{"notes":[{"content":{"submission_length":{"value":"Regular submission (no more than 12 pages of main content)"},"venue":{"value":"Accepted by TMLR"},"abstract":{"value":"Existing benchmarks for assessing the spatio-temporal understanding and reasoning abilities of video language models are susceptible to score inflation due to the presence of shortcut solutions based on superficial visual or textual cues. This paper mitigates the challenges in accurately assessing model performance by introducing the Minimal Video Pairs (MVP) benchmark, a simple shortcut-aware video QA benchmark for assessing the physical understanding of video language models. The benchmark is comprised of 55K high-quality multiple-choice video QA examples focusing on physical world understanding. Examples are curated from nine video data sources, spanning first-person egocentric and exocentric videos, robotic interaction data, and cognitive science intuitive physics benchmarks. To mitigate shortcut solutions that rely on superficial visual or textual cues and biases, each sample in MVP has a minimal-change pair — a visually similar video accompanied by an identical question but an opposing answer. To answer a question correctly, a model must provide correct answers for both examples in the minimal-change pair; as such, models that solely rely on visual or textual biases would achieve below random performance. Human performance on MVP is 92.9%, while the best open-source state-of-the- art video-language model achieves 40.2% compared to random performance at 25%."},"_bibtex":{"value":"@article{\nkrojer2025a,\ntitle={A Shortcut-aware Video-{QA} Benchmark for Physical Understanding via Minimal Video Pairs},\nauthor={Benno Krojer and Mojtaba Komeili and Candace Ross and Quentin Garrido and Koustuv Sinha and Nicolas Ballas and Mido Assran},\njournal={Transactions on Machine Learning Research},\nissn={2835-8856},\nyear={2025},\nurl={https://openreview.net/forum?id=gvFgNJcSw1},\nnote={}\n}"},"title":{"value":"A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs"},"pdf":{"value":"/pdf/f01dd06430bcd7b2aba068243d47011dcc96be6e.pdf"},"venueid":{"value":"TMLR"},"paperhash":{"value":"krojer|a_shortcutaware_videoqa_benchmark_for_physical_understanding_via_minimal_video_pairs"},"authorids":{"value":["~Benno_Krojer1","~Mojtaba_Komeili1","~Candace_Ross1","~Quentin_Garrido1","~Koustuv_Sinha1","~Nicolas_Ballas1","~Mido_Assran1"]},"assigned_action_editor":{"value":"~Tae-Hyun_Oh3"},"authors":{"value":["Benno Krojer","Mojtaba Komeili","Candace Ross","Quentin Garrido","Koustuv Sinha","Nicolas Ballas","Mido Assran"]}},"tmdate":1764585173082,"pdate":1764585173024,"tcdate":1752529561594,"writers":["TMLR"],"signatures":["TMLR/Paper5382/Authors"],"forum":"gvFgNJcSw1","license":"CC BY 4.0","number":5382,"cdate":1752529561594,"readers":["everyone"],"invitations":["TMLR/-/Submission","TMLR/-/Edit","TMLR/-/Under_Review","TMLR/Paper5382/-/Revision","TMLR/Paper5382/-/Camera_Ready_Revision","TMLR/-/Accepted"],"mdate":1764585173082,"odate":1753026882028,"domain":"TMLR","id":"gvFgNJcSw1","version":2},{"content":{"venue":{"value":"Video-Langauge Models Poster"},"pdf":{"value":"/pdf/cb240e31822630c8a8a6d43e475f1653b61857cc.pdf"},"keywords":{"value":["Vision","Language","Multimodal"]},"supplementary_material":{"value":"/attachment/24c6f01229c51f27a1b924b6cd0d390c0e5ebfb2.zip"},"venueid":{"value":"NeurIPS.cc/2024/Workshop/Video-Langauge_Models"},"paperhash":{"value":"mahmood|vurf_a_generalpurpose_reasoning_and_selfrefinement_framework_for_video_understanding"},"authorids":{"value":["~Ahmad_Mahmood1","~Ashmal_Vayani1","~Muzammal_Naseer1","~Salman_Khan4","~Fahad_Shahbaz_Khan1"]},"abstract":{"value":"Recent studies have demonstrated the effectiveness of Large Language Models (LLMs) as reasoning modules that can deconstruct complex tasks into more manageable sub-tasks, particularly when applied to visual reasoning tasks for images. In contrast, this paper introduces a Video Understanding and Reasoning Framework (VURF) based on the reasoning power of LLMs. Ours is a novel approach to extend the utility of LLMs in the context of video tasks, leveraging their capacity to generalize from minimal input and output demonstrations within a contextual framework. By presenting LLMs with pairs of instructions and their corresponding high-level programs, we harness their contextual learning capabilities to generate executable visual programs for video understanding.\nTo enhance program's accuracy and robustness, we implement two important strategies. Firstly, we employ a feedback-generation approach, powered by GPT-3.5, to rectify errors in programs utilizing unsupported functions. Secondly, taking motivation from recent works on self refinement of LLM outputs, we introduce an iterative procedure for improving the quality of the in-context examples by aligning the initial outputs to the outputs that would have been generated had the LLM not been bound by the structure of the in-context examples. Our results on several video-specific tasks, including visual QA, video anticipation, pose estimation and multi-video QA illustrate the efficacy of these enhancements in improving the performance of visual programming approaches for video tasks."},"_bibtex":{"value":"@inproceedings{\nmahmood2025vurf,\ntitle={{VURF}: A General-purpose Reasoning and Self-refinement Framework for Video Understanding},\nauthor={Ahmad Mahmood and Ashmal Vayani and Muzammal Naseer and Salman Khan and Fahad Shahbaz Khan},\nbooktitle={Workshop on Video-Language Models @ NeurIPS 2024},\nyear={2025},\nurl={https://openreview.net/forum?id=S92QnVEzQP}\n}"},"title":{"value":"VURF: A General-purpose Reasoning and Self-refinement Framework for Video Understanding"},"track":{"value":"Long Paper Track (up to 9 pages)"},"authors":{"value":["Ahmad Mahmood","Ashmal Vayani","Muzammal Naseer","Salman Khan","Fahad Shahbaz Khan"]}},"tmdate":1736861080305,"pdate":1730081752347,"tcdate":1725796759480,"writers":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission18/Authors"],"signatures":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission18/Authors"],"forum":"S92QnVEzQP","license":"CC BY 4.0","number":18,"cdate":1725796759480,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Post_Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission18/-/Full_Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Edit","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission18/-/Camera-Ready_Revision"],"mdate":1736861080305,"odate":1736861080293,"domain":"NeurIPS.cc/2024/Workshop/Video-Langauge_Models","id":"S92QnVEzQP","version":2},{"content":{"summary":{"value":"This work studies the instance-dependent regret guarantee in tabular Markov Decision Processes. The author focuses on the minimal sub-optimality gap structure and provides a logarithmic regret guarantee for two existing algorithms: UCB-Advantage and Q-EarlySettled-Advantage. Compared with previous instance-dependent guarantees, this work achieves a variance-aware regret bound that improves by a factor of H even under maximum variance. Additionally, when variance is low (e.g., deterministic transitions), the regret demonstrates improved dependency on the minimal sub-optimality gap. Furthermore, the author also proposes a gap-dependent policy-switching cost for the UCB-Advantage algorithm."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. In line 310, there is a typo fo$Q_h^k-Q_h^k$. \n\n2. In line 313, it seems questionable that the first term in G3 does not diminish to zero, while the regret should converge to zero as the episode k becomes sufficiently large.\n\n3. Lemma B.3 seems incorrect when $\\check{n}(s,a)=1$ immediately after a reset to 0.\n\n4. The $N(s,a)$ in Algorithm 1 should be $n(s,a)$."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"1. The author first proposes a novel algorithm that achieves a variance-aware regret bound with respect to the minimal sub-optimality gap.\n\n2. The author also proposes an instance-dependent policy-switching cost for the UCB-Advantage algorithm, which could be of independent interest.\n\n3. When variance is low (e.g., in deterministic transitions), the regret exhibits improved dependency on the minimal sub-optimality gap."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The main weakness is that the improvement in this work over existing results appears too limited.\n\n1. As discussed in line 264, the instance-dependent regret bound depends on the point-wise sub-optimality gap. In comparison, this work relies on the minimal sub-optimality gap across all state-action pairs. In most situations, the sub-optimality gap varies significantly across different state-action pairs, leading to a weaker performance in the regret guarantee presented in this work.\n\n2. For the instance-dependent guarantee with zero variance, this work achieves a sub-linear dependency on the sub-optimality gap. However, a similar result already exists without relying on the minimal sub-optimality gap assumption [1]. Compared with previous results, this work demonstrates worse dependency on the episode length H and the sub-optimality gap.\n[1] Sharp Variance-Dependent Bounds in Reinforcement Learning: Best of Both Worlds in Stochastic and Deterministic Environments\n\n3. Regarding the gap-dependent policy-switching cost, the claim in line 136 appears incorrect. When the optimal action set is small, the dominant term in equation (4) becomes the second term, resulting in an improvement of only \nlog T rather than a factor of A, which is minor.\n\n4. Regarding technical novelty, the author claims the introduction of a surrogate reference function; however, the importance of this reference function is not clearly explained in section 3.2. It would be helpful to further highlight its effect in the proof sketch."}},"nonreaders":[],"tmdate":1733214893956,"tcdate":1730598791434,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission953/Reviewer_Djdi"],"signatures":["ICLR.cc/2025/Conference/Submission953/Reviewer_Djdi"],"forum":"6tyPSkshtF","number":1,"license":"CC BY 4.0","cdate":1730598791434,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission953/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1733214893956,"domain":"ICLR.cc/2025/Conference","replyto":"6tyPSkshtF","id":"0hlGrRAPvs","forumContent":{"TLDR":{"value":"This paper analyzes the gap-dependent regrets and policy switching costs of two Q-Learning algorithms with variance reduction."},"venue":{"value":"ICLR 2025 Spotlight"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Reinforcement Learning","Q-Learning","Regret"]},"supplementary_material":{"value":"/attachment/db7bfbdbd30d841873c2053f544d8def8b09cf93.zip"},"primary_area":{"value":"reinforcement learning"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"We study the gap-dependent bounds of two important algorithms for on-policy $Q$-learning for finite-horizon episodic tabular Markov Decision Processes (MDPs): UCB-Advantage (Zhang et al. 2020) and Q-EarlySettled-Advantage (Li et al. 2021). UCB-Advantage and Q-EarlySettled-Advantage improve upon the results based on Hoeffding-type bonuses and achieve the {almost optimal} $\\sqrt{T}$-type regret bound in the worst-case scenario, where $T$ is the total number of steps. However, the benign structures of the MDPs such as a strictly positive suboptimality gap can significantly improve the regret. While gap-dependent regret bounds have been obtained for $Q$-learning with Hoeffding-type bonuses, it remains an open question to establish gap-dependent regret bounds for $Q$-learning using variance estimators in their bonuses and reference-advantage decomposition for variance reduction. We develop a novel error decomposition\nframework to prove gap-dependent regret bounds of UCB-Advantage and Q-EarlySettled-Advantage that are logarithmic in $T$ and improve upon existing ones for $Q$-learning algorithms. Moreover, we establish the gap-dependent bound for the policy switching cost of UCB-Advantage and improve that under the worst-case MDPs. To our knowledge, this paper presents the first gap-dependent regret analysis for $Q$-learning using variance estimators and reference-advantage decomposition and also provides the first gap-dependent analysis on policy switching cost for $Q$-learning."},"_bibtex":{"value":"@inproceedings{\nzheng2025gapdependent,\ntitle={Gap-Dependent Bounds for Q-Learning using Reference-Advantage Decomposition},\nauthor={Zhong Zheng and Haochen Zhang and Lingzhou Xue},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=6tyPSkshtF}\n}"},"title":{"value":"Gap-Dependent Bounds for Q-Learning using Reference-Advantage Decomposition"},"pdf":{"value":"/pdf/cf66213302f8b5edd07262254a0eb65b2b38001e.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"zheng|gapdependent_bounds_for_qlearning_using_referenceadvantage_decomposition"},"authorids":{"value":["~Zhong_Zheng3","~Haochen_Zhang8","~Lingzhou_Xue1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Zhong Zheng","Haochen Zhang","Lingzhou Xue"]}},"version":2},{"content":{"summary":{"value":"The paper proposes VIPO-R1, a training loop for video-reasoning MLLMs that alternates GRPO (online RL) with a rollout-aware verifier that curates contrastive CoT pairs, followed by DPO to refine the policy, then iterates. The verifier filters/labels rollouts using accuracy checks (e.g., Math-Verify or task metrics), reasoning–answer consistency checks, repetition and length checks, and builds several preference types (penalty, consistency, reflection). Compared to using GRPO alone, the loop aims to (i) stabilize optimization, (ii) lengthen CoTs without drifting, and (iii) reduce “right answer, wrong reasoning” inconsistencies. Experiments on VSI-Bench, Video-MMMU/MMMVU, TOMATO, and Video-MME report consistent gains over a Qwen2.5-VL-7B base and over strong baselines such as Video-R1/Kimi-VL-Thinking; the authors also highlight faster progress due to the DPO stage being ~7× cheaper per sample than GRPO."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. Which verifier LLM(s) are used in practice, and are they different from the base policy? Any evidence that results hold with multiple verifiers? \n\n2. How sensitive are results to the consistency/length/repetition thresholds and clustering settings used in data curation? Provide a sensitivity plot."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. Clear training recipe that combines known pieces in a useful way. The GRPO→Verifier→DPO cycle is well-motivated; the verifier’s multi-aspect filtering (accuracy, consistency, repetition, length) and construction of contrastive/reflection pairs is detailed and easy to reproduce at a high level.\n\n2. Empirical benefits across multiple video benchmarks, with especially large gains on hard video-reasoning tasks (e.g., +7.9% over GRPO on VSI-Bench; +5.6% over Video-R1 on Video-MMMU, as reported in the text) and reductions in reasoning-answer inconsistency. Case studies qualitatively show fewer spurious steps and better self-correction."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Verifier dependency & potential circularity:\\\nSeveral measurements of “consistency” and sample selection rely on LLM-based verification. If the same (or closely related) verifier is also used to curate training pairs, the evaluation may inherit its biases. A stronger case would use (i) external, task-specific non-LLM checks where possible and (ii) a held-out verifier for reporting consistency.\n\n2. Metric choice.\\\nThe paper emphasizes response length and Acc-Cons. Length is an input-dependent proxy; a human study or task-specific rubric (e.g., step validity on math/physics cases) would better validate that longer CoTs are meaningfully better. Suggest adding human evaluation on the necessity of the lengths.\n\n3. Suggested references.\\\nAuthors are also encouraged to discuss and reference the following models:\\\n[1] Wang et al. \"Video-RTS: Rethinking Reinforcement Learning and Test-Time Scaling for Efficient and Enhanced Video Reasoning\", https://arxiv.org/abs/2507.06485 \\\n[2] Sun et al. \"video-SALMONN-o1: Reasoning-enhanced Audio-visual Large Language Model\", https://arxiv.org/abs/2502.11775"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762943283438,"tcdate":1761342761768,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission25012/Reviewer_pKWZ"],"signatures":["ICLR.cc/2026/Conference/Submission25012/Reviewer_pKWZ"],"forum":"Zyy2wbKd8h","number":1,"license":"CC BY 4.0","cdate":1761342761768,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission25012/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762943283438,"domain":"ICLR.cc/2026/Conference","replyto":"Zyy2wbKd8h","id":"Me4N9PZIF1","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["video understanding","video question answering"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Applying Reinforcement Learning (RL) to Multimodal Large Language Models (MLLMs) shows significant promise for complex video reasoning. However, popular Reinforcement Fine-Tuning (RFT) methods, such as outcome-based Group Relative Policy Optimization (GRPO), are limited by data preparation bottlenecks (e.g., noise or high cost) and exhibit unstable improvements in the quality of long chain-of-thoughts (CoTs) and downstream performance. To address these limitations, we propose **VIPO-R1**, a **V**erifier-guided **I**terative **P**olicy **O**ptimization method designed to gradually enhance MLLMs' ability to generate long-term reasoning chains for challenging VideoQA. The core component is the Rollout-Aware Verifier, positioned between the GRPO and Direct Preference Optimization (DPO) training phases to form the GRPO-Verifier-DPO training loop. This verifier leverages small LLMs as a judge to assess the reasoning logic of rollouts, enabling the construction of high-quality contrastive data, including reflective and contextually consistent CoTs. These curated preference samples drive the efficient DPO stage (7x faster than GRPO), leading to marked improvements in reasoning chain quality, especially in terms of length and contextual consistency. This training loop benefits from GRPO's expansive search and DPO's targeted optimization. Experimental results demonstrate: 1) Faster and more effective optimization compared to standard GRPO variants, yielding superior performance; 2) Our trained models exceed the direct inference of large-scale instruction-tuned Video-LLMs, producing long and contextually consistent CoTs on diverse video reasoning tasks; and 3) Our model with one iteration outperforms powerful MLLMs (e.g., Kimi-VL) and thinking models (e.g., Video-R1), highlighting its effectiveness and stability."},"_bibtex":{"value":"@misc{\nli2026vipor,\ntitle={{VIPO}-R1: Cultivating Video Reasoning in {MLLM}s via Verifier-Guided Iterative Policy Optimization},\nauthor={yunxin li and Xinyu Chen and Zitao Li and zhenyu liu and Longyue Wang and Wenhan Luo and Baotian Hu and Min Zhang},\nyear={2026},\nurl={https://openreview.net/forum?id=Zyy2wbKd8h}\n}"},"title":{"value":"VIPO-R1: Cultivating Video Reasoning in MLLMs via Verifier-Guided Iterative Policy Optimization"},"pdf":{"value":"/pdf/6430a59edddcb9b8d53ba3e7b55fa2694a5f8e01.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"li|vipor1_cultivating_video_reasoning_in_mllms_via_verifierguided_iterative_policy_optimization"},"authorids":{"value":["~yunxin_li1","~Xinyu_Chen6","~Zitao_Li2","~zhenyu_liu4","~Longyue_Wang3","~Wenhan_Luo1","~Baotian_Hu1","~Min_Zhang9"]},"authors":{"value":["yunxin li","Xinyu Chen","Zitao Li","zhenyu liu","Longyue Wang","Wenhan Luo","Baotian Hu","Min Zhang"]}},"version":2},{"content":{"summary":{"value":"This paper introduces Chimera, a large-scale benchmark for evaluating visual-language models (VLMs) on diagram comprehension. The authors argue that current benchmarks overestimate model understanding because they fail to detect shortcut behaviors.  \nCHIMERA consists of 7,500 Wikipedia diagrams annotated with semantic triples and four levels of comprehension questions (entity recognition, relation understanding, knowledge grounding, and visual reasoning). Using this benchmark, the authors evaluate 15 open-weights VLMs across 7 model families, analyzing three shortcut types: visual-memorization, knowledge-recall, and Clever-Hans.  \nTheir findings suggest that Clever-Hans shortcuts significantly influence model performance, while visual-memorization and knowledge-recall shortcuts have smaller to moderate effects."},"soundness":{"value":1},"confidence":{"value":5},"questions":{"value":"**Questions for the Authors**\n\n1. Were any of the evaluated models fine-tuned on CHIMERA? If yes, please describe the training procedure and specify whether comparisons include both fine-tuned and zero-shot results.\n2. What was the inter-annotator agreement for the 300 manually evaluated diagrams?\n\n**Actionable Feedback**\n\n1. Conduct a human study to verify that ER questions are indeed easier than the other three tasks. This would strengthen claims about knowledge-recall shortcuts.\n2. Provide standard deviations or confidence intervals to assess reliability. Consider analyzing fully sufficient vs partially insufficient triples separately.\n3. Discuss potential biases introduced by automatically generated questions (Gemini) and how they may influence shortcut detection.\n4. Consider reorganizing the manuscript to separate introduction, related work, dataset description, and results for better readability and clarity."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"- Addresses an important gap in multimodal reasoning evaluation, distinguishing true comprehension from shortcut exploitation.\n\n- Proposes a structured framework for diagram understanding grounded in semiotic theory.\n    \n- Builds a large-scale dataset (Chimera) with hierarchical question design, enhancing diagnostic evaluation.\n    \n- Evaluates a broad range of open-weights VLMs, offering valuable comparative insights.\n    \n- The identification of three shortcut types provides a clear conceptual taxonomy."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- **Human evaluation and visual dependency**: The authors conduct a human evaluation on 300 diagrams to assess visual dependency, QA correctness, and triple completeness. However, they do not report inter-annotator agreement, and standard deviations across categories suggest significant variability. This raises concerns about the reliability of the visual-dependency measure, which is critical for assessing the Clever-Hans shortcut.\n    \n- **Ambiguity in fine-tuning**: The paper does not specify whether the 15 evaluated models were fine-tuned on CHIMERA. Fine-tuning could strongly influence performance and shortcut measurement, while zero-shot evaluation might reveal different behaviors.\n    \n- **Statistical rigor in shortcut measurement**: The visual-memorization shortcut is claimed based on a small mean difference (~2%) between original diagrams and visualized triples. No standard deviations or significance tests are reported. Given that human evaluation found triples to be only 74–86% fully sufficient, this small difference may reflect dataset quality rather than model memorization.\n    \n- **Task difficulty not human-validated**: Knowledge-recall shortcuts are inferred from ER (Entity Recognition) vs other tasks, assuming ER questions should be easier to respond for a VLM. However, no human study confirms relative difficulty. Automatically generated questions may introduce artifacts, so the 5% difference may not reliably indicate shortcut behavior. This can be assessed with a human evaluation which answer a sample of questions. If human performance doesn't align with VLM results on ER vs other tasks, it could reveal the influence of the knowledge-recall shortcut.\n    \n- **Clever-Hans overstatement**: Clever-Hans shortcuts are measured only on ER questions. Larger models show minimal differences when diagrams are removed, suggesting that the observed effect is primarily driven by smaller models. Therefore, the claim that all VLMs suffer from Clever-Hans may overgeneralize.\n    \n- **Unusual paper structure**: The paper merges introduction, related work, dataset description, and results into a single section, which makes it harder to clearly follow the flow of motivation, prior work, and methodology."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1764358395533,"tcdate":1761662723526,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission13473/Reviewer_JpPG"],"signatures":["ICLR.cc/2026/Conference/Submission13473/Reviewer_JpPG"],"forum":"q3eB3PhtqD","number":1,"license":"CC BY 4.0","cdate":1761662723526,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission13473/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1764358395533,"domain":"ICLR.cc/2026/Conference","replyto":"q3eB3PhtqD","id":"LPjBJZSV0s","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["vision-language model","diagram understanding","multimodal reasoning","visual question answering","shortcut learning","knowledge grounding","benchmark dataset","multimodal evaluation"]},"primary_area":{"value":"interpretability and explainable AI"},"abstract":{"value":"Diagrams convey symbolic information in a visual format rather than a linear stream of words, making them especially challenging for AI models to process. While recent evaluations suggest that vision-language models (VLMs) perform well on diagram-related benchmarks, their reliance on knowledge, reasoning, or modality shortcuts raises concerns about whether they genuinely understand and reason over diagrams.\nTo address this gap, we introduce Chimera, a comprehensive test suite comprising 7,500 high-quality diagrams sourced from Wikipedia; each diagram is annotated with its symbolic content represented by semantic triples along with multi-level questions designed to assess four fundamental aspects of diagram comprehension: entity recognition, relation understanding, knowledge grounding, and visual reasoning.\nWe use Chimera to measure the presence of three types of shortcuts in visual question answering: \n(1) the visual-memorization shortcut, where VLMs rely on memorized visual patterns;\n(2) the knowledge-recall shortcut, where models leverage memorized factual knowledge instead of interpreting the diagram; and\n(3) the Clever-Hans shortcut, where models exploit superficial language patterns or priors without true comprehension. We evaluate 15 open-source VLMs from 7 model families on Chimera and find that their seemingly strong performance largely stems from shortcut behaviors: visual-memorization shortcuts have slight impact, knowledge-recall shortcuts play a moderate role, and Clever-Hans shortcuts contribute significantly.\nThese findings expose critical limitations in current VLMs and underscore the need for more robust evaluation protocols that benchmark genuine comprehension of complex visual inputs (e.g., diagrams) rather than question-answering shortcuts."},"_bibtex":{"value":"@misc{\nchi2026chimera,\ntitle={Chimera: Diagnosing Shortcut Learning in Visual-Language Understanding},\nauthor={Ziheng Chi and Yifan Hou and Chenxi Pang and Shaobo Cui and Mubashara Akhtar and Mrinmaya Sachan},\nyear={2026},\nurl={https://openreview.net/forum?id=q3eB3PhtqD}\n}"},"title":{"value":"Chimera: Diagnosing Shortcut Learning in Visual-Language Understanding"},"pdf":{"value":"/pdf/f92e46808334228be1f33c8c27812a84e924f889.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"chi|chimera_diagnosing_shortcut_learning_in_visuallanguage_understanding"},"authorids":{"value":["~Ziheng_Chi1","~Yifan_Hou1","~Chenxi_Pang1","~Shaobo_Cui1","~Mubashara_Akhtar1","~Mrinmaya_Sachan3"]},"authors":{"value":["Ziheng Chi","Yifan Hou","Chenxi Pang","Shaobo Cui","Mubashara Akhtar","Mrinmaya Sachan"]}},"version":2},{"content":{"summary":{"value":"This paper proposes Self-Discriminative Optimization (SDO) to fine-tune video diffusion models. SDO first yielding degrades real samples by frequency-domain reweighting, and then uses these real/degraded pairs as\npositive and negative examples of DPO to optimize vido diffusion models. SDO demonstrate substantial gains\nin structural quality and semantic alignment than DPO and LoRA fine-tuning."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"V-bench is unable to evaluate the motion extent of generated videos. The  static videos usually get high score. How do authors measure the quality of video motion?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"1. This present clear motivation and the corresponding method  to deal  with the problem  of fine-tuning video diffusion models. \n\n2.  Self-degradation is a simple but effective method to generate positive and negative paris. \n\n3. Experimental results demonstrate SDO achieves superior performance than the existing\npost-training methods, with only a handful of high-quality\nsamples and minimal fine-tuning"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The generation of positive and negative pairs is core step of DPO. Although self-degradation is a effective method to yield these pairs, other approaches using reward model, e.g.,MLLM, ImageReward,HPSv2,  to classify the generative images as positive and negative samples   can also achieve this target. This process maybe cost more time, but many acceleration methods are able to quickly generate samples. I would like to see such comparison in the paper.\n\n2. The authors claim they use minimal fine-tuning cost, but the comparison with respect to  training hours and traning data is missing.\n\n3. How does SDO determine the papameters of self-degradation? What  if high-frequency components are reserved?"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762921419473,"tcdate":1762584548596,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission9992/Reviewer_qMKW"],"signatures":["ICLR.cc/2026/Conference/Submission9992/Reviewer_qMKW"],"forum":"I4jBCglUOI","number":4,"license":"CC BY 4.0","cdate":1762584548596,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission9992/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762921419473,"domain":"ICLR.cc/2026/Conference","replyto":"I4jBCglUOI","id":"EbKgLC3cqC","forumContent":{"venue":{"value":"ICLR 2026 Conference Desk Rejected Submission"},"keywords":{"value":["Video Generation; Post-training; Diffusion Models"]},"supplementary_material":{"value":"/attachment/4a0027ff1916872daebb868d488c775154c3baaa.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Recent preference alignment strategies have gained traction in large language models (LLMs) and are now being extended to broader generative domains. Approaches such as Direct Preference Optimization have been adapted to diffusion models by leveraging human-labeled preferences or auxiliary score models to distinguish ``winner'' from ``loser''. \nHowever, these methods face two key challenges: (1) the optimization process often overfits to the score model, resulting in suboptimal generation quality; and (2) the results generated from the same text prompt exhibit significant divergence, resulting in limited effective gradients and reduced training efficiency. These limitations are further exacerbated in video generation, where evaluation is more complex and inference is slower. In this work, we introduce Self-Discriminative Optimization that using only a handful of real samples, unlocks markedly higher-quality generation. First, we introduce self-degradation that applies frequency-domain reweighting to the latent representations from real samples, yielding degraded samples that more closely match the model’s original output distribution. This leads to controlled distortions such as low-quality, temporal inconsistency and object deformation.\nWe then use these real/degraded pairs as positive and negative examples to fine-tune the pretrained model discriminatively with automatically assigned, reliable labels. \nBy exploiting the richer gradients from these controllable degradation pairs, our experiments demonstrate substantial gains in structural quality and semantic alignment using only a handful of high-quality samples and minimal fine-tuning."},"_bibtex":{"value":"@misc{\nanonymous2026selfdiscriminative,\ntitle={Self-Discriminative Optimization for Video Diffusion Models},\nauthor={Anonymous},\nyear={2026},\nurl={https://openreview.net/forum?id=I4jBCglUOI}\n}"},"title":{"value":"Self-Discriminative Optimization for Video Diffusion Models"},"pdf":{"value":"/pdf/2c3b65a573548517220076d4f075149428b867a1.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Desk_Rejected_Submission"},"paperhash":{"value":"guo|selfdiscriminative_optimization_for_video_diffusion_models"},"authorids":{"value":["~Lanqing_Guo1","~Yichen_Liu17","~Yufei_Wang5","~Yingqing_He1","~Yan_Zheng4","~Hezhen_Hu2","~Junbo_Li3","~Xiaoyu_Wang1","~Zhangyang_Wang1"]},"authors":{"value":["Lanqing Guo","Yichen Liu","Yufei Wang","Yingqing He","Yan Zheng","Hezhen Hu","Junbo Li","Xiaoyu Wang","Zhangyang Wang"]}},"version":2},{"content":{"venue":{"value":"Video-Langauge Models Oral"},"pdf":{"value":"/pdf/93aef18fc1c56d2182f6cf1031598dc4ed7e7e9d.pdf"},"keywords":{"value":["video","benchmark","multimodal"]},"venueid":{"value":"NeurIPS.cc/2024/Workshop/Video-Langauge_Models"},"paperhash":{"value":"cai|temporalbench_benchmarking_finegrained_temporal_understanding_for_multimodal_video_models"},"authorids":{"value":["~Mu_Cai1","~Reuben_Tan1","~Jianrui_Zhang1","~Bocheng_Zou1","~Kai_Zhang10","~Yao_Feng2","~Fangrui_Zhu1","~Jing_Gu2","~Yiwu_Zhong1","~Yuzhang_Shang1","~Yao_Dou1","~Jaden_Park1","~Jianfeng_Gao1","~Yong_Jae_Lee2","~Jianwei_Yang1"]},"abstract":{"value":"Understanding fine-grained temporal dynamics is crucial for video understanding. Yet, popular video benchmarks, such as MSRVTT and TGIF, often fail to effectively evaluate AI models' temporal reasoning abilities due to the lack of fine-grained temporal annotations. As a result, text-based models, leveraging strong language priors, often perform comparably to video models, and image-trained models have been reported to outperform their video-trained counterparts on MSRVTT and TGIF. This paper introduces a new TemporalBench benchmark for fine-grained temporal event understanding in videos. TemporalBench, sourced from a diverse video datasets, consists of ∼10K pairs of video description questions, derived from ∼2K high-quality human-annotated video captions. Uniquely, our benchmark provides fine-grained temporal annotations to evaluate models' temporal reasoning abilities. Our results show that state-of-the-art models like GPT-4o achieve only 38.0% multiple binary QA accuracy on TemporalBench, demonstrating a significant human-AI gap in temporal understanding. We hope that TemporalBench is instrumental to fostering research on improving models' temporal reasoning capabilities."},"_bibtex":{"value":"@inproceedings{\ncai2025temporalbench,\ntitle={TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models},\nauthor={Mu Cai and Reuben Tan and Jianrui Zhang and Bocheng Zou and Kai Zhang and Yao Feng and Fangrui Zhu and Jing Gu and Yiwu Zhong and Yuzhang Shang and Yao Dou and Jaden Park and Jianfeng Gao and Yong Jae Lee and Jianwei Yang},\nbooktitle={Workshop on Video-Language Models @ NeurIPS 2024},\nyear={2025},\nurl={https://openreview.net/forum?id=bhuZn4URBH}\n}"},"title":{"value":"TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models"},"track":{"value":"Long Paper Track (up to 9 pages)"},"authors":{"value":["Mu Cai","Reuben Tan","Jianrui Zhang","Bocheng Zou","Kai Zhang","Yao Feng","Fangrui Zhu","Jing Gu","Yiwu Zhong","Yuzhang Shang","Yao Dou","Jaden Park","Jianfeng Gao","Yong Jae Lee","Jianwei Yang"]}},"tmdate":1736861081088,"pdate":1730081753606,"tcdate":1728932798357,"writers":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission56/Authors"],"signatures":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission56/Authors"],"forum":"bhuZn4URBH","license":"CC BY 4.0","number":56,"cdate":1728932798357,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Post_Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Edit","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission56/-/Camera-Ready_Revision"],"mdate":1736861081088,"odate":1736861081072,"domain":"NeurIPS.cc/2024/Workshop/Video-Langauge_Models","id":"bhuZn4URBH","version":2},{"content":{"summary":{"value":"This paper addresses the challenge of developing large multimodal models (LMMs) for video understanding, which has been limited by the scarcity of high-quality training data. To overcome this, the authors introduce LLaVA-Video-178K, a synthetic dataset designed for video instruction-following tasks such as detailed captioning, open-ended question-answering (QA), and multiple-choice QA. By training on this dataset alongside existing visual instruction tuning data, they develop LLaVA-Video, a new video LMM. Experiments demonstrate that LLaVA-Video performs strongly across various video benchmarks, underscoring the effectiveness of the synthetic dataset."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"- Could you elaborate on how your approach to using GPT-4o for generating video captions and QA pairs differs from prior work, such as ShareGPT4Video?\n- Additionally, would you consider exploring or providing insights on how non-GPT-based annotations might influence the model’s performance or diversity in understanding?\n- Your hierarchical captioning approach is similar to what has been previously proposed, such as in Video ReCap: Recursive Captioning of Hour-Long Videos (CVPR 2024). How does your method differ conceptually or practically from this prior work? Please clarify if I may have overlooked novel aspects in your captioning pipeline.\n- EgoSchema contains the videos from Ego4D, so the training set or even test set (not sure) can be observed in your training data mixture."},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- The recurrent, multi-level annotation strategy for generating detailed captions and question-answer pairs is effective.\n- The authors have conducted extensive experiments across multiple video benchmarks, demonstrating the effectiveness of the proposed LLaVA-Video models. The use of diverse datasets and thorough ablation studies supports the robustness of the results.\n- The paper is well-structured, clearly explaining the problem, methodology, and experimental design.\n- The significance of this work lies in its potential to advance video-language understanding models by providing a high-quality, open-source synthetic dataset. LLaVA-Video-178K has broad applicability in tasks such as video captioning and question-answering."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The primary concern lies in the heavy reliance on GPT-4o for video captioning and question-answer (QA) pair generation. This approach essentially distills GPT-4o’s video capabilities into a structured format, raising questions about the originality of the contribution. Moreover, This has been done in previous work: ShareGPT4Video: improving Video Understanding and Generation with Better Captions, NeurIPS 2024 D&B Track. This work is an incremental improvement over prior work by scaling the data from 40K QA pairs to about 1.3M QA pairs. \n- The hierarchical captioning strategy employed in the paper is not new. The concept of recursive or hierarchical video captioning has been previously explored, for instance, in Video ReCap: Recursive Captioning of Hour-Long Videos (CVPR 2024). The failure to cite and differentiate from this work undermines the perceived novelty.\n- The inclusion of multi-choice QA pairs appears to be tailored primarily for fitting into existing evaluation benchmarks rather than reflecting practical, real-world video understanding scenarios. This raises concerns about the broader utility of these QA pairs. \n- The model’s strong performance is largely due to fine-tuning from a powerful base model, LLaVA-OneVision. While the experimental results are compelling, the paper’s core contributions are somewhat overshadowed by the reliance on this pre-trained foundation."}},"nonreaders":[],"tmdate":1731427614849,"tcdate":1730718286598,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission2539/Reviewer_kqcW"],"signatures":["ICLR.cc/2025/Conference/Submission2539/Reviewer_kqcW"],"forum":"8Livf4oZxz","number":3,"license":"CC BY 4.0","cdate":1730718286598,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission2539/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427614849,"domain":"ICLR.cc/2025/Conference","replyto":"8Livf4oZxz","id":"ZWXtX1q7BN","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"TLDR":{"value":"The largest high-quality synthetic dataset specifically for video instruction-following"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video instruction dataset","video-language model"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"The development of video large multimodal models (LMMs) has been hindered by the difficulty of curating large amounts of high-quality raw data from the web. To address this, we consider an alternative approach, creating a high-quality synthetic dataset specifically for video instruction-following, namely LLaVA-Video-178K. This dataset includes key tasks such as detailed captioning, open-ended question-answering (QA), and multiple-choice QA. By training on this proposed dataset, in combination with existing visual instruction tuning data, we introduce LLaVA-Video, a new video LMM. Our experiments demonstrate that LLaVA-Video achieves strong performance across various video benchmarks, highlighting the effectiveness of our dataset. We plan to release the dataset, its generation pipeline, and the model checkpoints."},"_bibtex":{"value":"@misc{\nzhang2024video,\ntitle={Video Instruction Tuning with Synthetic Data},\nauthor={Yuanhan Zhang and Jinming Wu and Wei Li and Bo Li and Zejun MA and Ziwei Liu and Chunyuan Li},\nyear={2024},\nurl={https://openreview.net/forum?id=8Livf4oZxz}\n}"},"title":{"value":"Video Instruction Tuning with Synthetic Data"},"pdf":{"value":"/pdf/98c983fa704b94de7f370bfa5fcf2b7e8e7db9e8.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"zhang|video_instruction_tuning_with_synthetic_data"},"authorids":{"value":["~Yuanhan_Zhang1","~Jinming_Wu1","~Wei_Li78","~Bo_Li23","~Zejun_MA1","~Ziwei_Liu1","~Chunyuan_Li1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Yuanhan Zhang","Jinming Wu","Wei Li","Bo Li","Zejun MA","Ziwei Liu","Chunyuan Li"]}},"version":2},{"content":{"venue":{"value":"EMNLP 2024"},"pdf":{"value":"https://aclanthology.org/2024.emnlp-main.522.pdf"},"venueid":{"value":"dblp.org/conf/EMNLP/2024"},"paperhash":{"value":"taktasheva|rublimp_russian_benchmark_of_linguistic_minimal_pairs"},"authorids":{"value":["~Ekaterina_Taktasheva1","https://dblp.org/search/pid/api?q=author:Maxim_Bazhukov:","https://dblp.org/search/pid/api?q=author:Kirill_Koncha:","https://dblp.org/search/pid/api?q=author:Alena_Fenogenova:","https://dblp.org/search/pid/api?q=author:Ekaterina_Artemova:","https://dblp.org/search/pid/api?q=author:Vladislav_Mikhailov:"]},"html":{"value":"https://aclanthology.org/2024.emnlp-main.522"},"_bibtex":{"value":"@inproceedings{DBLP:conf/emnlp/TaktashevaBKFAM24,\n  author={Ekaterina Taktasheva and Maxim Bazhukov and Kirill Koncha and Alena Fenogenova and Ekaterina Artemova and Vladislav Mikhailov},\n  title={RuBLiMP: Russian Benchmark of Linguistic Minimal Pairs},\n  year={2024},\n  cdate={1704067200000},\n  pages={9268-9299},\n  url={https://aclanthology.org/2024.emnlp-main.522},\n  booktitle={EMNLP},\n  crossref={conf/emnlp/2024}\n}\n"},"abstract":{"value":"Minimal pairs are a well-established approach to evaluating the grammatical knowledge of language models. However, existing resources for minimal pairs address a limited number of languages and lack diversity of language-specific grammatical phenomena. This paper introduces the Russian Benchmark of Linguistic Minimal Pairs (RuBLiMP), which includes 45k pairs of sentences that differ in grammaticality and isolate a morphological, syntactic, or semantic phenomenon. In contrast to existing benchmarks of linguistic minimal pairs, RuBLiMP is created by applying linguistic perturbations to automatically annotated sentences from open text corpora and decontaminating test data. We describe the data collection protocol and present the results of evaluating 25 language models in various scenarios. We find that the widely used LMs for Russian are sensitive to morphological and agreement-oriented contrasts, but fall behind humans on phenomena requiring the understanding of structural relations, negation, transitivity, and tense. RuBLiMP, the codebase, and other materials are publicly available."},"title":{"value":"RuBLiMP: Russian Benchmark of Linguistic Minimal Pairs"},"authors":{"value":["Ekaterina Taktasheva","Maxim Bazhukov","Kirill Koncha","Alena Fenogenova","Ekaterina Artemova","Vladislav Mikhailov"]}},"tmdate":1739908555869,"pdate":1704067200000,"tcdate":1739908554597,"writers":["~"],"signatures":["~Ekaterina_Taktasheva1"],"forum":"YILBmoBv4D","license":"CC BY-SA 4.0","number":325353,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1739908555869,"domain":"DBLP.org","id":"YILBmoBv4D","version":2},{"content":{"venue":{"value":"CoRR 2024"},"pdf":{"value":"http://arxiv.org/pdf/2406.19232v3"},"venueid":{"value":"dblp.org/journals/CORR/2024"},"paperhash":{"value":"taktasheva|rublimp_russian_benchmark_of_linguistic_minimal_pairs"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Ekaterina_Taktasheva:","https://dblp.org/search/pid/api?q=author:Maxim_Bazhukov:","https://dblp.org/search/pid/api?q=author:Kirill_Koncha:","~Alena_Fenogenova1","https://dblp.org/search/pid/api?q=author:Ekaterina_Artemova:","https://dblp.org/search/pid/api?q=author:Vladislav_Mikhailov:"]},"html":{"value":"https://doi.org/10.48550/arXiv.2406.19232"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2406-19232,\n  publtype={informal},\n  author={Ekaterina Taktasheva and Maxim Bazhukov and Kirill Koncha and Alena Fenogenova and Ekaterina Artemova and Vladislav Mikhailov},\n  title={RuBLiMP: Russian Benchmark of Linguistic Minimal Pairs},\n  year={2024},\n  cdate={1704067200000},\n  journal={CoRR},\n  volume={abs/2406.19232},\n  url={https://doi.org/10.48550/arXiv.2406.19232}\n}\n"},"abstract":{"value":"Minimal pairs are a well-established approach to evaluating the grammatical knowledge of language models. However, existing resources for minimal pairs address a limited number of languages and lack diversity of language-specific grammatical phenomena. This paper introduces the Russian Benchmark of Linguistic Minimal Pairs (RuBLiMP), which includes 45k pairs of sentences that differ in grammaticality and isolate a morphological, syntactic, or semantic phenomenon. In contrast to existing benchmarks of linguistic minimal pairs, RuBLiMP is created by applying linguistic perturbations to automatically annotated sentences from open text corpora and carefully curating test data. We describe the data collection protocol and present the results of evaluating 25 language models in various scenarios. We find that the widely used language models for Russian are sensitive to morphological and agreement-oriented contrasts but fall behind humans on phenomena requiring understanding of structural relations, negation, transitivity, and tense. RuBLiMP, the codebase, and other materials are publicly available."},"title":{"value":"RuBLiMP: Russian Benchmark of Linguistic Minimal Pairs"},"authors":{"value":["Ekaterina Taktasheva","Maxim Bazhukov","Kirill Koncha","Alena Fenogenova","Ekaterina Artemova","Vladislav Mikhailov"]}},"tmdate":1728371774343,"pdate":1704067200000,"tcdate":1728371768191,"writers":["~"],"signatures":["~Alena_Fenogenova1"],"forum":"EDn1yimnxV","license":"CC BY-SA 4.0","number":144801,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1728371774343,"domain":"DBLP.org","id":"EDn1yimnxV","version":2},{"content":{"summary":{"value":"This paper introduces a reinforcement learning-based approach to task-aware video compression. The proposed method learns an RL agent to decide a QP map for each video frame coded by x264 in order to achieve task-aware video compression. Notably, the state signals are the MB-tree statistics with 10-frame look-ahead."},"soundness":{"value":1},"confidence":{"value":5},"questions":{"value":"(1) In Section 3, it is unclear what it means by “D is the precision between f(frame_rec) and f(frame_raw)”. Do you measure the distortion between frame_rec and frame_raw in the latent domain of the recognition network? Or this refers to the precision difference resulting from the replacement of the original image with the compressed image. If the task is ROI-based coding, how is D evaluated? \n\n(2) About the macroblock-level reward, it is unclear how the r_bit_rate is evaluated for one specific macroblock. \n\n(3) The action space is all the MB QP deltas within a frame, and is reduced by downsampling. I wonder if the agent can generalize to processing videos of different resolutions. \n\n(4) By looking at Tables 5 and 6 in the appendix, I realize that it is very costly to get the state signals since collecting some of them may require pre-encoding and pre-decoding the input video, e.g. intra encoding cost per MB, x264 calculated PSNR and SSIM, percentages of bits used for MV, etc. It is also mentioned that the MB-tree statistics were collected by allowing x264 to use look-ahead for 10 frames. In this case, how come the evaluation could achieve 30FPS? \n\n(5) The training and test split is kind of weird: 65 streams for training and 100 for testing. \n\n(6) In terms of PSNR, it is unclear why the rate-distortion performance of RL-RC-DoT is comparable to that of x264. From your training objective, there is no guarantee that this would be true.\n\n(7) Likewise, there appears to be no guarantee that the proposed agent would produce generalizable decoded videos. \n\n(8) It is unclear whether separate agents need to be trained in order to produce multiple rate-distortion points.  \n\n(9) For the ROI task, it appears that the agent must be very capable in order to predict the weighted sum of r_bit_rate and r_DT and to identify the ROI regions, or the state signal must be very informative. Is the same network architecture used across different tasks? Fig. 6 alone is not conclusive.\n\n(10)\tThere are missing references. \nHo et al., \"Neural Frank-Wolfe Policy Optimization for Region-of-Interest Intra-Frame Coding with HEVC/H.265,\"  VCIP’22.\nHo et al., \"A Dual-Critic Reinforcement Learning Framework for Frame-level Bit Allocation in HEVC/H.265 ,\" DCC'21."},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":1},"contribution":{"value":1},"strengths":{"value":"(1) The idea of using RL to perform task-aware video compression is interesting. \n\n(2) The proposed method is tested on two tasks: object detection and ROI-based coding."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"(1) The paper is lengthy and verbose. Lots of space is used to introduce some very basic concepts of video coding and quality metrics. \n\n(2) It appears that collecting the state signals is very costly. But, this part is not addressed at all. \n\n(3) The idea of introducing an RL agent for task-aware video compression is NOT new. \n\n(4) There are a few missing references. \n\n(5) There are many technical details that need further clarification."}},"nonreaders":[],"tmdate":1731428587816,"tcdate":1730703471411,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission7794/Reviewer_a2Py"],"signatures":["ICLR.cc/2025/Conference/Submission7794/Reviewer_a2Py"],"forum":"aQ7qYnY2nF","number":1,"license":"CC BY 4.0","cdate":1730703471411,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission7794/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428587816,"domain":"ICLR.cc/2025/Conference","replyto":"aQ7qYnY2nF","id":"VvVEeGsI6U","forumContent":{"TLDR":{"value":"RL based rate block-level control for task-aware video compression"},"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video compression","Rate control","Reinforcement Learning","Downstream task"]},"primary_area":{"value":"reinforcement learning"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Video encoders optimize compression for human perception by minimizing reconstruction error under bit-rate constraints. In many modern applications such as autonomous driving, an overwhelming majority of videos serve as input for AI systems performing tasks like object recognition or segmentation, rather than being watched by humans. It is therefore useful to optimize the encoder for a  downstream task instead of for perceptual image quality. However, a major challenge is how to combine such downstream optimization with existing standard video encoders, which are highly efficient and popular. Here, we address this challenge by controlling the Quantization Parameters (QPs) at the macro-block level to optimize the downstream task. This granular control allows us to prioritize encoding for task-relevant regions within each frame. We formulate this optimization problem as a Reinforcement Learning (RL) task, where the agent learns to balance long-term implications of choosing QPs on both task performance and bit-rate constraints. Notably, our policy does not require the downstream task as an input during inference, making it suitable for streaming applications and edge devices such as vehicles. We demonstrate significant improvements in two tasks, car detection, and ROI (saliency) encoding. Our approach improves task performance for a given bit rate compared to traditional task agnostic encoding methods, paving the way for more efficient task-aware video compression."},"_bibtex":{"value":"@misc{\ngadot2024real,\ntitle={Real Time Macro-Block Rate Control for Task-Aware Video Compression Using Reinforcement Learning},\nauthor={Uri Gadot and Assaf Shocher and Shie Mannor and Gal Chechik and Assaf Hallak},\nyear={2024},\nurl={https://openreview.net/forum?id=aQ7qYnY2nF}\n}"},"title":{"value":"Real Time Macro-Block Rate Control for Task-Aware Video Compression Using Reinforcement Learning"},"pdf":{"value":"/pdf/167170404677fb2439d1af258e70338b5af211dc.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"gadot|real_time_macroblock_rate_control_for_taskaware_video_compression_using_reinforcement_learning"},"authorids":{"value":["~Uri_Gadot1","~Assaf_Shocher1","~Shie_Mannor2","~Gal_Chechik1","~Assaf_Hallak1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Uri Gadot","Assaf Shocher","Shie Mannor","Gal Chechik","Assaf Hallak"]}},"version":2},{"content":{"summary":{"value":"The paper studies shortcut learning, i.e., why networks use to learn shortcut/spurious features over intended semantic features. The authors suggest that the network prefers a feature that is more available to quantify shortcut bias in terms of how a learned classifier deviates from an optimal classifier in its feature reliance. The authors study the neural network preference toward shortcut features using synthetic datasets with varying predictability and availability of the shortcut features. They empirically observed that networks prefer to learn shortcut features when they are more available, and ReLU is biased toward shortcuts. They also theoretically show using the NTK that linear networks are less biased towards the shortcut compared to the ReLU networks."},"presentation":{"value":"3 good"},"contribution":{"value":"3 good"},"soundness":{"value":"3 good"},"strengths":{"value":"* The paper is well-written and easy to follow. The problem of shortcut learning, which is not extensively studied, is important to understand.\n\n* The paper presents interesting insights into neural networks’ preference toward shortcuts. The paper studies shortcut learning empirically using controlled datasets and theoretically using the NTK.\n\n* The derivation for the bias of linear and ReLU networks using NTK would be helpful for future work in this domain."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"* The main observation that the model depth and non-linearity increase bias towards the shortcut features is intuitive. Both depth and non-linearity allows the model to learn a rich representation.\n\n* Theoretical analysis using NTK is problematic as they don't learn feature and thus cannot necessarily explain model's preference towards the shortcut feature. \n\n* Only vision tasks are explored in the paper, it is not clear if the observations will hold true for other domains.\n\n* It would be interesting to see experiments with the vision transformers."},"confidence":{"value":"3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked."},"questions":{"value":"* Do you think the observations will be similar for other non-linearities?"},"rating":{"value":"8: accept, good paper"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1700674221979,"tcdate":1698848169978,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission6363/Reviewer_KL3Q"],"signatures":["ICLR.cc/2024/Conference/Submission6363/Reviewer_KL3Q"],"forum":"Tj3xLVuE9f","number":3,"license":"CC BY 4.0","cdate":1698848169978,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission6363/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1700674221979,"domain":"ICLR.cc/2024/Conference","replyto":"Tj3xLVuE9f","id":"GiQWCqnNnH","forumContent":{"venue":{"value":"ICLR 2024 spotlight"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["shortcut learning","spurious correlations","architectural inductive bias"]},"primary_area":{"value":"visualization or interpretation of learned representations"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Deep-learning models can extract a rich assortment of features from data. Which features a model uses depends not only on *predictivity*---how reliably a feature indicates training-set labels---but also on *availability*---how easily the feature can be extracted from inputs. The literature on shortcut learning has noted examples in which models privilege one feature over another, for example texture over shape and image backgrounds over foreground objects. Here, we test hypotheses about which input properties are more available to a model, and systematically study how predictivity and availability interact to shape models' feature use. We construct a minimal, explicit generative framework for synthesizing classification datasets with two latent features that vary in predictivity and in factors we hypothesize to relate to availability, and we quantify a model's shortcut bias---its over-reliance on the shortcut (more available, less predictive) feature at the expense of the core (less available, more predictive) feature. We find that linear models are relatively unbiased, but introducing a single hidden layer with ReLU or Tanh units yields a bias. Our empirical findings are consistent with a theoretical account based on Neural Tangent Kernels. Finally, we study how models used in practice trade off predictivity and availability in naturalistic datasets, discovering availability manipulations which increase models' degree of shortcut bias. Taken together, these findings suggest that the propensity to learn shortcut features is a fundamental characteristic of deep nonlinear architectures warranting systematic study given its role in shaping how models solve tasks."},"_bibtex":{"value":"@inproceedings{\nhermann2024on,\ntitle={On the Foundations of Shortcut Learning},\nauthor={Katherine Hermann and Hossein Mobahi and Thomas FEL and Michael Curtis Mozer},\nbooktitle={The Twelfth International Conference on Learning Representations},\nyear={2024},\nurl={https://openreview.net/forum?id=Tj3xLVuE9f}\n}"},"title":{"value":"On the Foundations of Shortcut Learning"},"pdf":{"value":"/pdf/3f47b29f0e35691e7047d9fbfa0e4c47ea966e49.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference"},"paperhash":{"value":"hermann|on_the_foundations_of_shortcut_learning"},"authorids":{"value":["~Katherine_Hermann1","~Hossein_Mobahi2","~Thomas_FEL1","~Michael_Curtis_Mozer1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Katherine Hermann","Hossein Mobahi","Thomas FEL","Michael Curtis Mozer"]}},"version":2},{"content":{"summary":{"value":"This paper provides an in-depth analysis of the phenomenon of \"shortened learning\" in machine learning models, especially in the context of perceptual tasks. The authors confirm that basic empirical risk minimization (ERM) methods tend to prefer models that depend on shortcut features, even when models can achieve zero loss using only stable features. They attribute this to ERM's inductive bias to maximize margins across all samples. To address this, the authors propose an alternative loss function biased inductively with a uniform margin called MARG-CTRL. The paper demonstrates that MARG-CTRL mitigates shortcut learning on multiple vision and language tasks without the use of annotations of the shortcut feature in training or even validation. They also show that MARG-CTRL performs on par or better than the more complex and costly two-step shortcut mitigation method."},"presentation":{"value":"4 excellent"},"contribution":{"value":"4 excellent"},"soundness":{"value":"4 excellent"},"strengths":{"value":"-This paper presents a thorough analysis of the shortcut learning problem. \n-The authors propose a novel solution, MARG-CTRL, which has been shown to effectively mitigate shortcut learning in several tasks.\n-The paper is well-written and easy to understand."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"-Although MARG-CTRL is shown to perform well across different tasks, it is not clear how MARG-CTRL perform in scenarios where the shortcut features and the stable features are highly correlated.\n-The paper compares MARG-CTRL with two-stage shortcut-mitigating methods like JTT and CNC, it would be helpful to understand the specific scenarios where MARG-CTRL outperforms these methods and where it does not."},"confidence":{"value":"3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked."},"questions":{"value":"-How does MARG-CTRL perform in scenarios where the shortcut features and the stable features are highly correlated? Does this affect the effectiveness of MARG-CTRL?\n-What are the trade-offs involved in using MARG-CTRL versus two-stage shortcut-mitigating methods??\n-Are there specific scenarios or types of tasks for which MARG-CTRL may not be as effective?"},"rating":{"value":"5: Borderline accept: Technically solid paper where reasons to accept outweigh reasons to reject, e.g., limited evaluation. Please use sparingly."},"code_of_conduct":{"value":"Yes"},"limitations":{"value":"Please see the weaknesses part."}},"nonreaders":[],"tmdate":1702411520475,"tcdate":1688644377742,"writers":["NeurIPS.cc/2023/Conference","NeurIPS.cc/2023/Conference/Submission15594/Reviewer_PgqC"],"signatures":["NeurIPS.cc/2023/Conference/Submission15594/Reviewer_PgqC"],"forum":"zyZkaqNnpa","number":2,"license":"CC BY 4.0","cdate":1688644377742,"mdate":1702411520475,"readers":["everyone"],"invitations":["NeurIPS.cc/2023/Conference/Submission15594/-/Official_Review","NeurIPS.cc/2023/Conference/-/Edit"],"domain":"NeurIPS.cc/2023/Conference","replyto":"zyZkaqNnpa","id":"fEaGpgZJxH","forumContent":{"venue":{"value":"NeurIPS 2023 poster"},"keywords":{"value":["shortcut learning","spurious correlations","perfect stable feature","perception tasks","implicit bias in optimization","improving inductive biases"]},"_bibtex":{"value":"@inproceedings{\npuli2023dont,\ntitle={Don{\\textquoteright}t blame Dataset Shift! Shortcut Learning due to Gradients and Cross Entropy},\nauthor={Aahlad Manas Puli and Lily H Zhang and Yoav Wald and Rajesh Ranganath},\nbooktitle={Thirty-seventh Conference on Neural Information Processing Systems},\nyear={2023},\nurl={https://openreview.net/forum?id=zyZkaqNnpa}\n}"},"title":{"value":"Don’t blame Dataset Shift! Shortcut Learning due to Gradients and Cross Entropy"},"paperhash":{"value":"puli|dont_blame_dataset_shift_shortcut_learning_due_to_gradients_and_cross_entropy"},"TLDR":{"value":"Implicit biases toward maximizing margins induce shortcut learning in ERM even in tasks with perfect stable features, controlling margins mitigates shortcuts"},"abstract":{"value":"Common explanations for shortcut learning assume that the shortcut improves prediction only under the training distribution. Thus, models trained in the typical way by minimizing log-loss using gradient descent, which we call default-ERM, should utilize the shortcut. However, even when the stable feature determines the label in the training distribution and the shortcut does not provide any additional information, like in perception tasks, default-ERM exhibits shortcut learning. Why are such solutions preferred when the loss can be driven to zero when using the stable feature alone? By studying a linear perception task, we show that default-ERM’s preference for maximizing the margin, even without overparameterization, leads to models that depend more on the shortcut than the stable feature. This insight suggests that default-ERM’s implicit inductive bias towards max-margin may be unsuitable for perception tasks. Instead, we consider inductive biases toward uniform margins. We show that uniform margins guarantee sole dependence on the perfect stable feature in the linear perception task and suggest alternative loss functions, termed margin control (MARG-CTRL), that encourage uniform-margin solutions. MARG-CTRL techniques mitigate shortcut learning on a variety of vision and language tasks, showing that changing inductive biases can remove the need for complicated shortcut-mitigating methods in perception tasks."},"pdf":{"value":"/pdf/956a72600022df2e6738151ad2ff6a638cc32154.pdf"},"venueid":{"value":"NeurIPS.cc/2023/Conference"},"authorids":{"value":["~Aahlad_Manas_Puli1","~Lily_H_Zhang1","~Yoav_Wald1","~Rajesh_Ranganath2"]},"authors":{"value":["Aahlad Manas Puli","Lily H Zhang","Yoav Wald","Rajesh Ranganath"]}},"version":2},{"content":{"comment":{"value":"Dear AE,\n\nThank you for these good news and the additional higher-level feedback! We just submitted the camera-ready version of our paper with the minor revision you suggested:\n\n*For example, when introducing shortcuts, it would be helpful to include a concrete example with an illustration to aid the reader's understanding of what a shortcut is.*\n\nWe therefore added an example for each of the 4 shortcut baselines we study.\n\nOn a final note, we agree that \"there are tons of video QA benchmarks proposed nowadays and many of them also aim to mitigate the shortcuts\". However, I personally see the value of this work in providing a unified + diverse + large-scale collection of minimal video pairs. Many dataset do not try to curate minimal pairs of video or only do so in a smaller scale and not across many existing datasets. Let's see what datasets will end up being useful to the community!"},"title":{"value":"Camera Ready submitted"}},"parentInvitations":"TMLR/-/Official_Comment","tmdate":1763338483880,"tcdate":1763338483880,"writers":["TMLR","TMLR/Paper5382/Authors"],"signatures":["TMLR/Paper5382/Authors"],"forum":"gvFgNJcSw1","number":15,"license":"CC BY 4.0","cdate":1763338483880,"readers":["everyone"],"invitations":["TMLR/Paper5382/-/Official_Comment"],"mdate":1763338483880,"domain":"TMLR","replyto":"QQXf6FVeme","id":"lSuA96wYZ6","forumContent":{"submission_length":{"value":"Regular submission (no more than 12 pages of main content)"},"venue":{"value":"Accepted by TMLR"},"abstract":{"value":"Existing benchmarks for assessing the spatio-temporal understanding and reasoning abilities of video language models are susceptible to score inflation due to the presence of shortcut solutions based on superficial visual or textual cues. This paper mitigates the challenges in accurately assessing model performance by introducing the Minimal Video Pairs (MVP) benchmark, a simple shortcut-aware video QA benchmark for assessing the physical understanding of video language models. The benchmark is comprised of 55K high-quality multiple-choice video QA examples focusing on physical world understanding. Examples are curated from nine video data sources, spanning first-person egocentric and exocentric videos, robotic interaction data, and cognitive science intuitive physics benchmarks. To mitigate shortcut solutions that rely on superficial visual or textual cues and biases, each sample in MVP has a minimal-change pair — a visually similar video accompanied by an identical question but an opposing answer. To answer a question correctly, a model must provide correct answers for both examples in the minimal-change pair; as such, models that solely rely on visual or textual biases would achieve below random performance. Human performance on MVP is 92.9%, while the best open-source state-of-the- art video-language model achieves 40.2% compared to random performance at 25%."},"_bibtex":{"value":"@article{\nkrojer2025a,\ntitle={A Shortcut-aware Video-{QA} Benchmark for Physical Understanding via Minimal Video Pairs},\nauthor={Benno Krojer and Mojtaba Komeili and Candace Ross and Quentin Garrido and Koustuv Sinha and Nicolas Ballas and Mido Assran},\njournal={Transactions on Machine Learning Research},\nissn={2835-8856},\nyear={2025},\nurl={https://openreview.net/forum?id=gvFgNJcSw1},\nnote={}\n}"},"title":{"value":"A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs"},"pdf":{"value":"/pdf/f01dd06430bcd7b2aba068243d47011dcc96be6e.pdf"},"venueid":{"value":"TMLR"},"paperhash":{"value":"krojer|a_shortcutaware_videoqa_benchmark_for_physical_understanding_via_minimal_video_pairs"},"authorids":{"value":["~Benno_Krojer1","~Mojtaba_Komeili1","~Candace_Ross1","~Quentin_Garrido1","~Koustuv_Sinha1","~Nicolas_Ballas1","~Mido_Assran1"]},"assigned_action_editor":{"value":"~Tae-Hyun_Oh3"},"authors":{"value":["Benno Krojer","Mojtaba Komeili","Candace Ross","Quentin Garrido","Koustuv Sinha","Nicolas Ballas","Mido Assran"]}},"version":2},{"content":{"summary":{"value":"The paper aims to reduce the high inference cost of SAM2 in video segmentation. Specifically, they propose two modules: Object-aware Sparse Window Routing (SWR), which skips background windows in the image encoder based on object masks and saliency, and Object-aware Sparse Memory Retrieval (SMR), which selects only salient memory tokens and reuses their mask across frames. Together, these modules accelerate SAM2 by up to 1.75× with minimal accuracy loss on benchmarks such as SA-V, DAVIS, and MOSE."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"1. What is the specific layer index after which the router divides foreground and background tokens? How many subsequent modules benefit from reduced computation? It would be helpful to include an ablation varying the routing layer index to assess its impact.\n\n2. How is speedup defined in this paper? It would be clearer to disclose it in two aspects—FLOPs reduction and throughput improvement—to give a more complete view of efficiency.\n\n3. What is the intuition behind using two different temporal intervals (∆t =1,5) in the experiments? It seems the performance difference might largely stem from the frame rate (FPS) of the original benchmark. For high-FPS datasets like SA-V, increasing the interval could naturally yield greater gains since the memory bank contains more diverse frames.\n\n4. For SMR, have you evaluated a variant that selects memory frames only when object presence is confident (akin to SAM2Long’s strategy)? It would be informative to report the gain from this filtering."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The paper makes a solid technical contribution by streamlining the SAM2 model. The object-aware pruning of the image encoder and the introduction of a background shortcut for non-foreground patches are both clever ideas that substantially reduce computation.\n\n2. The ablation study is comprehensive. It not only analyzes the proposed components in isolation but also integrates other efficient methods (e.g., ToME) into their framework for comparison, which provides valuable insights."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The proposed routing mechanism heavily relies on the assumption of temporal consistency in video streams, meaning that no significant camera shaking or viewpoint shift occurs. This limits the method’s applicability in real-world scenarios with dynamic motion. It would be interesting to see comparisons with SAM2 on datasets such as MOSEv2[1] and SeCVOS[2], which feature frequent viewpoint transitions.\n\n\n\n\n2. Performing grid search on only one benchmark is not sufficient to demonstrate robustness. It would strengthen the paper to include additional grid search curves across multiple benchmarks (in Figure 5).\n\n\n\n\n\n\n\n[1] MOSEv2: A More Challenging Dataset for Video Object Segmentation in Complex Scenes\n[2] SeC: Advancing Complex Video Object Segmentation via Progressive Concept Construction"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762917786601,"tcdate":1762115363286,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission4947/Reviewer_DdRT"],"signatures":["ICLR.cc/2026/Conference/Submission4947/Reviewer_DdRT"],"forum":"HcivRSezJp","number":3,"license":"CC BY 4.0","cdate":1762115363286,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission4947/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762917786601,"domain":"ICLR.cc/2026/Conference","replyto":"HcivRSezJp","id":"m9cEBVF85c","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Segmemt Anything Model","Efficient Deep Learning","Model Acceleration"]},"supplementary_material":{"value":"/attachment/b3481d522f61dec0cb42900d807bcdca9c57c111.zip"},"primary_area":{"value":"infrastructure, software libraries, hardware, systems, etc."},"abstract":{"value":"Segment Anything Model 2 (SAM2) shows excellent performance in video object segmentation tasks; however, the heavy computational burden hinders its application in real-time video processing.\nAlthough there have been efforts to improve the efficiency of SAM2, most of them focus on retraining a lightweight backbone, with little exploration into post-training acceleration.\nIn this paper, we observe that SAM2 exhibits sparse perception pattern as biological vision, which provides opportunities for eliminating redundant computation and acceleration:\ni) In mask decoder, the attention primarily focuses on the foreground objects, whereas the image encoder in the earlier stage exhibits a broad attention span, which results in unnecessary computation to background regions.\nii) In memory bank, only a small subset of tokens in each frame contribute significantly to memory attention, and the salient regions exhibit temporal consistency, making full-token computation redundant.\nWith these insights, we propose Efficient-SAM2, which promotes SAM2 to adaptively focus on object regions while eliminating task-irrelevant computations, thereby significantly improving inference efficiency.\nSpecifically, for image encoder, we propose object-aware Sparse Window Routing (SWR), a window-level computation allocation mechanism that leverages the consistency and saliency cues from the previous-frame decoder to route background regions into a lightweight shortcut branch.\nMoreover, for memory attention, we propose object-aware Sparse Memory Retrieval (SMR), which allows only the salient memory tokens in each frame to participate in computation, \nwith the saliency pattern reused from their first recollection.\nWith negligible additional parameters and minimal training overhead, Efficient-SAM2 delivers 1.68$\\times$ speedup on SAM2.1-L model with only 1.0\\% accuracy drop on SA-V test set, where SWR and SMR provide 1.83$\\times$ and 1.78$\\times$ speedups, respectively."},"_bibtex":{"value":"@inproceedings{\nzhang2026efficientsam,\ntitle={Efficient-{SAM}2: Accelerating {SAM}2 with Object-Aware Visual Encoding and Memory Retrieval},\nauthor={Jing Zhang and Zhikai Li and Xuewen Liu and Qingyi Gu},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=HcivRSezJp}\n}"},"title":{"value":"Efficient-SAM2: Accelerating SAM2 with Object-Aware Visual Encoding and Memory Retrieval"},"pdf":{"value":"/pdf/b1801056e42ab7eb9562f53bd0d351cf557070d0.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"zhang|efficientsam2_accelerating_sam2_with_objectaware_visual_encoding_and_memory_retrieval"},"authorids":{"value":["~Jing_Zhang58","~Zhikai_Li1","~Xuewen_Liu1","~Qingyi_Gu1"]},"authors":{"value":["Jing Zhang","Zhikai Li","Xuewen Liu","Qingyi Gu"]}},"version":2},{"content":{"summary":{"value":"This paper proposes GateFlow, a transport-guided gating mechanism to mitigate shortcut learning in Vision-Language-Action (VLA) models. The authors argue that the optimization gap between ELBO and true NLL in flow-matching-based VLA training encourages spurious correlations between visual patterns and actions. GateFlow measures the Wasserstein distance between observation and action features to identify and suppress shortcut pathways, selectively amplifying features corresponding to genuine semantic understanding. The paper provides theoretical analysis and experiments on LIBERO benchmarks, reporting improved performance."},"soundness":{"value":1},"confidence":{"value":4},"questions":{"value":"see weakness 3, the definition of observation features and action features, and the calculation process is unclear."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. The paper highlights an important limitation of current VLA models: their tendency to overfit to spurious correlations (shortcut learning). The link between the ELBO–NLL gap and shortcut behavior is conceptually clear and well-motivated.\n2. Employing the Wasserstein distance as a measure of semantic alignment between different features is an interesting idea.\n3. The GateFlow module can be plugged into existing VLA architectures."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The paper’s statement about mitigating “shortcut learning” to enhance semantic understanding needs stronger argument. For VLAs with System1 and System2 structures, semantic reasoning primarily happens in a symbolic or high-level semantic layer (system 2). It is unclear whether deeper reasoning are needed for system1. Furthermore,  in practice, large-scale and diverse datasets, along with robust pretrained VLM features, already mitigate shortcuts substantially. It’s unclear why this architectural modification is necessary instead of improving data diversity or leveraging VLM pretraining better. A discussion comparing these two directions (data scaling vs. gating) is missing.\n2. The proposed alignment between observation features and actions is conceptually questionable. \n* There can be multiple valid actions for the same observation. Thus, their alignment may not represent true semantics. A more meaningful alignment might be between vision–language embeddings and actions, as that reflects task-level semantics. Also, since actions are per timestep, semantic correspondence is inherently weak.\n* The action features are not pretrained rather learned. Consider an experiment where you always put a green object for one specific task and no green object for the others, then when you train the model with such data, the action representation will be aligned with the observation feature during the training process. At inference time, when the green object is present, the model can just ignore the language instruction and perform the specific task. I would like to see such controlled experiments to validate the proposed method can suppress such shortcuts.\n3. It’s unclear what exactly constitutes “observation features”. Are they derived purely from the visual encoder, or do they include language and proprioceptive states? Similarly, “action features” are said to be used in the Wasserstein computation, but since actions are outputs, it is unclear how their features are extracted before prediction. Also, it’s unclear whether gating is applied per transformer layer.\n4. The overall performance gain over the strong π0.5 baseline is modest (e.g., +1.4% average, +4% on LIBERO-Long). Given that the method adds additional computations and hyperparameters, the improvement seems incremental rather than transformative.\n5. The use of LIBERO as the main benchmark is not ideal for demonstrating shortcut mitigation. Not only because model performance on this benchmark is saturated, but also because the simulated scenes are well controlled and less prone to the vision-action shortcut correlations. Experiments with real-world experiments with lighting changes or texture variations and simulation experiments with a dataset includes a spurious correlation (e.g., the experiment of a green object always associated with a specific task as aforementioned ) should be included to validate the method.\n6. Missing related works[1,2,3] on mitigating the shortcut.\n\n[1] Xing, Youguang, et al. \"Shortcut learning in generalist robot policies: The role of dataset diversity and fragmentation.\" arXiv preprint arXiv:2508.06426 (2025).\n\n[2] Lin, Fanqi, et al. \"OneTwoVLA: A Unified Vision-Language-Action Model with Adaptive Reasoning.\" arXiv preprint arXiv:2505.11917 (2025).\n\n[3] Huang, Huang, et al. \"Otter: A vision-language-action model with text-aware visual feature extraction.\" arXiv preprint arXiv:2503.03734 (2025)."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762918472290,"tcdate":1761959002890,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission6112/Reviewer_Renq"],"signatures":["ICLR.cc/2026/Conference/Submission6112/Reviewer_Renq"],"forum":"qOSy2PX4xS","number":3,"license":"CC BY 4.0","cdate":1761959002890,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission6112/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762918472290,"domain":"ICLR.cc/2026/Conference","replyto":"qOSy2PX4xS","id":"aRJW9fl0Nq","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"TLDR":{"value":"VLA models may easily learn shortcuts by exploiting spurious correlations. GateFlow uses transport distance to detect and suppress shortcuts while enhancing genuine understanding."},"keywords":{"value":["Vision-Language-Action","Flow Matching","Foundation Models","Embodied AI","Robotics"]},"primary_area":{"value":"applications to robotics, autonomy, planning"},"abstract":{"value":"Vision-Language-Action (VLA) models promise general-purpose robotic intelligence by leveraging pretrained vision-language representations. However, these models suffer from shortcut learning—exploiting spurious correlations between visual patterns and actions rather than developing semantic understanding. This occurs because VLA models optimize an Evidence Lower Bound (ELBO) proxy instead of the true likelihood, creating an optimization gap that enables memorized patterns to masquerade as genuine solutions. To mitigate this problem, we introduce GateFlow, a transport-guided gating mechanism that detects and suppresses shortcut learning by measuring the Wasserstein distance between observation and action representations. Low transport distance indicates semantic understanding and receives strong enhancement, while high distance reveals shortcuts and triggers suppression. This selective gating closes the ELBO-NLL gap by guiding optimization toward true likelihood minimization. We provide theoretical guarantees showing that GateFlow concentrates gradients on semantic features while eliminating spurious patterns. Empirically, GF-VLA achieves state-of-the-art performance on various tasks, with substantial improvements on long-range tasks or complex scenarios under non-stationary perturbations. GateFlow integrates seamlessly into existing VLA architectures with minimal computational overhead, offering a practical solution to more general robotic learning."},"_bibtex":{"value":"@misc{\nzhang2025gateflow,\ntitle={GateFlow: Mitigating Shortcut Learning in {VLA} Models via Gated Flow Matching},\nauthor={Wanpeng Zhang and Ye Wang and Hao Luo and Haoqi Yuan and Yicheng Feng and Sipeng Zheng and Qin Jin and Zongqing Lu},\nyear={2025},\nurl={https://openreview.net/forum?id=qOSy2PX4xS}\n}"},"title":{"value":"GateFlow: Mitigating Shortcut Learning in VLA Models via Gated Flow Matching"},"pdf":{"value":"/pdf/691cf8ebbc33f2ff0fab9ef100c422af853434ff.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"zhang|gateflow_mitigating_shortcut_learning_in_vla_models_via_gated_flow_matching"},"authorids":{"value":["~Wanpeng_Zhang1","~Ye_Wang12","~Hao_Luo6","~Haoqi_Yuan1","~Yicheng_Feng1","~Sipeng_Zheng1","~Qin_Jin1","~Zongqing_Lu2"]},"authors":{"value":["Wanpeng Zhang","Ye Wang","Hao Luo","Haoqi Yuan","Yicheng Feng","Sipeng Zheng","Qin Jin","Zongqing Lu"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a scalable synthetic framework for benchmarking video MLLMs. It decouples video content from query-response pairs and separates different aspects of video understanding skills. The authors construct a synthetic video benchmark with tasks like retrieval, ordering, and counting to evaluate video models' temporal perception, chronological order, and spatio-temporal coherence. The experiment on VNBench reveals performance gaps and weaknesses in current video models."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"Please see weakness."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. VideoNIAH offers a novel approach to evaluating video models by decoupling video content and query-response pairs.\n\n2. VNBench provides a scalable and flexible way to evaluate video understanding across different dimensions and video lengths, overcoming some limitations of traditional benchmarks.\n\n3. The paper conducts a thorough evaluation of multiple models, analyzing various factors such as haystack length, needle number, and model settings, providing valuable insights for model improvement."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. While VNBench assesses important aspects of video understanding, it may not cover all possible video-related tasks and scenarios.\n\n2. Despite the relatively low cost of the evaluation method proposed in this paper, it is vulnerable to being gamed by models. Models may develop strategies to perform well on the tasks without truly understanding the video content, thus compromising the reliability of the evaluation.\n\n3. The overall tasks in the proposed evaluation framework are relatively simplistic. After targeted training, models can achieve high scores with relative ease, which may not accurately reflect their ability to handle more complex and diverse real-world video understanding scenarios. This simplicity could lead to an overestimation of the models' capabilities and may not provide a comprehensive assessment of their true understanding and generalization ability in practical applications."}},"nonreaders":[],"tmdate":1731428630005,"tcdate":1730697938476,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission8848/Reviewer_aagQ"],"signatures":["ICLR.cc/2025/Conference/Submission8848/Reviewer_aagQ"],"forum":"ZJo6Radbqq","number":4,"license":"CC BY 4.0","cdate":1730697938476,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission8848/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428630005,"domain":"ICLR.cc/2025/Conference","replyto":"ZJo6Radbqq","id":"4saRzb1dhS","forumContent":{"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["video MLLM"]},"primary_area":{"value":"datasets and benchmarks"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Video understanding is a crucial next step for multimodal large language models (MLLMs).\nVarious benchmarks are introduced for better evaluating the MLLMs.\nNevertheless, current video benchmarks are still inefficient for evaluating video models during iterative development due to the high cost of constructing datasets and the difficulty in isolating specific skills.\nIn this paper, we propose VideoNIAH (Video Needle in A Haystack), a benchmark construction framework through synthetic video generation. \nVideoNIAH decouples video content from their query-responses by inserting unrelated visual 'needles' into original videos. \nThe framework automates the generation of query-response pairs using predefined rules, minimizing manual labor.  The queries focus on specific aspects of video understanding, enabling more skill-specific evaluations. The separation between video content and the queries also allow for increased video variety and evaluations across different lengths.\nUtilizing VideoNIAH, we compile a video benchmark, VNBench, which includes tasks such as retrieval, ordering, and counting to evaluate three key aspects of video understanding: temporal perception, chronological ordering, and spatio-temporal coherence. We conduct a comprehensive evaluation of both proprietary and open-source models, uncovering significant differences in their video understanding capabilities across various tasks. Additionally, we perform an in-depth analysis of the test results and model configurations. Based on these findings, we provide some advice for improving video MLLM training, offering valuable insights to guide future research and model development."},"_bibtex":{"value":"@inproceedings{\nzhao2025needle,\ntitle={Needle In A Video Haystack: A Scalable  Synthetic Evaluator for Video {MLLM}s},\nauthor={Zijia Zhao and Haoyu Lu and Yuqi Huo and Yifan Du and Tongtian Yue and Longteng Guo and Bingning Wang and weipeng chen and Jing Liu},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=ZJo6Radbqq}\n}"},"title":{"value":"Needle In A Video Haystack: A Scalable  Synthetic Evaluator for Video MLLMs"},"pdf":{"value":"/pdf/875d019fb5070873ce0564e8a175d24d3df4dc3f.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"zhao|needle_in_a_video_haystack_a_scalable_synthetic_evaluator_for_video_mllms"},"authorids":{"value":["~Zijia_Zhao1","~Haoyu_Lu1","~Yuqi_Huo1","~Yifan_Du1","~Tongtian_Yue1","~Longteng_Guo1","~Bingning_Wang3","~weipeng_chen2","~Jing_Liu1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Zijia Zhao","Haoyu Lu","Yuqi Huo","Yifan Du","Tongtian Yue","Longteng Guo","Bingning Wang","weipeng chen","Jing Liu"]}},"version":2},{"content":{"summary":{"value":"This paper proposed an augmentation technique in frequency domain, motivated by the phenomenon of frequency bias/frequency shortcut learning. The technique AAUA adversarially perturbs frequency components of images  which the models over-rely for classification. As AAUA might encourage shortcut learning in high-frequency, the authors further use AAD to drop out randomly frequency components that models highly depend on for classification. AAUA and AAD together can avoid frequency shortcut learning. It is shown that AAUA and AAD together can improve domain generalization capability of models."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"1. How are d_t^mu and the other initialized? Did the experiments consider different random seeds?\n\n2. Speaking of hardness of the augmented samples, how do AAUA and AAD compare with other adversarial techniques? \n\n3. Line 121-123: If correctly understood, it should be: responsive frequency describes the characteristics of mapping functions learned by networks and the mapping function transforms images into probability vectors. How does the smoothness represented by the probability vectors and how does it relate to frequency? The concept of 'smoothness' appears abruptly. \n\n4. What are the meaning of the two target domains in Fig. 2? Are they to show the distributions of augmented images are close to them? How are the density distributions of each domain computed? What do the peak values of each distribution mean?"},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"1. Well-written introduction. The introduction explains well the domain generalization problem and how it related to frequency bias / frequency shortcut learning. \n2. Interesting and insightful ablation study on the impact of AAUA and AAD on frequency shortcut learning. The authors observed that AAUA or AAD solely might encourage more shortcut learning, supporting the necessity to apply both AAUA and AAD together."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The first and second contribution points can be merged for a compact presentation. Moreover, whether this is the first work to combat frequency shortcuts for domain generalization is uncertain. Domain generalization is a broad topic, and it includes single-source domain generalization problem, e.g. training models on a clean dataset and evaluate them in a corrupted dataset. And the [1] focuses on reducing frequency shortcut learning to improve model robustness towards image corruptions, one of domain generalization problems. \n\n2. As this work proposed a frequency augmentation technique, there are more related work despite FACT and SADA mentioned in the paper, such as VIPAug [2], AFA [3], HybridAugment++ [4], etc. \n\n3. Vague statement: line 327-328: it is uncertain how the comparison is operated given the condition 'without the proposed AAD or AAUA module'. One has to calculate on their own to figure out the comparison is between AAD+L_JS and AAD+AAUA+L_JS, and AAUA+L_JS and AAD+AAUA+L_JS. \n\n4. Sec. 3.2 Theoretical justification is from others' work, motivating the design of the techniques but not having direct mathematical relationship with the proposed technique.  The input-input correlation matrix is introduced but not used.  Strictly speaking, Sec. 3 is more like a literature review explaining motivations instead of analysis carried out by the authors. \n\n5. Definition of the masks separating low and high frequency components: h_min of M_low is H/4, meaning that for frequency components close to the center will be considered as high-frequency, conflict to what is considered as low-frequency. This needs more clarification on why h_min is not 0. Same for w_min.  From the definition of M_low, AAUA seems to shift models' attention to partial mid-high frequencies. \n\n6. Experiments need to be comprehensive: \n\n*As AAUA and AAD augment images in the frequency domain, comparison with more other frequency augmentation techniques should be included, such as HybridAug++ [4], VIPAug [2], AFA [3], DFM-X [1], AugSVF [5]. \n\n*The experiments are limited to small datasets, with a few classes. It lacks comprehensive experiments on large datasets, e.g. ImageNet-1k with 1000 classes, to show the feasibility of the proposed technique as well as a fair comparison with other methods which provide results on ImageNet-1k. \n\n*Some results in Table 1 are missing and there lacks explanations. \n\n*In Sec. 3.1 the authors claimed that they focus on classification tasks. However, they abruptly show results on instance retrieval. It is understood that domain generalization is important to instance retrieval. But there needs more clarifications on how the proposed technique can address frequency shortcut learning in retrieval tasks when there are not many work analyzing shortcut learning in image retrieval. \n\n*Sec. 5.3 evaluates frequency shortcuts of trained models. Intuitively, the analyzed models are those trained in Sec. 5.2. However, the evaluation is carried out on ImageNet-10, which is not mentioned previously and not used for training. It is unclear at this stage whether the authors use the models trained in Sec. 5.2 for shortcut evaluation on the dataset ImageNet-10, or train new models on ImageNet-10 for shortcut evaluations. This needs clarifications.\n\n*Sec. 5.4.1 demonstrates that the proposed augmentation technique increases the frequency sensitivity of models compared to baseline. This seems to be a bad sign to model generalization capability according to [43]. However, there is no related discussions and Fig. 3a is not referred anywhere in the main paper. \n\n*Sec. 5.4.2 analyzes the hardness of augmented samples. However, the comparison is between AAUA and FACT, which seems unfair as FACT does not include adversarial setup while AAUA does. \n\nMinor: \nEach formula should end with period/comma. Line 271 'relative' should be 'related' if understood correctly. Table 3 not in the same page where mentioned. Table 4 exceeds width limit. \n\n[1] Wang, et al., 'DFM-X: Augmentation by Leveraging Prior Knowledge of Shortcut Learning', ICCVW2023.\n\n[2] Lee, et al., 'Domain Generalization with Vital Phase Augmentation', AAAI2024.\n\n[3] Vaish, et al., 'Fourier-basis Functions to Bridge Augmentation Gap: Rethinking Frequency Augmentation in Image Classification', CVPR2024.\n\n[4] Yucel, et al., 'HybridAugment++: Unified Frequency Spectra Perturbations for Model Robustness', ICCV2023.\n\n[5] Soklaski, et al., 'Fourier-Based Augmentations for Improved Robustness and Uncertainty Calibration', Neurips2021."},"limitations":{"value":"The authors addressed the limitations of their methods, regarding the infeasibility to directly locate frequency shortcuts."}},"nonreaders":[],"tmdate":1730878899720,"tcdate":1719831263490,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission4103/Reviewer_szDc"],"signatures":["NeurIPS.cc/2024/Conference/Submission4103/Reviewer_szDc"],"forum":"VMiLdBkCJM","number":1,"license":"CC BY 4.0","cdate":1719831263490,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission4103/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730878899720,"domain":"NeurIPS.cc/2024/Conference","replyto":"VMiLdBkCJM","id":"e5IYTaSweJ","forumContent":{"TLDR":{"value":"This paper proposes two data augmentation modules, which is also the first work, to tackle the topic of domain generalization from the perspective of preventing frequency shortcuts."},"venue":{"value":"NeurIPS 2024 poster"},"keywords":{"value":["Frequency Shortcut;Domain Generalization"]},"primary_area":{"value":"safety_in_machine_learning"},"abstract":{"value":"Domain generalization methods aim to learn transferable knowledge from source domains that can generalize well to unseen target domains. \nRecent studies show that neural networks frequently suffer from a simplicity-biased learning behavior which leads to over-reliance on specific frequency sets, namely as frequency shortcuts, instead of semantic information, resulting in poor generalization performance. \nDespite previous data augmentation techniques successfully enhancing generalization performances, they intend to apply more frequency shortcuts, thereby causing hallucinations of generalization improvement.\nIn this paper, we aim to prevent such learning behavior of applying frequency shortcuts from a data-driven perspective. Given the theoretical justification of models' biased learning behavior on different spatial frequency components, which is based on the dataset frequency properties, we argue that the learning behavior on various frequency components could be manipulated by changing the dataset statistical structure in the Fourier domain. \nIntuitively, as frequency shortcuts are hidden in the dominant and highly dependent frequencies of dataset structure, dynamically perturbating the over-reliance frequency components could prevent the application of frequency shortcuts.\nTo this end, we propose two effective data augmentation modules designed to collaboratively and adaptively adjust the frequency characteristic of the dataset, aiming to dynamically influence the learning behavior of the model and ultimately serving as a strategy to mitigate shortcut learning. Our code will be made publicly available."},"_bibtex":{"value":"@inproceedings{\nhe2024towards,\ntitle={Towards Combating Frequency Simplicity-biased Learning for Domain Generalization},\nauthor={Xilin He and Jingyu Hu and Qinliang Lin and Cheng Luo and Weicheng Xie and Siyang Song and Muhammad Haris Khan and Linlin Shen},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=VMiLdBkCJM}\n}"},"title":{"value":"Towards Combating Frequency Simplicity-biased Learning for Domain Generalization"},"pdf":{"value":"/pdf/1f0f5670f421586cd563925e2b6c31ce51929a7f.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"he|towards_combating_frequency_simplicitybiased_learning_for_domain_generalization"},"authorids":{"value":["~Xilin_He1","~Jingyu_Hu2","~Qinliang_Lin2","~Cheng_Luo4","~Weicheng_Xie1","~Siyang_Song1","~Muhammad_Haris_Khan3","~Linlin_Shen1"]},"authors":{"value":["Xilin He","Jingyu Hu","Qinliang Lin","Cheng Luo","Weicheng Xie","Siyang Song","Muhammad Haris Khan","Linlin Shen"]}},"version":2},{"content":{"summary":{"value":"The paper presents a DiT-based video restoration model trained with a novel concept-distillation strategy that uses a pre-trained T2V generator to produce aligned text–video pairs, which eliminates distribution drift. It also proposes a lightweight ControlNet projector and dual-branch connector that further suppress artifacts and enable dynamic control."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"I don’t have any further questions at the moment; I’m waiting to discuss with the other reviewers."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. It proposes concept distillation with a pre-trained T2V model to generate aligned text–video pairs.\n2. It markedly outperforms prior methods.\n3. A high-quality dataset is created that should significantly benefit the video-generation community."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Lack of video supplementary results: As a video-oriented work, without video results as supplementary material, it is difficult for the public to intuitively evaluate the model's performance, especially the quality of temporal consistency."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762916914439,"tcdate":1761926517129,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission3680/Reviewer_7MWW"],"signatures":["ICLR.cc/2026/Conference/Submission3680/Reviewer_7MWW"],"forum":"YV5Zgv8pdg","number":3,"license":"CC BY 4.0","cdate":1761926517129,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission3680/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762916914439,"domain":"ICLR.cc/2026/Conference","replyto":"YV5Zgv8pdg","id":"lmDNtunu1S","forumContent":{"TLDR":{"value":"We present Vivid-VR, a DiT-based generative video restoration method."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Video Restoration","Diffusion Transformer","Text-to-Video","ControlNet","Concept Distillation"]},"supplementary_material":{"value":"/attachment/1ac7838c26d644b9c27a33c04143e290587515fb.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"We present Vivid-VR, a DiT-based generative video restoration method built upon an advanced T2V foundation model, where ControlNet is leveraged to control the generation process, ensuring content consistency. However, conventional fine-tuning of such controllable pipelines frequently suffers from distribution drift due to limitations in imperfect multimodal alignment, resulting in compromised texture realism and temporal coherence. To tackle this challenge, we propose a concept distillation training strategy that utilizes the pretrained T2V model to synthesize training samples with embedded textual concepts, thereby distilling its conceptual understanding to preserve texture and temporal quality. To enhance generation controllability, we redesign the control architecture with two key components: 1) a control feature projector that filters degradation artifacts from input video latents to minimize their propagation through the generation pipeline, and 2) a new ControlNet connector employing a dual-branch design. This connector synergistically combines MLP-based feature mapping with cross-attention mechanism for dynamic control feature retrieval, enabling both content preservation and adaptive control signal modulation. Extensive experiments show that Vivid-VR performs favorably against existing approaches on both synthetic and real-world benchmarks, as well as AIGC videos, achieving impressive texture realism, visual vividness, and temporal consistency. The codes and checkpoints are publicly available at https://github.com/csbhr/Vivid-VR."},"_bibtex":{"value":"@inproceedings{\nbai2026vividvr,\ntitle={Vivid-{VR}: Distilling Concepts from Text-to-Video Diffusion Transformer for Photorealistic Video Restoration},\nauthor={Haoran Bai and Xiaoxu Chen and Canqian Yang and Zongyao He and Sibin Deng and Ying Chen},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=YV5Zgv8pdg}\n}"},"title":{"value":"Vivid-VR: Distilling Concepts from Text-to-Video Diffusion Transformer for Photorealistic Video Restoration"},"pdf":{"value":"/pdf/c0383acde3b4b19fa13cbf9ba5a770825b772ee9.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"bai|vividvr_distilling_concepts_from_texttovideo_diffusion_transformer_for_photorealistic_video_restoration"},"authorids":{"value":["~Haoran_Bai1","~Xiaoxu_Chen1","~Canqian_Yang1","~Zongyao_He1","~Sibin_Deng1","~Ying_Chen15"]},"authors":{"value":["Haoran Bai","Xiaoxu Chen","Canqian Yang","Zongyao He","Sibin Deng","Ying Chen"]}},"version":2},{"content":{"venue":{"value":"IEEE Trans. Circuits Syst. Video Technol. 2025"},"pdf":{"value":"https://ieeexplore.ieee.org/iel8/76/10989278/10786261.pdf"},"venueid":{"value":"dblp.org/journals/TCSV/2025"},"paperhash":{"value":"zeng|improving_video_moment_retrieval_by_auxiliary_momentquery_pairs_with_hyperinteraction"},"authorids":{"value":["~Runhao_Zeng1","https://dblp.org/search/pid/api?q=author:Yishen_Zhuo:","https://dblp.org/search/pid/api?q=author:Jialiang_Li:","https://dblp.org/search/pid/api?q=author:Yunjin_Yang:","https://dblp.org/search/pid/api?q=author:Huisi_Wu:","~Qi_Chen4","https://dblp.org/search/pid/api?q=author:Xiping_Hu_0001:","https://dblp.org/search/pid/api?q=author:Victor_C._M._Leung:"]},"html":{"value":"https://doi.org/10.1109/TCSVT.2024.3513633"},"_bibtex":{"value":"@article{DBLP:journals/tcsv/ZengZLYWCHL25,\n  author={Runhao Zeng and Yishen Zhuo and Jialiang Li and Yunjin Yang and Huisi Wu and Qi Chen and Xiping Hu and Victor C. M. Leung},\n  title={Improving Video Moment Retrieval by Auxiliary Moment-Query Pairs With Hyper-Interaction},\n  year={2025},\n  month={May},\n  cdate={1746057600000},\n  journal={IEEE Trans. Circuits Syst. Video Technol.},\n  volume={35},\n  number={5},\n  pages={3940-3954},\n  url={https://doi.org/10.1109/TCSVT.2024.3513633}\n}\n"},"abstract":{"value":"Most existing video moment retrieval (VMR) benchmark datasets face a common issue of sparse annotations-only a few moments being annotated. We argue that videos contain a broader range of meaningful moments that, if leveraged, could significantly enhance performance. Existing methods typically follow a generate-then-select paradigm, focusing primarily on generating moment-query pairs while neglecting the crucial aspect of selection. In this paper, we propose a new method, HyperAux, to yield auxiliary moment-query pairs by modeling the multi-modal hyper-interaction between video and language. Specifically, given a set of candidate moment-query pairs from a video, we construct a hypergraph with multiple hyperedges, each corresponding to a moment-query pair. Unlike traditional graphs where each edge connects only two nodes (frames or queries), each hyperedge connects multiple nodes, including all frames within a moment, semantically related frames outside the moment, and an input query. This design allows us to consider the frames within a moment as a whole, rather than modeling individual frame-query relationships separately. More importantly, constructing the relationships among all moment-query pairs within a video into a large hypergraph facilitates selecting higher-quality data from such pairs. On this hypergraph, we employ a hypergraph neural network to aggregate node information, update the hyperedge, and propagate video-language hyper-interactions to each connected node, resulting in context-aware node representations. This enables us to use node relevance to select high-quality moment-query pairs and refine the moments’ boundaries. We also exploit the discrepancy in semantic matching within and outside moments to construct a loss function for training the HGNN without human annotations. Our auxiliary data enhances the performance of twelve VMR models under fully-supervised, weakly-supervised, and zero-shot settings across three widely used VMR datasets: ActivityNet Captions, Charades-STA, and QVHighlights. We will release the source code and models publicly."},"title":{"value":"Improving Video Moment Retrieval by Auxiliary Moment-Query Pairs With Hyper-Interaction"},"authors":{"value":["Runhao Zeng","Yishen Zhuo","Jialiang Li","Yunjin Yang","Huisi Wu","Qi Chen","Xiping Hu","Victor C. M. Leung"]}},"tmdate":1768968238103,"pdate":1735689600000,"tcdate":1753144647185,"writers":["~"],"signatures":["~Qi_Chen4"],"forum":"ObmJTDxFp0","license":"CC BY-SA 4.0","number":579898,"cdate":1746057600000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1768968238103,"domain":"DBLP.org","id":"ObmJTDxFp0","version":2},{"content":{"venue":{"value":"ICCV 2023"},"pdf":{"value":"https://ieeexplore.ieee.org/iel7/10376473/10376477/10377725.pdf"},"venueid":{"value":"dblp.org/conf/ICCV/2023"},"paperhash":{"value":"chung|shortcutv2v_compression_framework_for_videotovideo_translation_based_on_temporal_redundancy_reduction"},"authorids":{"value":["~Chaeyeon_Chung2","https://dblp.org/search/pid/api?q=author:Yeojeong_Park:","~Seunghwan_Choi1","https://dblp.org/search/pid/api?q=author:Munkhsoyol_Ganbat:","~Jaegul_Choo1"]},"html":{"value":"https://doi.org/10.1109/ICCV51070.2023.00700"},"_bibtex":{"value":"@inproceedings{DBLP:conf/iccv/ChungPCGC23,\n  author={Chaeyeon Chung and Yeojeong Park and Seunghwan Choi and Munkhsoyol Ganbat and Jaegul Choo},\n  title={Shortcut-V2V: Compression Framework for Video-to-Video Translation based on Temporal Redundancy Reduction},\n  year={2023},\n  cdate={1672531200000},\n  pages={7578-7588},\n  url={https://doi.org/10.1109/ICCV51070.2023.00700},\n  booktitle={ICCV},\n  crossref={conf/iccv/2023}\n}\n"},"abstract":{"value":"Video-to-video translation aims to generate video frames of a target domain from an input video. Despite its usefulness, the existing networks require enormous computations, necessitating their model compression for wide use. While there exist compression methods that improve computational efficiency in various image/video tasks, a generally-applicable compression method for video-to-video translation has not been studied much. In response, we present Shortcut-V2V, a general-purpose compression framework for video-to-video translation. Shortcut-V2V avoids full inference for every neighboring video frame by approximating the intermediate features of a current frame from those of the previous frame. Moreover, in our framework, a newly-proposed block called AdaBD adaptively blends and deforms features of neighboring frames, which makes more accurate predictions of the intermediate features possible. We conduct quantitative and qualitative evaluations using well-known video-to-video translation models on various tasks to demonstrate the general applicability of our framework. The results show that Shortcut-V2V achieves comparable performance compared to the original video-to-video translation model while saving 3.2-5.7× computational cost and 7.8-44× memory at test time. Our code and videos are available at https://shortcut-v2v.github.io/."},"title":{"value":"Shortcut-V2V: Compression Framework for Video-to-Video Translation based on Temporal Redundancy Reduction"},"authors":{"value":["Chaeyeon Chung","Yeojeong Park","Seunghwan Choi","Munkhsoyol Ganbat","Jaegul Choo"]}},"tmdate":1760598100583,"pdate":1672531200000,"tcdate":1731330726366,"writers":["~"],"signatures":["~Jaegul_Choo1"],"forum":"KKb64ALnJU","license":"CC BY-SA 4.0","number":181500,"cdate":1672531200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1760598100583,"domain":"DBLP.org","id":"KKb64ALnJU","version":2},{"content":{"TLDR":{"value":"Video editing model for long video following open-domain instruction"},"venue":{"value":"Video-Langauge Models Poster"},"keywords":{"value":["Video Editing","Generative AI","Video Generation"]},"supplementary_material":{"value":"/attachment/60763abf2fca3c3190ce6f531a1e555a936a5e3e.pdf"},"abstract":{"value":"Video editing stands as a cornerstone of digital media, from entertainment and education to professional communication.\nHowever, previous methods often overlook the necessity of comprehensively understanding both global and local contexts, leading to inaccurate and inconsistency edits in the spatiotemporal dimension, especially for long videos.\nIn this paper, we introduce VIA, a unified spatiotemporal VIdeo Adaptation framework for global and local video editing, pushing the limits of consistently editing minute-long videos.\nFirst, to ensure local consistency within individual frames, the foundation of VIA is a novel test-time editing adaptation method, which adapts a pre-trained image editing model for improving consistency between potential editing directions and the text instruction, and adapts masked latent variables for precise local control.\nFurthermore, to maintain global consistency over the video sequence, we introduce spatiotemporal adaptation that adapts consistent attention variables in key frames and strategically applies them across the whole sequence to realize the editing effects.\nExtensive experiments demonstrate that, compared to baseline methods, our VIA approach produces edits that are more faithful to the source videos, more coherent in the spatiotemporal context, and more precise in local control. More importantly, we show that VIA can achieve consistent long video editing in minutes, unlocking the potentials for advanced video editing tasks over long video sequences."},"_bibtex":{"value":"@inproceedings{\ngu2025via,\ntitle={{VIA}: A Spatiotemporal Video Adaptation Framework for Global and Local Video Editing},\nauthor={Jing Gu and Yuwei Fang and Ivan Skorokhodov and Peter Wonka and Xinya Du and Sergey Tulyakov and Xin Eric Wang},\nbooktitle={Workshop on Video-Language Models @ NeurIPS 2024},\nyear={2025},\nurl={https://openreview.net/forum?id=dOdPMh8ZnU}\n}"},"title":{"value":"VIA: A Spatiotemporal Video Adaptation Framework for Global and Local Video Editing"},"pdf":{"value":"/pdf/fa3a120b83b9efe82fae424225495a1c28763650.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Workshop/Video-Langauge_Models"},"paperhash":{"value":"gu|via_a_spatiotemporal_video_adaptation_framework_for_global_and_local_video_editing"},"authorids":{"value":["~Jing_Gu2","~Yuwei_Fang1","~Ivan_Skorokhodov1","~Peter_Wonka1","~Xinya_Du1","~Sergey_Tulyakov1","~Xin_Eric_Wang2"]},"track":{"value":"Long Paper Track (up to 9 pages)"},"authors":{"value":["Jing Gu","Yuwei Fang","Ivan Skorokhodov","Peter Wonka","Xinya Du","Sergey Tulyakov","Xin Eric Wang"]}},"tmdate":1736861080653,"pdate":1730081752980,"tcdate":1726002550027,"writers":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission36/Authors"],"signatures":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission36/Authors"],"forum":"dOdPMh8ZnU","license":"CC BY 4.0","number":36,"cdate":1726002550027,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Post_Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission36/-/Full_Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Edit","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission36/-/Camera-Ready_Revision"],"mdate":1736861080653,"odate":1736861080628,"domain":"NeurIPS.cc/2024/Workshop/Video-Langauge_Models","id":"dOdPMh8ZnU","version":2},{"content":{"TLDR":{"value":"Linguistic minimal pairs unveil the internal linguistic representations of large language models, revealing consistency in high-resource languages and alignment with fine-grained theoretical categories."},"venue":{"value":"MINT@NeurIPS2024"},"pdf":{"value":"/pdf/a12574405a273c7e93f483a56d76f6d8c72a754c.pdf"},"keywords":{"value":["Large Language Model","Linguistic minimal pairs"]},"venueid":{"value":"NeurIPS.cc/2024/Workshop/MINT"},"paperhash":{"value":"zhou|linguistic_minimal_pairs_elicit_linguistic_similarity_in_large_language_models"},"authorids":{"value":["~Xinyu_Zhou11","~Delong_Chen1","~Samuel_Cahyawijaya1","~Xufeng_Duan1","~Zhenguang_Cai1"]},"abstract":{"value":"We introduce a novel analysis that leverages linguistic minimal pairs to probe the internal linguistic representations of Large Language Models (LLMs). By measuring the similarity between LLM activation differences across minimal pairs, we quantify linguistic similarity and gain insight into the linguistic knowledge captured by LLMs. Our large-scale experiments, spanning over 100 LLMs and 150,000 minimal pairs in three languages, reveal that linguistic similarity is more consistent in high-resource languages, influenced by training data composition, and strongly aligned with fine-grained theoretical linguistic categories, but weakly aligned with broader categories. This work demonstrates the potential of minimal pairs as a window into the neural representations of language, shedding light on the relationship between LLMs and linguistic theory."},"email_of_author_nominated_as_reviewer":{"value":"delong.chen@connect.ust.hk"},"_bibtex":{"value":"@inproceedings{\nzhou2024linguistic,\ntitle={Linguistic Minimal Pairs Elicit Linguistic Similarity in Large Language Models},\nauthor={Xinyu Zhou and Delong Chen and Samuel Cahyawijaya and Xufeng Duan and Zhenguang Cai},\nbooktitle={🍃 MINT: Foundation Model Interventions},\nyear={2024},\nurl={https://openreview.net/forum?id=BfLvu8JERw}\n}"},"title":{"value":"Linguistic Minimal Pairs Elicit Linguistic Similarity in Large Language Models"},"authors":{"value":["Xinyu Zhou","Delong Chen","Samuel Cahyawijaya","Xufeng Duan","Zhenguang Cai"]}},"tmdate":1734279411748,"pdate":1728466361816,"tcdate":1726310765041,"writers":["NeurIPS.cc/2024/Workshop/MINT","NeurIPS.cc/2024/Workshop/MINT/Submission36/Authors"],"signatures":["NeurIPS.cc/2024/Workshop/MINT/Submission36/Authors"],"forum":"BfLvu8JERw","license":"CC BY 4.0","number":36,"cdate":1726310765041,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Workshop/MINT/-/Submission","NeurIPS.cc/2024/Workshop/MINT/-/Post_Submission","NeurIPS.cc/2024/Workshop/MINT/-/Edit","NeurIPS.cc/2024/Workshop/MINT/Submission36/-/Camera_Ready"],"mdate":1734279411748,"odate":1734279411732,"domain":"NeurIPS.cc/2024/Workshop/MINT","id":"BfLvu8JERw","version":2},{"content":{"venue":{"value":"CoRR 2025"},"pdf":{"value":"https://arxiv.org/pdf/2504.02768v4"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"jumelet|multiblimp_10_a_massively_multilingual_benchmark_of_linguistic_minimal_pairs"},"html":{"value":"https://doi.org/10.48550/arXiv.2504.02768"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2504-02768,\n  publtype={informal},\n  author={Jaap Jumelet and Leonie Weissweiler and Arianna Bisazza},\n  title={MultiBLiMP 1.0: A Massively Multilingual Benchmark of Linguistic Minimal Pairs},\n  year={2025},\n  month={April},\n  cdate={1743465600000},\n  journal={CoRR},\n  volume={abs/2504.02768},\n  url={https://doi.org/10.48550/arXiv.2504.02768}\n}\n"},"abstract":{"value":"We introduce MultiBLiMP 1.0, a massively multilingual benchmark of linguistic minimal pairs, covering 101 languages and 2 types of subject-verb agreement, containing more than 128,000 minimal pairs. Our minimal pairs are created using a fully automated pipeline, leveraging the large-scale linguistic resources of Universal Dependencies and UniMorph. MultiBLiMP 1.0 evaluates abilities of LLMs at an unprecedented multilingual scale, and highlights the shortcomings of the current state-of-the-art in modelling low-resource languages."},"title":{"value":"MultiBLiMP 1.0: A Massively Multilingual Benchmark of Linguistic Minimal Pairs"},"authors":{"value":[{"fullname":"Jaap Jumelet","username":""},{"fullname":"Leonie Weissweiler","username":"~Leonie_Weissweiler1"},{"fullname":"Arianna Bisazza","username":""}]}},"tmdate":1790663520861,"pdate":1767139200000,"externalIds":["dblp:journals/corr/abs-2504-02768"],"tcdate":1790663516243,"writers":["~","OpenReview.net/Public_Article/DBLP.org","OpenReview.net/Support"],"signatures":["~Leonie_Weissweiler1"],"forum":"ov28tNNdgT","license":"CC BY-SA 4.0","number":158853,"cdate":1743465600000,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/DBLP.org/-/Record","OpenReview.net/Public_Article/-/Edit"],"mdate":1790663520861,"domain":"OpenReview.net/Public_Article","id":"ov28tNNdgT","version":2},{"content":{"summary":{"value":"This paper aims to leverage the pre-trained single-modal diffusion models for generating aligned video and audio. To achieve this objective and involve minimal computing costs, the authors proposed a lightweight joint guidance module to adjust the independent distribution from base models to match the joint distribution over audio and video by designing and training a discriminator. Based on the experimental results on the benchmark datasets, it seems that the proposed method improves the alignment score of generated audio and video while involving reduced trainable parameters."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. Did you adopt other evaluation metrics to evaluate the proposed method on the performance of temporal alignment between generated audio and video? Because it seems that the ImageBind score is just for semantic alignment.\n2. Did you try different discriminator structures or larger trainable parameters to see if there will be future performance improvement? \n3. Why the metric scores of MM-Diffusion in Table 2 are different from the ones of the original paper? \n4. It is better to compare the proposed method with \"Seeing and Hearing\" discussed in the related works by using the same pre-trained base models."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. It is interesting to solve the joint audio-video generative problems by adjusting the pre-trained modality-specific generators with minimal trainable parameters rather than fully training the joint audio-video generation models.\n2. Sufficient mathematic equations and proofs are given to support that the discriminator-based guidance can adjust the estimated scores from audio diffusion and video diffusion to approximately match the score of joint audio-video distribution.\n3. The experimental results shown in the table present the proposed method can improve the alignment performance of generated audio and video in both in-domain and out-of-domain settings.\n4. The paper is generally well-written."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The description that training a discriminator can adjust the pre-trained modality-specific diffusion models towards a joint audio-video generator is not very clear. For example, how does the discriminator-based guidance achieve both semantic and temporal alignment?  It is better to make it more clear. \n2. The improvement of metric score is limited especially for the OOD setting. In addition, from Table 4, the FVD and IB-TV scores of AudioLDM/AnimateDiff with the proposed guidance module are even worse. \n3. Based on the provided video files, although the performance in the IND is improved by the proposed method, the OOD setting performs poorly in generating high-quality video with the audio track. \n4. It is meaningful to see how different base models affect the performance. Whether more powerful pre-trained video and audio generators further improve the performance when equipped with the same guidance module?"}},"nonreaders":[],"tmdate":1733116500673,"tcdate":1730581206839,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission5876/Reviewer_Kk8X"],"signatures":["ICLR.cc/2025/Conference/Submission5876/Reviewer_Kk8X"],"forum":"agbiPPuSeQ","number":2,"license":"CC BY 4.0","cdate":1730581206839,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission5876/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1733116500673,"domain":"ICLR.cc/2025/Conference","replyto":"agbiPPuSeQ","id":"1zgsG6lnaR","forumContent":{"TLDR":{"value":"Building audio-video joint generative model with minimal computational cost via multimodal discriminator on the top of pre-trained single-modal generative models"},"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Diffusion models","Multi-modal data","Audio-visual generative models"]},"supplementary_material":{"value":"/attachment/002a0a08a630cdbec4ef8021c406d0e8b4216e96.zip"},"primary_area":{"value":"generative models"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"This study aims to construct an audio-video generative model with minimal computational cost by leveraging pre-trained single-modal generative models for audio and video.\nTo achieve this, we propose a novel method that guides single-modal models to cooperatively generate well-aligned samples across modalities. \nSpecifically, given two pre-trained base diffusion models, we train a lightweight joint guidance module to adjust scores separately estimated by the base models to match the score of joint distribution over audio and video. \nWe show that this guidance can be computed using the gradient of the optimal discriminator, which distinguishes real audio-video pairs from fake ones independently generated by the base models. \nBased on this analysis, we construct a joint guidance module by training this discriminator.\nAdditionally, we adopt a loss function to stabilize the discriminator's gradient and make it work as a noise estimator, as in standard diffusion models. \nEmpirical evaluations on several benchmark datasets demonstrate that our method improves both single-modal fidelity and multimodal alignment with relatively few parameters.\nThe code is available at: [https://github.com/SonyResearch/MMDisCo](https://github.com/SonyResearch/MMDisCo)."},"_bibtex":{"value":"@inproceedings{\nhayakawa2025mmdisco,\ntitle={{MMD}isCo: Multi-Modal Discriminator-Guided Cooperative Diffusion for Joint Audio and Video Generation},\nauthor={Akio Hayakawa and Masato Ishii and Takashi Shibuya and Yuki Mitsufuji},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=agbiPPuSeQ}\n}"},"title":{"value":"MMDisCo: Multi-Modal Discriminator-Guided Cooperative Diffusion for Joint Audio and Video Generation"},"pdf":{"value":"/pdf/d2c6e85ea8b816d429d4916770364984d9ec58b1.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"hayakawa|mmdisco_multimodal_discriminatorguided_cooperative_diffusion_for_joint_audio_and_video_generation"},"authorids":{"value":["~Akio_Hayakawa1","~Masato_Ishii1","~Takashi_Shibuya1","~Yuki_Mitsufuji1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Akio Hayakawa","Masato Ishii","Takashi Shibuya","Yuki Mitsufuji"]}},"version":2},{"content":{"summary":{"value":"The paper introduces Video-EM, a training-free framework designed to improve the performance of Video Large Language Models in understanding long-form videos by overcoming context window limitations and frame redundancy. Video-EM reformulates long-form video question answering by treating isolated keyframes as temporally ordered episodic events, capturing essential spatio-temporal relationships often missed by traditional static sampling methods. The framework involves three core components: Key Event Selection, Episodic Memory Representation (which encodes dynamic scene narratives and relationships), and a Chain-of-Thought reasoning module that iteratively selects a minimal yet highly informative subset of memories. Extensive experiments across four long-video benchmarks demonstrate that Video-EM enhances the accuracy and efficiency of some Video-LLM backbones using fewer frames on average."},"soundness":{"value":1},"confidence":{"value":5},"questions":{"value":"See weaknesses plus\n- The paper highlights the reduction in frames input to the final Video-LLM (e.g., 41 frames down to an average of 9 on EgoSchema). What is the complete end-to-end inference latency (or total computational cost) for the full Video-EM pipeline (including Key Event Selection, Episodic Memory Representation, and Chain-of-Thought steps using all five foundation models and Qwen3-8B)? How does this total cost compare to the baseline model running the maximum allowed frame input?\n- The paper acknowledges that the method is \"limited by the accuracy of captioners and object detectors\". What testing or simulation was performed to quantify how a decrease in accuracy (e.g., failure rate) in a crucial upstream component (such as Grounding-DINO missing key objects or Tarsier2-7B generating an inaccurate Dynamic Scene Narrative) propagates and impacts the final Video-LLM performance?\n- Given the strong claims of superiority over prior methods, why were empirical comparisons against other existing plug-and-play, training-free long-video understanding frameworks with similar goals, such as HERMES (which also uses episodes and semantics), omitted? Providing context for these comparisons is crucial for substantiating Video-EM's novelty and competitive edge in the crowded field of V-LLM accelerators.\n- In the description of the multi-grained semantic retrieval (L193 onwards), is the summation of equation (1) over the set Q={q1,q2,q3} or Q={q,qo,qs}? In other words, what is qi and why do we have Wq1, Wq2 and Wq3 but no Wq, Wqo and Wqs?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"- Instead of treating selected frames as disconnected images (a stated limitation of previous keyframe retrieval methods), Video-EM reformulates them as temporally ordered episodic events, avoiding the temporal discontinuities that often disrupt the semantic narrative of events in traditional methods.\n- Video-EM leverages a Chain-of-Thought (CoT) thinking strategy to iteratively identify and retrieve a minimal yet informative subset of episodic memories.\n- Video-EM is a training-free framework that can be integrated with off-the-shelf Video-LLM backbones without requiring retraining or architectural modification."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Video-EM is a complex, multi-stage pipeline that relies heavily on the quality and coordination of several external, specialized foundation models.\n- The concept of episodic memory has been explored by HERMES [1], also a plug-and-play model, with similar claims as Video-EM, yet the differences/similarities between the two are not specified, nor were the results of HERMES discussed in the manuscript.\n- Several plug-and-play modules for Video-LLM accuracy/efficiency improvements have been published in recent years such as FastV [2], VisionZip [3], VFlowOpt [4] in addition to the aforementioned HERMES [1]. I am curious about the comparison results with theses other plug-and-play frameworks in terms of accuracy/efficiency tradeoffs, and also in terms of methodology.\n- While Video-EM successfully reduces the number of frames processed by the final Video-LLM (from 41 frames down to an average of 9 on EgoSchema, for example), the preceding processing steps require extensive computation across multiple large models (CLIP, DINOv2, RAFT, Grounding-DINO, Tarsier2-7B, Qwen3-8B). I thus believe, the slight accuracy improvement does not justify the upstream cost of putting such a system together.\n- I also think such a system is very fragile. A deficiency in the initial retrieval stage (Key Event Selection) or the intermediate processing stages directly impacts the quality of the final input provided to the Video-LLM. It follows that these results would be a headache to replicate.\n- The author’s efficiency claims are not substantiated. Fewer frames do not equal more efficient.\n- Ambiguous variable definition: In section 3.2, the \"Adaptive Event Expansion\" paragraph, the authors define alpha as a variable with a value between 0 and 1, yet immediately after that, the paper states that alpha is set to 2. I am quite confused by that. \n- I think Figure 2 has too much text, is quite convoluted, and the bright red color is not easy on the eye. \n\n\n[1] Faure, Gueter Josmy, et al. \"Hermes: temporal-coherent long-form understanding with episodes and semantics.\" Proceedings of the IEEE/CVF International Conference on Computer Vision. 2025.\n\n[2] Chen, Liang, et al. \"An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models.\" European Conference on Computer Vision. Cham: Springer Nature Switzerland, 2024.\n\n[3] Yang, Senqiao, et al. \"Visionzip: Longer is better but not necessary in vision language models.\" Proceedings of the Computer Vision and Pattern Recognition Conference. 2025.\n\n[4] Yang, Sihan, et al. \"Vflowopt: A token pruning framework for lmms with visual information flow-guided optimization.\" Proceedings of the IEEE/CVF International Conference on Computer Vision. 2025."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762920278870,"tcdate":1761642898588,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission8370/Reviewer_g7A1"],"signatures":["ICLR.cc/2026/Conference/Submission8370/Reviewer_g7A1"],"forum":"aLQsPnVNnk","number":2,"license":"CC BY 4.0","cdate":1761642898588,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission8370/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762920278870,"domain":"ICLR.cc/2026/Conference","replyto":"aLQsPnVNnk","id":"CzzjPcs9v1","forumContent":{"TLDR":{"value":"We introduce Video-EM, a training‑free framework that treats long video question answering as an episodic memory retrieval-and‑reasoning problem inspired by human cognitive psychology."},"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Multi-modal Vision","Video Understanding"]},"primary_area":{"value":"foundation or frontier models, including LLMs"},"abstract":{"value":"Video Large Language Models (Video-LLMs) excel at general video understanding but struggle with long-form videos due to context-window limits. Consequently, recent approaches focus on keyframe retrieval, condensing lengthy videos into a small set of informative frames. Despite their practicality, these methods simplify the problem to static text-image matching, overlooking spatio-temporal relationships crucial for capturing scene transitions and contextual continuity, and may yield redundant keyframes with limited information, diluting salient cues essential for accurate video question answering. To address these limitations, we introduce Video-EM, a training-free framework inspired by the principles of human episodic memory, designed to facilitate robust and contextually grounded reasoning. Rather than treating keyframes as isolated visual entities, Video-EM explicitly models them as temporally ordered episodic events, capturing both spatial relationships and temporal dynamics necessary for accurately reconstructing the underlying narrative. Furthermore, the framework leverages chain-of-thought (CoT) thinking with LLMs to iteratively identify a minimal yet highly informative subset of episodic memories, enabling efficient and accurate question answering by Video-LLMs. Extensive evaluations on multiple mainstream long-video benchmarks demonstrate the superiority of Video-EM, which achieves highly competitive results while using fewer frames."},"_bibtex":{"value":"@misc{\nwang2026episodic,\ntitle={Episodic Memory Representation for Long Video Understanding},\nauthor={Yun wang and Long Zhang and Jingren Liu and Jiaqi Yan and Zhanjie Zhang and Jiahao Zheng and Ao Ma and Xun Yang and Dapeng Wu and Xiangyu Chen and Xuelong Li},\nyear={2026},\nurl={https://openreview.net/forum?id=aLQsPnVNnk}\n}"},"title":{"value":"Episodic Memory Representation for Long Video Understanding"},"pdf":{"value":"/pdf/f020aad44e91a19ec0c3c30fd66c852f9da5374f.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"wang|episodic_memory_representation_for_long_video_understanding"},"authorids":{"value":["~Yun_wang13","~Long_Zhang4","~Jingren_Liu1","~Jiaqi_Yan3","~Zhanjie_Zhang2","~Jiahao_Zheng1","~Ao_Ma2","~Xun_Yang1","~Dapeng_Wu1","~Xiangyu_Chen5","~Xuelong_Li2"]},"authors":{"value":["Yun wang","Long Zhang","Jingren Liu","Jiaqi Yan","Zhanjie Zhang","Jiahao Zheng","Ao Ma","Xun Yang","Dapeng Wu","Xiangyu Chen","Xuelong Li"]}},"version":2},{"content":{"summary":{"value":"This paper introduces SparseCut, a general multimodal fusion framework for MLLMs, which enhances cross-modal understanding by establishing sparse shortcut connections between multiple layers of the vision encoder and the language model. These shortcuts allow hierarchical and multi-grained visual information to be integrated efficiently without extending the LLM’s input length. By fusing high- and low-resolution visual features through cross-attention, SparseCut preserves rich semantics while maintaining computational efficiency. Experiments on various benchmarks show consistent improvements over LLaVA and DeepStack, achieving better performance with minimal additional cost and strong scalability across different base LLMs."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. The choice of shortcut pattern (density, distribution) may require manual tuning.\n2. The method relies on a frozen vision encoder, potentially limiting deeper cross-modal alignment."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":3},"strengths":{"value":"1. This paper proposes a method that can efficiently integrates multi-level and multi-resolution visual features without increasing computational cost.\n2. Through shortcut connections, SparseCut effectively incorporates multi-granularity visual features into the LLM while preserving its original context length and computational efficiency.\n3. The experimental results demonstrate strong generalization and scalability across different base LLMs."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The choice of shortcut pattern (density, distribution) may require manual tuning.\n2. The method relies on a frozen vision encoder, potentially limiting deeper cross-modal alignment."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762931428629,"tcdate":1761894698985,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission19542/Reviewer_opN4"],"signatures":["ICLR.cc/2026/Conference/Submission19542/Reviewer_opN4"],"forum":"p9Hc1o6By5","number":2,"license":"CC BY 4.0","cdate":1761894698985,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission19542/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762931428629,"domain":"ICLR.cc/2026/Conference","replyto":"p9Hc1o6By5","id":"073PQnsfzn","forumContent":{"TLDR":{"value":"A general cross-modal fusion architecture of MLLM with shortcut connnections for multi-level multi-grained feature integration."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Multimodal Large Language Models；Cross-modal Alignment；Multi-granularity Fusion"]},"supplementary_material":{"value":"/attachment/efffac5c3913013c5a3c47ef52ddcb88a0807f4a.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"With the remarkable success of large language models (LLMs) in natural language understanding and generation, multimodal large language models (MLLMs) have rapidly advanced in their ability to process data across multiple modalities.\nWhile most existing efforts focus on scaling up language models or constructing higher-quality training data, limited attention has been paid to effectively integrating cross-modal knowledge into the language space.\nIn vision-language models, for instance, aligning modalities using only high-level visual features often discards the rich semantic information present in mid- and low-level features, limiting the model’s ability of cross-modality understanding.\nTo address this issue, we propose SparseCut, a general cross-modal fusion architecture for MLLMs, introducing sparse shortcut connections between the cross-modal encoder and the LLM. These shortcut connections enable the efficient and hierarchical integration of visual features at multiple levels, facilitating richer semantic fusion without increasing computational overhead.\nWe further introduce an efficient multi-grained feature fusion module, which performs the fusion of visual features before routing them through the shortcuts.\nThis preserves the original language context and does not increase the overall input length, thereby avoiding an increase in computational complexity for the LLM.\nWe systematically evaluate the performance of various shortcut patterns and demontrate that SparseCut can enhance the performance of MLLMs across various multimodal benchmarks with high training stability. It is also compatible with different base LLMs."},"_bibtex":{"value":"@misc{\nzhang2026sparse,\ntitle={Sparse Shortcuts: Facilitating Efficient Fusion in Multimodal Large Language Models},\nauthor={Jingrui Zhang and Yong Zhang and Feng Liang and Wei Wang and Runhao Zeng and Xiping Hu},\nyear={2026},\nurl={https://openreview.net/forum?id=p9Hc1o6By5}\n}"},"title":{"value":"Sparse Shortcuts: Facilitating Efficient Fusion in Multimodal Large Language Models"},"pdf":{"value":"/pdf/997a29bc50763364cde0f32664fad0d9da59c797.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"zhang|sparse_shortcuts_facilitating_efficient_fusion_in_multimodal_large_language_models"},"authorids":{"value":["~Jingrui_Zhang1","~Yong_Zhang23","~Feng_Liang5","~Wei_Wang80","~Runhao_Zeng1","~Xiping_Hu1"]},"authors":{"value":["Jingrui Zhang","Yong Zhang","Feng Liang","Wei Wang","Runhao Zeng","Xiping Hu"]}},"version":2},{"content":{"summary":{"value":"This paper proposes Flow Uniqueness Models (FUM), which splits the overall generative trajectory $P(\\epsilon \\rightarrow x_0)$ into two sub-paths, $P(\\epsilon \\rightarrow x_s)$ and $P(x_s \\rightarrow x_0)$. The first sub-path is to enforce velocity uniqueness via strictly one-to-one sample paris, while the second is aligned through a flow consistency constraint. To achieve this, this paper introduces two variants of flow consistency: a Shortcut Models-based strategy and a MeanFlow-based strategy. In addition, this paper reports empirical results on three benchmark datasets to demonstrate the performance of FUM."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"- How is R, the so-called reversible range, quantitatively determined? Please provide a reproducible criterion and sensitivity analysis.\n- Why is the sub-path division necessary at all? Show an ablation comparing full-path vs two-path training.\n- There are several typos and inconsistencies. For example, in Eq.(4), $v_{\\theta}(x_i, 0, i)$ should be $v_{\\theta}(x_i, i, s)$.\n- For clarity, it may be better to use $u$ to denote average velocity instead of $v$, to distinguish it from instantaneous velocity."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"+ This paper addresses the important direction of one-step generation and reports results on multiple datasets.\n+ This paper unifies ideas from Shortcut and MeanFlow models under a unified framework, with reproducible experiment settings."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Unconvincing motivation and weak necessity of sub-path division. Velocity ambiguity in flow matching originates from marginalization over random $(x_0, \\epsilon)$ pairs, not from the temporal structure itself. Simply splitting the trajectory into segments cannot reduce or eliminate this ambiguity. \n- Sampling scheme restricts model capacity. As shown in Eq.(3)/(4), the method samples $i \\sim U(0,s)$ and $j \\sim U(s,1)$, forcing training pairs to be drawn only across the two sub-paths. Unlike MeanFlow or Shortcut, which allow arbitrary $(i,j)\\in[0,1]^2$, this design reduces temporal coverage and learning diversity. Moreover, it potentially breaks smoothness of the learned continuous flow. \n- Limited experiments and modest performance. There is no ablation study demonstrating that the sub-path division improves training stability or performance. On ImageNet-64 (Table 2), FUM underperforms several strong baselines (e.g., sCT, ECM)."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762919152346,"tcdate":1761893336573,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission6913/Reviewer_VjJE"],"signatures":["ICLR.cc/2026/Conference/Submission6913/Reviewer_VjJE"],"forum":"ZMqIgONdJZ","number":1,"license":"CC BY 4.0","cdate":1761893336573,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission6913/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762919152346,"domain":"ICLR.cc/2026/Conference","replyto":"ZMqIgONdJZ","id":"W2PUQI6RJ5","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Flow matching","one-step generation","flow consistency"]},"primary_area":{"value":"generative models"},"abstract":{"value":"Recent advances in generative modeling frameworks, such as diffusion models and flow matching, have achieved record-breaking performance.\nNevertheless, these approaches involve iterative sampling procedures across many neural network passes, which severely limits their practical deployment, particularly in domains demanding real-time interaction.\nAlthough considerable effort has been devoted to accelerating sampling, achieving high-quality one-step generation remains an open challenge, motivating research into a new era of generative modeling.\nMotivated by this, we put forward a novel and effective framework, termed \\textit{Flow Uniqueness Models} (\\textbf{FUM}).\nThe core idea of FUM is to construct strictly one-to-one image pairs, thereby enforcing velocity uniqueness along the entire sampling path, which forms as the foundation for few-step sampling.\nBy leveraging this modeling mechanism, FUM not only achieves remarkable one-step generative performance but also provides the flexibility to balance image quality against the number of sampling steps.\nExtensive experiments on three benchmark datasets comprehensively validate the superiority of our proposed FUM."},"_bibtex":{"value":"@misc{\nzhang2026building,\ntitle={Building Flow Uniqueness in One-step Generative Modeling},\nauthor={Junyu Zhang and Daochang Liu and Liu.Liu and Zhizhong Su and Jong Hwan Ko and Shichao Zhang and Eunbyung Park and Chang Xu},\nyear={2026},\nurl={https://openreview.net/forum?id=ZMqIgONdJZ}\n}"},"title":{"value":"Building Flow Uniqueness in One-step Generative Modeling"},"pdf":{"value":"/pdf/e1f62ddaec95e6411b4597cb38cc6a0699d5c444.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"zhang|building_flow_uniqueness_in_onestep_generative_modeling"},"authorids":{"value":["~Junyu_Zhang2","~Daochang_Liu1","~Liu.Liu19","~Zhizhong_Su3","~Jong_Hwan_Ko2","~Shichao_Zhang3","~Eunbyung_Park1","~Chang_Xu4"]},"authors":{"value":["Junyu Zhang","Daochang Liu","Liu.Liu","Zhizhong Su","Jong Hwan Ko","Shichao Zhang","Eunbyung Park","Chang Xu"]}},"version":2},{"content":{"summary":{"value":"This work proposes a novel video-language framework, HAWK, aiming at understanding video anomalies, which incorporates motion modality to enhance its capability. The work generates rich language descriptions for seven different video anomaly datasets, and also generates question-answer pairs to tackle potential user inquiries. The proposed framework demonstrates SOTA performance for video anomaly understanding and question-answering across multiple scenarios, which will advance the open-world anomaly understanding field."},"soundness":{"value":4},"confidence":{"value":5},"questions":{"value":"See the Weakness part."},"rating":{"value":7},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":4},"strengths":{"value":"The work proposes a novel vision-language framework to address open-world video anomaly understanding, which is a very different pipeline from previous classification-based anomaly detection pipelines. To build a vision-language model for open-world video anomaly understanding, the work adopts seven different video anomaly datasets for training, where rich language descriptions and question-answer pairs are generated. Experiments for video anomaly understanding and question-answering across multiple scenarios demonstrate the effectiveness of the proposed framework. I believe the work will advance the field of open-world video anomaly understanding."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The work adopts Gunnar Farneback’s algorithm to obtain the motion modality. Is this algorithm efficient? \n2. The Ubnormal dataset consists of virtual anomaly videos. Do the authors consider the gap between real and virtual anomaly videos, which is mentioned in previous works [1, 2]?\n\n[1] Ubnormal: New benchmark for supervised open-set video anomaly detection, CVPR 2022\n\n[2] Generating Anomalies for Video Anomaly Detection with Prompt-based Feature Mapping, CVPR 2023"},"limitations":{"value":"The paper has discussed the limitations and potential impacts of the work."}},"nonreaders":[],"tmdate":1730878748889,"tcdate":1720409654341,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission2096/Reviewer_kzaw"],"signatures":["NeurIPS.cc/2024/Conference/Submission2096/Reviewer_kzaw"],"forum":"vBKoEZ1PG3","number":1,"license":"CC BY 4.0","cdate":1720409654341,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission2096/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730878748889,"domain":"NeurIPS.cc/2024/Conference","replyto":"vBKoEZ1PG3","id":"HdBDxzzW4d","forumContent":{"venue":{"value":"NeurIPS 2024 poster"},"TLDR":{"value":"A large vision-language model for understanding open-world video anomalies."},"keywords":{"value":["Video Anomalies Understanding"]},"primary_area":{"value":"machine_vision"},"abstract":{"value":"Video Anomaly Detection (VAD) systems can autonomously monitor and identify disturbances, reducing the need for manual labor and associated costs. However, current VAD systems are often limited by their superficial semantic understanding of scenes and minimal user interaction. Additionally, the prevalent data scarcity in existing datasets restricts their applicability in open-world scenarios.\nIn this paper, we introduce HAWK, a novel framework that leverages interactive large Visual Language Models (VLM) to interpret video anomalies precisely. Recognizing the difference in motion information between abnormal and normal videos, HAWK explicitly integrates motion modality to enhance anomaly identification. To reinforce motion attention, we construct an auxiliary consistency loss within the motion and video space, guiding the video branch to focus on the motion modality. Moreover, to improve the interpretation of motion-to-language, we establish a clear supervisory relationship between motion and its linguistic representation. Furthermore, we have annotated over 8,000 anomaly videos with language descriptions, enabling effective training across diverse open-world scenarios, and also created 8,000 question-answering pairs for users' open-world questions. The final results demonstrate that HAWK achieves SOTA performance, surpassing existing baselines in both video description generation and question-answering. Our codes/dataset/demo will be released at https://github.com/jqtangust/hawk."},"_bibtex":{"value":"@inproceedings{\ntang2024hawk,\ntitle={{HAWK}: Learning to Understand Open-World Video Anomalies},\nauthor={Jiaqi Tang and Hao LU and RUIZHENG WU and Xiaogang Xu and Ke Ma and Cheng Fang and Bin Guo and Jiangbo Lu and Qifeng Chen and Ying-Cong Chen},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=vBKoEZ1PG3}\n}"},"title":{"value":"HAWK: Learning to Understand Open-World Video Anomalies"},"pdf":{"value":"/pdf/72f5b57d7e449d765338f2a60df13977e0717488.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"tang|hawk_learning_to_understand_openworld_video_anomalies"},"authorids":{"value":["~Jiaqi_Tang1","~Hao_LU8","~RUIZHENG_WU1","~Xiaogang_Xu2","~Ke_Ma11","~Cheng_Fang6","~Bin_Guo3","~Jiangbo_Lu1","~Qifeng_Chen1","~Ying-Cong_Chen1"]},"authors":{"value":["Jiaqi Tang","Hao LU","RUIZHENG WU","Xiaogang Xu","Ke Ma","Cheng Fang","Bin Guo","Jiangbo Lu","Qifeng Chen","Ying-Cong Chen"]}},"version":2},{"content":{"research_area_keywords":{"value":"linguistic theories, benchmarking, language resources, evaluation, less-resourced languages"},"languages_studied":{"value":"Turkish"},"venue":{"value":"ACL ARR 2025 May Submission"},"_bibtex":{"value":"@inproceedings{\nanonymous2025turblimp,\ntitle={Tur{BL}i{MP}: A Turkish Benchmark of Linguistic Minimal Pairs},\nauthor={Anonymous},\nbooktitle={Submitted to ACL Rolling Review - May 2025},\nyear={2025},\nurl={https://openreview.net/forum?id=caQim4TYzK},\nnote={under review}\n}"},"title":{"value":"TurBLiMP: A Turkish Benchmark of Linguistic Minimal Pairs"},"contribution_types":{"value":["Data resources","Data analysis"]},"abstract":{"value":"We introduce TurBLiMP, the first Turkish benchmark of linguistic minimal pairs, designed to evaluate the linguistic abilities of monolingual and multilingual language models (LMs). Covering 16 linguistic phenomena with 1000 minimal pairs each, TurBLiMP fills an important gap in linguistic evaluation resources for Turkish. In designing the benchmark, we give extra attention to two properties of Turkish that remain understudied in current syntactic evaluations of LMs, namely word order flexibility and subordination through morphological processes. Our experiments on a wide range of LMs and a newly collected set of human acceptability judgments reveal that even cutting-edge Large LMs still struggle with grammatical phenomena that are not challenging for humans, and may also exhibit different sensitivities to word order and morphological complexity compared to humans."},"paper_type":{"value":"Long"},"pdf":{"value":"/pdf/8c3a997944ebf6f236b512d5cdb66f431ca087ea.pdf"},"research_area":{"value":"Linguistic theories, Cognitive Modeling and Psycholinguistics"},"venueid":{"value":"aclweb.org/ACL/ARR/2025/May/Submission"}},"tmdate":1782962219656,"tcdate":1747733254445,"writers":["aclweb.org/ACL/ARR/2025/May","aclweb.org/ACL/ARR/2025/May/Submission6357/Authors"],"signatures":["aclweb.org/ACL/ARR/2025/May/Submission6357/Authors"],"forum":"caQim4TYzK","license":"CC BY 4.0","number":6357,"cdate":1747733254445,"readers":["everyone"],"invitations":["aclweb.org/ACL/ARR/2025/May/-/Submission","aclweb.org/ACL/ARR/2025/May/-/Edit","aclweb.org/ACL/ARR/2025/May/-/Post_Submission","aclweb.org/ACL/ARR/2025/May/-/Preprint_Release_Submission","aclweb.org/ACL/ARR/2025/May/-/Preprint_Post_Submission"],"mdate":1782962219656,"odate":1753766020949,"domain":"aclweb.org/ACL/ARR/2025/May","id":"caQim4TYzK","version":2},{"content":{"venue":{"value":"CoRR 2025"},"pdf":{"value":"https://arxiv.org/pdf/2506.13487v2"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"basar|turblimp_a_turkish_benchmark_of_linguistic_minimal_pairs"},"html":{"value":"https://doi.org/10.48550/arXiv.2506.13487"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2506-13487,\n  publtype={informal},\n  author={Ezgi Basar and Francesca Padovani and Jaap Jumelet and Arianna Bisazza},\n  title={TurBLiMP: A Turkish Benchmark of Linguistic Minimal Pairs},\n  year={2025},\n  month={June},\n  cdate={1748736000000},\n  journal={CoRR},\n  volume={abs/2506.13487},\n  url={https://doi.org/10.48550/arXiv.2506.13487}\n}\n"},"abstract":{"value":"We introduce TurBLiMP, the first Turkish benchmark of linguistic minimal pairs, designed to evaluate the linguistic abilities of monolingual and multilingual language models (LMs). Covering 16 linguistic phenomena with 1000 minimal pairs each, TurBLiMP fills an important gap in linguistic evaluation resources for Turkish. In designing the benchmark, we give extra attention to two properties of Turkish that remain understudied in current syntactic evaluations of LMs, namely word order flexibility and subordination through morphological processes. Our experiments on a wide range of LMs and a newly collected set of human acceptability judgments reveal that even cutting-edge Large LMs still struggle with grammatical phenomena that are not challenging for humans, and may also exhibit different sensitivities to word order and morphological complexity compared to humans."},"title":{"value":"TurBLiMP: A Turkish Benchmark of Linguistic Minimal Pairs"},"authors":{"value":[{"fullname":"Ezgi Basar","username":"~Ezgi_Başar1"},{"fullname":"Francesca Padovani","username":""},{"fullname":"Jaap Jumelet","username":""},{"fullname":"Arianna Bisazza","username":""}]}},"tmdate":1779739776098,"pdate":1767139200000,"externalIds":["dblp:journals/corr/abs-2506-13487"],"tcdate":1779739772728,"writers":["~","OpenReview.net/Public_Article/DBLP.org","OpenReview.net/Support"],"signatures":["~Ezgi_Başar1"],"forum":"OmFixzlrOi","license":"CC BY-SA 4.0","number":24867,"cdate":1748736000000,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/DBLP.org/-/Record","OpenReview.net/Public_Article/-/Edit"],"mdate":1779739776098,"domain":"OpenReview.net/Public_Article","id":"OmFixzlrOi","version":2},{"content":{"venue":{"value":"EMNLP 2025"},"pdf":{"value":"https://aclanthology.org/2025.emnlp-main.834.pdf"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"basar|turblimp_a_turkish_benchmark_of_linguistic_minimal_pairs"},"html":{"value":"https://doi.org/10.18653/v1/2025.emnlp-main.834"},"_bibtex":{"value":"@inproceedings{DBLP:conf/emnlp/BasarPJB25,\n  author={Ezgi Basar and Francesca Padovani and Jaap Jumelet and Arianna Bisazza},\n  title={TurBLiMP: A Turkish Benchmark of Linguistic Minimal Pairs},\n  year={2025},\n  cdate={1735689600000},\n  pages={16495-16510},\n  url={https://doi.org/10.18653/v1/2025.emnlp-main.834},\n  booktitle={EMNLP},\n  crossref={conf/emnlp/2025}\n}\n"},"abstract":{"value":"We introduce TurBLiMP, the first Turkish benchmark of linguistic minimal pairs, designed to evaluate the linguistic abilities of monolingual and multilingual language models (LMs). Covering 16 linguistic phenomena with 1000 minimal pairs each, TurBLiMP fills an important gap in linguistic evaluation resources for Turkish. In designing the benchmark, we give extra attention to two properties of Turkish that remain understudied in current syntactic evaluations of LMs, namely word order flexibility and subordination through morphological processes. Our experiments on a wide range of LMs and a newly collected set of human acceptability judgments reveal that even cutting-edge Large LMs still struggle with grammatical phenomena that are not challenging for humans, and may also exhibit different sensitivities to word order and morphological complexity compared to humans."},"title":{"value":"TurBLiMP: A Turkish Benchmark of Linguistic Minimal Pairs"},"authors":{"value":[{"fullname":"Ezgi Basar","username":"~Ezgi_Başar1"},{"fullname":"Francesca Padovani","username":""},{"fullname":"Jaap Jumelet","username":""},{"fullname":"Arianna Bisazza","username":""}]}},"tmdate":1779739775812,"pdate":1767139200000,"externalIds":["dblp:conf/emnlp/BasarPJB25"],"tcdate":1779739772745,"writers":["~","OpenReview.net/Public_Article/DBLP.org","OpenReview.net/Support"],"signatures":["~Ezgi_Başar1"],"forum":"JJ0eq9QeOW","license":"CC BY-SA 4.0","number":24868,"cdate":1735689600000,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/DBLP.org/-/Record","OpenReview.net/Public_Article/-/Edit"],"mdate":1779739775812,"domain":"OpenReview.net/Public_Article","id":"JJ0eq9QeOW","version":2},{"content":{"venue":{"value":"CoRR 2025"},"pdf":{"value":"http://arxiv.org/pdf/2506.13487v1"},"venueid":{"value":"dblp.org/journals/CORR/2025"},"paperhash":{"value":"basar|turblimp_a_turkish_benchmark_of_linguistic_minimal_pairs"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Ezgi_Basar:","~Francesca_Padovani1","https://dblp.org/search/pid/api?q=author:Jaap_Jumelet:","https://dblp.org/search/pid/api?q=author:Arianna_Bisazza:"]},"html":{"value":"https://doi.org/10.48550/arXiv.2506.13487"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2506-13487,\n  publtype={informal},\n  author={Ezgi Basar and Francesca Padovani and Jaap Jumelet and Arianna Bisazza},\n  title={TurBLiMP: A Turkish Benchmark of Linguistic Minimal Pairs},\n  year={2025},\n  month={June},\n  cdate={1748736000000},\n  journal={CoRR},\n  volume={abs/2506.13487},\n  url={https://doi.org/10.48550/arXiv.2506.13487}\n}\n"},"abstract":{"value":"We introduce TurBLiMP, the first Turkish benchmark of linguistic minimal pairs, designed to evaluate the linguistic abilities of monolingual and multilingual language models (LMs). Covering 16 linguistic phenomena with 1000 minimal pairs each, TurBLiMP fills an important gap in linguistic evaluation resources for Turkish. In designing the benchmark, we give extra attention to two properties of Turkish that remain understudied in current syntactic evaluations of LMs, namely word order flexibility and subordination through morphological processes. Our experiments on a wide range of LMs and a newly collected set of human acceptability judgments reveal that even cutting-edge Large LMs still struggle with grammatical phenomena that are not challenging for humans, and may also exhibit different sensitivities to word order and morphological complexity compared to humans."},"title":{"value":"TurBLiMP: A Turkish Benchmark of Linguistic Minimal Pairs"},"authors":{"value":["Ezgi Basar","Francesca Padovani","Jaap Jumelet","Arianna Bisazza"]}},"tmdate":1753813738150,"pdate":1735689600000,"tcdate":1753813733484,"writers":["~"],"signatures":["~Francesca_Padovani1"],"forum":"ptqAcrCRYl","license":"CC BY-SA 4.0","number":600385,"cdate":1748736000000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1753813738150,"domain":"DBLP.org","id":"ptqAcrCRYl","version":2},{"content":{"summary":{"value":"This paper proposes Efficient-SAM2, a post-training acceleration framework to address the significant computational bottleneck of SAM2 in real-time video object segmentation. The authors identify a mismatch between SAM2's dense computation and its inherently sparse perception, highlighting redundancy in the image encoder's background processing and in full-token memory retrieval. To exploit this, the framework introduces two key components: object-aware Sparse Window Routing (SWR) and object-aware Sparse Memory Retrieval (SMR). SWR dynamically routes irrelevant background windows in the encoder to a lightweight shortcut branch, guided by saliency and consistency cues from the previous frame's decoder. SMR leverages temporal consistency by identifying a sparse set of salient memory tokens during their first recollection and reusing this pattern for subsequent frames, drastically reducing memory attention computations."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"While the paper presents a valuable approach to making SAM2 more efficient, the design of the SWR and SMR modules appears highly heuristic and is tightly coupled to assumptions about temporal continuity. This raises a significant question: Is SAM2's strong, general-purpose segmentation performance fully preserved? The evaluation is currently confined to standard VOS benchmarks. To truly validate that these heuristics do not compromise the model's robustness, I would ask the authors to provide evaluations on more diverse datasets, particularly on challenging \"in-the-wild\" videos, which would better test the limits of these heuristic assumptions."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The motivation for reducing SAM2's computational overhead is well-grounded and intuitive for video object segmentation\n2. The post-training approach is practical, enabling efficient adaptation by leveraging the generalized parameters of the pre-trained SAM2.\n3. The method achieves a good speed-performance trade-off, delivering a speedup of nearly 2x while incurring only a minimal and acceptable performance degradation of approximately 1%."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The SWR component is heavily dependent on the previous frame's prediction and salient mask. I think this may cause challenges in some cases, such as rapid motion, abrupt scene cuts, or severe occlusions, where this temporal assumption would be violated.\n2. For a video domain paper, the qualitative results with static images are insufficient. Supplemental videos would be significantly stronger to properly demonstrate temporal consistency, failure modes (especially in scenarios mentioned in point 1), and the practical impact of the optimizations.\n3. I think an evaluation is needed to determine if the SMR module, which uses a cached saliency pattern, maintains its effectiveness in long video scenarios where significant appearance and context drift are likely.\n4. The paper lacks a dedicated discussion of its limitations. This omission leaves the impression that the proposed efficiencies might be confined to easy (or trained) scenarios."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762917786905,"tcdate":1761975880754,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission4947/Reviewer_gSJw"],"signatures":["ICLR.cc/2026/Conference/Submission4947/Reviewer_gSJw"],"forum":"HcivRSezJp","number":2,"license":"CC BY 4.0","cdate":1761975880754,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission4947/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762917786905,"domain":"ICLR.cc/2026/Conference","replyto":"HcivRSezJp","id":"3gBwG9PwUq","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Segmemt Anything Model","Efficient Deep Learning","Model Acceleration"]},"supplementary_material":{"value":"/attachment/b3481d522f61dec0cb42900d807bcdca9c57c111.zip"},"primary_area":{"value":"infrastructure, software libraries, hardware, systems, etc."},"abstract":{"value":"Segment Anything Model 2 (SAM2) shows excellent performance in video object segmentation tasks; however, the heavy computational burden hinders its application in real-time video processing.\nAlthough there have been efforts to improve the efficiency of SAM2, most of them focus on retraining a lightweight backbone, with little exploration into post-training acceleration.\nIn this paper, we observe that SAM2 exhibits sparse perception pattern as biological vision, which provides opportunities for eliminating redundant computation and acceleration:\ni) In mask decoder, the attention primarily focuses on the foreground objects, whereas the image encoder in the earlier stage exhibits a broad attention span, which results in unnecessary computation to background regions.\nii) In memory bank, only a small subset of tokens in each frame contribute significantly to memory attention, and the salient regions exhibit temporal consistency, making full-token computation redundant.\nWith these insights, we propose Efficient-SAM2, which promotes SAM2 to adaptively focus on object regions while eliminating task-irrelevant computations, thereby significantly improving inference efficiency.\nSpecifically, for image encoder, we propose object-aware Sparse Window Routing (SWR), a window-level computation allocation mechanism that leverages the consistency and saliency cues from the previous-frame decoder to route background regions into a lightweight shortcut branch.\nMoreover, for memory attention, we propose object-aware Sparse Memory Retrieval (SMR), which allows only the salient memory tokens in each frame to participate in computation, \nwith the saliency pattern reused from their first recollection.\nWith negligible additional parameters and minimal training overhead, Efficient-SAM2 delivers 1.68$\\times$ speedup on SAM2.1-L model with only 1.0\\% accuracy drop on SA-V test set, where SWR and SMR provide 1.83$\\times$ and 1.78$\\times$ speedups, respectively."},"_bibtex":{"value":"@inproceedings{\nzhang2026efficientsam,\ntitle={Efficient-{SAM}2: Accelerating {SAM}2 with Object-Aware Visual Encoding and Memory Retrieval},\nauthor={Jing Zhang and Zhikai Li and Xuewen Liu and Qingyi Gu},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=HcivRSezJp}\n}"},"title":{"value":"Efficient-SAM2: Accelerating SAM2 with Object-Aware Visual Encoding and Memory Retrieval"},"pdf":{"value":"/pdf/b1801056e42ab7eb9562f53bd0d351cf557070d0.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"zhang|efficientsam2_accelerating_sam2_with_objectaware_visual_encoding_and_memory_retrieval"},"authorids":{"value":["~Jing_Zhang58","~Zhikai_Li1","~Xuewen_Liu1","~Qingyi_Gu1"]},"authors":{"value":["Jing Zhang","Zhikai Li","Xuewen Liu","Qingyi Gu"]}},"version":2},{"content":{"summary":{"value":"The paper introduces LongViTU, a large-scale dataset designed for long-form video understanding, featuring approximately 121k question-answer pairs across 900 hours of video content. It addresses challenges in long-form video understanding by offering a dataset with diverse real-world scenarios, explicit timestamp labels, long certificate lengths, fine-grained categorization, and open-ended precise QA pairs. LongViTU is curated to facilitate instruction tuning for long-form videos, involving the organization of video content into a hierarchical tree and incorporating self-revision mechanisms to ensure high-quality QA pairs. The authors primarily validate the effectiveness of LongViTU through experiments conducted on two different models."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1.\tIs it possible that using GPT-4 for evaluation may struggle to distinguish fine-grained semantics? For instance, if sentences differ by only one or two keywords but convey significantly different meanings, how would GPT-4 rate them in such cases?\n\n2.\tCan LongViTU still deliver substantial performance improvements on models that perform better?"},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1.\tThe approach of organizing video content into a hierarchical tree structure is innovative. This method allows for the generation of question-answer pairs that capture both spatial and temporal details, which is a creative extension of existing video understanding frameworks.\n2.\tThe dataset provides fine-grained categorization of questions, which is crucial for advancing the understanding of complex video content and adds depth to the quality of the dataset."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1.\tIn Table 2, it can be observed that there is a lack of differentiation in the benchmark. The performance gap between the best-performing Gemini-1.5-Pro and the other models is not evident. According to the reviewer, in most existing benchmarks, Gemini-1.5-Pro demonstrates a significant performance advantage over Video-LLaVA.\n\n2.\tThe proposed benchmark employs GPT-4 for assessment, which may introduce additional bias. \n\n3.\tThe validation method employed was released some time ago, and its baseline performance is no longer highly competitive compared to more recent models. It remains unclear whether it can still deliver significant performance improvements on more recently proposed models."}},"nonreaders":[],"tmdate":1731428926224,"tcdate":1730603941940,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission7604/Reviewer_fVx1"],"signatures":["ICLR.cc/2025/Conference/Submission7604/Reviewer_fVx1"],"forum":"4j9plQoOH1","number":2,"license":"CC BY 4.0","cdate":1730603941940,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission7604/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428926224,"domain":"ICLR.cc/2025/Conference","replyto":"4j9plQoOH1","id":"unj6rj7Dif","forumContent":{"TLDR":{"value":"We propose a large-scale instruction-tuning dataset for long-form video understanding."},"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["vision language models","instruction-tuning","long-form video understanding"]},"primary_area":{"value":"datasets and benchmarks"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"This paper presents LongViTU, a large-scale (~121k QA pairs, ~900h videos), automatically generated dataset for long-form video understanding. Our key idea is inspired by the success of Large Language Models (LLMs) and Multimodal Language Models (MLMs) that are fueled by machine-generated instruction-following data (*e.g.*, InstructGPT, LLaVA). We developed a *systematic* approach to produce massive question-answeringing pairs tailored to virtually unbounded long videos by organizing them into a ***hierarchical tree***, incorporating ***self-revision*** mechanisms to guarantee high quality. We curate LongViTU for each QA pair: 1) involves a long context (average *certificate length* of 4.6 minutes); 2) requires rich knowledge and condensed reasoning (commonsense, causality, planning, *etc.*); 3) explicit labels the timestamps of relevant events throughout the entire video. Furthermore, LongViTU provides a benchmark to facilitate future research in instruction-following for long-form videos. Our experiments first reveal the performance gap between open-source video MLMs and their commercial counterparts (*e.g.*, Gemini-1.5-Pro) on this benchmark. Supervised Fine-Tuning (SFT) on open-source models led to Video-LLaVA achieving the best performance, with a GPT-4 score of $50.7$, closely following $52.3$ by the leading closed-source model Gemini-1.5-Pro, underscoring the substantial challenge posed by our benchmark. Further SFT on LongViTU with Video-LLaVA resulted in improvements of $30.7$% on the In-Distribution (ID) benchmark EgoSchema; $12.9$% and $0.6$% on the Out-of-Distribution (OOD) benchmarks WorldQA and VideoMME, respectively. These outcomes demonstrate the effectiveness and robust OOD generalizability of our proposed instruction-tuning scheme for long-form video understanding. The dataset, SFT models, and code are publicly available on the anonymous page [LongViTU](https://longvitu.github.io)."},"_bibtex":{"value":"@misc{\nwu2024longvitu,\ntitle={LongVi{TU}: Instruction Tuning for Long-Form Video Understanding},\nauthor={Rujie Wu and Xiaojian Ma and Hai Ci and Yue Fan and Yuxuan Wang and Haozhe Zhao and Qing Li and Yizhou Wang},\nyear={2024},\nurl={https://openreview.net/forum?id=4j9plQoOH1}\n}"},"title":{"value":"LongViTU: Instruction Tuning for Long-Form Video Understanding"},"pdf":{"value":"/pdf/e663a2eb9e041444826a666f95acc8764c6e736b.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"wu|longvitu_instruction_tuning_for_longform_video_understanding"},"authorids":{"value":["~Rujie_Wu2","~Xiaojian_Ma1","~Hai_Ci1","~Yue_Fan2","~Yuxuan_Wang4","~Haozhe_Zhao1","~Qing_Li1","~Yizhou_Wang1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Rujie Wu","Xiaojian Ma","Hai Ci","Yue Fan","Yuxuan Wang","Haozhe Zhao","Qing Li","Yizhou Wang"]}},"version":2},{"content":{"venue":{"value":"Video-Langauge Models Poster"},"TLDR":{"value":"We examine the consistency of Video-LLMs in understanding temporal moments within videos, a key factor for robust video language comprehension, with various probes.."},"keywords":{"value":["Video Large Language Models; Video Temporal Reasoning"]},"supplementary_material":{"value":"/attachment/80994dc324244a360e5afc269371f487e1bfe353.zip"},"abstract":{"value":"Recent advancements in video large language models (Video-LLMs) have shown capabilities of temporally-grounding language queries or retrieving video moments in videos. However, such capabilities have not been thoroughly verified to be robust and trustable. In this study, we explore the consistency of Video-LLMs in grasping temporal moments within videos — a critical indicator for robust and trustworthy video language comprehension. Specifically, we devise different probes where Video-LLMs first predict temporal moments based on language queries, followed by verification questions to assess whether the predicted moments accurately reflect the queries. Our results show that current Video-LLMs respond unintuitively to such assessment; they often fail to provide consistent answers upon re-evaluation and even get near chance-level performance. This reveals the significant shortcomings in the current capabilities of Video-LLMs for reliable video temporal understanding, underscoring the need for further research and development in this field."},"_bibtex":{"value":"@inproceedings{\njung2025can,\ntitle={Can Video Large Language Models Comprehend Language in Videos?},\nauthor={Minjoon Jung and Junbin Xiao and Byoung-Tak Zhang and Angela Yao},\nbooktitle={Workshop on Video-Language Models @ NeurIPS 2024},\nyear={2025},\nurl={https://openreview.net/forum?id=HxBGjM9OGm}\n}"},"title":{"value":"Can Video Large Language Models Comprehend Language in Videos?"},"pdf":{"value":"/pdf/3631a49355d2485ee6a622838e9d1244170cc773.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Workshop/Video-Langauge_Models"},"paperhash":{"value":"jung|can_video_large_language_models_comprehend_language_in_videos"},"authorids":{"value":["~Minjoon_Jung2","~Junbin_Xiao1","~Byoung-Tak_Zhang1","~Angela_Yao1"]},"track":{"value":"Long Paper Track (up to 9 pages)"},"authors":{"value":["Minjoon Jung","Junbin Xiao","Byoung-Tak Zhang","Angela Yao"]}},"tmdate":1736861080175,"pdate":1730081752175,"tcdate":1725614345384,"writers":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission14/Authors"],"signatures":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission14/Authors"],"forum":"HxBGjM9OGm","license":"CC BY 4.0","number":14,"cdate":1725614345384,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Post_Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission14/-/Full_Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Edit","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission14/-/Camera-Ready_Revision"],"mdate":1736861080175,"odate":1736861080160,"domain":"NeurIPS.cc/2024/Workshop/Video-Langauge_Models","id":"HxBGjM9OGm","version":2},{"content":{"summary":{"value":"The paper studies training-free acceleration for autoregressive video diffusion. The authors empirically demonstrate that different chunks should have independent feature caching policies rather than a single global caching policy. A redundancy-aware KV-cache compression scheme is also adopted for long-video generation.\nExperiments on MAGI-1 and SkyReels-V2 demonstrate a noticeable speedup with a minor drop in VBench score."},"soundness":{"value":2},"confidence":{"value":5},"questions":{"value":"- line 99-101: Please clarify the experimental/testing setup for this result (e.g., GPU type, memory, batch size, video length). Without this, it’s hard to judge how generalizable the observation is.\n- The paper claims to “provide insights into memory–quality trade-offs,” but the experiments do not actually show memory vs. quality curves/tables (e.g., VBench vs. cache size/compression ratio). This weakens the contribution.\n- How is the chunkwise caching policy obtained in practice? Is it derived offline from a calibration set, or learned/heuristic? When generating videos of different lengths or motion patterns, does the policy need to be recomputed or adapted?\n- Do you have numerical benchmarks (peak memory, KV size per frame/chunk, vs. baseline) to substantiate the claim of reducing the memory with KV-cache compression?\n- misc:\n  - The citation of [1] is not about diffusion and seems unrelated to the context in which it is cited.\n  - VBench is a scaled score, not a percentage\n\n[1] Mengwei Xu, et la, Deepcache: Principled cache for mobile deep vision. 2018"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":1},"contribution":{"value":2},"strengths":{"value":"- The paper proposed training-free, plug-in acceleration for causal video diffusion. Treating each video chunk as its own, with an independent reuse policy, is well motivated by the observed heterogeneity across chunks at the same timesteps.\n- Experiment isolates the benefit of chunkwise feature reuse over full reuse and shows that kv-compression has a small impact."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Memory claims lack evidence. Memory usage is stated to be fixed, but no benchmark on memory usage is exhibited in the paper.\n- Lack of quality comparison. Only a few images are displayed in the paper. No video clip was provided, making it hard to evaluate the visual quality.\n- Lack of evaluation on long-video benchmark, e.g., VBench-long, since the method is claimed to be helpful for long-video generation.\n- MAGI-1 already applied window attention (8-second preceding video content). Weakening the motivation for KV-cache compression.\n- MAGI-1 has a shortcut step-distill version. The proposed method does not compare with it, nor apply the feature reuse to it.\n- Feature reuse necessarily increases memory consumption because the cache has to be retained. This seems to conflict with the stated motivation of KVcache compression, which is to reduce memory usage."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762942681652,"tcdate":1761979775040,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission23495/Reviewer_LS46"],"signatures":["ICLR.cc/2026/Conference/Submission23495/Reviewer_LS46"],"forum":"vko4DuhKbh","number":3,"license":"CC BY 4.0","cdate":1761979775040,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission23495/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762942681652,"domain":"ICLR.cc/2026/Conference","replyto":"vko4DuhKbh","id":"BiapuQ0QA6","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Autoregressive video generation","chunkwise caching","KV cache compression","ultra-long video synthesis","video acceleration"]},"primary_area":{"value":"generative models"},"abstract":{"value":"Autoregressive models, often built on Transformer architectures, represent a powerful paradigm for generating ultra-long videos by synthesizing content in sequential chunks. However, this sequential generation process is notoriously slow. While caching strategies have proven effective for accelerating traditional video diffusion models, existing methods assume uniform denoising across all frames—an assumption that breaks down in autoregressive models where different video chunks exhibit varying similarity patterns at identical timesteps.\nIn this paper, we present \\textbf{FlowCache}, the first caching framework specifically designed for autoregressive video generation. Our key insight is that each video chunk should maintain independent caching policies, allowing fine-grained control over which chunks require recomputation at each timestep. We introduce a chunkwise caching strategy that dynamically adapts to the unique denoising characteristics of each chunk, complemented by a joint importance–redundancy optimized KV cache compression mechanism that maintains fixed memory bounds while preserving generation quality.\nOur method achieves remarkable speedups of $\\textbf{2.38}\\times$ on MAGI-1 and $\\textbf{6.7}\\times$ on SkyReels-V2, with negligible quality degradation (VBench: $0.87\\uparrow$ and $0.79\\downarrow$ respectively). These results demonstrate that FlowCache, successfully unlocks the potential of autoregressive models for real-time, ultra-long video generation—establishing a new benchmark for efficient video synthesis at scale. The code is available at https://github.com/mikeallen39/FlowCache."},"_bibtex":{"value":"@inproceedings{\nma2026flow,\ntitle={Flow Caching for Autoregressive Video Generation},\nauthor={Yuexiao Ma and Xuzhe Zheng and Jing Xu and Xiwei Xu and Feng Ling and Xiawu Zheng and Huafeng Kuang and Huixia Li and XING WANG and Xuefeng Xiao and Fei Chao and Rongrong Ji},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=vko4DuhKbh}\n}"},"title":{"value":"Flow Caching for Autoregressive Video Generation"},"pdf":{"value":"/pdf/47168dd2cd36d5843cbd694a7589d9794c419f62.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"ma|flow_caching_for_autoregressive_video_generation"},"authorids":{"value":["~Yuexiao_Ma1","~Xuzhe_Zheng2","~Jing_Xu27","~Xiwei_Xu3","~Feng_Ling1","~Xiawu_Zheng1","~Huafeng_Kuang2","~Huixia_Li2","~XING_WANG3","~Xuefeng_Xiao1","~Fei_Chao1","~Rongrong_Ji5"]},"authors":{"value":["Yuexiao Ma","Xuzhe Zheng","Jing Xu","Xiwei Xu","Feng Ling","Xiawu Zheng","Huafeng Kuang","Huixia Li","XING WANG","Xuefeng Xiao","Fei Chao","Rongrong Ji"]}},"version":2},{"content":{"summary":{"value":"This paper presents a new video-to-audio model called SALSA-V that efficiently generates long-form audio given a target video. To perform long-form audio generation, the proposed method iteratively outpaints the previously generated audio, and for that purpose, the model is trained with masked ground-truth audio as conditional input. For efficient generation, the training objective includes a shortcut loss, which enables few-step generation during inference. The experimental results show that the proposed model outperforms existing models, especially when the duration of the target video is long or the number of sampling steps is small."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"- Is there any reason why the authors did not try SALSA-V with 1B parameters?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- The proposed method is simple and should be easy to implement.\n- The advantage of the proposed method over MMAudio is significant when the duration of the target video is long or the number of sampling steps is small."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The novelty of the methodology is marginal.\n  - The main modifications from MMAudio are training with masked audio and the use of shortcut loss, and both techniques appear to be adopted in a straightforward manner.\n  - It would be beneficial to clearly describe the particular challenges in applying these techniques to video-to-audio models and how the proposed method addresses them. Alternatively, the authors can mention any empirical insights that could be helpful for future studies in the field of video-to-audio generation.\n- The comparison in the experiments needs additional baselines.\n  - For efficient generation, Frieren (or applying reflow to the proposed model with a flow matching formulation), as mentioned in Section 2.3, would be a good baseline.\n  - For long-form generation, V-AURA (mentioned in Section 2.1) and LoVA (mentioned in Section 2.4) could be used as baselines.\n- The benefit of audio-conditional generation in SALSA-V is not clear in Figure 4.\n  - I understand that some audio patterns in the conditional audio faithfully appear in the generated audio, but it seems that the ground-truth audio does not actually contain such patterns within the generated timeline. Thus, it is not clear if the conditional audio contributes to accurate audio generation.\n- I could not access the demo page due to a timeout error."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762922766490,"tcdate":1761609364519,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission11727/Reviewer_gTEV"],"signatures":["ICLR.cc/2026/Conference/Submission11727/Reviewer_gTEV"],"forum":"3FcjKYmNY5","number":1,"license":"CC BY 4.0","cdate":1761609364519,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission11727/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762922766490,"domain":"ICLR.cc/2026/Conference","replyto":"3FcjKYmNY5","id":"TbAG3g4qmn","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["video-to-audio","audio generation","diffusion"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"We propose SALSA-V, a multimodal video-to-audio generation model capable of synthesizing highly synchronized, high-fidelity long-form audio from silent video content. Our approach introduces a masked diffusion objective, enabling audio-conditioned generation and the seamless synthesis of audio sequences of unconstrained length. Additionally, by integrating a shortcut loss into our training process, we achieve rapid generation of high-quality audio samples in as few as eight sampling steps, paving the way for near-real-time applications without requiring dedicated fine-tuning or retraining. We demonstrate that SALSA-V significantly outperforms existing state-of-the-art methods in both audiovisual alignment and synchronization with video content in quantiative evaluation and a human listening study. Furthermore, our use of random masking during training enables our model to match spectral characteristics of reference audio samples, broadening its applicability to professional audio synthesis tasks such as Foley generation and sound design."},"_bibtex":{"value":"@misc{\ndellali2026salsav,\ntitle={{SALSA}-V: Shortcut-Augmented Long-form Synchronized Audio from Videos},\nauthor={Amir Dellali and Luca A Lanzend{\\\"o}rfer and Florian Gr{\\\"o}tschla and Roger Wattenhofer},\nyear={2026},\nurl={https://openreview.net/forum?id=3FcjKYmNY5}\n}"},"title":{"value":"SALSA-V: Shortcut-Augmented Long-form Synchronized Audio from Videos"},"pdf":{"value":"/pdf/780501365dd6cdc6dc1c38131015c6ed7be54b52.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"dellali|salsav_shortcutaugmented_longform_synchronized_audio_from_videos"},"authorids":{"value":["~Amir_Dellali1","~Luca_A_Lanzendörfer1","~Florian_Grötschla1","~Roger_Wattenhofer1"]},"authors":{"value":["Amir Dellali","Luca A Lanzendörfer","Florian Grötschla","Roger Wattenhofer"]}},"version":2},{"content":{"venue":{"value":"COLING 2025"},"pdf":{"value":"https://aclanthology.org/2025.coling-main.459.pdf"},"venueid":{"value":"dblp.org/conf/COLING/2025"},"paperhash":{"value":"zhou|linguistic_minimal_pairs_elicit_linguistic_similarity_in_large_language_models"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Xinyu_Zhou:","https://dblp.org/search/pid/api?q=author:Delong_Chen:","https://dblp.org/search/pid/api?q=author:Samuel_Cahyawijaya:","~Xufeng_Duan1","https://dblp.org/search/pid/api?q=author:Zhenguang_G._Cai:"]},"html":{"value":"https://aclanthology.org/2025.coling-main.459/"},"_bibtex":{"value":"@inproceedings{DBLP:conf/coling/ZhouCCDC25,\n  author={Xinyu Zhou and Delong Chen and Samuel Cahyawijaya and Xufeng Duan and Zhenguang G. Cai},\n  title={Linguistic Minimal Pairs Elicit Linguistic Similarity in Large Language Models},\n  year={2025},\n  cdate={1735689600000},\n  pages={6866-6888},\n  url={https://aclanthology.org/2025.coling-main.459/},\n  booktitle={COLING},\n  crossref={conf/coling/2025}\n}\n"},"abstract":{"value":"We introduce a novel analysis that leverages linguistic minimal pairs to probe the internal linguistic representations of Large Language Models (LLMs). By measuring the similarity between LLM activation differences across minimal pairs, we quantify the linguistic similarity and gain insight into the linguistic knowledge captured by LLMs. Our large-scale experiments, spanning 100+ LLMs and 150k minimal pairs in three languages, reveal properties of linguistic similarity from four key aspects: consistency across LLMs, relation to theoretical categorizations, dependency to semantic context, and cross-lingual alignment of relevant phenomena. Our findings suggest that 1) linguistic similarity is significantly influenced by training data exposure, leading to higher cross-LLM agreement in higher-resource languages. 2) Linguistic similarity strongly aligns with fine-grained theoretical linguistic categories but weakly with broader ones. 3) Linguistic similarity shows a weak correlation with semantic similarity, showing its context-dependent nature. 4) LLMs exhibit limited cross-lingual alignment in their understanding of relevant linguistic phenomena. This work demonstrates the potential of minimal pairs as a window into the neural representations of language in LLMs, shedding light on the relationship between LLMs and linguistic theory."},"title":{"value":"Linguistic Minimal Pairs Elicit Linguistic Similarity in Large Language Models"},"authors":{"value":["Xinyu Zhou","Delong Chen","Samuel Cahyawijaya","Xufeng Duan","Zhenguang G. Cai"]}},"tmdate":1747802407999,"pdate":1735689600000,"tcdate":1747802393630,"writers":["~"],"signatures":["~Xufeng_Duan1"],"forum":"ggPRPFsST7","license":"CC BY-SA 4.0","number":546344,"cdate":1735689600000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1747802407999,"domain":"DBLP.org","id":"ggPRPFsST7","version":2},{"content":{"venue":{"value":"CoRR 2024"},"pdf":{"value":"http://arxiv.org/pdf/2409.12435v2"},"venueid":{"value":"dblp.org/journals/CORR/2024"},"paperhash":{"value":"zhou|linguistic_minimal_pairs_elicit_linguistic_similarity_in_large_language_models"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Xinyu_Zhou:","https://dblp.org/search/pid/api?q=author:Delong_Chen:","https://dblp.org/search/pid/api?q=author:Samuel_Cahyawijaya:","~Xufeng_Duan1","https://dblp.org/search/pid/api?q=author:Zhenguang_G._Cai:"]},"html":{"value":"https://doi.org/10.48550/arXiv.2409.12435"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2409-12435,\n  publtype={informal},\n  author={Xinyu Zhou and Delong Chen and Samuel Cahyawijaya and Xufeng Duan and Zhenguang G. Cai},\n  title={Linguistic Minimal Pairs Elicit Linguistic Similarity in Large Language Models},\n  year={2024},\n  cdate={1704067200000},\n  journal={CoRR},\n  volume={abs/2409.12435},\n  url={https://doi.org/10.48550/arXiv.2409.12435}\n}\n"},"abstract":{"value":"We introduce a novel analysis that leverages linguistic minimal pairs to probe the internal linguistic representations of Large Language Models (LLMs). By measuring the similarity between LLM activation differences across minimal pairs, we quantify the and gain insight into the linguistic knowledge captured by LLMs. Our large-scale experiments, spanning 100+ LLMs and 150k minimal pairs in three languages, reveal properties of linguistic similarity from four key aspects: consistency across LLMs, relation to theoretical categorizations, dependency to semantic context, and cross-lingual alignment of relevant phenomena. Our findings suggest that 1) linguistic similarity is significantly influenced by training data exposure, leading to higher cross-LLM agreement in higher-resource languages. 2) Linguistic similarity strongly aligns with fine-grained theoretical linguistic categories but weakly with broader ones. 3) Linguistic similarity shows a weak correlation with semantic similarity, showing its context-dependent nature. 4) LLMs exhibit limited cross-lingual alignment in their understanding of relevant linguistic phenomena. This work demonstrates the potential of minimal pairs as a window into the neural representations of language in LLMs, shedding light on the relationship between LLMs and linguistic theory. Codes and data are available at https://github.com/ChenDelong1999/Linguistic-Similarity"},"title":{"value":"Linguistic Minimal Pairs Elicit Linguistic Similarity in Large Language Models"},"authors":{"value":["Xinyu Zhou","Delong Chen","Samuel Cahyawijaya","Xufeng Duan","Zhenguang G. Cai"]}},"tmdate":1747802406215,"pdate":1704067200000,"tcdate":1747802393592,"writers":["~"],"signatures":["~Xufeng_Duan1"],"forum":"xmE7nJb96q","license":"CC BY-SA 4.0","number":546342,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1747802406215,"domain":"DBLP.org","id":"xmE7nJb96q","version":2},{"content":{"venue":{"value":"BlackboxNLP 2024"},"pdf":{"value":"/pdf/b5b542a8a3435ac3dac827cf31abeb3b47dcf940.pdf"},"keywords":{"value":["Linguistic minimal pairs","Large language models"]},"venueid":{"value":"EMNLP/2024/Workshop/BlackBoxNLP"},"paperhash":{"value":"zhou|linguistic_minimal_pairs_elicit_linguistic_similarity_in_large_language_models"},"authorids":{"value":["xinyuzhou314@gmail.com","~Delong_Chen1","~Samuel_Cahyawijaya1","~Xufeng_Duan1","~Zhenguang_Cai1"]},"abstract":{"value":"We introduce a novel analysis that leverages linguistic minimal pairs to probe the internal linguistic representations of Large Language Models (LLMs). By measuring the similarity between LLM activation differences across minimal pairs, we quantify the linguistic similarity and gain insight into the linguistic knowledge captured by LLMs. Our large-scale experiments, spanning 100+ LLMs and 150k minimal pairs in three languages, reveal properties of linguistic similarity from four key aspects: consistency across LLMs, relation to theoretical categorizations, dependency to semantic context, and cross-lingual alignment of relevant phenomena.\nOur findings suggest that 1)linguistic similarity is significantly influenced by training data exposure, leading to higher cross-LLM agreement in higher-resource languages. 2)Linguistic similarity strongly aligns with fine-grained theoretical linguistic categories but weakly with broader ones. 3)Linguistic similarity shows a weak correlation with semantic similarity, showing its context-dependent nature. 4)LLMs exhibit limited cross-lingual alignment in their understanding of relevant linguistic phenomena. This work demonstrates the potential of minimal pairs as a window into the neural representations of language in LLMs, shedding light on the relationship between LLMs and linguistic theory."},"_bibtex":{"value":"@inproceedings{\nzhou2024linguistic,\ntitle={Linguistic Minimal Pairs Elicit Linguistic Similarity in Large Language Models},\nauthor={Xinyu Zhou and Delong Chen and Samuel Cahyawijaya and Xufeng Duan and Zhenguang Cai},\nbooktitle={The 7th BlackboxNLP Workshop},\nyear={2024},\nurl={https://openreview.net/forum?id=sPtGEfcCbd}\n}"},"title":{"value":"Linguistic Minimal Pairs Elicit Linguistic Similarity in Large Language Models"},"track":{"value":"Extended abstract"},"copyright_PDF":{"value":"/attachment/bafcd66676d4690e4e37186780a142500f4f4fa8.pdf"},"authors":{"value":["Xinyu Zhou","Delong Chen","Samuel Cahyawijaya","Xufeng Duan","Zhenguang Cai"]}},"tmdate":1728212755996,"pdate":1726912746738,"tcdate":1724154487996,"writers":["EMNLP/2024/Workshop/BlackBoxNLP","EMNLP/2024/Workshop/BlackBoxNLP/Submission90/Authors"],"signatures":["EMNLP/2024/Workshop/BlackBoxNLP/Submission90/Authors"],"forum":"sPtGEfcCbd","license":"CC BY 4.0","number":90,"cdate":1724154487996,"readers":["everyone"],"invitations":["EMNLP/2024/Workshop/BlackBoxNLP/-/Submission","EMNLP/2024/Workshop/BlackBoxNLP/-/Post_Submission","EMNLP/2024/Workshop/BlackBoxNLP/-/Edit","EMNLP/2024/Workshop/BlackBoxNLP/Submission90/-/Camera-Ready_Revision"],"mdate":1728212755996,"odate":1726912746738,"domain":"EMNLP/2024/Workshop/BlackBoxNLP","id":"sPtGEfcCbd","version":2},{"content":{"summary":{"value":"This paper tackles the problem of generating physically plausible videos using text-to-video diffusion models The proposed method, DiffPhy, fine-tunes a pretrained video diffusion model with guidance from LLMs and Multimodal LLMs. LLM first infers physical context from the text prompt, and an MLLM evaluates the generated video against these physical rules. The MLLM’s feedback is converted into a continuous differentiable score that supervises the diffusion model, encouraging alignment with physical laws. A failure-aware refinement module further injects attention based on detected physical inconsistencies. Experiments on physics-oriented benchmarks show improved physical plausibility and semantic alignment compared to existing models."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"See the weakness."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"The paper proposes a reasonable integration of textual reasoning and multimodal verification to encourage physically consistent video generation, extending prior work on language-guided diffusion with an additional physics-aware supervision signal. The proposed continuous score estimator provides a practical way to incorporate non-differentiable LVLM feedback into gradient-based fine-tuning. The failure-aware refinement further introduces an interesting idea of using detected failure cases as textual cues to guide attention."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"**1) Incomplete Related Work & Baseline Coverage**  \nWhile the “Video Physics Reasoning” subsection is conceptually relevant, the cited works mainly focus on traditional physics simulation or early neural reasoning models, rather than recent diffusion-based approaches that explicitly address physical consistency in video generation. Incorporating more recent studies such as **PhyT2V [1]** and **Yang et al. [2]** would strengthen the connection between DiffPhy and the current landscape of physics-aware video diffusion. Also, the baseline comparison could be expanded to include these contemporary methods.\n\n**2)  Insufficient Comparison with Alignment-Based Methods**  \nThe paper does not adequately position its approach within the growing literature on alignment-based fine-tuning frameworks that leverage LVLM feedback (or more broadly, reward-model signals) for diffusion model adaptation (e.g., VADER [3]). While the proposed continuous score estimator provides a differentiable way to incorporate feedback, it remains unclear how this approach compares to or improves upon existing reward-based and preference-alignment methods in terms of training stability, supervision efficiency, or performance. A more systematic discussion or experimental comparison would help clarify the method’s advantages and situate it more clearly within this line of work.\n\n**3) Uncertain Effectiveness of Failure-Aware Refinement**  \nThe proposed failure-aware refinement introduces an interesting idea of using MLLM-identified failure cases as textual feedback to guide attention. However, this mechanism ultimately relies on the possibility that conditioning on such failure cues helps the model focus on physically inconsistent regions, rather than providing any guarantee of actual correction. Since the feedback is binary and lacks spatial-temporal grounding, it is unclear whether this refinement truly fixes the underlying physical inconsistencies or simply injects additional noise. Moreover, the paper does not include an ablation isolating this component during training, making it even more difficult to assess its actual impact. It would also be helpful to include inference-time ablations comparing this refinement against simpler alternatives, such as seed resampling, to demonstrate that the proposed strategy offers tangible advantages beyond random variation.\n\n**4) Citation Format**  \nCitation references are not consistently enclosed in parentheses throughout the paper.\n\n[1] *PhyT2V: LLM-Guided Iterative Self-Refinement for Physics-Grounded Text-to-Video Generation*  \n[2] *Towards Physically Plausible Video Generation via VLM Planning*  \n[3] *Video Diffusion Alignment via Reward Gradients*"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762918558642,"tcdate":1761858258103,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission6229/Reviewer_pvJs"],"signatures":["ICLR.cc/2026/Conference/Submission6229/Reviewer_pvJs"],"forum":"lPKsPBstHg","number":3,"license":"CC BY 4.0","cdate":1761858258103,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission6229/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762918558642,"domain":"ICLR.cc/2026/Conference","replyto":"lPKsPBstHg","id":"tTq6QI3GLt","forumContent":{"TLDR":{"value":"We propose DiffPhy, a generic framework that enables physically-correct and semantically coherent video generation by fine-tuning a pre-trained video diffusion model."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Video Diffusion","Video Generation","Physical Commensense"]},"supplementary_material":{"value":"/attachment/dd73e8deab01feb18bdae06f4f6ecc4dca24d57c.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Recent video diffusion models have demonstrated their great capability in generating visually-pleasing results, while synthesizing the correct physical effects in generated videos remains challenging. The complexity of real-world motions, interactions, and dynamics introduce great difficulties when learning physics from data. In this work, we propose DiffPhy, a generic framework that enables physically-correct and photo-realistic video generation by fine-tuning a pre-trained video diffusion model. Our method leverages large language models (LLMs) to infer rich physical context from the text prompt. To incorporate this context into the video diffusion model, we use a multimodal large language model (MLLM) to verify intermediate latent variables against the inferred physical rules, guiding the model’s gradient updates accordingly. MLLM’s textual output is transformed into continuous signals. We then formulate a set of training objectives that jointly ensure physical accuracy and semantic alignment with the input text.  Additionally, failure facts of physical phenomena are corrected via attention injection. We also establish a high-quality physical video dataset containing diverse phyiscal actions and events to facilitate effective finetuning. Extensive experiments on public benchmarks demonstrate that DiffPhy is able to produce state-of-the-art results across diverse physics-related scenarios. Code and data will be made available post-review."},"_bibtex":{"value":"@misc{\nzhang2026think,\ntitle={Think Before You Diffuse: Infusing Physical Rules into Video Diffusion},\nauthor={Ke Zhang and Cihan Xiao and Jiacong Xu and Yiqun Mei and Vishal M. Patel},\nyear={2026},\nurl={https://openreview.net/forum?id=lPKsPBstHg}\n}"},"title":{"value":"Think Before You Diffuse: Infusing Physical Rules into Video Diffusion"},"pdf":{"value":"/pdf/abed66725f92f4403a0e6585ebdac7432376d20d.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"zhang|think_before_you_diffuse_infusing_physical_rules_into_video_diffusion"},"authorids":{"value":["~Ke_Zhang17","~Cihan_Xiao1","~Jiacong_Xu1","~Yiqun_Mei1","~Vishal_M._Patel1"]},"authors":{"value":["Ke Zhang","Cihan Xiao","Jiacong Xu","Yiqun Mei","Vishal M. Patel"]}},"version":2},{"content":{"TLDR":{"value":"This paper proposes RACCooN, a versatile and user-friendly video-to-paragraph-to-video generative framework that supports multiple video editing capabilities such as removal, addition, and modification, through a unified pipeline."},"venue":{"value":"Video-Langauge Models Poster"},"keywords":{"value":["Video Editing","Video Inpainting","Multimodal Large Language Model"]},"supplementary_material":{"value":"/attachment/866b28f4db9e63e7fda44b55efa32ffbe632aaf5.zip"},"abstract":{"value":"Recent video generative models primarily rely on carefully written text prompts for specific tasks, like inpainting or style editing. They require labor-intensive textual descriptions for input videos, hindering their flexibility to adapt personal/raw videos to user specifications. This paper proposes RACCooN, a versatile and user-friendly video-to-paragraph-to-video generative framework that supports multiple video editing capabilities such as removal, addition, and modification, through a unified pipeline. RACCooN consists of two principal stages: Video-to-Paragraph (V2P) and Paragraph-to-Video (P2V). In the V2P stage, we automatically describe video scenes in well-structured natural language, capturing both the holistic context and focused object details. Subsequently, in the P2V stage, users can optionally refine these descriptions to guide the video diffusion model, enabling various modifications to the input video, such as removing, changing subjects, and/or adding new objects. The proposed approach stands out from other methods through several significant contributions: (1) RACCooN suggests a multi-granular spatiotemporal pooling strategy to generate well-structured video descriptions, capturing both the broad context and object details without requiring complex human annotations, simplifying precise video content editing based on text for users. (2) Our video generative model incorporates auto-generated narratives or instructions to enhance the quality and accuracy of the generated content. It supports the addition of video objects, inpainting, and attribute modification within a unified framework, surpassing existing video editing and inpainting benchmarks. The proposed framework demonstrates impressive versatile capabilities in video-to-paragraph generation, video content editing, and can be incorporated into other SoTA video generative models for further enhancement."},"_bibtex":{"value":"@inproceedings{\nyoon2025raccoon,\ntitle={{RACC}ooN: Remove, Add, and Change Video Content with Auto-Generated Narratives},\nauthor={Jaehong Yoon and Shoubin Yu and Mohit Bansal},\nbooktitle={Workshop on Video-Language Models @ NeurIPS 2024},\nyear={2025},\nurl={https://openreview.net/forum?id=oWEbHXXDNs}\n}"},"title":{"value":"RACCooN: Remove, Add, and Change Video Content with Auto-Generated Narratives"},"pdf":{"value":"/pdf/2dd1f756b4f70629270f4fafa54602a9f2d01198.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Workshop/Video-Langauge_Models"},"paperhash":{"value":"yoon|raccoon_remove_add_and_change_video_content_with_autogenerated_narratives"},"authorids":{"value":["~Jaehong_Yoon1","~Shoubin_Yu1","~Mohit_Bansal2"]},"track":{"value":"Short Paper Track (up to 3 pages)"},"authors":{"value":["Jaehong Yoon","Shoubin Yu","Mohit Bansal"]}},"tmdate":1736861080828,"pdate":1730081753174,"tcdate":1726014224030,"writers":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission43/Authors"],"signatures":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission43/Authors"],"forum":"oWEbHXXDNs","license":"CC BY 4.0","number":43,"cdate":1726014224030,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Post_Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Edit"],"mdate":1736861080828,"odate":1736861080813,"domain":"NeurIPS.cc/2024/Workshop/Video-Langauge_Models","id":"oWEbHXXDNs","version":2},{"content":{"summary":{"value":"This paper reframes the ambiguous sparse-view reconstruction task as a temporal generation task and unleashes the strong generative prior of large pre-trained video diffusion models. To maintain consistency while generating the video clip, this work proposes to first construct a global point cloud using DUSt3R and then encode it into a contextual space to provide the 3D structure condition. Afterward, it will set the first frame as well as the last frame and repurpose the video diffusion model as an interpolator to synthesize plausible frames. Finally, it will incorporate a confidence-aware 3DGS optimization for reconstruction. Experiments show the paper's superiority in terms of quality and generalization capability to larger viewpoint change compared with previous work."},"soundness":{"value":4},"confidence":{"value":4},"questions":{"value":"Please refer to the weakness part. I will be glad to raise my rate if my concerns can be addressed!"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The idea of repurposing a video diffusion model as a spatial interpolator sounds reasonable and can provide generative priors to unseen regions.\n2. How this work injects structure-aware guidance by using the point cloud from DUSt3R seems critical and can provide insights for future work.\n3. This work achieves visually better results than previous methods and seems to generalize better to large viewpoint changes.\n4. The paper is well-written and should be easy to follow as long as the code is released."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. One strength of directly applying a feed-forward reconstruction like pixelSplat[1] or MVSplat[2] is the inference speed. The proposed method needs to incorporate DUSt3R, a video diffusion model, and a confidence-aware optimization to get the final result. I wonder about the total time consumption during inference compared with baseline models.\n2. From the reviewer's perspective, there seem to be multiple solutions given the first frame and the last frame of a video, would it be better to incorporate camera pose information during the training and inference of the video diffusion model? Or on the other side, is the video diffusion model trying to learn a certain type of camera motion due to the limitation of fine-tuning dataset? \n3. This paper proposes to utilize a global point cloud as the 3D structure guidance. From the ablation study in Fig.5, the quality seems to be poor without this guidance, showing the importance of incorporating this point cloud. While DUSt3R itself can provide a global point cloud as the reconstruction result, I would like to see a comparison with DUSt3R apart from MVSplat and pixelSplat. Moreover, pixelSplat and MVSplat do not incorporate a global point cloud as input. \n4. This work mainly deals with static scene reconstruction, it would be interesting to reconstruct a dynamic scene in future work, which is also mentioned by the authors in the future work part. \n\n[1] pixelSplat: 3D Gaussian Splats from Image Pairs for Scalable Generalizable 3D Reconstruction, CVPR 2024 \n\n[2] MVSplat: Efficient 3D Gaussian Splatting from Sparse Multi-View Images, ECCV 2024\n\n[3] DUSt3R: Geometric 3D Vision Made Easy, CVPR 2024"}},"nonreaders":[],"tmdate":1731428639681,"tcdate":1729325390024,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission9317/Reviewer_JdpZ"],"signatures":["ICLR.cc/2025/Conference/Submission9317/Reviewer_JdpZ"],"forum":"Z30Mdbv5jO","number":1,"license":"CC BY 4.0","cdate":1729325390024,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission9317/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428639681,"domain":"ICLR.cc/2025/Conference","replyto":"Z30Mdbv5jO","id":"6w3VJRvMzw","forumContent":{"TLDR":{"value":"ReconX is a novel sparse-view 3D scene reconstruction paradigm that reframes the ambiguous reconstruction challenge as a temporal generation task."},"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Sparse-view Reconstruction","Video Diffusion","Gaussian Splatting"]},"supplementary_material":{"value":"/attachment/85bb44579fb246807dc36f1fcaffbce8ea01e93d.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Advancements in 3D scene reconstruction have transformed 2D images from the real world into 3D models, producing realistic 3D results from hundreds of input photos. Despite great success in dense-view reconstruction scenarios, rendering a detailed scene from sparse views is still an ill-posed optimization problem, often resulting in artifacts and distortions in unseen areas. In this paper, we propose ReconX, a novel 3D scene reconstruction paradigm that reframes the ambiguous reconstruction problem as a temporal generation task. The key insight is to unleash the strong generative prior of large pre-trained video diffusion models for sparse-view reconstruction. Nevertheless, it is challenging to preserve 3D view consistency when directly generating video frames from pre-trained models. To address this issue, given limited input views, the proposed ReconX first constructs a global point cloud and encodes it into a contextual space as the 3D structure condition. Guided by the condition, the video diffusion model then synthesizes video frames that are detail-preserved and exhibit a high degree of 3D consistency, ensuring the coherence of the scene from various perspectives. Finally, we recover the 3D scene from the generated video through a confidence-aware 3D Gaussian Splatting optimization scheme. Extensive experiments on various real-world datasets show the superiority of ReconX over state-of-the-art methods in terms of quality and generalizability."},"_bibtex":{"value":"@misc{\nliu2025reconx,\ntitle={ReconX: Reconstruct Any Scene from Sparse Views with Video Diffusion Model},\nauthor={Fangfu Liu and Wenqiang Sun and Hanyang Wang and Yikai Wang and Haowen Sun and Junliang Ye and Jun Zhang and Yueqi Duan},\nyear={2025},\nurl={https://openreview.net/forum?id=Z30Mdbv5jO}\n}"},"title":{"value":"ReconX: Reconstruct Any Scene from Sparse Views with Video Diffusion Model"},"pdf":{"value":"/pdf/9364ea472d15f95d4ba49d9d3ab667616846ed10.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"liu|reconx_reconstruct_any_scene_from_sparse_views_with_video_diffusion_model"},"authorids":{"value":["~Fangfu_Liu2","~Wenqiang_Sun1","~Hanyang_Wang4","~Yikai_Wang2","~Haowen_Sun1","~Junliang_Ye1","~Jun_Zhang25","~Yueqi_Duan1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Fangfu Liu","Wenqiang Sun","Hanyang Wang","Yikai Wang","Haowen Sun","Junliang Ye","Jun Zhang","Yueqi Duan"]}},"version":2},{"content":{"summary":{"value":"The paper presents \"VideoTetris,\" a new framework designed to improve text-to-video generation in complex scenarios with dynamic changes and multiple objects. It introduces spatio-temporal compositional diffusion techniques for better alignment with textual semantics and integrates a dynamic-aware data processing and consistency regularization to enhance video consistency. The results from extensive experiments demonstrate significant qualitative and quantitative improvements in T2V generation."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"How is continuity ensured between different shots?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The paper is well-written and easy to understand.\n2. The motivations for spatio-temporal compositional diffusion and Dynamic-Aware Video Data Processing are clearly explained, with supportive experimental evidence provided.\n\n3. The experiments are sufficiently thorough, demonstrating that VideoTetris achieves impressive results in both short and long video production scenarios."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1.Although Spatio-Temporal Compositional Diffusion directly adjusts cross-attention, it still segments videos into different shots, each with a specific layout, suggesting it's essentially a layout-based method. This approach contradicts the motivation outlined in section 3.1.\n\n2.The paper presents limited technical innovations, focusing more on engineering implementation. In Spatio-Temporal Compositional Diffusion, it uses LLMs to pre-process prompts and region masks, which are then sequentially applied during generation. Meanwhile, Dynamic-Aware Video Data Processing enhances data preprocessing for higher quality."},"limitations":{"value":"The manuscript includes discussions on limitations and social impacts."}},"nonreaders":[],"tmdate":1730879408605,"tcdate":1720715313000,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission10727/Reviewer_uyep"],"signatures":["NeurIPS.cc/2024/Conference/Submission10727/Reviewer_uyep"],"forum":"RPM7STrnVz","number":2,"license":"CC BY 4.0","cdate":1720715313000,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission10727/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730879408605,"domain":"NeurIPS.cc/2024/Conference","replyto":"RPM7STrnVz","id":"BAYreCzjpq","forumContent":{"venue":{"value":"NeurIPS 2024 poster"},"keywords":{"value":["Text-to-Video Generation","Video Diffusion Models"]},"supplementary_material":{"value":"/attachment/f72c15c52e21ddf70f17cb82f2289107037072b9.zip"},"primary_area":{"value":"generative_models"},"abstract":{"value":"Diffusion models have demonstrated great success in text-to-video (T2V) generation. However, existing methods may face challenges when handling complex (long) video generation scenarios that involve multiple objects or dynamic changes in object numbers. To address these limitations, we propose VideoTetris, a novel framework that enables compositional T2V generation. Specifically, we propose spatio-temporal compositional diffusion to precisely follow complex textual semantics by manipulating and composing the attention maps of denoising networks spatially and temporally. Moreover, we propose a new dynamic-aware data processing pipeline and a consistency regularization method to enhance the consistency of auto-regressive video generation. Extensive experiments demonstrate that our VideoTetris achieves impressive qualitative and quantitative results in compositional T2V generation. Code is available at: https://github.com/YangLing0818/VideoTetris"},"_bibtex":{"value":"@inproceedings{\ntian2024videotetris,\ntitle={VideoTetris: Towards Compositional Text-to-Video Generation},\nauthor={Ye Tian and Ling Yang and Haotian Yang and Yuan Gao and Yufan Deng and Xintao Wang and Zhaochen Yu and Xin Tao and Pengfei Wan and Di ZHANG and Bin CUI},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=RPM7STrnVz}\n}"},"title":{"value":"VideoTetris: Towards Compositional Text-to-Video Generation"},"pdf":{"value":"/pdf/bb68997ead13efc218660269a1fa4d189f0588a2.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"tian|videotetris_towards_compositional_texttovideo_generation"},"authorids":{"value":["~Ye_Tian15","~Ling_Yang1","~Haotian_Yang1","~Yuan_Gao32","~Yufan_Deng2","~Xintao_Wang1","~Zhaochen_Yu2","~Xin_Tao3","~Pengfei_Wan1","~Di_ZHANG3","~Bin_CUI2"]},"authors":{"value":["Ye Tian","Ling Yang","Haotian Yang","Yuan Gao","Yufan Deng","Xintao Wang","Zhaochen Yu","Xin Tao","Pengfei Wan","Di ZHANG","Bin CUI"]}},"version":2},{"content":{"venue":{"value":"Video-Langauge Models Oral"},"pdf":{"value":"/pdf/b2613ded1b4f3f8d2ef8047079500b9aabf089cf.pdf"},"keywords":{"value":["text-to-video generation","physical commonsense","human and automatic evaluation"]},"venueid":{"value":"NeurIPS.cc/2024/Workshop/Video-Langauge_Models"},"paperhash":{"value":"bansal|videophy_evaluating_physical_commonsense_for_video_generation"},"authorids":{"value":["~Hritik_Bansal2","~Zongyu_Lin1","~Tianyi_Xie1","~Zeshun_Zong1","~Michal_Yarom1","~Yonatan_Bitton1","~Chenfanfu_Jiang3","~Yizhou_Sun1","~Kai-Wei_Chang1","~Aditya_Grover1"]},"abstract":{"value":"Recent advances in internet-scale video data pretraining have led to the development of text-to-video generative models that can create high-quality videos across a broad range of visual concepts and styles. Due to their ability to synthesize realistic motions and render complex objects, these generative models have the potential to become general-purpose simulators of the physical world. However, it is unclear how far we are from this goal with the existing text-to-video generative models. To this end, we present VideoPhy, a benchmark designed to assess whether the generated videos follow physical commonsense for real-world activities (e.g. marbles will roll down when placed on a slanted surface). Specifically, we curate a list of 688 captions that involve interactions between various material types in the physical world (e.g., solid-solid, solid-fluid, fluid-fluid). We then generate videos conditioned on these captions from diverse state-of-the-art text-to-video generative models, including open models (e.g., VideoCrafter2) and closed models (e.g., Lumiere from Google, Pika). Further, our human evaluation reveals that the existing models severely lack the ability to generate videos adhering to the given text prompts, while also lack physical commonsense. Specifically, the best performing model, Pika, generates videos that adhere to the caption and physical laws for only 19.7% of the instances. VideoPhy thus highlights that the video generative models are far from accurately simulating the physical world. Finally, we also supplement the dataset with an auto-evaluator, \\model{}, to assess semantic adherence and physical commonsense at scale."},"_bibtex":{"value":"@inproceedings{\nbansal2025videophy,\ntitle={VideoPhy: Evaluating Physical Commonsense for Video Generation},\nauthor={Hritik Bansal and Zongyu Lin and Tianyi Xie and Zeshun Zong and Michal Yarom and Yonatan Bitton and Chenfanfu Jiang and Yizhou Sun and Kai-Wei Chang and Aditya Grover},\nbooktitle={Workshop on Video-Language Models @ NeurIPS 2024},\nyear={2025},\nurl={https://openreview.net/forum?id=xMlYKYFd03}\n}"},"title":{"value":"VideoPhy: Evaluating Physical Commonsense for Video Generation"},"track":{"value":"Long Paper Track (up to 9 pages)"},"authors":{"value":["Hritik Bansal","Zongyu Lin","Tianyi Xie","Zeshun Zong","Michal Yarom","Yonatan Bitton","Chenfanfu Jiang","Yizhou Sun","Kai-Wei Chang","Aditya Grover"]}},"tmdate":1736861080077,"pdate":1730081752113,"tcdate":1725588445817,"writers":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission13/Authors"],"signatures":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission13/Authors"],"forum":"xMlYKYFd03","license":"CC BY 4.0","number":13,"cdate":1725588445817,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Post_Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Edit","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission13/-/Camera-Ready_Revision"],"mdate":1736861080077,"odate":1736861080058,"domain":"NeurIPS.cc/2024/Workshop/Video-Langauge_Models","id":"xMlYKYFd03","version":2},{"content":{"summary":{"value":"This paper proposes a framework named VISTA to address the issue of shortcut learning in vision-language models (VLMs), where models tend to rely on superficial visual cues rather than developing a deep understanding of the logical relationships between questions and visual inputs. The VISTA framework explicitly decomposes the reasoning process into a visual sensor (VLM) for perception and a reasoning module (LLM) for logical inference, thereby mitigating the influence of shortcut learning. Experimental results on two benchmarks, MMVP and SeedBench-500, demonstrate the effectiveness of the proposed approach."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"1. Does the VLM component itself still suffer from shortcut learning? In the proposed agent system, the VLM appears to be reduced to a perceptual module, leaving its inherent shortcut learning issues unaddressed.\n\n2. How does the method generalize to more comprehensive benchmarks? Evaluation on challenging benchmarks such as MMMU and MMMU-Pro would better demonstrate its generalization capability."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. The authors propose the VISTA framework to address shortcut learning in VLMs, which explicitly separates visual perception (sensor) from logical reasoning (reasoner) to mitigate reliance on spurious visual cues."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. While the VISTA framework attempts to address shortcut learning by employing a dual-agent architecture (VLM + LLM), this approach does not fundamentally solve the underlying issue within the VLM itself. The VLM component remains susceptible to shortcut learning, merely transferring rather than resolving this critical limitation.\n\n2. The evaluation is currently limited to established benchmarks. To better demonstrate the method's robustness and generalizability, performance should be validated on more recent and challenging VQA benchmarks such as MMMU and MMMU-Pro."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762923846667,"tcdate":1762016779949,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission13122/Reviewer_y7z3"],"signatures":["ICLR.cc/2026/Conference/Submission13122/Reviewer_y7z3"],"forum":"F1pWCHoSSA","number":3,"license":"CC BY 4.0","cdate":1762016779949,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission13122/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762923846667,"domain":"ICLR.cc/2026/Conference","replyto":"F1pWCHoSSA","id":"CqUZz8rb2Q","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Visual Question Answering","Reasoning"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"End-to-end Vision-language models (VLMs) often rely on spurious visual cues, conflating perception with decision-making. We introduce VISTA (Visual Information Separation for Text-based Analysis), which enforces an explicit information bottleneck between a text-only reasoner and a stateless VLM sensor. The LLM reasoner decomposes each question and iteratively queries a VLM for visual facts; the VLM is instructed to reject queries that require high-level inference, creating an explicit information bottleneck. Trained on only 641 questions, VISTA yields large robustness gains on SpuriVerse across two vision backbones (+16.29\\% with Qwen-2.5-VL-7B and +6.77\\% with Llama-3.2-Vision-11B), while direct SFT or RL on the VLM fails to remedy spuriosity and can even exacerbate it. Despite never exposing the reasoner to raw pixels, VISTA slightly improves or remains on par with VLMs on everyday-scene benchmarks, including MMVP and SeedBench. Our learned reasoners transfer across sensors, indicating algorithmic rather than model-specific generalization. Together, VISTA enables spurious-resistant VQA by upgrading the brain, not the eyes."},"_bibtex":{"value":"@misc{\nli2026unbiased,\ntitle={Unbiased Visual Reasoning with Controlled Visual Inputs},\nauthor={Zhaonan Li and Shijie Lu and Fei Wang and Jacob Dineen and Xiao Ye and Zhikun Xu and Siyi Liu and Young Min Cho and Bangzheng Li and Daniel Chang and Kenny Nguyen and Qizheng Yang and Muhao Chen and Ben Zhou},\nyear={2026},\nurl={https://openreview.net/forum?id=F1pWCHoSSA}\n}"},"title":{"value":"Unbiased Visual Reasoning with Controlled Visual Inputs"},"pdf":{"value":"/pdf/2893d25c2c2e312c0781097170a2b21c8c26d296.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"li|unbiased_visual_reasoning_with_controlled_visual_inputs"},"authorids":{"value":["~Zhaonan_Li1","~Shijie_Lu1","~Fei_Wang12","~Jacob_Dineen1","~Xiao_Ye1","~Zhikun_Xu1","~Siyi_Liu2","~Young_Min_Cho1","~Bangzheng_Li1","~Daniel_Chang1","~Kenny_Nguyen1","~Qizheng_Yang2","~Muhao_Chen1","~Ben_Zhou1"]},"authors":{"value":["Zhaonan Li","Shijie Lu","Fei Wang","Jacob Dineen","Xiao Ye","Zhikun Xu","Siyi Liu","Young Min Cho","Bangzheng Li","Daniel Chang","Kenny Nguyen","Qizheng Yang","Muhao Chen","Ben Zhou"]}},"version":2},{"content":{"venue":{"value":"Video-Langauge Models Poster"},"pdf":{"value":"/pdf/527a7a89db04942668ec9cf410e27b9c39d5970b.pdf"},"keywords":{"value":["Datasets and benchmarking","Video understanding","Multi-modal learning","Visual question answering","Long-form video","Metrics and benchmarks"]},"supplementary_material":{"value":"/attachment/62c19c6f366bd2293b97d1a39bb658db79088de0.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Workshop/Video-Langauge_Models"},"paperhash":{"value":"rawal|cinepile_a_long_video_question_answering_dataset_and_benchmark"},"authorids":{"value":["~Ruchit_Rawal1","~Khalid_Saifullah1","~Ronen_Basri1","~David_Jacobs1","~Gowthami_Somepalli1","~Tom_Goldstein1"]},"abstract":{"value":"Current long-form video understanding datasets often fail to provide genuine comprehension challenges, as many tasks can be solved by analyzing only a few random frames. To address this issue, we present a novel dataset and benchmark, CinePile, specifically designed for authentic long-form video understanding. This paper details our innovative approach for creating a question-answer dataset, utilizing advanced LLMs with human-in-the-loop and building upon human-generated raw data. Our comprehensive dataset comprises 305,000 multiple-choice questions (MCQs), covering various visual and multimodal aspects, including temporal comprehension, understanding human-object interactions, and reasoning about events or actions within a scene. Additionally, we evaluate recent video-centric LLMs, both open-source and proprietary, on the test split of our dataset. The findings reveal that even state-of-the-art video-centric LLMs significantly lag behind human performance in these tasks, highlighting the complexity and challenge inherent in video understanding."},"_bibtex":{"value":"@inproceedings{\nrawal2025cinepile,\ntitle={CinePile: A Long Video Question Answering Dataset and Benchmark},\nauthor={Ruchit Rawal and Khalid Saifullah and Ronen Basri and David Jacobs and Gowthami Somepalli and Tom Goldstein},\nbooktitle={Workshop on Video-Language Models @ NeurIPS 2024},\nyear={2025},\nurl={https://openreview.net/forum?id=V3VNqzVFxF}\n}"},"title":{"value":"CinePile: A Long Video Question Answering Dataset and Benchmark"},"track":{"value":"Short Paper Track (up to 3 pages)"},"authors":{"value":["Ruchit Rawal","Khalid Saifullah","Ronen Basri","David Jacobs","Gowthami Somepalli","Tom Goldstein"]}},"tmdate":1736861080709,"pdate":1730081753077,"tcdate":1726002829678,"writers":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission38/Authors"],"signatures":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission38/Authors"],"forum":"V3VNqzVFxF","license":"CC BY 4.0","number":38,"cdate":1726002829678,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Post_Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission38/-/Full_Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Edit","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission38/-/Camera-Ready_Revision"],"mdate":1736861080709,"odate":1736861080691,"domain":"NeurIPS.cc/2024/Workshop/Video-Langauge_Models","id":"V3VNqzVFxF","version":2},{"content":{"TLDR":{"value":"We propose a concise, interpretable language representation that helps long video reasoning, while utilizing the context of LLMs efficiently and effectively."},"venue":{"value":"Video-Langauge Models Poster"},"keywords":{"value":["long video reasoning; large language models"]},"supplementary_material":{"value":"/attachment/ef5284240ea313c4c32878693e41518ec37ff285.pdf"},"abstract":{"value":"Language has become a prominent modality in computer vision with the rise of multi-modal LLMs. Despite supporting long context-lengths, their effectiveness in handling long-term information gradually declines with input length. This becomes critical, especially in applications such as long-form video understanding. In this paper, we introduce a Language Repository (LangRepo) for LLMs, that maintains concise and structured information as an interpretable (i.e., all-textual) representation. It consists of write and read operations that focus on pruning redundancies in text, and extracting information at various temporal scales. The proposed framework is evaluated on zero-shot video VQA benchmarks, showing state-of-the-art performance at its scale. Our code is available at https://github.com/kkahatapitiya/LangRepo."},"_bibtex":{"value":"@inproceedings{\nkahatapitiya2025language,\ntitle={Language Repository for Long Video Understanding},\nauthor={Kumara Kahatapitiya and Kanchana Ranasinghe and Jongwoo Park and Michael S Ryoo},\nbooktitle={Workshop on Video-Language Models @ NeurIPS 2024},\nyear={2025},\nurl={https://openreview.net/forum?id=CTaWnvcTNz}\n}"},"title":{"value":"Language Repository for Long Video Understanding"},"pdf":{"value":"/pdf/af0afd482f12be1f445568eb3b087717dd61d3a7.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Workshop/Video-Langauge_Models"},"paperhash":{"value":"kahatapitiya|language_repository_for_long_video_understanding"},"authorids":{"value":["~Kumara_Kahatapitiya1","~Kanchana_Ranasinghe1","~Jongwoo_Park1","~Michael_S_Ryoo1"]},"track":{"value":"Short Paper Track (up to 3 pages)"},"authors":{"value":["Kumara Kahatapitiya","Kanchana Ranasinghe","Jongwoo Park","Michael S Ryoo"]}},"tmdate":1736861080560,"pdate":1730081752936,"tcdate":1725990878169,"writers":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission35/Authors"],"signatures":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission35/Authors"],"forum":"CTaWnvcTNz","license":"CC BY 4.0","number":35,"cdate":1725990878169,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Post_Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission35/-/Full_Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Edit","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission35/-/Camera-Ready_Revision"],"mdate":1736861080560,"odate":1736861080544,"domain":"NeurIPS.cc/2024/Workshop/Video-Langauge_Models","id":"CTaWnvcTNz","version":2},{"content":{"TLDR":{"value":"We provide a large-scale study on evaluating current vision-language foundation models focusing on transfer-learning onto video understanding task."},"venue":{"value":"Video-Langauge Models Poster"},"keywords":{"value":["Video understanding","video-language foundation model","action recognition","multi-modal learning"]},"supplementary_material":{"value":"/attachment/596dfa45066b7caf25e79f1956b125f9f91c7bad.pdf"},"abstract":{"value":"Vision-Language foundation models, including vision-language models (VLMs) and vision-large language models (VLLMs), have been evolving rapidly and have shown good performance on different downstream video understanding tasks, especially on web datasets. However, it is still an open question how much these VLMs and VLLMs perform in more challenging scenarios like Activities of Daily Living (ADL). To answer this, we provide a comprehensive study of VLMs and VLLMs by comparing their zero-shot transfer ability to five downstream tasks including action classification, video retrieval, video description, action forecasting, and frame-wise action segmentation. Extensive experiments are conducted on eleven real-world, human-centric video understanding datasets (e.g., Toyota Smarthome, Penn Action, UAV-Human, EgoExo4D, TSU, Charades) to study these tasks with our insights into the strengths and limitations of these models in zero-shot settings. Moreover, we provide in-deep analysis to find the best setting to improve the model performance in zero-shot action classification tasks. Based on our experiments, we find that these models are still far away from satisfactory performance in all evaluated tasks, particularly in densely labeled and long video datasets."},"_bibtex":{"value":"@inproceedings{\nali2025quo,\ntitle={Quo Vadis, Video Understanding with Vision-Language Foundation Models?},\nauthor={Mahmoud ALI and Di Yang and Arkaprava Sinha and Dominick Reilly and Srijan Das and Gianpiero Francesca and Francois Bremond},\nbooktitle={Workshop on Video-Language Models @ NeurIPS 2024},\nyear={2025},\nurl={https://openreview.net/forum?id=JWwCTmVkYo}\n}"},"title":{"value":"Quo Vadis, Video Understanding with Vision-Language Foundation Models?"},"pdf":{"value":"/pdf/b97d0e1f1619839e0fa7589e9ce31c0ebc5a2401.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Workshop/Video-Langauge_Models"},"paperhash":{"value":"ali|quo_vadis_video_understanding_with_visionlanguage_foundation_models"},"authorids":{"value":["~Mahmoud_ALI1","~Di_Yang4","asinha13@charlotte.edu","dreilly1@charlotte.edu","~Srijan_Das1","~Gianpiero_Francesca1","~Francois_Bremond1"]},"track":{"value":"Long Paper Track (up to 9 pages)"},"authors":{"value":["Mahmoud ALI","Di Yang","Arkaprava Sinha","Dominick Reilly","Srijan Das","Gianpiero Francesca","Francois Bremond"]}},"tmdate":1736861080525,"pdate":1730081752853,"tcdate":1725984163450,"writers":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission33/Authors"],"signatures":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission33/Authors"],"forum":"JWwCTmVkYo","license":"CC BY 4.0","number":33,"cdate":1725984163450,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Post_Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission33/-/Full_Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Edit","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission33/-/Camera-Ready_Revision"],"mdate":1736861080525,"odate":1736861080510,"domain":"NeurIPS.cc/2024/Workshop/Video-Langauge_Models","id":"JWwCTmVkYo","version":2},{"content":{"venue":{"value":"Video-Langauge Models Poster"},"pdf":{"value":"/pdf/8de3b165a4c91d140d992d7ff16d9d55acff2e2d.pdf"},"keywords":{"value":["Audio Generation","Multimodal Generation"]},"venueid":{"value":"NeurIPS.cc/2024/Workshop/Video-Langauge_Models"},"paperhash":{"value":"jeong|read_watch_and_scream_sound_generation_from_text_and_video"},"authorids":{"value":["~Yujin_Jeong2","~Yunji_Kim1","~Sanghyuk_Chun1","~Jiyoung_Lee2"]},"abstract":{"value":"Despite the impressive progress of multimodal generative models, generating sound solely from text poses challenges in ensuring comprehensive scene depiction and temporal alignment.\nMeanwhile, video-to-audio generation limits the flexibility to prioritize sound synthesis for specific objects within the scene.\nTo tackle these challenges, we propose a novel video-and-text-to-audio generation method, called \\ours, where video serves as a conditional control for a text-to-audio generation model.\nEspecially, our method estimates the structural information of sound (namely, energy) from the video while receiving key content cues from a user prompt.\nWe employ a well-performing text-to-audio model to consolidate the video control, which is much more efficient for training multimodal diffusion models with massive triplet-paired (audio-video-text) data.\nIn addition, by separating the generative components of audio, it becomes a more flexible system that allows users to freely adjust the energy, surrounding environment, and primary sound source according to their preferences.\nExperimental results demonstrate that our method shows superiority in terms of quality, controllability, and training efficiency.\nOur demo is available at https://naver-ai.github.io/rewas."},"_bibtex":{"value":"@inproceedings{\njeong2025read,\ntitle={Read, Watch and Scream! Sound Generation from Text and Video},\nauthor={Yujin Jeong and Yunji Kim and Sanghyuk Chun and Jiyoung Lee},\nbooktitle={Workshop on Video-Language Models @ NeurIPS 2024},\nyear={2025},\nurl={https://openreview.net/forum?id=wo3xj8vL1M}\n}"},"title":{"value":"Read, Watch and Scream! Sound Generation from Text and Video"},"track":{"value":"Long Paper Track (up to 9 pages)"},"authors":{"value":["Yujin Jeong","Yunji Kim","Sanghyuk Chun","Jiyoung Lee"]}},"tmdate":1736861079854,"pdate":1730081751764,"tcdate":1724677653194,"writers":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission2/Authors"],"signatures":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission2/Authors"],"forum":"wo3xj8vL1M","license":"CC BY 4.0","number":2,"cdate":1724677653194,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Post_Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Edit","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission2/-/Camera-Ready_Revision"],"mdate":1736861079854,"odate":1736861079746,"domain":"NeurIPS.cc/2024/Workshop/Video-Langauge_Models","id":"wo3xj8vL1M","version":2},{"content":{"summary":{"value":"The paper proposes CAT-Video, a corruption-aware training framework for latent video diffusion models (LVDMs). It introduces two structured, low-rank embedding perturbations: Batch-Centered Noise Injection (BCNI) and Spectrum-Aware Contextual Noise (SACN). The idea is to inject data-aligned noise during training to improve robustness to noisy text/multimodal conditioning and reduce semantic drift across timesteps. The method is supported by a theoretical sketch (entropy/Wasserstein/score-drift bounds) and experiments on WebVid-2M, MSR-VTT, MSVD, and UCF-101, with reported FVD gains over Gaussian/Uniform corruption and some large diffusion baselines; there are also brief extensions to autoregressive models and multimodal video understanding"},"soundness":{"value":2},"confidence":{"value":5},"questions":{"value":"See weaknesses."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"Robustness of T2V models to noisy/ambiguous conditioning is important and under-tested. The paper articulates temporal error accumulation in diffusion and motivates structured corruption specifically for video.\n\nBCNI perturbs along deviations from batch mean; SACN perturbs along principal spectral modes with exponentially decayed variances. Both add minimal overhead and have a single scale hyper-parameter."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Novelty is limited relative to prior “noisy/structured conditioning” literature; positioning is dated.\nInjecting noise into conditioning embeddings (Gaussian/Uniform, token-level swaps/replacements) and exploiting structure/low-rank directions have been explored extensively in image diffusion and multimodal finetuning (e.g., corruption-aware pretraining, token-level perturbations, NEFTune-style noisy embeddings). The paper cites several such works but doesn’t clearly differentiate BCNI/SACN as more than “reasonable engineering variants” adapted to video. A crisper technical delta vs. structured conditioning noise in recent video works is missing. Suggest: provide head-to-head against stronger structured baselines, not only isotropic Gaussian/Uniform or simple temporal gradients (TANI/HSCAN as defined). For instance, compare to learned/adaptive corruption schedules, prompt-space adversarial perturbations with temporal regularizers, or curriculum-style noise aligned to caption syntax/scene cuts\n\n2. Experimental scope feels behind current 2025-26 standards; datasets/architectures and metrics are narrow.\nCore results rely on WebVid-2M/MSR-VTT/MSVD/UCF-101 with FVD-centric reporting. Modern T2V evaluation has moved toward longer videos, higher resolutions, compositional control, and stronger perception-aligned metrics (e.g., VBench subsets with motion/physics consistency, human studies with calibrated protocols). The paper mentions VBench/EvalCrafter in passing, but full tables are pushed to the appendix; the main text should surface these with stronger analysis and significance tests. Also, many baselines listed are from earlier generations of models; direct comparisons to current-gen LVDMs/DiT-style video transformers trained at scale are missing. Actionable:\nAdd long-horizon (≥10–16 s) and high-res evaluations;\nInclude modern SOTA baselines trained on 2024–2026 corpora (and/or reproduce baseline training with matched compute);\nReport human preference and motion/consistency metrics prominently in main text.\n\n3. Claims vs. baselines are hard to validate due to scale mismatch and unclear fairness.\nSeveral “beats larger models with 5× less data” statements are made, but the compared methods differ in architecture, parameterization, pretraining recipes, and dataset curation. Without matched compute / strong-to-strong comparisons (e.g., same backbone with/without CAT; or retrained SOTA with the authors’ code), it’s difficult to attribute gains to BCNI/SACN rather than other confounders. Please (i) include paired-control runs on the same modern backbone at matched data/compute; (ii) show scaling curves (data, steps, guidance, steps vs. FVD) for clean vs. CAT; (iii) report statistical significance (mean±std is in the appendix, but main-text needs tests across seeds).\n\n4. Theory is mostly appendix-level and not tightly coupled to practice.\nThe paper sketches entropy/Wasserstein/score-drift bounds and a D/d “complexity gap,” but practical instantiations (how d is chosen/estimated online, how SVD in SACN scales with sequence length, and why the assumed low-frequency dominance universally holds) are not convincingly validated. The main text should tie specific theorems to measurable proxies (e.g., Lipschitz estimates, score-norm smoothness, sensitivity slopes) and show correlations across datasets/noise levels. Provide ablations on rank d, spectral weighting schedules, and wall-clock overhead broken down by operator.\n\n5. The multimodal video understanding experiment (AVSD) is conducted at 0.5B scale only (LLaVA-OV-0.5B-FT and PAVE-0.5B). By ICLR-26 standards, this falls short of contemporary MLLM practice (7B/13B and stronger video-language backbones). For a credible “model-agnostic” claim, please include ≥7B variants (e.g., PAVE-7B/13B or comparable video-LLaVA baselines\n\n6. The paper claims “we also validate scalability by extending CAT to autoregressive video generation (NOVA)…,” but NOVA is not tabulated in the main paper. Table 3(a) discusses scalability to AR using MAGVIT/CogVideo as references, while NOVA-specific results are not shown; the authors point to Appendix Tables 14–16, yet Table 14 aggregates AR results by corruption type (BCNI/SACN/Gaussian/Uniform) without a clear NOVA row, making the claim hard to verify from the main text. Please provide an explicit NOVA baseline line with matched settings (params, data, steps) and report AR scaling curves (data/compute vs. FVD) to substantiate “model-agnostic” robustness."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762931540294,"tcdate":1762423717259,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission19699/Reviewer_Nqx2"],"signatures":["ICLR.cc/2026/Conference/Submission19699/Reviewer_Nqx2"],"forum":"unZhwukf0T","number":4,"license":"CC BY 4.0","cdate":1762423717259,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission19699/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762931540294,"domain":"ICLR.cc/2026/Conference","replyto":"unZhwukf0T","id":"xyjAauAItA","forumContent":{"TLDR":{"value":"We introduce CAT-Video, a corruption-aware training framework that improves robustness and temporal coherence in video diffusion models through structured, data-aligned noise injection."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["video diffusion","corruption-aware training","robust video generation","structured noise injection","multimodal robustness","temporal coherence"]},"supplementary_material":{"value":"/attachment/c99cdbeea034fc1e74ad38310569a2906228cc13.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Latent Video Diffusion Models (LVDMs) have achieved state-of-the-art generative quality for image and video generation; however, they remain brittle under noisy conditioning, where small perturbations in text or multimodal embeddings can cascade over timesteps and cause semantic drift. Existing corruption strategies from image diffusion (Gaussian, Uniform) fail in video settings because static noise disrupts temporal fidelity. In this paper, we propose **CAT-Video**, a corruption-aware training framework with structured, data-aligned noise injection tailored for video diffusion. Our two operators—*Batch-Centered Noise Injection (BCNI)* and *Spectrum-Aware Contextual Noise (SACN)*—align perturbations with batch semantics or spectral dynamics to preserve coherence. CAT-Video yields substantial gains: BCNI reduces FVD by **31.9%** on WebVid-2M, MSR-VTT, and MSVD, while SACN improves UCF-101 by **12.3%**, outperforming Gaussian, Uniform, and even large diffusion baselines like DEMO (2.3B) and Lavie (3B) despite training on $\\mathbf{5}\\times$ less data. Ablations confirm the unique value of low-rank, data-aligned noise, and theory establishes why these operators tighten robustness and generalization bounds. CAT-Video thus sets a new framework for robust video diffusion, and our experiments show that it can also be extended to autoregressive generation and multimodal video understanding LLMs."},"_bibtex":{"value":"@misc{\nmaduabuchi2026catvideo,\ntitle={{CAT}-{VIDEO}: {CORRUPTION}-{AWARE} {TRAINING} {FOR} {ROBUST} {VIDEO} {DIFFUSION} {MODELS}},\nauthor={Chika Maduabuchi and Hao Chen and Yujin Han and Jindong Wang},\nyear={2026},\nurl={https://openreview.net/forum?id=unZhwukf0T}\n}"},"title":{"value":"CAT-VIDEO: CORRUPTION-AWARE TRAINING FOR ROBUST VIDEO DIFFUSION MODELS"},"pdf":{"value":"/pdf/36b779e3109b1503084b84e3c1c30d2bc4985918.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"maduabuchi|catvideo_corruptionaware_training_for_robust_video_diffusion_models"},"authorids":{"value":["~Chika_Maduabuchi1","~Hao_Chen15","~Yujin_Han1","~Jindong_Wang4"]},"authors":{"value":["Chika Maduabuchi","Hao Chen","Yujin Han","Jindong Wang"]}},"version":2},{"content":{"summary":{"value":"This paper studies why shortcut learning happens. It focuses on two characteristics that input features may have: availability and predictivity. Availability refers to how frequently that type of feature is available in the data, while predictivity means how useful it is in predicting the target. Shortcuts are described as overly relying on available features that are less predictive. Based on this interpretation of shortcuts, the paper experimentally shows on synthetic and natural image datasets that availability explains shortcut learning where predictivity cannot. These findings are supplemented with a theoretical analysis that concludes that adding a single hidden layer already biases the model to rely on shortcuts."},"presentation":{"value":"3 good"},"contribution":{"value":"3 good"},"soundness":{"value":"3 good"},"strengths":{"value":"The paper is well-written and the main message that availability is a key driver to shortcut learning is convincingly conveyed.\n\nBoth empirical support and theoretical support are provided to underline the effect of availability on shortcut learning."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The theoretical analysis is limited to a single hidden layer MLP which does not reflect the type of architectures used in practice. It is difficult to conclude whether this also holds for other types of architectures like Transformers or CNNs.\n\nExperiments are limited to supervised classification settings. Self-supervised training or other tasks like object detection would be interesting to consider in the context of shortcut learning."},"confidence":{"value":"2: You are willing to defend your assessment, but it is quite likely that you did not understand the central parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked."},"questions":{"value":"The paper points out that availability can explain (some) occurrences of shortcut learning. The availability can be determined by looking at the training data distribution, but could one also identify these shortcuts stemming from availability by directly looking at the train neural network weights? For example, if multiple filters in the same CNN layer share the same pattern?"},"rating":{"value":"6: marginally above the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636702191,"tcdate":1699217505744,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission6363/Reviewer_QCsi"],"signatures":["ICLR.cc/2024/Conference/Submission6363/Reviewer_QCsi"],"forum":"Tj3xLVuE9f","number":5,"license":"CC BY 4.0","cdate":1699217505744,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission6363/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636702191,"domain":"ICLR.cc/2024/Conference","replyto":"Tj3xLVuE9f","id":"XSM1qXN2iv","forumContent":{"venue":{"value":"ICLR 2024 spotlight"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["shortcut learning","spurious correlations","architectural inductive bias"]},"primary_area":{"value":"visualization or interpretation of learned representations"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Deep-learning models can extract a rich assortment of features from data. Which features a model uses depends not only on *predictivity*---how reliably a feature indicates training-set labels---but also on *availability*---how easily the feature can be extracted from inputs. The literature on shortcut learning has noted examples in which models privilege one feature over another, for example texture over shape and image backgrounds over foreground objects. Here, we test hypotheses about which input properties are more available to a model, and systematically study how predictivity and availability interact to shape models' feature use. We construct a minimal, explicit generative framework for synthesizing classification datasets with two latent features that vary in predictivity and in factors we hypothesize to relate to availability, and we quantify a model's shortcut bias---its over-reliance on the shortcut (more available, less predictive) feature at the expense of the core (less available, more predictive) feature. We find that linear models are relatively unbiased, but introducing a single hidden layer with ReLU or Tanh units yields a bias. Our empirical findings are consistent with a theoretical account based on Neural Tangent Kernels. Finally, we study how models used in practice trade off predictivity and availability in naturalistic datasets, discovering availability manipulations which increase models' degree of shortcut bias. Taken together, these findings suggest that the propensity to learn shortcut features is a fundamental characteristic of deep nonlinear architectures warranting systematic study given its role in shaping how models solve tasks."},"_bibtex":{"value":"@inproceedings{\nhermann2024on,\ntitle={On the Foundations of Shortcut Learning},\nauthor={Katherine Hermann and Hossein Mobahi and Thomas FEL and Michael Curtis Mozer},\nbooktitle={The Twelfth International Conference on Learning Representations},\nyear={2024},\nurl={https://openreview.net/forum?id=Tj3xLVuE9f}\n}"},"title":{"value":"On the Foundations of Shortcut Learning"},"pdf":{"value":"/pdf/3f47b29f0e35691e7047d9fbfa0e4c47ea966e49.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference"},"paperhash":{"value":"hermann|on_the_foundations_of_shortcut_learning"},"authorids":{"value":["~Katherine_Hermann1","~Hossein_Mobahi2","~Thomas_FEL1","~Michael_Curtis_Mozer1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Katherine Hermann","Hossein Mobahi","Thomas FEL","Michael Curtis Mozer"]}},"version":2},{"content":{"summary":{"value":"This paper designs an evaluation protocol and surveys a number of video generation papers published over the last few years.  The authors point out shortcomings of evaluation protocols in these priors works (e.g. they often limit studies to video quality comparisons with no clear definitions of video quality, small number of videos, annotators etc).  The authors promise to publish their trainings for annotators along with an interface.  They show that these trainings help significantly to improve reliability of annotator scores.\n\nAnother contribution is a method that doesn’t require all pairs of models to be compared and dynamically is able to select pairs to be compared and show experimentally that this scales better than the naive quadratic approach. \n\nFinally the authors use their methodology to compare 5 recent models and show among other things that Runway and Pika outperform their open source competitors."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"See weaknesses above."},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":2},"strengths":{"value":"Overall the list of surveyed papers in this paper is quite impressive and comprehensive. Moreover, having an evaluation method that can be used without all-pairs of comparisons to be run is clearly important particularly as the number of video generation models increases.  I also absolutely agree with the weaknesses of other video evaluations and agree that human evals are the golden standard and so making a public one available would be very beneficial for the community."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Weaknesses include the following:\n* The authors cite works like Evalcrafter/Vbench but are vague about what value this current submission brings over these prior approaches. Specifically they claim that these approaches “may lack diversity to cover all real-world scenarios” but there are no details given.  As I’ve started to see a number of papers published using the Evalcrafter methodology, I would like to see a clear explanation of the pros and cons.\n* Moreover the number of models actually evaluated in this paper is fairly small (only 5) — given that this is an evaluation methodology paper, it would be much stronger to include more models and perhaps even to show clearly that these human evaluations do not correlate well with things like FID or FVD.\n* I recommend citing and considering using ELO scores which similarly allow for us to score and rank models based on not-all-pairs comparisons and also allow for dynamic selection of model pairs to be compared side by side.  See e.g., https://lmsys.org/blog/2023-05-03-arena/ for a recent example using ELO to rank LLMs by quality as perceived by humans.  Related is the Microsoft TrueSkill work which has also been used widely.\n* Though the authors make a big deal of the fact that few prior works reveal the details of their interface and how the videos are displayed, I found very few details actually revealed in this paper.  The one clear detail I can see is that the authors mention standardizing the height of the video across models.  But this may not be the right thing to do!  Consider the VideoPoet work which generate videos in portrait aspect ratios — resizing so that heights are fixed would be unfair to these vertical videos.  Other details that are important are how to best show videos that are generated at different framerates and even different lengths (for example, how does one compare a 5 second clip from e.g. VideoCrafter to a 1 minute clip from Sora?).\n* Finally some things that are marked as objective clearly are not objective like aesthetic quality.  It would be helpful / more convincing to explain clearly what the training looks like for such a dimension.  I would presume that one example showing what aesthetic quality means is not nearly enough…\n\n\nMinor quibbles:\n* I recommend being clear that “higher is better” in many plots\n* The authors also use a lot of acronyms which impedes readability in various parts of the paper (e.g. last subsection of 4.3)\n* I somewhat disagree with the characterization of FVD being about temporal consistency… even though this is part of it, the statements about the other metrics also apply — it compares feature representations from real and generated images and assesses diversity and clarity…\n* I recommend seeing/citing the Videopoet paper that had very similar comparisons."},"limitations":{"value":"yes"}},"nonreaders":[],"tmdate":1730878939792,"tcdate":1720800830784,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission4622/Reviewer_2pJH"],"signatures":["NeurIPS.cc/2024/Conference/Submission4622/Reviewer_2pJH"],"forum":"0AwMciNShl","number":4,"license":"CC BY 4.0","cdate":1720800830784,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission4622/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730878939792,"domain":"NeurIPS.cc/2024/Conference","replyto":"0AwMciNShl","id":"jzW81U4Zvk","forumContent":{"venue":{"value":"NeurIPS 2024 poster"},"keywords":{"value":["Human evaluation","text-to-video models"]},"primary_area":{"value":"evaluation"},"abstract":{"value":"Recent text-to-video (T2V) technology advancements, as demonstrated by models such as Gen2, Pika, and Sora, have significantly broadened its applicability and popularity. \nDespite these strides, evaluating these models poses substantial challenges. \nPrimarily, due to the limitations inherent in automatic metrics, manual evaluation is often considered a superior method for assessing T2V generation. However, existing manual evaluation protocols face reproducibility, reliability, and practicality issues.\nTo address these challenges, this paper introduces the Text-to-Video Human Evaluation (T2VHE) protocol, a comprehensive and standardized protocol for T2V models. \nThe T2VHE protocol includes well-defined metrics, thorough annotator training, and an effective dynamic evaluation module. \nExperimental results demonstrate that this protocol not only ensures high-quality annotations but can also reduce evaluation costs by nearly 50\\%.\nWe will open-source the entire setup of the T2VHE protocol, including the complete protocol workflow, the dynamic evaluation component details, and the annotation interface code. This will help communities establish more sophisticated human assessment protocols."},"_bibtex":{"value":"@inproceedings{\nzhang2024rethinking,\ntitle={Rethinking Human Evaluation Protocol for Text-to-Video Models: Enhancing Reliability, Reproducibility, and Practicality},\nauthor={Tianle Zhang and Langtian Ma and Yuchen Yan and Yuchen Zhang and Yue Yang and Ziyao Guo and Wenqi Shao and Kai Wang and Yang You and Yu Qiao and Ping Luo and Kaipeng Zhang},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=0AwMciNShl}\n}"},"title":{"value":"Rethinking Human Evaluation Protocol for Text-to-Video Models: Enhancing Reliability, Reproducibility, and Practicality"},"pdf":{"value":"/pdf/aa115a2ffe88cc5707a0f711b0ee921175fa9141.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"zhang|rethinking_human_evaluation_protocol_for_texttovideo_models_enhancing_reliability_reproducibility_and_practicality"},"authorids":{"value":["~Tianle_Zhang4","~Langtian_Ma1","~Yuchen_Yan5","~Yuchen_Zhang8","~Yue_Yang6","~Ziyao_Guo1","~Wenqi_Shao2","~Kai_Wang8","~Yang_You1","~Yu_Qiao1","~Ping_Luo2","~Kaipeng_Zhang1"]},"authors":{"value":["Tianle Zhang","Langtian Ma","Yuchen Yan","Yuchen Zhang","Yue Yang","Ziyao Guo","Wenqi Shao","Kai Wang","Yang You","Yu Qiao","Ping Luo","Kaipeng Zhang"]}},"version":2},{"content":{"summary":{"value":"The paper introduces TemporalBench, a new benchmark for evaluating fine-grained temporal understanding in multimodal video models. It addresses the limitation of current video benchmarks like MSRVTT and TGIF, which fail to effectively evaluate AI models' temporal reasoning due to the lack of detailed temporal annotations. TemporalBench consists of approximately 10K pairs of video description questions derived from around 2K high-quality human-annotated video captions. It provides fine-grained temporal annotations to assess models' ability to reason about time-sensitive events. The benchmark tests models on tasks like video QA, captioning, and long video understanding, and shows that even state-of-the-art models like GPT-4o perform poorly, highlighting a significant gap between human and AI capabilities in temporal understanding."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. When will the dataset be open-sourced?\n\n2. How to balance the evaluation cost and evaluation effect?"},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. TemporalBench improves the evaluation of multimodal video models' spatio-temporal comprehension by providing approximately 10K pairs of video description questions with fine-grained temporal annotations that capture subtleties in video actions and events.\n\n2. The benchmark supports various tasks such as video QA and video captioning and sources data from a range of video domains, ensuring a broad and diverse set of testing scenarios.\n\n3. By benchmarking against human performance, TemporalBench exposes the significant gaps in current state-of-the-art models' fine-grained temporal understanding, indicating areas for future research and development."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. One major concern is about the video data sources used in this benchmark. From what I can see, the authors only used existing academic datasets such as COIN, Activity, and EgoExo4D to source their videos, rather than collecting any new video data from more recent real-world scenarios. This is quite problematic because these academic datasets are already widely known and used in the research community. What this means is that many models may have already been trained or fine-tuned on these videos in some way. As a result, model developers could potentially find ways to optimize their models specifically for these familiar videos without too much effort, leading to artificially inflated performance scores that don't actually reflect how well the models would work on truly novel, real-world video content.\n\n2. I found it quite frustrating that the paper doesn't include any comprehensive statistical comparison tables between TemporalBench and other existing video benchmarks. Without such detailed comparisons showing metrics like dataset size, video duration distributions, annotation density, and task complexity, it's really hard for readers like me to get a clear picture of exactly what makes TemporalBench special or better than previous benchmarks. Some simple comparison tables would have made it much easier to understand TemporalBench's unique contributions and advantages.\n\n3. Regarding the MBA evaluation method proposed in the paper - while I appreciate that it creates a more challenging and thorough evaluation scenario, I have some practical concerns. The method requires significantly more computational resources and evaluation time compared to traditional approaches. This raises an important question that the authors haven't fully addressed: how do we find the right balance between having a rigorous evaluation process and keeping the computational costs manageable? This is especially relevant for researchers with limited computing resources.\n\n4. A significant limitation of this work is that the dataset hasn't been made publicly available yet. This is quite disappointing because it prevents other researchers from independently verifying the results or building upon this work. The research community would benefit tremendously if the authors could release the complete dataset, including all the videos, annotations, and evaluation scripts. This would allow for more transparent validation of the benchmark's claims and accelerate progress in temporal understanding research. I strongly encourage the authors to consider making their dataset open-source as soon as possible."}},"nonreaders":[],"tmdate":1731427456006,"tcdate":1730611337006,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission1635/Reviewer_s93Z"],"signatures":["ICLR.cc/2025/Conference/Submission1635/Reviewer_s93Z"],"forum":"Wto5U7q6I2","number":3,"license":"CC BY 4.0","cdate":1730611337006,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission1635/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427456006,"domain":"ICLR.cc/2025/Conference","replyto":"Wto5U7q6I2","id":"oDRNXVxAik","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"TLDR":{"value":"We introduce a finegrained multimodal video understanding benchmark"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["video","benchmark","multimodel"]},"supplementary_material":{"value":"/attachment/1b00a92cc0682497e1612273bbcb5db6abfd1201.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Understanding fine-grained temporal dynamics is crucial for video understanding. Yet, popular video benchmarks, such as MSRVTT and TGIF, often fail to effectively evaluate AI models' temporal reasoning abilities due to the lack of fine-grained temporal annotations. \nAs a result, text-based models, leveraging strong language priors, often perform comparably to video models, and image-trained models have been reported to outperform their video-trained counterparts on MSRVTT and TGIF. This paper introduces a new TemporalBench benchmark for fine-grained temporal event understanding in videos. TemporalBench, sourced from a diverse video datasets, consists of $\\sim$10K pairs of video description questions, derived from $\\sim$2K high-quality human-annotated video captions.  Uniquely, our benchmark provides fine-grained temporal annotations to evaluate models' temporal reasoning abilities. Our results show that state-of-the-art models like GPT-4o achieve only 38.0\\% multiple binary QA accuracy on TemporalBench, demonstrating a significant human-AI gap in temporal understanding. We hope that TemporalBench is instrumental to fostering research on improving models' temporal reasoning capabilities."},"_bibtex":{"value":"@misc{\ncai2024temporalbench,\ntitle={TemporalBench: Towards Fine-grained Temporal Understanding for  Multimodal Video  Models},\nauthor={Mu Cai and Reuben Tan and Jianrui Zhang and Bocheng Zou and Kai Zhang and Feng Yao and Fangrui Zhu and Jing Gu and Yiwu Zhong and Yuzhang Shang and Yao Dou and Jaden Park and Jianfeng Gao and Yong Jae Lee and Jianwei Yang},\nyear={2024},\nurl={https://openreview.net/forum?id=Wto5U7q6I2}\n}"},"title":{"value":"TemporalBench: Towards Fine-grained Temporal Understanding for  Multimodal Video  Models"},"pdf":{"value":"/pdf/433bcaa84f8ff694d5c6245d586377638b770240.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"cai|temporalbench_towards_finegrained_temporal_understanding_for_multimodal_video_models"},"authorids":{"value":["~Mu_Cai1","~Reuben_Tan1","~Jianrui_Zhang1","~Bocheng_Zou1","~Kai_Zhang10","~Feng_Yao1","~Fangrui_Zhu1","~Jing_Gu2","~Yiwu_Zhong1","~Yuzhang_Shang1","~Yao_Dou1","~Jaden_Park1","~Jianfeng_Gao1","~Yong_Jae_Lee2","~Jianwei_Yang1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Mu Cai","Reuben Tan","Jianrui Zhang","Bocheng Zou","Kai Zhang","Feng Yao","Fangrui Zhu","Jing Gu","Yiwu Zhong","Yuzhang Shang","Yao Dou","Jaden Park","Jianfeng Gao","Yong Jae Lee","Jianwei Yang"]}},"version":2},{"content":{"summary":{"value":"The paper proposes PanoWorldX, a novel framework for high-fidelity and controllable panoramic video generation with diverse camera trajectories. PanoWorldX consists of a novel pipeline using the Unreal Engine to procedurally generate a large, diverse dataset of panoramic video-trajectory pairs, and a Sphere-Aware Diffusion Transformer that reprojects equirectangular features onto the spherical surface to accurately model geometric adjacency in the latent space. By specifically addressing the geometric misalignment between the spherical geometry of panoramic data and conventional diffusion model priors, PanoWorldX demonstrates significant improvements over existing methods in terms of motion range, control precision, and visual quality."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"- Modifying the standard DiT architecture with additional attention branches (Sphere-Aware and Exploration-Aware) inherently increases computational load during training and inference. What is the latency or memory overhead compared to the baseline DiT model?\n- While trajectory complexity was discussed, can the model handle cases like abrupt changes in elevation, sharp turns in confined spaces, or rapid changes in camera velocity without noticeable artifacts?\n- What is the individual contribution of the Exploration-Aware Attention in the ablation study?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"- The panorama data synthesis pipeline for large-scale and accurately labeled panoramic video-trajectory pairs is a valuable contribution, as the scarcity of data has limited the development of this task.\n- The Sphere-Aware Attention and Exploration-Aware Attention are theoretically sound and practical for integration into existing DiT architectures. Explicitly modeling for the spherical topology is effective for high-fidelity panoramic generation,\n- Experimental results demonstrate that PanoWorldX outperforms baseline methods and achieves good panoramic video generation performance."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- While I agree that synthesized panoramic data is necessary, the domain gap between synthetic and real-world data remains significant. Complex real-world scenarios, such as variations in lighting, non-rigid deformation, and motion blur and noise discrepancies that occur during shooting, remain challenging to represent in synthetic data. This paper's analysis of the data lacks a detailed discussion of these cases, which may be beneficial for future high-quality panoramic data synthesis.\n- PanoWorldX focuses heavily on static scene generation and camera movement. It is unclear how PanoWorld-X handles highly dynamic, non-structural elements, such as moving characters, weather changes, or complex large motions. Further discussion or ablation studies on the temporal stability of non-rigid objects would enhance the review of temporal coherence."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762922916679,"tcdate":1761934307456,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission11906/Reviewer_3MJN"],"signatures":["ICLR.cc/2026/Conference/Submission11906/Reviewer_3MJN"],"forum":"iZyBEbq6jR","number":2,"license":"CC BY 4.0","cdate":1761934307456,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission11906/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762922916679,"domain":"ICLR.cc/2026/Conference","replyto":"iZyBEbq6jR","id":"L3VYS9spRe","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["immersive content"]},"supplementary_material":{"value":"/attachment/5773686dc6b17b2632d3dc874878b8b3fbb061b7.zip"},"primary_area":{"value":"other topics in machine learning (i.e., none of the above)"},"abstract":{"value":"Generating a complete and explorable 360-degree visual world enables a wide range of downstream applications. While prior works have advanced the field, they remain constrained by either narrow field-of-view limitations, which hinder the synthesis of continuous and holistic scenes, or insufficient controllability that restricts free exploration by users or autonomous agents. To address this, we propose PanoWorld-X, a novel framework for high-fidelity and controllable panoramic video generation with diverse camera trajectories. \n First, we propose a novel pipeline for synthesizing panoramic video-trajectory dataset pairs in virtual 3D environments via Unreal Engine. This pipeline consists of four main steps and enables the collection of a large-scale dataset with rich scene diversity and accurate trajectory annotations.\n To achieve precise panoramic video generation, we identify that the bottleneck arises from the misalignment between the spherical geometry of panoramic data and the inductive priors of conventional video diffusion models. To address this, we leverage the spherical connectivity characteristics of panorama data, and propose a Sphere-Aware Diffusion Transformer that reprojects equirectangular features onto the spherical surface, thereby capturing geometric adjacency in the latent space. This design significantly improves both visual fidelity and spatiotemporal continuity.\n   Extensive experiments demonstrate that our PanoWorld-X achieves superior performance in various aspects, including motion range, control precision, and visual quality, underscoring its potential for real-world applications."},"_bibtex":{"value":"@misc{\nyin2026panoworldx,\ntitle={PanoWorld-X: Generating Explorable Panoramic Worlds via Sphere-Aware Video Diffusion},\nauthor={Yuyang Yin and Hao-Xiang Guo and Fangfu Liu and Mengyu Wang and Hanwen Liang and Eric Li and Yikai Wang and Xiaojie Jin and Yao Zhao and Yunchao Wei},\nyear={2026},\nurl={https://openreview.net/forum?id=iZyBEbq6jR}\n}"},"title":{"value":"PanoWorld-X: Generating Explorable Panoramic Worlds via Sphere-Aware Video Diffusion"},"pdf":{"value":"/pdf/8d2464ff9acb82ee4bdbbfda2c8b3031ed0755c7.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"yin|panoworldx_generating_explorable_panoramic_worlds_via_sphereaware_video_diffusion"},"authorids":{"value":["~Yuyang_Yin1","~Hao-Xiang_Guo1","~Fangfu_Liu2","~Mengyu_Wang5","~Hanwen_Liang2","~Eric_Li8","~Yikai_Wang2","~Xiaojie_Jin1","~Yao_Zhao1","~Yunchao_Wei1"]},"authors":{"value":["Yuyang Yin","Hao-Xiang Guo","Fangfu Liu","Mengyu Wang","Hanwen Liang","Eric Li","Yikai Wang","Xiaojie Jin","Yao Zhao","Yunchao Wei"]}},"version":2},{"content":{"summary":{"value":"This paper observes that video frames have strong local spatial correlation compared to distant locations, and thus propose to perform parallel frame decoding following the diagonal trajectory. The proposed method can be applied on pre-trained AR video models, and addresses the time bottleneck of their original sequential decoding paradigm. It can be adapted training-free or with finetuning. The proposed method is tested on multiple video frame prediction benchmarks, achieving large speed-up with minimal quality degradation."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"- How the VRAM will change when applied the parallel diagonal decoding? This is not a major issue, but would help comprehensive understanding the full trade-off."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"- This work has good motivation and observation that video content is spatially locally correlated across frames, and thus proposes to decode only dependent on the spatially local content of previous frames, i.e. from a corner and following diagonals.\n\n- The proposed method is (almost) training-free and can be handily adapted to pre-trained AR models to gain big speed-up with minimal quality drop."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The proposed method is only evaluated on frame prediction task, where a domain-specific model continue the video given a leading clip. Not only the task complexity but also the core observation might depend on the video types, e.g. egocentric videos vs action videos. It would be further highlighted to test with open-domain T2V AR models on general video generation only conditioned on non-visual input, especially on general videos like human actions, object moving with fixed camera view, etc. to solidate the observation and proposed method's generability.\n\n- The quantitative metrics need to be upgrade. FVD is highly sensitive to statistic data and thus needs to be calculated on at least thousands of samples to be stable enough. For general open-domain T2V cases, more comprehensive benchmarks like VBench etc. should be involved."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762923370417,"tcdate":1761981462722,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission12498/Reviewer_9SGt"],"signatures":["ICLR.cc/2026/Conference/Submission12498/Reviewer_9SGt"],"forum":"hMZbIIHMDI","number":4,"license":"CC BY 4.0","cdate":1761981462722,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission12498/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762923370417,"domain":"ICLR.cc/2026/Conference","replyto":"hMZbIIHMDI","id":"dQAqMDILon","forumContent":{"TLDR":{"value":"Diagonal Decoding introduces a training-free method to accelerate autoregressive video generation models, achieving up to 10x speedup."},"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Autoregressive Video Generation","Accelerate"]},"supplementary_material":{"value":"/attachment/dca1188b6a0626f1c795320e92d4144a8f119b33.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Autoregressive Transformer has demonstrated impressive performance in generation models. However, their sequential, token-by-token decoding becomes a severe bottleneck for video generation, which may require generating tens of thousands of tokens sequentially. In this paper, we introduce Diagonal Decoding (DiagD), a training-free inference acceleration algorithm that exploits spatiotemporal correlations to speed up autoregressively pre-trained models. DiagD generates tokens simultaneously along diagonal trajectories in the spatial-temporal token grid, enabling parallel decoding within frames and partial overlap across successive frames. The proposed algorithm is versatile and adaptive to various generative models and tasks and offers adjustable trade-offs between speed and visual quality. Furthermore, we propose a cost-effective fine-tuning strategy that aligns the attention patterns of the model with the new decoding order to show the potential of training with DiagD. Experiments on several autoregressive video generation models and datasets demonstrate that DiagD achieves up to **10x** speed-up over naive sequential decoding, while preserving comparable visual fidelity."},"_bibtex":{"value":"@misc{\nye2025fast,\ntitle={Fast Autoregressive Video Generation with Diagonal Decoding},\nauthor={Yang Ye and Junliang Guo and Haoyu Wu and Tianyu He and Tim Pearce and Tabish Rashid and Katja Hofmann and Jiang Bian},\nyear={2025},\nurl={https://openreview.net/forum?id=hMZbIIHMDI}\n}"},"title":{"value":"Fast Autoregressive Video Generation with Diagonal Decoding"},"pdf":{"value":"/pdf/1bfdf77fa15553c8b8a78a14ccebb3d89e5bcab3.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"ye|fast_autoregressive_video_generation_with_diagonal_decoding"},"authorids":{"value":["~Yang_Ye2","~Junliang_Guo1","~Haoyu_Wu3","~Tianyu_He1","~Tim_Pearce1","~Tabish_Rashid1","~Katja_Hofmann1","~Jiang_Bian1"]},"authors":{"value":["Yang Ye","Junliang Guo","Haoyu Wu","Tianyu He","Tim Pearce","Tabish Rashid","Katja Hofmann","Jiang Bian"]}},"version":2},{"content":{"venue":{"value":"LATA 2016"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-319-30000-9_14.pdf"},"venueid":{"value":"dblp.org/conf/LATA/2016"},"paperhash":{"value":"smetsers|minimal_separating_sequences_for_all_pairs_of_states"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Rick_Smetsers:","https://dblp.org/search/pid/api?q=author:Joshua_Moerman:","~David_N._Jansen1"]},"html":{"value":"https://doi.org/10.1007/978-3-319-30000-9_14"},"_bibtex":{"value":"@inproceedings{DBLP:conf/lata/SmetsersMJ16,\n  author={Rick Smetsers and Joshua Moerman and David N. Jansen},\n  title={Minimal Separating Sequences for All Pairs of States},\n  year={2016},\n  cdate={1451606400000},\n  pages={181-193},\n  url={https://doi.org/10.1007/978-3-319-30000-9_14},\n  booktitle={LATA},\n  crossref={conf/lata/2016}\n}\n"},"abstract":{"value":"Finding minimal separating sequences for all pairs of inequivalent states in a finite state machine is a classic problem in automata theory. Sets of minimal separating sequences, for instance, play a central role in many conformance testing methods. Moore has already outlined a partition refinement algorithm that constructs such a set of sequences in \\(\\mathcal {O}(mn)\\) time, where m is the number of transitions and n is the number of states. In this paper, we present an improved algorithm based on the minimization algorithm of Hopcroft that runs in \\(\\mathcal {O}(m \\log n)\\) time. The efficiency of our algorithm is empirically verified and compared to the traditional algorithm."},"title":{"value":"Minimal Separating Sequences for All Pairs of States"},"authors":{"value":["Rick Smetsers","Joshua Moerman","David N. Jansen"]}},"tmdate":1741262351203,"pdate":1451606400000,"tcdate":1741262311810,"writers":["~"],"signatures":["~David_N._Jansen1"],"forum":"EkunSKPVMq","license":"CC BY-SA 4.0","number":361921,"cdate":1451606400000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1741262351203,"domain":"DBLP.org","id":"EkunSKPVMq","version":2},{"content":{"summary":{"value":"The authors combine base diffusion models for audio and video into a unified framework, training it to generate both modalities simultaneously. They introduce two mechanisms to enhance the alignment of audio-video pairs: 1) Timestep Adjustment, which supplies distinct timestep information to each base model; and 2) Cross-Modal Conditioning as Positional Encoding (CMC-PE),.a design for additional modules that integrates cross-modal information as temporal position data, akin to positional encoding. This approach offers an inductive bias for achieving temporal alignment in the generated data compared to the traditional cross-attention mechanism."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1. The FVD score on the Landscape dataset is much worse than that on the VGGSound dataset. However, the Landscape dataset is a quite small and domain dataset. Intuitively, Landscape should be a more simple dataset than VGGSound, thus is expected to be easier for modeling.\n2. In Line 404-406, the authors claimed that the CMC-PE improves FVD than cross-attention. However, the Table 1 demonstrates poor FVD when using CMC-PE, which seems to be conflict."},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. The paper proposes a unified model for joint audio and video generation, integrating base diffusion models for both modalities.It introduces two designed mechanisms, Timestep Adjustment and CMC-PE, to enhance the alignment between audio and video pairs.\n2. The proposed CMC-PE method can explicitly align audio and video signal by incorporating cross-modal information like temporal position embedding. The AV-align socres in Table.1 demonstate the effectiveness of CMC-PE in improving audio-video alignment."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Lacking comparison with other sounding video generative models based on off-the-shelf audio or video generation models like [1]. The authors only reported quantitative comparison with models trained from scratch, which is unfair since the proposed method utilized pretrained models.\n\n2. Lacking qualitative comparison between the proposed method and baseline methods. The authors only present audio-video samples generated by the proposed method. However, qualitative comparison with other methods is necessary to better explain how the proposed method can improve temporal alignment of audio and video.\n\n\n[1] Xing Y, He Y, Tian Z, et al. Seeing and hearing: Open-domain visual-audio generation with diffusion latent aligners[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2024: 7151-7161."}},"nonreaders":[],"tmdate":1731427491942,"tcdate":1730639353112,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission1798/Reviewer_NBr8"],"signatures":["ICLR.cc/2025/Conference/Submission1798/Reviewer_NBr8"],"forum":"WqL4wOU3tw","number":2,"license":"CC BY 4.0","cdate":1730639353112,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission1798/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427491942,"domain":"ICLR.cc/2025/Conference","replyto":"WqL4wOU3tw","id":"YD2ZWkn77z","forumContent":{"TLDR":{"value":"We have built a simple but strong baseline method for sounding video generation."},"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["sounding video generation","diffusion models","audio-visual","audio generation","video generation"]},"supplementary_material":{"value":"/attachment/2d82eb8e0d2d743d33fcbec5b978470a6b5164c0.zip"},"primary_area":{"value":"generative models"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"In this work, we build a simple but strong baseline for sounding video generation. Given base diffusion models for audio and video, we integrate them with additional modules into a single model and train it to make the model jointly generate audio and video. To enhance alignment between audio-video pairs, we introduce two novel mechanisms in our model. The first one is timestep adjustment, which provides different timestep information to each base model. It is designed to align how samples are generated along with timesteps across modalities. The second one is a new design of the additional modules, termed Cross-Modal Conditioning as Positional Encoding (CMC-PE). In CMC-PE, cross-modal information is embedded as if it represents temporal position information, and the embeddings are fed into the model like positional encoding. Compared with the popular cross-attention mechanism, CMC-PE provides a better inductive bias for temporal alignment in the generated data. Experimental results validate the effectiveness of the two newly introduced mechanisms and also demonstrate that our method outperforms existing methods. The source code will be released upon acceptance."},"_bibtex":{"value":"@misc{\nishii2024a,\ntitle={A Simple but Strong Baseline for Sounding Video Generation: Effective Adaptation of Audio and Video Diffusion Models for Joint Generation},\nauthor={Masato Ishii and Akio Hayakawa and Takashi Shibuya and Yuki Mitsufuji},\nyear={2024},\nurl={https://openreview.net/forum?id=WqL4wOU3tw}\n}"},"title":{"value":"A Simple but Strong Baseline for Sounding Video Generation: Effective Adaptation of Audio and Video Diffusion Models for Joint Generation"},"pdf":{"value":"/pdf/1dfdf2580c1048c62f8ef127890726c23f0bf60b.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"ishii|a_simple_but_strong_baseline_for_sounding_video_generation_effective_adaptation_of_audio_and_video_diffusion_models_for_joint_generation"},"authorids":{"value":["~Masato_Ishii1","~Akio_Hayakawa1","~Takashi_Shibuya1","~Yuki_Mitsufuji1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Masato Ishii","Akio Hayakawa","Takashi Shibuya","Yuki Mitsufuji"]}},"version":2},{"content":{"summary":{"value":"The manuscript proposes a new framework that uses an interactive large-scale visual language model (VLM) to accurately explain video anomalies. The framework can explicitly integrate motion modalities to enhance anomaly identification. The manuscript also constructs an auxiliary consistency loss to guide the video branch to focus on motion modalities. The manuscript annotates more than 8,000 anomaly videos with language descriptions, enabling effective training in different open-world scenarios, and creates 8,000 question-answer pairs for users' open-world questions. The final results show that HAWK achieves SOTA performance, surpassing existing baselines in both video description generation and question answering."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"As mentioned above."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1.This manuscript proposes a novel video language framework HAWK, which aims to understand video anomalies, and combines motion modalities to enhance its capabilities.\n2.This manuscript collects seven video anomaly datasets from different scenarios and generates rich language descriptions for each video. At the same time, considering the diversity of open-world questions, question-answer pairs are generated to solve potential user queries.\n3.The proposed method achieves state-of-the-art performance on three public anomaly detection datasets."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. There are already many studies focusing on the background information in abnormal events. Have the authors considered describing the background information when describing the action information?\n[1] Scene-aware context reasoning for unsupervised abnormal event detection in videos[C]//Proceedings of the 28th ACM international conference on multimedia. 2020: 184-192.\n[2] Few-shot scene-adaptive anomaly detection[C]//Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part V 16. Springer International Publishing, 2020: 125 -141.\n[3] Hierarchical scene normality-binding modeling for anomaly detection in surveillance videos[C]//Proceedings of the 30th ACM international conference on multimedia. 2022: 6103-6112.\n2.The dataset proposed in the manuscript also supports question answering in open scenes, but the method does not reflect how to use these questions and answers to improve the model's anomaly detection and understanding capabilities.\n3.When constructing a dataset, if a description is generated for each second in the dataset, it may result in a lot of repeated descriptions. In order to avoid this redundancy problem, have you considered using keyframes instead of second-by-second descriptions to reduce the complexity and computational cost of data processing?\n4.Most previous studies pay equal attention to various parts of the video, such as background, motion information, character appearance, etc. This paper focuses on motion information. Section 4.3 mainly extracts language descriptions related to motion. Whether background information is not considered. Moreover, the 7 datasets proposed in the manuscript contain more abnormal scenes. I am very curious why the scene information is not paid attention to."},"limitations":{"value":"Yes, the authors address possible limitations of their study."}},"nonreaders":[],"tmdate":1730878748782,"tcdate":1720521719065,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission2096/Reviewer_Cy7X"],"signatures":["NeurIPS.cc/2024/Conference/Submission2096/Reviewer_Cy7X"],"forum":"vBKoEZ1PG3","number":2,"license":"CC BY 4.0","cdate":1720521719065,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission2096/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730878748782,"domain":"NeurIPS.cc/2024/Conference","replyto":"vBKoEZ1PG3","id":"MMPddwrTNN","forumContent":{"venue":{"value":"NeurIPS 2024 poster"},"TLDR":{"value":"A large vision-language model for understanding open-world video anomalies."},"keywords":{"value":["Video Anomalies Understanding"]},"primary_area":{"value":"machine_vision"},"abstract":{"value":"Video Anomaly Detection (VAD) systems can autonomously monitor and identify disturbances, reducing the need for manual labor and associated costs. However, current VAD systems are often limited by their superficial semantic understanding of scenes and minimal user interaction. Additionally, the prevalent data scarcity in existing datasets restricts their applicability in open-world scenarios.\nIn this paper, we introduce HAWK, a novel framework that leverages interactive large Visual Language Models (VLM) to interpret video anomalies precisely. Recognizing the difference in motion information between abnormal and normal videos, HAWK explicitly integrates motion modality to enhance anomaly identification. To reinforce motion attention, we construct an auxiliary consistency loss within the motion and video space, guiding the video branch to focus on the motion modality. Moreover, to improve the interpretation of motion-to-language, we establish a clear supervisory relationship between motion and its linguistic representation. Furthermore, we have annotated over 8,000 anomaly videos with language descriptions, enabling effective training across diverse open-world scenarios, and also created 8,000 question-answering pairs for users' open-world questions. The final results demonstrate that HAWK achieves SOTA performance, surpassing existing baselines in both video description generation and question-answering. Our codes/dataset/demo will be released at https://github.com/jqtangust/hawk."},"_bibtex":{"value":"@inproceedings{\ntang2024hawk,\ntitle={{HAWK}: Learning to Understand Open-World Video Anomalies},\nauthor={Jiaqi Tang and Hao LU and RUIZHENG WU and Xiaogang Xu and Ke Ma and Cheng Fang and Bin Guo and Jiangbo Lu and Qifeng Chen and Ying-Cong Chen},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=vBKoEZ1PG3}\n}"},"title":{"value":"HAWK: Learning to Understand Open-World Video Anomalies"},"pdf":{"value":"/pdf/72f5b57d7e449d765338f2a60df13977e0717488.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"tang|hawk_learning_to_understand_openworld_video_anomalies"},"authorids":{"value":["~Jiaqi_Tang1","~Hao_LU8","~RUIZHENG_WU1","~Xiaogang_Xu2","~Ke_Ma11","~Cheng_Fang6","~Bin_Guo3","~Jiangbo_Lu1","~Qifeng_Chen1","~Ying-Cong_Chen1"]},"authors":{"value":["Jiaqi Tang","Hao LU","RUIZHENG WU","Xiaogang Xu","Ke Ma","Cheng Fang","Bin Guo","Jiangbo Lu","Qifeng Chen","Ying-Cong Chen"]}},"version":2},{"content":{"summary":{"value":"Video Retrieval is a very challenging task, hence the proposal of composite video retrieval, which uses images and text as retrieval signals to search for video content. Current content retrieval faces two issues:\n(1). A lack of video-retrieval datasets with fine-grained descriptions.\n(2). A lack of effective solutions to implement good video-retrieval. \nIn response to the first issue, the authors proposed a fine-grained large-scale video retrieval dataset, FineCVR-1M, which includes one million video-text pairs. Addressing the second model-side problem, the authors proposed a decoupled text representation-cross modality alignment model. However, the methodology part requires more ablations and the related works are incomplete. It would be better to address these issues for further ranking improvement."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"1. During the construction of the WebVid-CoVR dataset, did the authors consider the possibility of retrieval mismatch due to video length when they mixed short videos like MSRVTT with longer ones like ActivityNet?"},"rating":{"value":8},"details_of_ethics_concerns":{"value":"None"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The authors have provided a high-quality, fine-grained dataset in the field of video retrieval, which has made significant contributions to the community.\n\n2. The authors analyzed the advantages of this new dataset over the WebVid-CoVR dataset, which is also proposed for the composed video retrieval task. CoVR uses LLM to describe the differences between two videos and does not support fine-grained details. What we generate are fine-grained text descriptions.\n\n3. The authors' dataset structure is quite ingenious, as it introduces LLM to label the three important components between two video pairs: retained component, injected component, and excluded component, to generate fine-grained text descriptions."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. In the abstract section, the challenges of this composed video retrieval task are described too simplistically. The difficulty of fine-grained video-text modeling is a long-standing and intractable problem. The authors could provide a more nuanced description of the challenges.\n\n2. The author should conduct a more detailed exploration of the temporal encoder. Dicosa[1] found that using the text-side [cls] token to perform cosine similarity with each video frame feature, and then passing through softmax to convert into a probability distribution to guide the weighted summation of video features is a better compression method. Could this process bring performance improvements in composed video retrieval?\n\n3. The author lacks citations of relevant papers. Specifically, DPC-KNN has also been used in other video retrieval works [2], and relevant studies include Dicosa [2] and FreestyleRet [3]. FreestyleRet contains the exploration of composed image retrieval.\n\n[1]. Text-video retrieval with Disentangled Conceptualization and Set-to-Set Alignment\n[2]. Video-Text as Game Players: Hierarchical Banzhaf Interaction for Cross-Modal Representation Learning\n[3]. FreestyleRet: Retrieving Images from Style-Diversified Queries"}},"nonreaders":[],"tmdate":1732549806695,"tcdate":1730446299305,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission1597/Reviewer_FBRb"],"signatures":["ICLR.cc/2025/Conference/Submission1597/Reviewer_FBRb"],"forum":"wGa2plE8ka","number":3,"license":"CC BY 4.0","cdate":1730446299305,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission1597/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732549806695,"domain":"ICLR.cc/2025/Conference","replyto":"wGa2plE8ka","id":"QKHmli6EPP","forumContent":{"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Composed Video Retrieval; Fine-grained Representation; Feature Disentanglement"]},"supplementary_material":{"value":"/attachment/d9a2ca7ff50404ea1fcda5d16fd9bfe8b30251b0.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"With the explosive growth of video data, finding videos that meet detailed requirements in large datasets has become a challenge. To address this, the composed video retrieval task has been introduced, enabling users to retrieve videos using complex queries that involve both visual and textual information. However, the inherent heterogeneity between the modalities poses significant challenges. Textual data are highly abstract, while video content contains substantial redundancy. The modality gap in information representation makes existing methods struggle with the modality fusion and alignment required for fine-grained composed retrieval. To overcome these challenges, we first introduce FineCVR-1M, a fine-grained composed video retrieval dataset containing 1,010,071 video-text triplets with detailed textual descriptions. This dataset is constructed through an automated process that identifies key concept changes between video pairs to generate textual descriptions for both static and action concepts. For fine-grained retrieval methods, the key challenge lies in understanding the detailed requirements. Text description serves as clear expressions of intent, but it requires models to distinguish subtle differences in the description of video semantics. Therefore, we propose a textual Feature Disentanglement and Cross-modal Alignment framework (FDCA) that disentangles features at both the sentence and token levels. At the sequence level, we separate text features into retained and injected features. At the token level, an Auxiliary Token Disentangling mechanism is proposed to disentangle texts into retained, injected, and excluded tokens. The disentanglement at both levels extracts fine-grained features, which are aligned and fused with the reference video to extract global representations for video retrieval. Experiments on FineCVR-1M dataset demonstrate the superior performance of FDCA. Our code and dataset are available at: https://may2333.github.io/FineCVR/."},"_bibtex":{"value":"@inproceedings{\nwu2025learning,\ntitle={Learning Fine-Grained Representations through Textual Token Disentanglement in Composed Video Retrieval},\nauthor={Yue WU and Zhaobo Qi and Yiling Wu and Junshu Sun and Yaowei Wang and Shuhui Wang},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=wGa2plE8ka}\n}"},"title":{"value":"Learning Fine-Grained Representations through Textual Token Disentanglement in Composed Video Retrieval"},"pdf":{"value":"/pdf/920758adbc591c30ca384d1334e95f204d0e7658.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"wu|learning_finegrained_representations_through_textual_token_disentanglement_in_composed_video_retrieval"},"authorids":{"value":["~Yue_WU22","~Zhaobo_Qi1","~Yiling_Wu1","~Junshu_Sun1","~Yaowei_Wang1","~Shuhui_Wang1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Yue WU","Zhaobo Qi","Yiling Wu","Junshu Sun","Yaowei Wang","Shuhui Wang"]}},"version":2},{"content":{"summary":{"value":"This paper argues that the development and application of Chinese VLP and multimodal LLM are lagging behind the English counterpart, due to the lack of a large-scale Chinese video-language dataset. Thus, they propose a new dataset Youku-mPLUG, which consists of 10 million Chinese video-text pairs for pertaining, and a dataset with 0.3 million videos for downstream benchmarks, including video-text retrieval, video captioning, and video category classification. Meanwhile, they investigate popular video-language models (e.g., ALPRO, mPLUG-2), and the new proposed model mPLUG-video. The model mPLUG-video consists of a trainable video encoder, a visual abstractor module, and a frozen pre-trained LLM decoder. Experiments show that models pre-trained on Youku-mPLUG gain on multiple tasks. Furthermore, by building on top of Bloomz, mPLUG-video can achieve impressive zero-shot performance with very few trainable parameter."},"presentation":{"value":"2 fair"},"contribution":{"value":"2 fair"},"soundness":{"value":"3 good"},"strengths":{"value":"+ This paper proposes a large-scale dataset with 10 million Chinese video-text pairs for pertaining, and a dataset with 0.3 million videos for downstream benchmarks. Several off-the-shelf techniques have been used to ensure the high-quality of training videos."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"+ The novelty of the new model mPLUG-video is limited. The proposed three modules, and partially efficient tuning are all well studied techniques in this area.\n\n+ The improvements brought by the proposed mPLUG-video are limited.\n\n+ One of the key contributions in this paper is the proposed new dataset. It would be better to demonstrate the high quality of the newly collected data. Based on the example shown in Figure 9, the text annotations look very noisy."},"confidence":{"value":"5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully."},"questions":{"value":"In the first paragraph of the introduction section, the authors argue that existing methods of translating English to Chinese suffer intrinsic linguistic and cultural gaps. Could you give more explicit examples to show the harmfulness of these methods?"},"rating":{"value":"5: marginally below the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636455398,"tcdate":1698839190544,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission4734/Reviewer_bL3j"],"signatures":["ICLR.cc/2024/Conference/Submission4734/Reviewer_bL3j"],"forum":"mzxKLZNbrQ","number":2,"license":"CC BY 4.0","cdate":1698839190544,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission4734/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636455398,"domain":"ICLR.cc/2024/Conference","replyto":"mzxKLZNbrQ","id":"g1ZNJaArIX","forumContent":{"venue":{"value":"Submitted to ICLR 2024"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Chinese","Video-language pre-training","video-language benchmarks","mPLUG","Youku","video captioning","video classification"]},"primary_area":{"value":"datasets and benchmarks"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"We firstly release the largest public Chinese high-quality video-language dataset named Youku-mPLUG, which is collected from Youku, a well-known Chinese video-sharing website, with strict criteria of safety, diversity, quality, and copyright. Youku-mPLUG contains 10 million Chinese video-text pairs filtered from 400 million raw videos across a wide range of 45 diverse categories for large-scale pre-training. In addition, to facilitate a comprehensive evaluation of video-language models, we carefully build the largest human-annotated Chinese benchmarks covering three popular video-language tasks across cross-modal retrieval, video captioning, and video category classification. \nWe also provide comprehensive benchmark evaluations of models across different architectures including encoder-only (i.e., ALPRO), encoder-decoder (i.e., mPLUG-2), and decoder-only (i.e., mPLUG-Video) for comparison. Especially, we train the first Chinese Multimodal LLM with only 1.7% trainable parameters for video understanding. Experiments show that models pre-trained on Youku-mPLUG gain up to 23.1% improvement in video category classification. Besides, mPLUG-video achieves a new state-of-the-art result on these benchmarks with 80.5% top-1 accuracy in video category classification and 68.9 CIDEr score in video captioning, respectively. Finally, the 2.7B version of mPLUG-video demonstrates impressive instruction and video understanding ability. The zero-shot instruction understanding experiment indicates that pretraining with Youku-mPLUG can enhance the ability to comprehend overall and detailed visual semantics, recognize scene text, and leverage open-domain knowledge."},"_bibtex":{"value":"@misc{\nxu2024youkumplug,\ntitle={Youku-m{PLUG}: A 10 Million Large-scale Chinese Video-Language Dataset for Pre-training and Benchmarks},\nauthor={Haiyang Xu and Qinghao Ye and Xuan Wu and Ming Yan and Yuan Miao and Jiabo Ye and Anwen Hu and Guohai Xu and Yaya Shi and Guangwei Xu and Chenliang Li and Qi Qian and Maofei Que and Ji Zhang and Xiao Zeng and Fei Huang},\nyear={2024},\nurl={https://openreview.net/forum?id=mzxKLZNbrQ}\n}"},"title":{"value":"Youku-mPLUG: A 10 Million Large-scale Chinese Video-Language Dataset for Pre-training and Benchmarks"},"pdf":{"value":"/pdf/614c835115f1443393e0caec779fcaaf14e25811.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Rejected_Submission"},"paperhash":{"value":"xu|youkumplug_a_10_million_largescale_chinese_videolanguage_dataset_for_pretraining_and_benchmarks"},"authorids":{"value":["~Haiyang_Xu1","~Qinghao_Ye1","~Xuan_Wu5","~Ming_Yan2","~Yuan_Miao1","~Jiabo_Ye1","~Anwen_Hu1","~Guohai_Xu1","~Yaya_Shi1","~Guangwei_Xu2","~Chenliang_Li2","~Qi_Qian1","~Maofei_Que1","~Ji_Zhang3","~Xiao_Zeng4","~Fei_Huang1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Haiyang Xu","Qinghao Ye","Xuan Wu","Ming Yan","Yuan Miao","Jiabo Ye","Anwen Hu","Guohai Xu","Yaya Shi","Guangwei Xu","Chenliang Li","Qi Qian","Maofei Que","Ji Zhang","Xiao Zeng","Fei Huang"]}},"version":2},{"content":{"summary":{"value":"This submission addresses the challenge of explainability in graph neural networks (GNNs). The authors introduce a new method called Shortcut-guided Graph Rationalization (SGR), designed to provide rationale by identifying significant nodes or edges in a graph. The SGR method has two stages: 1) it employs a shortcut guide trained with an early stop strategy to capture shortcut information; 2) it separates the graph into rationale and non-rationale parts and uses mutual information (MI) minimization and maximization to refine the rationale subgraph representations. The method is tested on both synthetic and real-world datasets. The experimental results indicate the effectiveness of the proposed method. The core idea is to allow GNNs to learn from these shortcuts to improve their prediction accuracy and reliability, especially under shifts in the data environment."},"presentation":{"value":"3 good"},"contribution":{"value":"2 fair"},"soundness":{"value":"3 good"},"strengths":{"value":"1. Based on some observations from previous work, the proposed Shortcut-guided Graph Rationalization (SGR) method is new as it identifies significant nodes or edges for rationale.\n\n2. SGR demonstrates robustness across various datasets, including synthetic and real-world scenarios, by filtering out misleading correlations. Its ability to refine rationale subgraph representations through mutual information optimization is beneficial under shifts in the data environment.\n\n3. The paper presents extensive experimental results to validate the effectiveness of the SGR method."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The proposed method assumes that shortcut representations are available, which may not always be true in real-world applications. This assumption could limit the generalizability of the approach to scenarios where such information is not readily identifiable. Furthermore, it is unclear how to construct these shortcuts effectively.\n\n2. The early stop strategy is the key to success for shortcut representations. In Appendix B, the authors provided some simple tricks, but the validation is unclear in a general setting.\n\n3. The shortcuts are challenging to obtain in a general setting since there is a lack of clear definition. The presence and nature of shortcuts can be highly dependent on the specifics of the dataset. In some cases, what constitutes a shortcut might be subtle or complex, making it hard to pinpoint without in-depth analysis. \n\n4. The authors assume the shortcut representation is available. The shortcut feature is easy to learn. These assumptions may be too strong for some applications."},"confidence":{"value":"4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work."},"questions":{"value":"In addition to the above weaknesses, I have the following questions:\n\n1. I may be wrong, but I was wondering whether the superior performance of the proposed method is due to extra training shortcuts datasets or other factors. If the shortcuts training data contributes more, then it is a little bit unfair to compare the proposed method with non-shortcut methods. The reason I am asking is that it seems the key challenge of the problem is to identify effective shortcuts in a general setting. Once the shortcuts are identified, you can have some advantages if you use them for training the GNN model. Is it the logic?\n\n2. Identifying shortcuts often requires domain knowledge or not."},"rating":{"value":"5: marginally below the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636812054,"tcdate":1698806323383,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission6956/Reviewer_gXJ9"],"signatures":["ICLR.cc/2024/Conference/Submission6956/Reviewer_gXJ9"],"forum":"XcwHDoKvVg","number":3,"license":"CC BY 4.0","cdate":1698806323383,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission6956/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636812054,"domain":"ICLR.cc/2024/Conference","replyto":"XcwHDoKvVg","id":"RQflNk7lz1","forumContent":{"venue":{"value":"ICLR 2024 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["graph rationalization","shortcut learning"]},"primary_area":{"value":"representation learning for computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"The remarkable success in graph neural networks (GNNs) promotes the Graph Rationalization methods that aim to provide explanations to support the prediction results by identifying a small subset of the original graph (i.e., rationale). Although existing methods have achieved promising results, recent studies have proved that these methods still suffer from  exploiting shortcuts in the data to yield task results and compose rationales. Different from previous methods plagued by shortcuts, in this paper, we propose a Shortcut-guided Graph Rationalization (SGR) method, which identifies rationales by learning from shortcuts. Specifically, SGR consists of two training stages. In the first stage, we train a shortcut guider with an early stop strategy to obtain shortcut information. During the second stage, SGR separates the graph into the rationale and non-rationale subgraphs and lets them learn from the shortcut information generated by the frozen shortcut guider to identify which information belongs to shortcuts and which does not. Finally, we employ the non-rationale subgraphs as environments and identify the invariant rationales which filter out the shortcuts under environment shifts. Extensive experimental results on both synthetic and real-world datasets clearly validate the effectiveness of our proposed method."},"_bibtex":{"value":"@misc{\nyue2024learning,\ntitle={Learning from Shortcut: A Shortcut-guided Approach for Graph Rationalization},\nauthor={Linan Yue and Qi Liu and Ye Liu and Weibo Gao and Chao Song},\nyear={2024},\nurl={https://openreview.net/forum?id=XcwHDoKvVg}\n}"},"title":{"value":"Learning from Shortcut: A Shortcut-guided Approach for Graph Rationalization"},"pdf":{"value":"/pdf/40f49261208b6613b6cba47b97c7d708165772cf.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Withdrawn_Submission"},"paperhash":{"value":"yue|learning_from_shortcut_a_shortcutguided_approach_for_graph_rationalization"},"authorids":{"value":["~Linan_Yue1","~Qi_Liu3","~Ye_Liu10","~Weibo_Gao1","~Chao_Song2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Linan Yue","Qi Liu","Ye Liu","Weibo Gao","Chao Song"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a training-free method for light trajectory editing in videos. The approach extends the single-image relighting model, IC-Light, to the video domain by incorporating priors from video diffusion models to ensure temporal consistency. The core technical contributions include a light map injection module, which introduces trajectory-aware noise into the VDM's latent space, and a geometry-aware relighting module. This second module processes both RGB frames and corresponding normal maps estimated via StableNormal to guide the relighting process. The results are visually compelling and demonstrate good adherence to the specified lighting trajectories."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. In Eq. (3), what exact $\\omega$ values are used across videos and trajectories? Is it fixed or tuned per video?\n2. Could you elaborate on the temporal stability of the StableNormal predictions? Are the normal maps computed independently per frame and fed directly into the video VAE, or is some form of temporal smoothing or other consistency enforcement applied?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":3},"strengths":{"value":"1. The paper tackles a relatively unexplored task of lighting control in video diffusion models rather than just global style. The direction, ideas, and results are promising and the problem is quite interesting.\n2. By building on existing, well-known T2I and T2V models, the method is accessible and its components are easier to understand and potentially replicate.\n3. The strategy of injecting trajectory-aware noise into the VDM's initial latent space appears effective for guiding the lighting. This concept may be adaptable to other conditional video generation tasks.\n4. The design of the geometry-aware relighting module, which dynamically blends RGB and normal map information, is technically sound. It correctly reflects that surface geometry should remain invariant during a relighting task."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The use of PSNR_y against a pure white reference is not convincing. Relighting is a complex task involving light, geometry, and material, and does not necessarily mean \"the brighter the better\". This metric mainly captures pixel-wise brightness in 2D space and fails to account for the geometric and directional correctness of the illumination. It cannot distinguish excessive lighting and ignores geometry and material reflectance.\n2. The evaluation is somewhat limited. The dataset contains only 50 videos. This small scale and likely limited diversity are insufficient to robustly validate the method's generalizability. The paper would significantly benefit from showcasing more results across a wider variety of video content and lighting conditions.\n\nMinor issues:\n\nThere are plenty of papers about generative portrait relighting. Though they deal with portrait videos, I think they are still related to this topic. It is better to cite and discuss these papers. Following are some examples:\n\n+ Lumos: Learning to Relight Portrait Images via a Virtual Light Stage and Synthetic-to-Real Adaptation\n\n+ Neural Video Portrait Relighting in Real-time via Consistency Modeling\n\n+ Real-time 3D-aware Portrait Video Relighting"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762923010359,"tcdate":1761914909125,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission12029/Reviewer_fsKh"],"signatures":["ICLR.cc/2026/Conference/Submission12029/Reviewer_fsKh"],"forum":"5ft8vd9rwc","number":1,"license":"CC BY 4.0","cdate":1761914909125,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission12029/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762923010359,"domain":"ICLR.cc/2026/Conference","replyto":"5ft8vd9rwc","id":"tz9qphozCa","forumContent":{"TLDR":{"value":"LightCtrl relight a given video via a user-specific light trajectory in a zero-shot manner."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["video relighting; controllable video editing"]},"supplementary_material":{"value":"/attachment/9620cf406218d8104168386c0036f3d45fbf2f7a.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Recent diffusion models have achieved remarkable success in image relighting, and this success has quickly been reproduced in video relighting. Although these methods can relight videos under various conditions, their ability to explicitly control the illumination in the relighted video remains limited. Therefore, we present \\name, the first controllable video relighting method that offers explicit control over the video illumination through a user-supplied light trajectory in a training-free manner. This is essentially achieved by leveraging a hybrid approach that combines pre-trained diffusion models: a pre-trained image relighting diffusion model is used to relight each frame individually, followed by a video diffusion prior that enhances the temporal consistency of the relighted sequence. In particular, to enable explicit control over dynamically varying lighting in the relighted video, we introduce two key components. \nFirst, the Light Map Injection module samples light trajectory-specific noise and injects it into the latent representation of the source video, significantly enhancing illumination coherence with respect to the conditional light trajectory. \nSecond, the Geometry-Aware Relighting module dynamically combines RGB and normal map latents in the frequency domain to suppress the influence of the original lighting in the input video, thereby further improving the relighted video's adherence to the input light trajectory. \nOur experiments demonstrate that \\name can generate high-quality video results with diverse illumination changes closely following the light trajectory condition, indicating improved controllability over baseline methods. The code will be released at: https://github.com/GVCLab/LightCtrl."},"_bibtex":{"value":"@inproceedings{\npeng2026lightctrl,\ntitle={LightCtrl: Training-free Controllable Video Relighting},\nauthor={Yizuo Peng and Xuelin Chen and Kai Zhang and Xiaodong Cun},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=5ft8vd9rwc}\n}"},"title":{"value":"LightCtrl: Training-free Controllable Video Relighting"},"pdf":{"value":"/pdf/ebecc07688f72c477de1cca270c7a0498aa35be7.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"peng|lightctrl_trainingfree_controllable_video_relighting"},"authorids":{"value":["~Yizuo_Peng1","~Xuelin_Chen1","~Kai_Zhang16","~Xiaodong_Cun1"]},"authors":{"value":["Yizuo Peng","Xuelin Chen","Kai Zhang","Xiaodong Cun"]}},"version":2},{"content":{"venue":{"value":"Video-Langauge Models Oral"},"pdf":{"value":"/pdf/12d3803519276abafedf4201a5d5b8e2cc39df1e.pdf"},"keywords":{"value":["text-to-video generation","multi-scene","time-aligned captions"]},"venueid":{"value":"NeurIPS.cc/2024/Workshop/Video-Langauge_Models"},"paperhash":{"value":"bansal|talc_timealigned_captions_for_multiscene_texttovideo_generation"},"authorids":{"value":["~Hritik_Bansal2","~Yonatan_Bitton1","~Michal_Yarom1","~Idan_Szpektor1","~Aditya_Grover1","~Kai-Wei_Chang1"]},"abstract":{"value":"Most of these text-to-video (T2V) generative models often produce single-scene video clips that depict an entity performing a particular action (e.g., 'a red panda climbing a tree'). However, it is pertinent to generate multi-scene videos since they are ubiquitous in the real-world (e.g., 'a red panda climbing a tree' followed by 'the red panda sleeps on the top of the tree'). To generate multi-scene videos from the pretrained T2V model, we introduce a simple and effective Time-Aligned Captions (TALC) framework. Specifically, we enhance the text-conditioning mechanism in the T2V architecture to recognize the temporal alignment between the video scenes and scene descriptions. For instance, we condition the visual features of the earlier and later scenes of the generated video with the representations of the first scene description (e.g., 'a red panda climbing a tree') and second scene description (e.g., 'the red panda sleeps on the top of the tree'), respectively. As a result, we show that the T2V model can generate multi-scene videos that adhere to the multi-scene text descriptions and be visually consistent (e.g., entity and background). Further, we finetune the pretrained T2V model with multi-scene video-text data using the TALC framework. We show that the TALC-finetuned model outperforms the baseline by achieving a relative gain of 29% in the overall score, which averages visual consistency and text adherence using human evaluation."},"_bibtex":{"value":"@inproceedings{\nbansal2025talc,\ntitle={{TALC}: Time-Aligned Captions for Multi-Scene Text-to-Video Generation},\nauthor={Hritik Bansal and Yonatan Bitton and Michal Yarom and Idan Szpektor and Aditya Grover and Kai-Wei Chang},\nbooktitle={Workshop on Video-Language Models @ NeurIPS 2024},\nyear={2025},\nurl={https://openreview.net/forum?id=HDnb9nhN45}\n}"},"title":{"value":"TALC: Time-Aligned Captions for Multi-Scene Text-to-Video Generation"},"track":{"value":"Long Paper Track (up to 9 pages)"},"authors":{"value":["Hritik Bansal","Yonatan Bitton","Michal Yarom","Idan Szpektor","Aditya Grover","Kai-Wei Chang"]}},"tmdate":1736861080026,"pdate":1730081752047,"tcdate":1725588274699,"writers":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission12/Authors"],"signatures":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission12/Authors"],"forum":"HDnb9nhN45","license":"CC BY 4.0","number":12,"cdate":1725588274699,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Post_Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Edit","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission12/-/Camera-Ready_Revision"],"mdate":1736861080026,"odate":1736861080010,"domain":"NeurIPS.cc/2024/Workshop/Video-Langauge_Models","id":"HDnb9nhN45","version":2},{"content":{"TLDR":{"value":"We train agent monitors on minimal counterfactual pairs, where the malicious twin differs from a real benign trajectory only by a bounded local edit, raising a 4B monitor's AUROC from 0.56 to 0.76 on an out-of-domain 968-trajectory benchmark."},"venue":{"value":"ICLR 2027 Conference Submission"},"keywords":{"value":["AI safety","LLM agents","agent monitoring","misbehavior detection","counterfactual data generation"]},"primary_area":{"value":"alignment, fairness, safety, privacy, and societal considerations"},"abstract":{"value":"Agent monitors support safer deployment by detecting malicious actions within legitimate work, sometimes before task completion. Yet separately collected benign and malicious runs can differ throughout execution, making it difficult to construct matched contrasts for training and evaluation. We propose a framework for training agent monitors with controlled minimal counterfactual pairs. We construct each pair by producing a malicious twin from a real benign trajectory through bounded local edits. An adversarial editor inserts actions for a malicious side task, while mechanical checks and behavioral judges assess structural validity and coherence. The construction also supports feedback-guided collection, using monitor responses to target subsequent edits. The pairs support direct monitor training and controlled trajectory-level, paired, and prefix-level evaluation. On CUA-SHADE-Arena, an external benchmark of computer-use agent trajectories, direct training on our pairs raises Qwen3.5-35B-A3B AUROC from 0.665 prompted to 0.866, above 0.753 with STRIDE and Gloom training alone. In a separate four-round collection study, feedback-guided generation reaches 0.881 AUROC versus 0.770 for random collection at the same training-set size. Our trained monitors achieve Pareto-optimal cost–detection trade-offs among the evaluated configurations, with Qwen3.5-397B-A17B reaching detection performance comparable to Gemini 3.1 Pro at about 3.65-fold lower modeled inference cost. This framework connects failure diagnosis, data construction, and monitor training through controlled counterfactual pairs, providing a practical way to improve detection with agent monitors."},"_bibtex":{"value":"@inproceedings{\nanonymous2026training,\ntitle={Training Agent Monitors with Minimal Counterfactual Pairs},\nauthor={Anonymous},\nbooktitle={Submitted to The Fifteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=Ts9bZceJ5W},\nnote={under review}\n}"},"title":{"value":"Training Agent Monitors with Minimal Counterfactual Pairs"},"pdf":{"value":"/pdf/56f42b842ef50bf0bf4998a990098c3aa7e8d890.pdf"},"venueid":{"value":"ICLR.cc/2027/Conference/Submission"}},"tmdate":1791233761302,"tcdate":1789749029378,"writers":["ICLR.cc/2027/Conference","ICLR.cc/2027/Conference/Submission45912/Authors"],"signatures":["ICLR.cc/2027/Conference/Submission45912/Authors"],"forum":"Ts9bZceJ5W","license":"CC BY 4.0","number":45912,"cdate":1789749029378,"readers":["everyone"],"invitations":["ICLR.cc/2027/Conference/-/Submission","ICLR.cc/2027/Conference/-/Edit","ICLR.cc/2027/Conference/-/Submission_Change_Before_Bidding","ICLR.cc/2027/Conference/Submission45912/-/Full_Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Reviewing"],"mdate":1791233761302,"odate":1791031328286,"domain":"ICLR.cc/2027/Conference","id":"Ts9bZceJ5W","version":2},{"content":{"venue":{"value":"Video-Langauge Models Poster"},"pdf":{"value":"/pdf/ed68f286a914a46d56c18409650fcc0232486566.pdf"},"keywords":{"value":["Large Language Vision Model","ADL"]},"venueid":{"value":"NeurIPS.cc/2024/Workshop/Video-Langauge_Models"},"paperhash":{"value":"chakraborty|llavidal_benchmarking_large_language_vision_models_for_daily_activities_of_living"},"authorids":{"value":["~Rajatsubhra_Chakraborty1","~Arkaprava_Sinha1","~Dominick_Reilly1","~Manish_Kumar_Govind3","~Pu_Wang1","~Francois_Bremond1","~Srijan_Das1"]},"abstract":{"value":"With the increasing pervasiveness of video content throughout society, the demand for robust video-language models is increasingly urgent. In this work we introduce LLAVIDAL, a Large Language Vision Model tailored for Activities of Daily Living (ADL). Unlike existing models primarily trained on curated web videos, LLAVIDAL leverages a novel multiview RGB-D dataset, ADL-X, which includes 100K untrimmed video-instruction pairs, enriched with 3D skeletons and object trajectories to mimic real-world complexities. The model integrates these features to effectively understand intricate human behaviors and spatiotemporal dynamics typical of daily activities. We also introduce ADLMCQ, a new benchmark designed to evaluate the proficiency of video-language models in interpreting ADL content. Our evaluations demonstrate that LLAVIDAL significantly outperforms existing models, showcasing superior ability to process and reason about real-life video scenarios. The insights gained underscore the necessity for advanced processing techniques to handle the scale and multimodality of video data, alongside a need for comprehensive benchmarks that reflect real-world use cases more accurately. The instruction tuning data is available at https://adl-x.github.io"},"_bibtex":{"value":"@inproceedings{\nchakraborty2025llavidal,\ntitle={{LLAVIDAL}: Benchmarking Large Language Vision Models for Daily Activities of Living},\nauthor={Rajatsubhra Chakraborty and Arkaprava Sinha and Dominick Reilly and Manish Kumar Govind and Pu Wang and Francois Bremond and Srijan Das},\nbooktitle={Workshop on Video-Language Models @ NeurIPS 2024},\nyear={2025},\nurl={https://openreview.net/forum?id=llegT88ct7}\n}"},"title":{"value":"LLAVIDAL: Benchmarking Large Language Vision Models for Daily Activities of Living"},"track":{"value":"Short Paper Track (up to 3 pages)"},"authors":{"value":["Rajatsubhra Chakraborty","Arkaprava Sinha","Dominick Reilly","Manish Kumar Govind","Pu Wang","Francois Bremond","Srijan Das"]}},"tmdate":1736861080713,"pdate":1730081753130,"tcdate":1726003852123,"writers":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission39/Authors"],"signatures":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission39/Authors"],"forum":"llegT88ct7","license":"CC BY 4.0","number":39,"cdate":1726003852123,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Post_Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission39/-/Full_Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Edit","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission39/-/Camera-Ready_Revision"],"mdate":1736861080713,"odate":1736861080694,"domain":"NeurIPS.cc/2024/Workshop/Video-Langauge_Models","id":"llegT88ct7","version":2},{"content":{"summary":{"value":"This paper introduces Astra, a framework for interactive world modeling that generates long and temporal coherence video sequences across diverse scenarios. Astra enhances a pre-trained video model with an light-weight action-aware adapter for precise action conditioning, a noise-augmented history memory during training to ensure long-term consistency, and a mixture of action experts to effectively handle diverse action inputs."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"1. In Section 3.3, the explanation of the types of random noise and blur used is unclear. Please provide a more detailed description.\n2. It was not properly analyzed the lightweight action-aware adapter model complexity. Please provide a more detailed description and comparison.\n3. In Figure 6, compared to YUME, the results from Astra appear to exhibit some color shift. Could the authors explain the cause of this phenomenon?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. This paper uses a lightweight action-aware adapter for precise action conditioning.\n2. Astra achieves good responsiveness and is able to generate long, temporally coherent video sequences by employing a noise-as-mask strategy during training.\n3. Astra employs a mixture of action experts to effectively adapt to diverse scenarios and handle various types of action inputs."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Mixture of action experts idea is similar to [1, 2] and action-aware adapter is similar to [3, 4]. Please provide a conceptual comparison with these reference.\n2. The paper does not thoroughly analyze the underlying reasons why the noise-as-mask strategy enables the generation of long, temporally coherent video sequences.\n3. The paper does not explain why the router network performs so well across diverse scenarios and with various types of action inputs.\n\n[1] Mixture of Action Expert Embeddings: Multi-Task ACT\n\n[2] DriveMoE: Mixture-of-Experts for Vision-Language-Action Model in End-to-End Autonomous Driving\n\n[3] Long-Context Autoregressive Video Modeling with Next-Frame Prediction\n\n[4] Epona: Autoregressive Diffusion World Model for Autonomous Driving"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762921032086,"tcdate":1762404118366,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission9431/Reviewer_AW9Q"],"signatures":["ICLR.cc/2026/Conference/Submission9431/Reviewer_AW9Q"],"forum":"8UZpmrxoLG","number":3,"license":"CC BY 4.0","cdate":1762404118366,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission9431/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762921032086,"domain":"ICLR.cc/2026/Conference","replyto":"8UZpmrxoLG","id":"5hwiiiZd76","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["world model","video generation"]},"supplementary_material":{"value":"/attachment/ea12713313b6fe2f964a565e8846e0c7015a7a2e.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Recent advances in diffusion transformers have empowered video generation models to generate high-quality video clips from texts or images. However, world models with the ability to predict long-horizon futures from past observations and actions remain underexplored, especially for general-purpose scenarios and various forms of actions. To bridge this gap, we introduce Astra, an interactive general world model that generates real-world futures for diverse scenarios (e.g., autonomous driving, robot grasping) with precise action interactions (e.g., camera motion, robot action). We propose an autoregressive denoising architecture and use temporal causal attention to aggregate past observations and support streaming outputs. We use a noise-augmented history memory to avoid over-reliance on past frames to balance responsiveness with temporal coherence. For precise action control, we introduce an action-aware adapter that directly injects action signals into the denoising process. We further develop a mixture of action experts that dynamically route heterogeneous action modalities, enhancing versatility across diverse real-world tasks such as exploration, manipulation, and camera control. Astra achieves interactive, consistent, and general long-term video prediction and supports various forms of interactions. Experiments across multiple datasets demonstrate the improvements of Astra in fidelity, long-range prediction, and action alignment over existing state-of-the-art world models."},"_bibtex":{"value":"@inproceedings{\nzhu2026astra,\ntitle={Astra: General Interactive World Model with Autoregressive Denoising},\nauthor={Yixuan Zhu and Feng Jiaqi and Wenzhao Zheng and Yuan Gao and Xin Tao and Pengfei Wan and Jiwen Lu and Jie Zhou},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=8UZpmrxoLG}\n}"},"title":{"value":"Astra: General Interactive World Model with Autoregressive Denoising"},"pdf":{"value":"/pdf/9a15a13d6ccbd962dfcd42fc2ce3419ba1550a4f.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"zhu|astra_general_interactive_world_model_with_autoregressive_denoising"},"authorids":{"value":["~Yixuan_Zhu1","~Feng_Jiaqi1","~Wenzhao_Zheng1","~Yuan_Gao32","~Xin_Tao3","~Pengfei_Wan1","~Jiwen_Lu1","~Jie_Zhou3"]},"authors":{"value":["Yixuan Zhu","Feng Jiaqi","Wenzhao Zheng","Yuan Gao","Xin Tao","Pengfei Wan","Jiwen Lu","Jie Zhou"]}},"version":2},{"content":{"summary":{"value":"The paper proposes a two-stage latent diffusion model, which contains keyframe generation and frame interpolation, for text-to-video generation tasks. Starting from a pretrained text-to-image model, separate temporal blocks are used to model the temporal information. A video decoder based on MoVQGAN is also adopted to improve the generation quality."},"presentation":{"value":"3 good"},"contribution":{"value":"2 fair"},"soundness":{"value":"2 fair"},"strengths":{"value":"1. The topic of this paper is significant.\n2. The paper is overall clear and well-formulated.\n3. The training of the proposed model seems resource-efficient compared to most text-to-video methods, as only 8-16 A100 GPUs and 120k data pairs were adopted in training."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. There needs to be more explanations and analysis about the proposed methods and results. The authors should illustrate the possible reason why separate temporal blocks perform better than temporal layers.\n2. As the limitations of quantitative metrics, more visual results should be provided and compared with other text-to-video methods.\n3. The paper doesn't mention the number of parameters of the video generation model, and only 120k internal data pairs are utilized for training. The model maybe overfit to the training data. As the data domain is also agnostic, it's hard to decide whether the experiment results are evidential enough or not."},"confidence":{"value":"4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work."},"questions":{"value":"1. Could the authors give more explanations about the separate temporal blocks over temporal layers?\n2. As mentioned in the introduction part, the scarcity of open-source text-video datasets impedes the development of video generation. Is it possible to contribute the internal training data to the community?"},"rating":{"value":"5: marginally below the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699637152640,"tcdate":1698736400062,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission9160/Reviewer_qUnt"],"signatures":["ICLR.cc/2024/Conference/Submission9160/Reviewer_qUnt"],"forum":"JBLgjRRuHG","number":3,"license":"CC BY 4.0","cdate":1698736400062,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission9160/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699637152640,"domain":"ICLR.cc/2024/Conference","replyto":"JBLgjRRuHG","id":"16wokFoLtX","forumContent":{"venue":{"value":"ICLR 2024 Conference Withdrawn Submission"},"TLDR":{"value":"In this paper we propose a two-stage latent diffusion video generation architecture and a new MoVQ video decoding scheme. We also conduct experimental research to compare temporal blocks and temporal layers for architecture design."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["text-to-video","video generation","temporal consistency","frames interpolation","Inception Score","CLIPSIM","MoVQ video decoder"]},"supplementary_material":{"value":"/attachment/85e824fa7cd2f57c7ab32187b44fd2b90363c8a2.zip"},"primary_area":{"value":"generative models"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Multimedia generation approaches occupy a prominent place in artificial intelligence research. Text-to-image models achieved high quality results over the last years, however video synthesis methods recently started to develop. In this paper we present a new two-stage latent diffusion video generation architecture using a new MoVQ video decoding scheme. The first stage concerns keyframes synthesis, while the second one is devoted to interpolated frames generation. We compare two temporal conditioning approaches during evaluation and show the improvement of using temporal blocks over temporal layers in terms of IS and CLIPSIM metrics reflecting video generation quality aspects. We also evaluate different configurations of MoVQ-based video decoding scheme to achieve higher PSNR, SSIM, MSE and LPIPS scores. Finally, we compare our pipeline with existing solutions and achieve top-3 CLIPSIM metric score (0.2976)."},"_bibtex":{"value":"@misc{\nvladimir2024efficient,\ntitle={Efficient architectural aspects for text-to-video generation pipeline},\nauthor={Arkhipkin Sergeevich Vladimir and Zein Shaheen and Viacheslav Vasilev and Denis Valerievich Dimitrov and Andrey Kuznetsov},\nyear={2024},\nurl={https://openreview.net/forum?id=JBLgjRRuHG}\n}"},"title":{"value":"Efficient architectural aspects for text-to-video generation pipeline"},"pdf":{"value":"/pdf/de7dea6dc2c02c9d3dc83b1f407bdf5087d624b7.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Withdrawn_Submission"},"paperhash":{"value":"vladimir|efficient_architectural_aspects_for_texttovideo_generation_pipeline"},"authorids":{"value":["~Arkhipkin_Sergeevich_Vladimir1","~Zein_Shaheen1","~Viacheslav_Vasilev1","~Denis_Valerievich_Dimitrov1","~Andrey_Kuznetsov2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Arkhipkin Sergeevich Vladimir","Zein Shaheen","Viacheslav Vasilev","Denis Valerievich Dimitrov","Andrey Kuznetsov"]}},"version":2},{"content":{"summary":{"value":"1. The paper introduces PanoWorld-X, a framework for generating controllable 360-degree panoramic videos with diverse exploration paths. 2. The paper proposes the PanoExplorer dataset, a large-scale synthetic dataset of panoramic video-trajectory pairs.\n3. The paper proposes sphere-aware DiT block. This block consists of explorable-aware attention to handle trajectory control and a sphere-aware attention to improve spatiotemporal consistency.\n4. The experimental results demonstrate that PanoWorld-X achieves superior performance compared to previous panoramic video generation and trajectory-based methods."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. Given the final generated resolution is 480×960, what is the native resolution of the PanoExplorer dataset? It connects with the utility of the dataset for future high-resolution panoramic research.\n2. How capable is the model of inferring an exploration path purely based on text instructions?\n3. The current evaluation focuses heavily on static, synthetic scenes. How does PanoWorld-X handle out-of-distribution real-world panoramic images even with moving objects?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"1. The paper introduces PanoWorld-X, which can generate high-fidelity, explorable panoramic videos controlled by trajectories.\n2. The proposed Sphere-Aware DiT is the main technical contribution. It is well-motivated by the need to address the geometric misalignment between spherical panorama data and standard diffusion models, explicitly capturing spherical connectivity.\n3. The PanoExplorer dataset is a valuable contribution to the field, which contains 110k panoramic videos paired with trajectory annotations from Unreal Engine."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The architectural contribution of the trajectory control mechanism(ControlNet) appears limited and there are many previous works about it in the field of general video generation. Another novelty is sphere-aware attention, where the similar spherical representation usage is discussed in previous panoramic image generation works likes PanFusion and Text2Light.\n2. The explanation of the Exploration Route Representation in Section 3.3 lacks clarity. The paper introduces a 6-DoF signal $Er_i$​ but then discusses converting it into extrinsic matrices and using Plücker embeddings. The precise relationship between this initial 6-DoF vector and the final $R^{F×H×W×6}$ embedding tensor depicted in Figure 2 is ambiguous and needs to be clarified.\n3. The practical utility of the framework is severely constrained by its low output resolution of 480×960. For panoramic content, this resolution is insufficient for reality usage, especially given the base model supports higher resolutions. This issue is acknowledged in the supplementary materials, where perspective crops are described as being \"less than 50 pixels\".\n4. While quantitative metrics on the evaluation of trajectory adherence are provided, the paper do not offer side-by-side visualization of the desired camera path overlaid on the generated video to demonstrate the model's trajectory-conditioned effect on the generated video.\n\nI will consider adjusting my score if my concerns are well-resolved."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762922916198,"tcdate":1761998885830,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission11906/Reviewer_8Add"],"signatures":["ICLR.cc/2026/Conference/Submission11906/Reviewer_8Add"],"forum":"iZyBEbq6jR","number":3,"license":"CC BY 4.0","cdate":1761998885830,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission11906/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762922916198,"domain":"ICLR.cc/2026/Conference","replyto":"iZyBEbq6jR","id":"0pPFuXGLD5","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["immersive content"]},"supplementary_material":{"value":"/attachment/5773686dc6b17b2632d3dc874878b8b3fbb061b7.zip"},"primary_area":{"value":"other topics in machine learning (i.e., none of the above)"},"abstract":{"value":"Generating a complete and explorable 360-degree visual world enables a wide range of downstream applications. While prior works have advanced the field, they remain constrained by either narrow field-of-view limitations, which hinder the synthesis of continuous and holistic scenes, or insufficient controllability that restricts free exploration by users or autonomous agents. To address this, we propose PanoWorld-X, a novel framework for high-fidelity and controllable panoramic video generation with diverse camera trajectories. \n First, we propose a novel pipeline for synthesizing panoramic video-trajectory dataset pairs in virtual 3D environments via Unreal Engine. This pipeline consists of four main steps and enables the collection of a large-scale dataset with rich scene diversity and accurate trajectory annotations.\n To achieve precise panoramic video generation, we identify that the bottleneck arises from the misalignment between the spherical geometry of panoramic data and the inductive priors of conventional video diffusion models. To address this, we leverage the spherical connectivity characteristics of panorama data, and propose a Sphere-Aware Diffusion Transformer that reprojects equirectangular features onto the spherical surface, thereby capturing geometric adjacency in the latent space. This design significantly improves both visual fidelity and spatiotemporal continuity.\n   Extensive experiments demonstrate that our PanoWorld-X achieves superior performance in various aspects, including motion range, control precision, and visual quality, underscoring its potential for real-world applications."},"_bibtex":{"value":"@misc{\nyin2026panoworldx,\ntitle={PanoWorld-X: Generating Explorable Panoramic Worlds via Sphere-Aware Video Diffusion},\nauthor={Yuyang Yin and Hao-Xiang Guo and Fangfu Liu and Mengyu Wang and Hanwen Liang and Eric Li and Yikai Wang and Xiaojie Jin and Yao Zhao and Yunchao Wei},\nyear={2026},\nurl={https://openreview.net/forum?id=iZyBEbq6jR}\n}"},"title":{"value":"PanoWorld-X: Generating Explorable Panoramic Worlds via Sphere-Aware Video Diffusion"},"pdf":{"value":"/pdf/8d2464ff9acb82ee4bdbbfda2c8b3031ed0755c7.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"yin|panoworldx_generating_explorable_panoramic_worlds_via_sphereaware_video_diffusion"},"authorids":{"value":["~Yuyang_Yin1","~Hao-Xiang_Guo1","~Fangfu_Liu2","~Mengyu_Wang5","~Hanwen_Liang2","~Eric_Li8","~Yikai_Wang2","~Xiaojie_Jin1","~Yao_Zhao1","~Yunchao_Wei1"]},"authors":{"value":["Yuyang Yin","Hao-Xiang Guo","Fangfu Liu","Mengyu Wang","Hanwen Liang","Eric Li","Yikai Wang","Xiaojie Jin","Yao Zhao","Yunchao Wei"]}},"version":2},{"content":{"summary":{"value":"This paper aims to address the challenge of reasoning shortcuts in VLMs, where models often rely on textual priors instead of visual evidence. To overcome this, the authors propose Visionary R1, a visual reasoning–aware training framework that integrates GRPO with visual chain-of-thought supervision. The method encourages models to reason explicitly about visual cues and penalizes text-only heuristics, producing a more grounded reasoning process. Experiments across benchmarks demonstrate that Visionary R1 substantially reduces shortcut reliance and improves multimodal reasoning accuracy, outperforming both supervised fine-tuning and conventional RLHF approaches. The results highlight the effectiveness of R1-style visual reasoning training for mitigating spurious shortcut behavior in VLMs."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"Please refer to the weaknesses."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1)\tVisionary-R1 introduces a conceptually simple but powerful “caption-before-reason” strategy that forces the model to understand the image context before reasoning.\n2)\tUnlike prior VLMs requiring large-scale GPT-4-generated CoT supervision, Visionary-R1 is trained purely from question–answer pairs, significantly improving scalability and autonomy.\n3)\tThe inclusion of an auxiliary caption reward explicitly reduces the tendency to rely on superficial visual cues, encouraging deeper, generalizable reasoning."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1)\tBecause the R1 training process lacks explicit CoT supervision, it is uncertain whether the observed problems stem solely from shortcut learning. Other possible causes, such as unstable reward optimization or limited exploration, may also contribute. The paper should analyze these alternatives more carefully or justify why shortcut bias is the only plausible explanation.\n2)\tIn Figure 1, the paper mentions “shortcut” but does not clearly define what it refers to. In this paper, it seems related to models exploiting textual cues instead of visual reasoning, yet the description lacks formal or measurable criteria. The authors should explicitly clarify what constitutes a shortcut and how it is identified or quantified.\n3)\tIn line 92, the paper claims that the captioning step enables deeper image analysis rather than reliance on superficial cues. However, captioning typically focuses on describing visible, surface-level content rather than abstract reasoning. The paper does not clearly explain why or how generating captions contributes to improved reasoning ability. A stronger empirical or theoretical justification is needed to show that captioning truly enhances reasoning depth instead of merely restating image descriptions.\n4)\tFigure 2 shows that longer reasoning traces may improve performance, but some prior studies suggest that overthinking can harm efficiency and even accuracy [1]. The paper should clarify why increased reasoning length is considered beneficial here and whether there is evidence that such extended reasoning reflects genuine improvement rather than redundancy or noise. A discussion comparing “productive reasoning” versus “overthinking” would make the interpretation more balanced and convincing.    \n[1] More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models.\n5)\tThe paper claims that Visionary-R1 mitigates shortcut reasoning and reduces hallucination, yet it is not evaluated on dedicated hallucination detection or grounding benchmarks. Testing on such datasets would provide stronger evidence that the method truly decreases visual hallucinations rather than merely improving task accuracy.\n6.    Since the proposed Visionary-R1 framework introduces a new reinforcement training paradigm, open-sourcing the code is essential for reproducibility, fair comparison, and future research extensions."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762925260674,"tcdate":1762098158077,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission14918/Reviewer_bsz4"],"signatures":["ICLR.cc/2026/Conference/Submission14918/Reviewer_bsz4"],"forum":"bya3KOdLeS","number":5,"license":"CC BY 4.0","cdate":1762098158077,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission14918/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762925260674,"domain":"ICLR.cc/2026/Conference","replyto":"bya3KOdLeS","id":"i65i4AeFiG","forumContent":{"TLDR":{"value":"We addressed a shortcut problem in applying reinforcement learning to VLMs with Visionary-R1, trained on 273K CoT-free visual question-answer pairs  using only reinforcement learning."},"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Large Vision Language Model","Visual Reasoning","Reinforcement Learning"]},"primary_area":{"value":"foundation or frontier models, including LLMs"},"abstract":{"value":"Learning general-purpose reasoning capabilities has long been a challenging problem in AI. Recent research in large language models (LLMs), such as DeepSeek-R1, has shown that reinforcement learning techniques like GRPO can enable pre-trained LLMs to develop reasoning capabilities using simple question-answer pairs. In this paper, we aim to train visual language models (VLMs) to perform reasoning on image data through reinforcement learning and visual question-answer pairs, without any explicit chain-of-thought (CoT) supervision. Our findings indicate that simply applying reinforcement learning to a VLM---by prompting the model to produce a reasoning chain before providing an answer---can lead the model to develop shortcuts from easy questions, thereby reducing its ability to generalize across unseen data distributions. We argue that the key to mitigating shortcut learning is to encourage the model to interpret images prior to reasoning. Therefore, we train the model to adhere to a caption-reason-answer output format: initially generating a detailed caption for an image, followed by constructing an extensive reasoning chain. When trained on 273K CoT-free visual question-answer pairs and using only reinforcement learning, our model, named Visionary-R1, outperforms strong multimodal models, such as GPT-4o, Claude3.5-Sonnet, and Gemini-1.5-Pro, on multiple visual reasoning benchmarks. Code and models will be publicly released."},"_bibtex":{"value":"@misc{\nxia2025visionaryr,\ntitle={Visionary-R1: Mitigating Shortcuts in Visual Reasoning with Reinforcement Learning},\nauthor={Jiaer Xia and Yuhang Zang and Peng Gao and Yixuan Li and Kaiyang Zhou},\nyear={2025},\nurl={https://openreview.net/forum?id=bya3KOdLeS}\n}"},"title":{"value":"Visionary-R1: Mitigating Shortcuts in Visual Reasoning with Reinforcement Learning"},"pdf":{"value":"/pdf/3dd28e833f99b895e121c749459d21c527822646.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"xia|visionaryr1_mitigating_shortcuts_in_visual_reasoning_with_reinforcement_learning"},"authorids":{"value":["~Jiaer_Xia1","~Yuhang_Zang1","~Peng_Gao3","~Yixuan_Li1","~Kaiyang_Zhou1"]},"authors":{"value":["Jiaer Xia","Yuhang Zang","Peng Gao","Yixuan Li","Kaiyang Zhou"]}},"version":2},{"content":{"venue":{"value":"Pattern Recognit. 2025"},"venueid":{"value":"dblp.org/journals/PR/2025"},"paperhash":{"value":"park|fast_video_anomaly_detection_via_contextaware_shortcut_exploration_and_abnormal_feature_distance_learning"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Chaewon_Park:","~Donghyeong_Kim2","https://dblp.org/search/pid/api?q=author:MyeongAh_Cho:","~Minjung_Kim4","https://dblp.org/search/pid/api?q=author:Minseok_Lee:","https://dblp.org/search/pid/api?q=author:Seungwook_Park:","https://dblp.org/search/pid/api?q=author:Sangyoun_Lee:"]},"html":{"value":"https://doi.org/10.1016/j.patcog.2024.110877"},"_bibtex":{"value":"@article{DBLP:journals/pr/ParkKCKLPL25,\n  author={Chaewon Park and Donghyeong Kim and MyeongAh Cho and Minjung Kim and Minseok Lee and Seungwook Park and Sangyoun Lee},\n  title={Fast video anomaly detection via context-aware shortcut exploration and abnormal feature distance learning},\n  year={2025},\n  cdate={1735689600000},\n  journal={Pattern Recognit.},\n  volume={157},\n  pages={110877},\n  url={https://doi.org/10.1016/j.patcog.2024.110877}\n}\n"},"abstract":{"value":"Highlights•Patch anomaly generation enforces normality learning, in terms of appearance and motion.•Anomaly distance learning enlarges the feature distance of normal and abnormal frames.•Context-aware shortcut widens the quality gap between outputs for normal and abnormal samples.•Our method is fast, accurate, and free from pre-trained networks."},"title":{"value":"Fast video anomaly detection via context-aware shortcut exploration and abnormal feature distance learning"},"authors":{"value":["Chaewon Park","Donghyeong Kim","MyeongAh Cho","Minjung Kim","Minseok Lee","Seungwook Park","Sangyoun Lee"]}},"tmdate":1742364648728,"pdate":1735689600000,"tcdate":1731474800285,"writers":["~"],"signatures":["~DongHyeong_Kim1"],"forum":"iv6DGkctBZ","license":"CC BY-SA 4.0","number":198252,"cdate":1735689600000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1742364648728,"domain":"DBLP.org","id":"iv6DGkctBZ","version":2},{"content":{"summary":{"value":"This paper introduces CAT-Video, a corruption-aware training framework that improves robustness in latent video diffusion models (LVDMs) under noisy or imperfect conditioning.\nThe core contribution lies in two structured corruption strategies: Batch-Centered Noise Injection (BCNI), which injects noise along the deviation from the batch mean, and Spectrum-Aware Contextual Noise (SACN), which perturbs only the low-frequency spectral components of conditioning embeddings. Theoretical analysis supports tighter generalization bounds and improved temporal fidelity, and extensive experiments across multiple benchmarks demonstrate empirical gains over conventional Gaussian and Uniform corruption techniques as well as large-scale diffusion models."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1.\tWould combining multiple corruption techniques during training (e.g., applying BCNI followed by SACN) be beneficial, or would such composite perturbations destabilize the training process? Some insight into this interaction could be valuable for readers.\n2.\tHow are subsets of training data (e.g., 2 M vs. 10 M in the original DEMO paper) selected? Are they random samples or curated based on specific criteria?\n3.\tFor very long or high-resolution videos, does computing BCNI or SACN introduce significant additional computational overhead during training?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1.\tThe paper presents a clear exploration of corruption-aware training for video diffusion, supported by solid theoretical derivations that extend existing analyses to the temporal setting.\n2.\tExtensive experiments across four standard benchmarks and additional evaluations on autoregressive and multimodal models confirm broad empirical robustness.\n3.\tAblation studies, sensitivity analyses, and detailed appendices ensure reproducibility; released code and metrics further strengthen the work’s transparency."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1.\tAlthough the authors claim model-agnostic applicability, most analyses are conducted on the DEMO backbone with OpenCLIP-based encoders. This limits the universality claim, while evaluating BCNI/SACN on one or two additional diffusion backbones would strengthen the evidence.\n2.\tWhile CAT-Video demonstrates strong robustness on standard datasets, further evaluation under more realistic, noisy, or weakly aligned conditions would better reflect its practical robustness.\n3.\tRelated work on corruption-aware methods in image diffusion and other noisy-input domains is not sufficiently discussed. This section primarily centers on LVDMs, which are not the main focus of this work.\n4.\tAdditional qualitative visualizations would help clarify how CAT-Video improves visual coherence compared with other corruption strategies."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762931540896,"tcdate":1762055994447,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission19699/Reviewer_S25s"],"signatures":["ICLR.cc/2026/Conference/Submission19699/Reviewer_S25s"],"forum":"unZhwukf0T","number":3,"license":"CC BY 4.0","cdate":1762055994447,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission19699/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762931540896,"domain":"ICLR.cc/2026/Conference","replyto":"unZhwukf0T","id":"gx0765T5kL","forumContent":{"TLDR":{"value":"We introduce CAT-Video, a corruption-aware training framework that improves robustness and temporal coherence in video diffusion models through structured, data-aligned noise injection."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["video diffusion","corruption-aware training","robust video generation","structured noise injection","multimodal robustness","temporal coherence"]},"supplementary_material":{"value":"/attachment/c99cdbeea034fc1e74ad38310569a2906228cc13.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Latent Video Diffusion Models (LVDMs) have achieved state-of-the-art generative quality for image and video generation; however, they remain brittle under noisy conditioning, where small perturbations in text or multimodal embeddings can cascade over timesteps and cause semantic drift. Existing corruption strategies from image diffusion (Gaussian, Uniform) fail in video settings because static noise disrupts temporal fidelity. In this paper, we propose **CAT-Video**, a corruption-aware training framework with structured, data-aligned noise injection tailored for video diffusion. Our two operators—*Batch-Centered Noise Injection (BCNI)* and *Spectrum-Aware Contextual Noise (SACN)*—align perturbations with batch semantics or spectral dynamics to preserve coherence. CAT-Video yields substantial gains: BCNI reduces FVD by **31.9%** on WebVid-2M, MSR-VTT, and MSVD, while SACN improves UCF-101 by **12.3%**, outperforming Gaussian, Uniform, and even large diffusion baselines like DEMO (2.3B) and Lavie (3B) despite training on $\\mathbf{5}\\times$ less data. Ablations confirm the unique value of low-rank, data-aligned noise, and theory establishes why these operators tighten robustness and generalization bounds. CAT-Video thus sets a new framework for robust video diffusion, and our experiments show that it can also be extended to autoregressive generation and multimodal video understanding LLMs."},"_bibtex":{"value":"@misc{\nmaduabuchi2026catvideo,\ntitle={{CAT}-{VIDEO}: {CORRUPTION}-{AWARE} {TRAINING} {FOR} {ROBUST} {VIDEO} {DIFFUSION} {MODELS}},\nauthor={Chika Maduabuchi and Hao Chen and Yujin Han and Jindong Wang},\nyear={2026},\nurl={https://openreview.net/forum?id=unZhwukf0T}\n}"},"title":{"value":"CAT-VIDEO: CORRUPTION-AWARE TRAINING FOR ROBUST VIDEO DIFFUSION MODELS"},"pdf":{"value":"/pdf/36b779e3109b1503084b84e3c1c30d2bc4985918.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"maduabuchi|catvideo_corruptionaware_training_for_robust_video_diffusion_models"},"authorids":{"value":["~Chika_Maduabuchi1","~Hao_Chen15","~Yujin_Han1","~Jindong_Wang4"]},"authors":{"value":["Chika Maduabuchi","Hao Chen","Yujin Han","Jindong Wang"]}},"version":2},{"content":{"summary":{"value":"This paper conducts video scene graph generation in a weakly supervised manner. Several key steps are employed to generate scene graph labels from video-text annotations. The authors first design text prompts for multimodal large language models to parse the video captions into sentences along a time horizon. Then, a clustering-based caption-frame alignment framework is designed to align video frames with the segmented sentences. With images and their aligned texts, the authors use existing scene graph parsing methods and scene grounding methods to generate scene graphs. The authors also conduct experiments on the video scene graph generation benchmark , the Action Genome (AG) dataset, to validate the effectiveness the proposed method."},"soundness":{"value":4},"confidence":{"value":3},"questions":{"value":"I have no further questions. Please see the weakness part."},"rating":{"value":6},"details_of_ethics_concerns":{"value":"None"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":3},"strengths":{"value":"1. This paper proposes a weakly supervised pipeline for video scene graph generation, an important research topic in multimodal learning. Researchers can build a larger scene graph generation dataset with this pipeline from annotated video-text pairs. Researchers can also leverage the MLLMs to generate scene graphs without any human labels in further days.\n2. The design of the temporality-aware caption segmentation module is interesting; it can convert the video's text caption into multiple sentence segments. MLLM may find it challenging to finish such tasks. It would be helpful to add more details about the segment process. \n3. Aligning the segmented sentences with the video frames is critical for generating a scene graph. The authors provide many details about conducting the clustering-based alignment."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Although the proposed VSNLS pipeline can achieve good results on VidSGG tasks, its innovation still needs to be enhanced and emphasized. The VSNLS pipeline has four critical steps: the temporality-aware caption segmentation part, the action duration variability-aware caption-frame alignment part, the generation of pseudo-localized scene graphs part, and the adverse action classes generation part. Except for the action duration variability-aware caption-frame alignment part, I am unclear about the novelty and technical insights in other sections. \n1. In section 2.2, the authors design prompts to segment the video captions into sentences. Since there are a lot of prompt engineering works, what are the new insights in this part that can motivate the following researchers to conduct their work? Are there some key points that need to be noticed in design prompts?\n2. In section 2.4, the scene graph parsing and grounding utilize state-of-the-art methods to generate graph nodes. I guess readers are more willing to see how the authors develop some techniques to make these state-of-the-art methods conduct parsing or grounding more accurately.\n3. The experimental setting is organized in a bit of chaos. How the training, validation, and testing sets are split is not introduced. From line 316 to line 325. the reader can hardly understand whether the model is trained on the AG dataset or the combination of the AG caption and the MSVD caption dataset.\n4. There is no ablation study on the effectiveness of the microdesign in the ADV part. How different clustering strategies influence the final results is important to understand the rationale behind the VSNLS framework."}},"nonreaders":[],"tmdate":1731427804359,"tcdate":1730633403770,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission3326/Reviewer_6GYf"],"signatures":["ICLR.cc/2025/Conference/Submission3326/Reviewer_6GYf"],"forum":"GQgPj1H4pO","number":2,"license":"CC BY 4.0","cdate":1730633403770,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission3326/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427804359,"domain":"ICLR.cc/2025/Conference","replyto":"GQgPj1H4pO","id":"SOWr5rA7Gd","forumContent":{"TLDR":{"value":"We propose  a weakly-supervised video scene graph generation framework that aims to relieve the annotation costs by training a model using natural language supervision."},"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video Scene Understanding","Weakly Supervised Learning","Large Language Model"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Existing Video Scene Graph Generation (VidSGG) studies are trained in a fully supervised manner, which requires all frames in a video to be annotated, thereby incurring high annotation cost compared to Image Scene Graph Generation (ImgSGG). Although the annotation cost of VidSGG can be alleviated by adopting a weakly supervised approach commonly used for ImgSGG (WS-ImgSGG) that uses image captions, there are two key reasons that hinder such a naive adoption: 1) Temporality within video captions, i.e., unlike image captions, video captions include temporal markers (e.g., before, while, then, after) that indicate time-related details, and 2) Variability in action duration, i.e., unlike human actions in image captions, human actions in video captions unfold over varying duration. To address these issues, we propose a Natural Language-based Video Scene Graph Generation (NL-VSGG) framework that only utilizes the readily available video captions for training a VidSGG model. NL-VSGG consists of two key modules: Temporality-aware Caption Segmentation (TCS) module and Action Duration Variability-aware caption-frame alignment (ADV) module. Specifically, TCS segments the video captions into multiple sentences in a temporal order based on a Large Language Model (LLM), and ADV aligns each segmented sentence with appropriate frames considering the variability in action duration. Our approach leads to a significant enhancement in performance compared to simply applying the WS-ImgSGG pipeline to VidSGG on the Action Genome dataset. As a further benefit of utilizing the video captions as weak supervision, we show that the VidSGG model trained by NL-VSGG is able to predict a broader range of action classes that are not included in the training data, which makes our framework practical in reality."},"_bibtex":{"value":"@inproceedings{\nkim2025weakly,\ntitle={Weakly Supervised Video Scene Graph Generation via Natural Language Supervision},\nauthor={Kibum Kim and Kanghoon Yoon and Yeonjun In and Jaehyeong Jeon and Jinyoung Moon and Donghyun Kim and Chanyoung Park},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=GQgPj1H4pO}\n}"},"title":{"value":"Weakly Supervised Video Scene Graph Generation via Natural Language Supervision"},"pdf":{"value":"/pdf/6082acd7ebb0364ab36782cd36c1bb838a5bbf63.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"kim|weakly_supervised_video_scene_graph_generation_via_natural_language_supervision"},"authorids":{"value":["~Kibum_Kim1","~Kanghoon_Yoon2","~Yeonjun_In1","~Jaehyeong_Jeon1","~Jinyoung_Moon1","~Donghyun_Kim2","~Chanyoung_Park1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Kibum Kim","Kanghoon Yoon","Yeonjun In","Jaehyeong Jeon","Jinyoung Moon","Donghyun Kim","Chanyoung Park"]}},"version":2},{"content":{"venue":{"value":"Video-Langauge Models Poster"},"TLDR":{"value":"We propose a multimodal in-context ensemble learning with pseudo labels for better step-by-step, low-level video understanding."},"pdf":{"value":"/pdf/a5a4443ee4a6e432dcd7a86d45a04cab81c19949.pdf"},"keywords":{"value":["video-language model","in-context learning","ensemble learning","SOP","pseudo-labels","temporal."]},"venueid":{"value":"NeurIPS.cc/2024/Workshop/Video-Langauge_Models"},"paperhash":{"value":"xu|incontext_ensemble_learning_from_pseudo_labels_improves_videolanguage_models_in_lowlevel_workflow_understanding"},"authorids":{"value":["~Moucheng_Xu1","~Evangelos_Chatzaroulas1","~Luc_McCutcheon1","~Abdul_Ahad1","~Hamzah_Azeem1","~Janusz_Marecki1","~Ammar_Anwar1"]},"abstract":{"value":"A Standard Operating Procedure (SOP) defines a step-by-step written guide for a business software workflow. SOP generation is a crucial step towards automating end-to-end software workflows. Manually creating SOPs can be time-consuming. Recent advancements in large video-language models offer the potential for automating SOP generation by analyzing recordings of human demonstrations. However, current large video-language models face challenges with zero-shot SOP generation. In this work, we first explore in-context learning with video-language models for SOP generation. We then propose In-Context Ensemble Learning, to aggregate pseudo labels of SOPs. The proposed in-context ensemble learning increases test-time compute and enables the models to learn beyond its context window limit with an implicit consistency regularisation. We report that in-context learning helps video-language models to generate more temporally accurate SOPs, and the proposed in-context ensemble learning can consistently enhance the capabilities of the video-language models in SOP generation."},"_bibtex":{"value":"@inproceedings{\nxu2025incontext,\ntitle={In-Context Ensemble Learning from Pseudo Labels Improves Video-Language Models in Low-Level Workflow Understanding},\nauthor={Moucheng Xu and Evangelos Chatzaroulas and Luc McCutcheon and Abdul Ahad and Hamzah Azeem and Janusz Marecki and Ammar Anwar},\nbooktitle={Workshop on Video-Language Models @ NeurIPS 2024},\nyear={2025},\nurl={https://openreview.net/forum?id=xWM6MZJUv8}\n}"},"title":{"value":"In-Context Ensemble Learning from Pseudo Labels Improves Video-Language Models in Low-Level Workflow Understanding"},"track":{"value":"Short Paper Track (up to 3 pages)"},"authors":{"value":["Moucheng Xu","Evangelos Chatzaroulas","Luc McCutcheon","Abdul Ahad","Hamzah Azeem","Janusz Marecki","Ammar Anwar"]}},"tmdate":1736861080457,"pdate":1730081752794,"tcdate":1725980754209,"writers":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission32/Authors"],"signatures":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission32/Authors"],"forum":"xWM6MZJUv8","license":"CC BY 4.0","number":32,"cdate":1725980754209,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Post_Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Edit","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission32/-/Camera-Ready_Revision"],"mdate":1736861080457,"odate":1736861080441,"domain":"NeurIPS.cc/2024/Workshop/Video-Langauge_Models","id":"xWM6MZJUv8","version":2},{"content":{"summary":{"value":"This paper introduces VideoMathQA, a benchmark designed to evaluate multimodal mathematical reasoning in real-world videos, addressing the limitations of existing static image/text-based math benchmarks that lack support for temporal extension, dynamic visuals, and cross-modal integration. The benchmark comprises 420 manually curated video-question pairs spanning 10 mathematical domains and video durations from 10 seconds to over 1 hour. Each pair includes expert-annotated multi-step reasoning (2,945 total steps) with timestamps, enabling fine-grained evaluation of intermediate inference."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. Given the high annotation cost, what semi-automatic or crowdsourcing strategies are you exploring to scale VideoMathQA? Could synthetic data or data augmentation techniques be integrated without compromising the \"real-world\" essence of the benchmark?\n2. Models perform poorly on some categories, such as topology and graph theory. Could you provide qualitative examples of why these domains are more challenging for current models?"},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. VideoMathQA fills a gap by focusing on temporally extended cross-modal reasoning for math, an underexplored area in existing video benchmarks.  \n2. The benchmark leverages graduate-level experts to create detailed step-wise reasoning with timestamps. \n3. The four evaluation strategies address limitations of traditional MCQ and provide nuanced insights.    \n4. The authors systematically investigate factors impacting performance (model size, video duration, subtitles, frame sampling) and conduct error analysis across 7 categories, offering actionable guidance for model improvement."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. With only 420 video-question pairs, the benchmark may lack sufficient diversity to generalize across all real-world math instructional scenarios. The high annotation cost raises concerns about scalability."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762916663078,"tcdate":1761754191376,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission3317/Reviewer_bHWm"],"signatures":["ICLR.cc/2026/Conference/Submission3317/Reviewer_bHWm"],"forum":"VI4kGUfPio","number":2,"license":"CC BY 4.0","cdate":1761754191376,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission3317/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762916663078,"domain":"ICLR.cc/2026/Conference","replyto":"VI4kGUfPio","id":"vzdS0TQeXg","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"TLDR":{"value":"Mathematical Reasoning using Video MLLM Benchmark"},"keywords":{"value":["Multimodal Reasoning","Video Question Answering","Mathematical Understanding","Temporal Reasoning","Visual Grounding"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Mathematical reasoning in real-world video presents a fundamentally different challenge than static images or text. It requires interpreting fine-grained visual information, accurately reading handwritten or digital text, and integrating spoken cues, often dispersed non-linearly over time. In such multimodal contexts, success hinges not just on perception, but on selectively identifying and integrating the right details from a rich and noisy stream of content. To this end, we introduce VideoMathQA, a benchmark designed to evaluate whether models can perform such temporally extended cross-modal reasoning on videos. The benchmark spans 10 diverse mathematical domains, covering videos from 10 seconds to over 1 hour. We employ graduate-level experts to ensure high quality, for over 920 man-hours of annotation. To reflect real-world scenarios, questions are designed around three core reasoning challenges: direct problem solving, conceptual transfer, which requires applying learned methods to new problems; and deep instructional comprehension, involving multi-step reasoning over extended explanations and partially worked-out solutions. Each question includes multi-step reasoning annotations, enabling fine-grained diagnosis of model capabilities. Through this benchmark, we establish an evaluation framework for models that must reason, rather than merely perceive, jointly ground concepts across visual, audio, and textual modalities, across temporally extended mathematical problem settings."},"_bibtex":{"value":"@inproceedings{\nrasheed2026videomathqa,\ntitle={VideoMath{QA}: Benchmarking Mathematical Reasoning via Multimodal Understanding in Video},\nauthor={Hanoona Abdul Rasheed and Abdelrahman M Shaker and Anqi Tang and Muhammad Maaz and Ming-Hsuan Yang and Salman Khan and Fahad Shahbaz Khan},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=VI4kGUfPio}\n}"},"title":{"value":"VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Video"},"pdf":{"value":"/pdf/ecfd49591f38470840444b642dfebfb19c5d0c39.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"rasheed|videomathqa_benchmarking_mathematical_reasoning_via_multimodal_understanding_in_video"},"authorids":{"value":["~Hanoona_Abdul_Rasheed1","~Abdelrahman_M_Shaker1","~Anqi_Tang2","~Muhammad_Maaz1","~Ming-Hsuan_Yang1","~Salman_Khan4","~Fahad_Shahbaz_Khan1"]},"authors":{"value":["Hanoona Abdul Rasheed","Abdelrahman M Shaker","Anqi Tang","Muhammad Maaz","Ming-Hsuan Yang","Salman Khan","Fahad Shahbaz Khan"]}},"version":2},{"content":{"summary":{"value":"The paper tackles the problem Class Incremental Learning (CIL), focusing on the shortcut learning problem. To do so, the authors give a theoretical perspective justifying the existence of short learning in Continual Leaning and propose to leverage a large pool of data called library. The library is then used to detect potential shortcut learning issues when learning the current task and allows for relearning. The authors conduct experiments on various relevant CIL datasets and show a considerable improvement compared to state-of-the-art methods."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"See weaknesses"},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- The information bottleneck perspective of the shortcut learning problem is appreciate\n- The performances are compelling\n- The core idea is easy to follow\n- The code is shared\n- The paper is easy to follow"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"## Weaknesses\n - l 44. \"the replay buffer is fundamentally limited given how it is constructed because the model may be biased when selecting the buffer without accessing the future task data\". The replay buffer construction can vary from one paper to the other. No method is specified here, hindering the validity of such argument. In addition, the most common memory building strategy remains the reservoir sampling, which does not rely on the model performances at all, making this argument completely invalid in this context. The authors should be more precise in this statement.\n- The entire introduction, while making various claims about the existing strategies, contains only one single reference.\n- l52 \"Most previous work attributes the inability to retain the previous task data performance as the forgetting issue\". This is the actual definition of forgetting.\n- l53. The described case of wrong learning or shortcut learning in the case of Continual Learning is not new, see [1, 2]. Such work should be considered for comparison. \n- Work on learning representation with mutual information between task have been proposed before and are not mentioned in the paper [3, 4]. The usage of contrastive learning as a mutual information proxy is not new, see also [5].\n- l262. What about regularization strategies to mitigate shortcut learning?\n- l 311 \"Furthermore, just increasing replay-buffer size is not feasible as the computation cost of training on the replay-buffer will increase with the size of replay-buffer.\" I disagree with this statement. The batch size of data extracted from the buffer will still be the same, so the computation overhead does not depend on the buffer size. However, the space complexity will indeed increase.\n- Following on previous point, I do not understand how the authors justify the usage of library while denying large memory size. To me the usage of library is completely similar to using an infinite-size memory buffer and the proposed methods should be compared with larger memory size. Similarly, is computing the difficulty score for each sample in the library computationally intensive?\n- The authors claims an improved computation efficiency compared to large memory buffer, but such information is not provided in the paper.\n- How do you choose the library size for a given dataset? \n- While the paper claims to solve shortcut learning in this context, I do not see muuch evidence of that in the paper.\n- Why not compare to [4] ? It is cited in the paper and also addresses shortcut learning. A comparison of various CAM with compared methods, on top of performances, would be required to prove that the proposed method solves shortcut learning.\n### citations\n[1] Wang, Maorong, Nicolas Michel, Ling Xiao, and Toshihiko Yamasaki. \"Improving Plasticity in Online Continual Learning via Collaborative Learning.\" In _Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition_, pp. 23460-23469. 2024.\n\n[2] Yujie Wei, Jiaxin Ye, Zhizhong Huang, Junping Zhang, and\nHongming Shan. Online prototype learning for online continual learning. In ICCV, 2023.\n\n[3] Gu, Y., Yang, X., Wei, K., & Deng, C. (2022). Not just selection, but exploration: Online class-incremental continual learning via dual view consistency. In _Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition_ (pp. 7442-7451).\n\n[4] Guo, Y., Liu, B., & Zhao, D. (2022, June). Online continual learning through mutual information maximization. In _International conference on machine learning_ (pp. 8109-8126). PMLR.\n\n[5] Mai, Z., Li, R., Kim, H., & Sanner, S. (2021). Supervised contrastive replay: Revisiting the nearest class mean classifier in online class-incremental continual learning. In _Proceedings of the IEEE/CVF conference on computer vision and pattern recognition_ (pp. 3589-3599)."}},"nonreaders":[],"tmdate":1731427982807,"tcdate":1730718059588,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission3923/Reviewer_gKAw"],"signatures":["ICLR.cc/2025/Conference/Submission3923/Reviewer_gKAw"],"forum":"gCYFtUKXSc","number":4,"license":"CC BY 4.0","cdate":1730718059588,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission3923/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427982807,"domain":"ICLR.cc/2025/Conference","replyto":"gCYFtUKXSc","id":"FkIkil4MAf","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Continual learning","Data-Efficient Learning","Information Theory"]},"primary_area":{"value":"transfer learning, meta learning, and lifelong learning"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Replay-based methods provide a promising solution to address catastrophic forgetting issue in continual learning. They try to retain previous knowledge by using a small amount of data from previous tasks stored in a fix-sized buffer. In this work, we invoke the information bottleneck principles and reveal some fundamental limitations of those methods on their effectiveness in capturing the truly important features from the prior tasks by relying on the buffer data selected according to the model's performance on known tasks. Since future tasks are not accessible during model training and buffer construction, the trained model and the buffer data tend to be biased towards making accurate predictions on the labels of known tasks. However, when new task samples are introduced along with labels, the biased model and the buffer data become less effective in differentiating samples of the old tasks from those of the new ones. Inspired by the way humans learn over time, we propose a novel relearning technique that makes use of additional past data, referred to as the library, to test how much information the model loses after learning the new task. We then realign the model towards those forgotten samples by training on a carefully selected small subset samples from the library for a few epochs with comparable computational cost as existing replay-based models. The experimental results on multiple real-world datasets demonstrate that the proposed relearning process can improve the performance of the state-of-the-art continual learning methods by a large margin."},"_bibtex":{"value":"@misc{\nacharya2024avoid,\ntitle={Avoid Being a Shortcut Learner through Library-Based Re-Learning},\nauthor={Abhinab Acharya and Dayou Yu and Yuansheng Zhu and Qi Yu and Xumin Liu},\nyear={2024},\nurl={https://openreview.net/forum?id=gCYFtUKXSc}\n}"},"title":{"value":"Avoid Being a Shortcut Learner through Library-Based Re-Learning"},"pdf":{"value":"/pdf/f04bb1b55721f773882cd541101931bfa8d127a7.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"acharya|avoid_being_a_shortcut_learner_through_librarybased_relearning"},"authorids":{"value":["~Abhinab_Acharya1","~Dayou_Yu1","~Yuansheng_Zhu1","~Qi_Yu1","~Xumin_Liu1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Abhinab Acharya","Dayou Yu","Yuansheng Zhu","Qi Yu","Xumin Liu"]}},"version":2},{"content":{"summary":{"value":"The paper proposes a semi-supervised preference learning framework that infers pseudo-preference labels for many unlabeled video segment pairs from a small set of labeled pairs. The method encodes segments with a video foundation model (ViFM) into latent embeddings, computes an optimal transport (OT) plan between labeled and unlabeled segment sets in the latent space under uniform marginals, aggregates labeled pairwise relations through the transport couplings with the preference score matrix to score unlabeled pairs, and then normalizes/thresholds these scores to produce pseudo-labels. A reward model trained on real + pseudo labels is used for offline RL. Experiments span D4RL, MetaWorld, Robomimic, and two real-robot tasks"},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. Assumption of the marginal. The paper has an assumption about optimal transport where the marginals are uniform. This assumption may not handle imbalanced or redundant segment distributions and am wondering if the authors have considered this situation or if they have encountered this in practice. \n\n2. More qualitative results: It would help the readers better understand how the pseudo labels are produced by providing some qualitative examples of the labeled segments and unlabeled ones. In Fig.1, the paper only has one example of preference matrix and transport plan but they are not connected with the corresponding video segments. Therefore, it would be more clear to show the corresponding segments too. For example,  you can show (i) labeled segments and their preference matrix, (ii) several unlabeled pairs, and (iii) corresponding transport plans/couplings and the computed score. \n\n3. Cost metric choice: the paper mentions Euclidean/cosine/others are potential metrics for the distance function but uses Euclidean in experiments.  If the authors can provide some explanations about why euclidean distance is chosen and include some examples of the distance with similar pairs of video segments vs different pairs, it would help readers understand more about the metrics and pretrained video representations.\n\n4. The paper is mostly focused on offline rl settings but the proposed way to learn reward function can also be used for online preference-based RL like PEBBLE[3]. I am curious to see if the inferred preference from a few labeled samples can be robust to the on-policy distribution shift in the sampled segment pairs. \n\n5. In the experiments, the paper has compared among (i) learning with task rewards (ii) learning with only small N real labels and (iii) learning with real + pseudo labels (N+M). It would also provide more information about the quality of the pseudo labels if the method is also compared against IQL learned from reward function trained with (N+M) real labels. This can be obtained from ground-truth task reward labeling and can serve as an upper bound. The gap is how far the pseudo labels are from the real labels. \n\n[3] Lee, Kimin, Laura M. Smith, and Pieter Abbeel. \"PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training.\" International Conference on Machine Learning. PMLR, 2021."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The paper has clear problem setup and writing; the proposed method and pipeline is easy to follow.\n\n2. The paper proposed a novel application of Optimal Transport to propagate preferences from limited supervision to unlabeled data to reduce the labeling costs without touching the pretrained video representation itself, inspired by prior work to use optimal transport in reward learning. \n\n3. The experiments covers a wide range of tasks with both simulation and real-robot evaluations.\n\n4. It also includes detailed ablations of design choices (video encoders, thresholds, number of preferences, etc.)."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. One of the very-related work to this paper is [1] which is also cited in this paper. This prior work also proposes to use optimal transport to compute the reward but it is based on a adapted visual representation learned via small number of preferences while this work is fixing the pre-trained representations and only propagating the preferences from labeled ones to unlabeled ones to optimal transport. Since one of the followup work [2] of [1] also uses learned reward function from a few preference labels to create pseudo labels, it is worth comparing to this baseline to see whether the proposed method outperforms this prior work. \n\n2. The evaluated tasks in robomimic, metaworld and real-world are relatively short horizon. The scalability of the proposed method to long-horizon, high-precision tasks (e.g., Square, Tool-Hang in robomimic) is unclear.\n\n3. One of the benefit to use large scale pretrained video encoder is their generalization capabilities across tasks. However, in this paper, it still assumes the labeled video segments and the unlabeled ones are in the same domain, which may weaken the motivation of using those pretrained video representations. \n\n\n[1] Tian, Thomas, et al. \"What Matters to You? Towards Visual Representation Alignment for Robot Learning.\" The Twelfth International Conference on Learning Representations.\n[2] Tian, Ran, et al. \"Maximizing alignment with minimal feedback: Efficiently learning rewards for visuomotor robot policy alignment.\" arXiv preprint arXiv:2412.04835 (2024)."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762931616647,"tcdate":1762401789052,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission19772/Reviewer_EaLy"],"signatures":["ICLR.cc/2026/Conference/Submission19772/Reviewer_EaLy"],"forum":"wWvrC9oajI","number":3,"license":"CC BY 4.0","cdate":1762401789052,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission19772/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762931616647,"domain":"ICLR.cc/2026/Conference","replyto":"wWvrC9oajI","id":"fNsv7xz3vv","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"TLDR":{"value":"A semi-supervised preference learning method that use optimal transport over embeddings of the video foundation model to perform pseudo-labeling."},"keywords":{"value":["Preference-based Reinforcement Learning","Offline Reinforcement Learning","Feedback Efficiency","Optimal Transport","Video Foundation Models","Semi-supervised Learning","Robotics"]},"supplementary_material":{"value":"/attachment/0900ade87e89d9caa591eb55def3c30043c6aeb1.zip"},"primary_area":{"value":"reinforcement learning"},"abstract":{"value":"Conveying complex objectives to reinforcement learning (RL) agents often requires meticulous reward engineering. Preference-based RL offers a promising alternative by learning reward functions from human feedback, but its scalability is hindered by the large amount of feedback required. Inspired by recent advances in Video Foundation Models (ViFMs), we present Video-based Optimal Transport Preference (VOTP), a semi-supervised preference learning framework that can learn effective reward functions from only a handful of preference labels. By leveraging optimal transport in the representation space of ViFMs for pseudolabeling, VOTP can utilize large amounts of unlabeled data for reward learning, substantially reducing the need for human supervision. Extensive experiments across locomotion and manipulation tasks show that VOTP outperforms existing PbRL methods under limited feedback. We further validate VOTP on real robotic tasks, demonstrating its ability to learn useful rewards with minimal human input."},"_bibtex":{"value":"@misc{\nluu2026videobased,\ntitle={Video-Based Optimal Transport for Feedback-Efficient Offline Preference-Based Reinforcement Learning},\nauthor={Tung Minh Luu and Hwanhee Kim and Younghwan Lee and Chang D. Yoo},\nyear={2026},\nurl={https://openreview.net/forum?id=wWvrC9oajI}\n}"},"title":{"value":"Video-Based Optimal Transport for Feedback-Efficient Offline Preference-Based Reinforcement Learning"},"pdf":{"value":"/pdf/6049eb8a04439e084c5b27b711f23b1dd27f79e7.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"luu|videobased_optimal_transport_for_feedbackefficient_offline_preferencebased_reinforcement_learning"},"authorids":{"value":["~Tung_Minh_Luu2","~Hwanhee_Kim2","~Younghwan_Lee1","~Chang_D._Yoo1"]},"authors":{"value":["Tung Minh Luu","Hwanhee Kim","Younghwan Lee","Chang D. Yoo"]}},"version":2},{"content":{"summary":{"value":"This paper proposes VideoDetective, a multimodal large language model designed to address long video question-answering (LVQA) tasks through an efficient question-aware memory mechanism.\n\nThe key contributions include:\n\n  1. Question-aware Memory Mechanism: The method processes long videos by dividing them into sub-segments and employs learnable memory tokens to compress visual information in a\n  question-aware manner.\n  2. Recurrent Critical Clue Seeking: A memory bank stores and aggregates semantic representations across video segments to maintain historical context for subsequent processing.\n  3. GLVC Dataset: A new long video QA dataset featuring concrete temporal clues scattered throughout entire videos, designed to better evaluate models' ability to ground critical\n  information.\n  4. Efficiency Claims: The method reportedly enables processing of 100K tokens (3600 frames) with only 32K context length, requiring 2 minutes inference time and 37GB GPU memory.\n\n  The approach is motivated by human cognitive processes of \"thinking while watching\" and aims to seek small amounts of crucial clues rather than processing entire video content at once."},"soundness":{"value":1},"confidence":{"value":5},"questions":{"value":"**Training Process Mechanics**\n  - Please provide the complete training algorithm with explicit gradient computation formulas: $\\frac{\\partial L}{\\partial M_i} = ?$ for $i = 1,2,...,S$\n  - Explain exactly what \"memory tokens do not participate in loss calculation\" means technically\n  - Provide pseudocode showing how computational graphs are maintained across recurrent segments\n  - Clarify the specific loss function and optimization target for the warmup stage\n\n**Experimental Methodology**\n  - Can you provide results on GLVC dataset using zero-shot evaluation (without training on GLVC)?\n  - What are the exact hardware specifications and software environments for efficiency comparisons?\n  - Can you include ablation studies removing question-aware compression and recurrent memory components?\n\n**Sampling and Data Processing**\n  - How exactly are videos of different lengths processed during training and inference?\n  - What is the strategy for handling the last segment when video length is not divisible by 32?\n  - How do you maintain semantic coherence when using fixed-size segments?\n  - Can you provide analysis showing the method actually captures \"critical clues\" rather than random information?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":1},"strengths":{"value":"**Originality**\n\n  - Dataset Contribution: GLVC provides temporal grounding annotations that could benefit the community for more rigorous evaluation of long video understanding capabilities.\n\n**Quality**\n\n  - Well-Motivated Approach: The method addresses a real limitation of current MLLMs in handling long video contexts due to memory constraints.\n  - Comprehensive Evaluation: The paper evaluates on multiple established benchmarks (VideoMME, MLVU, LongVideoBench, etc.) covering both short and long video scenarios.\n\n**Clarity**\n\n  - Clear Problem Statement: The paper clearly articulates the challenges of long video QA and memory limitations.\n  - Good Visual Presentation: Figures effectively illustrate the overall architecture and key concepts.\n\n**Significance**\n\n  - Important Problem: Long video understanding is a significant challenge for current MLLMs with practical applications.\n  - Memory Efficiency: If the technical claims are validated, the approach could enable broader deployment of video understanding models."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"**Fundamental Training Process Contradictions**\n The paper contains a critical technical inconsistency in Section 3.3:\n  - Claims \"memory tokens do not participate in loss calculation\" yet the loss function $L = \\sum_{j=1}^l -\\log p(x_j|V_1, M_1, Q, \\cdots, V_S, M_S, Q, x_0, ..., x_{j-1})$ explicitly depends\n  on memory tokens $M_i$\n  - If memory tokens don't participate in loss calculation, how do they receive gradients for optimization?\n  - This creates a fundamental contradiction that questions the technical feasibility of the approach\n\n**Gradient Backpropagation Problems**\n  The recurrent processing design raises serious gradient flow issues:\n  - Each segment's memory tokens depend on historical context from previous segments\n  - Maintaining computational graphs across all segments would require enormous memory (contradicting efficiency claims)\n  - The paper provides no explanation of how gradients backpropagate through the memory bank updates\n  - Missing details on whether detach() operations are used and where\n\n**Incomplete Training Strategy Description**\n  The two-stage training process lacks crucial technical details:\n  - Warmup stage: Uses video-caption pairs but provides no loss function or training objective\n  - Compression ratio inconsistency: Warmup uses ratio 32, main training uses ratio 16 without justification\n  - Learning rate jump: 100× increase from warmup (1e-6) to main training (1e-4) lacks theoretical and emperical basis\n\n**Insufficient Ablation Studies**\n  Current ablations only test compression ratios, missing critical components:\n  - No validation of question-aware compression effectiveness\n  - Missing ablation on recurrent memory mechanism vs. simple aggregation\n  - No verification that the model actually \"seeks critical clues\" as claimed\n\n**Unfair Experimental Comparisons**\n  - Data leakage: VideoDetective trained on GLVC dataset but evaluated on it (Table 2)\n\n**Sampling Strategy Inconsistencies**\n  Section 5.1 reveals problematic data handling:\n  - Fixed 32-frame segments ignore semantic boundaries\n  - No strategy for handling videos shorter/longer than expected lengths\n  - Compression ratio $k = N_i/\\alpha$ undefined for segments with $N_i < \\alpha$\n  - Training-inference mismatch in handling variable-length sequences\n\n**Limited Performance Gains**\n  Results show concerning patterns:\n  - Large gaps with SOTA: 10-15 point deficits compared to GPT-4o, Gemini-1.5-Pro\n  - Failure on key benchmarks: Acknowledged poor performance on LongVideoBench\n  - Marginal improvements: Small gains over same-scale models don't justify complexity\n  - Short video regression: Performance drops on short videos suggest fundamental limitations"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762918253184,"tcdate":1761376567264,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission5773/Reviewer_Y4eZ"],"signatures":["ICLR.cc/2026/Conference/Submission5773/Reviewer_Y4eZ"],"forum":"9glgMTTZb9","number":1,"license":"CC BY 4.0","cdate":1761376567264,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission5773/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762918253184,"domain":"ICLR.cc/2026/Conference","replyto":"9glgMTTZb9","id":"1pheHhSNgI","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["MLLM; LongVideo"]},"supplementary_material":{"value":"/attachment/2e42a1c09f6d6f7b74ce5b21f718da8e25d638dd.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Long Video Question-Answering (LVQA) presents a significant challenge for Multi-modal Large Language Models (MLLMs) due to immense context and overloaded information, which could also lead to prohibitive memory consumption. \nWhile existing methods attempt to address these issues by reducing visual tokens or extending model's context length, they may miss useful information or take considerable computation.\nIn fact, when answering given questions, only a small amount of crucial information is required.\nTherefore, we propose an efficient question-aware memory mechanism, enabling MLLMs to recurrently seek these critical clues. Our approach, named VideoDetective, simplifies this task by iteratively processing video sub-segments. For each sub-segment, a question-aware compression strategy is employed by introducing a few special memory tokens to achieve purposefully compression. This allows models to effectively seek critical clues while reducing visual tokens.\nThen, due to history context could have a significant impact, we recurrently aggregate and store these memory tokens to update history context, which would be reused for subsequent sub-segments. \nFurthermore, to more effectively measure model's long video understanding ability, we introduce GLVC (Grounding Long Video Clues), a long video question-answering dataset, which features grounding critical and concrete clues scattered throughout entire videos.\nExperimental results demonstrate our method enables MLLMs with limited context length of 32K to efficiently process 100K tokens (3600 frames, an hour-long video sampled at 1fps), requiring only 2 minutes and 37GB GPU memory usage. Evaluation results across multiple long video benchmarks illustrate our method can more effectively seek critical clues from massive information."},"_bibtex":{"value":"@misc{\ndu2026video,\ntitle={Video Detective: Seek Critical Clues Recurrently to Answer Question from Long Videos},\nauthor={Henghui Du and Chang Zhou and Chunjie Zhang and Xi Chen and Di Hu},\nyear={2026},\nurl={https://openreview.net/forum?id=9glgMTTZb9}\n}"},"title":{"value":"Video Detective: Seek Critical Clues Recurrently to Answer Question from Long Videos"},"pdf":{"value":"/pdf/bf8baa9af35aa0e72e7a5b3d40c3c4f35d3b6d6c.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"du|video_detective_seek_critical_clues_recurrently_to_answer_question_from_long_videos"},"authorids":{"value":["~Henghui_Du1","~Chang_Zhou3","~Chunjie_Zhang6","~Xi_Chen21","~Di_Hu1"]},"authors":{"value":["Henghui Du","Chang Zhou","Chunjie Zhang","Xi Chen","Di Hu"]}},"version":2},{"content":{"summary":{"value":"TemporalBench is a benchmark focused on fine-grained temporal dynamics in video understanding. It contains ~2K videos with human-authored, temporally detailed captions, yielding ~15K QA pairs (≈10K short, ≈5K long). Negatives are crafted as minimal temporal perturbations (word-level and event-level), intended to force models to reason over counts, order, duration, and motion direction—not just single-frame semantics. The authors identify a structural bias in standard MCQ (a “centralized” distractor effect) and propose Multiple Binary Accuracy (MBA)—factorizing each (M+1)-way MCQ into M binary decisions—to reduce option-structure shortcuts. Experiments span generative video LMMs and video embedding models, with a human AMT reference. Results show a substantial gap: e.g., Gemini-2.5-Pro 43.6% MBA on short-video QA vs humans 67.9%, with further drops on long-video tasks and near-chance performance for embedding models."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"* Does the number of negatives per item (M) vary across your short-video set? If so, how do you normalize MBA so scores are comparable across items/models?\n\n* To what extent do lexical/style artifacts drive performance differences between positives and negatives? What are text-only LLM results under MBA, broken down by negative type (word vs. event), domain, and video length?\n\n* How robust are results to the source of negatives (human vs. LLM) and to cross-LLM generation (e.g., train/eval splits with different negative generators)?\n\n* How do you guarantee unique alignment of each dense sub-caption within concatenated long videos, and can models show where the evidence lies (frame indices)?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"* Clear problem framing. Pinpoints why many “video” benchmarks collapse to static recognition and single-frame shortcuts; motivates a benchmark where temporal signals are necessary.\n\n* Fine-grained supervision. Human-refined captions capture action frequency, ordering, duration, motion polarity—details that cannot be recovered from a single frame.\n\n* Minimal, targeted negatives. Word-level and event-level edits (e.g., “three slices”→“two slices”) create controlled contrasts that probe specific temporal skills.\n\n* Broad evaluation scope. Covers short (<20s) and long (≤20 min) settings and evaluates both generative and embedding model families with frame-budget studies.\n\n* Metric contribution (MBA). Identifies an MCQ option-structure bias and offers a principled factorization into multiple binary decisions; shows large MCQ vs MBA gaps (e.g., 78.7% vs 43.6%).\n\n* Human baseline & worker hygiene. Excludes caption annotators from evaluation; uses onboarding and curation to ensure reliable human reference."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"* MBA comparability & calibration. MBA multiplies per-option accuracies; difficulty depends on M. If M varies across items/splits, raw MBA isn’t directly comparable. The long-video setup fixes the chance (~9.5%), but short-video details on M distribution and normalization are unclear.\n\n* Residual text-pattern cues. Although MBA mitigates centralized-option bias, lexical/semantic artifacts in positives vs. LLM-generated negatives may remain (style, length, rare tokens). A stronger audit of text-only performance per item type would help.\n\n* Negative generation provenance. Negatives are primarily LLM-generated and then author-filtered; this may inject LLM-specific artifacts that certain models can learn to exploit or avoid. Human-authored adversarial negatives or cross-LLM generation ablations would strengthen claims.\n\n* Annotation consistency & reliability. The two-stage AMT→author pipeline is sound, but the paper lacks inter-annotator agreement, edit-rate statistics, and error taxonomies for refined captions (especially beyond FineGym).\n\n* Embedding model diagnosis. “Near-chance due to small embedding size (768–2048)” is speculative; the deficit may stem from temporal aggregation design, training data, or loss formulations. A dimension-controlled ablation is needed.\n\n* Long-video construction bias. Long items are formed by concatenating dense clip captions; correctness hinges on exact sub-caption alignment. This composition could advantage strategies that match local segments rather than truly reasoning over long temporal narratives.\n\n* Scope & coverage. Audio is removed during annotation (good for visual reliance), but the evaluation setup for audio/subtitles isn’t fully specified. If subtitles/text are ever provided at test time, it risks re-introducing non-visual shortcuts."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762915557924,"tcdate":1762826993749,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission588/Reviewer_TWji"],"signatures":["ICLR.cc/2026/Conference/Submission588/Reviewer_TWji"],"forum":"XQfRnmOzY8","number":4,"license":"CC BY 4.0","cdate":1762826993749,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission588/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762915557924,"domain":"ICLR.cc/2026/Conference","replyto":"XQfRnmOzY8","id":"BzcvgNYDAh","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"TLDR":{"value":"TemporalBench: Evaluating Fine-Grained Temporal Dynamics Understanding for Multimodal Models"},"keywords":{"value":["Video; Benchmark; Temporal"]},"supplementary_material":{"value":"/attachment/9160fda348c6f1df228c11cd91238e3ac9d7b8df.zip"},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Understanding fine-grained temporal dynamics is crucial for multimodal video comprehension and generation. Due to the lack of fine-grained temporal annotations, existing video benchmarks mostly resemble static image benchmarks and are insufficient at evaluating models for temporal understanding.  In this paper, we introduce *TemporalBench*, a benchmark dedicated to evaluating **fine-grained temporal understanding** in videos. *TemporalBench* consists of $\\sim$15K video question-answer pairs, derived from $\\sim$2K high-quality human annotations detailing the temporal dynamics. As a result, our benchmark provides a unique testbed for evaluating various temporal understanding and reasoning abilities such as *action frequency, motion magnitude, event order*, etc.  Moreover, it enables evaluations on various tasks such as both short and long video understanding, as well as different models including multimodal embedding models and text generation models.  Furthermore, we notice a critical pitfall for multi-choice QA where LLMs can detect the subtle changes in negative captions and find a \"centralized” description as a cue for its prediction. To correct such bias, we propose **Multiple Binary Accuracy (MBA)**, a new metric for dense temporal understanding.  Results show that state-of-the-art models like GPT-4o achieve only **38.5%** short video QA, demonstrating a significant gap ($\\sim$30%) between humans and AI in temporal understanding.  We hope that *TemporalBench* can foster research on improving models' temporal reasoning capabilities. Both dataset and code will be available."},"_bibtex":{"value":"@misc{\ncai2025temporalbench,\ntitle={TemporalBench: Evaluating Fine-Grained Temporal Dynamics Understanding for Multimodal Models},\nauthor={Mu Cai and Reuben Tan and Jianrui Zhang and Bocheng Zou and Kai Zhang and Feng Yao and Fangrui Zhu and Jing Gu and Yiwu Zhong and Yuzhang Shang and Yao Dou and Jaden Park and Jianfeng Gao and Yong Jae Lee and Jianwei Yang},\nyear={2025},\nurl={https://openreview.net/forum?id=XQfRnmOzY8}\n}"},"title":{"value":"TemporalBench: Evaluating Fine-Grained Temporal Dynamics Understanding for Multimodal Models"},"pdf":{"value":"/pdf/99e03a75194d1ccd20c1d2ebaf5eb24601970f43.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"cai|temporalbench_evaluating_finegrained_temporal_dynamics_understanding_for_multimodal_models"},"authorids":{"value":["~Mu_Cai1","~Reuben_Tan1","~Jianrui_Zhang1","~Bocheng_Zou1","~Kai_Zhang10","~Feng_Yao1","~Fangrui_Zhu1","~Jing_Gu2","~Yiwu_Zhong1","~Yuzhang_Shang1","~Yao_Dou1","~Jaden_Park1","~Jianfeng_Gao1","~Yong_Jae_Lee2","~Jianwei_Yang1"]},"authors":{"value":["Mu Cai","Reuben Tan","Jianrui Zhang","Bocheng Zou","Kai Zhang","Feng Yao","Fangrui Zhu","Jing Gu","Yiwu Zhong","Yuzhang Shang","Yao Dou","Jaden Park","Jianfeng Gao","Yong Jae Lee","Jianwei Yang"]}},"version":2},{"content":{"summary":{"value":"This paper introduces LongViTU, a novel large-scale dataset (~121k QA pairs, ~900h videos) for long-form video understanding.  The authors address the limitations of existing video question-answering (VQA) datasets by focusing on several key aspects:  diverse real-world scenarios (leveraging Ego4D), explicit timestamp labels for QA-related events, long average certificate length (4.6 minutes), fine-grained categorization of QA pairs (spatiotemporal understanding, episodic reasoning, commonsense inference), and open-ended, precise QA generation.  A hierarchical pipeline, employing LLMs (primarily GPT-4) at multiple stages (hierarchical video tree construction, long-form QA generation, self-revision), is used for automatic dataset creation.  Experiments demonstrate the challenges posed by LongViTU to existing video language models (VLMs), showing a performance gap even between open-source and commercial models. Fine-tuning on LongViTU improves performance on both in-distribution and out-of-distribution benchmarks."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. How were the specific parameters for the sliding window (five segments) determined? What is the sensitivity of the results to changes in this parameter?\n2. What is the inter-annotator agreement (IAA) for the human annotations used in the Ego4D dataset, and how does this affect the quality of LongViTU?\n3. What are the computational costs associated with generating and processing LongViTU?\n4. Can you provide a more detailed analysis of the biases present in the generated QA pairs?\n5. How does the performance of the fine-tuned models change with different sizes of the LongViTU training set?"},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. LongViTU explicitly addresses the limitations of temporal context, length, and fine-grained question types from the perspective of sft.  The hierarchical pipeline for automatic dataset generation is a sound procedure to create long-form annotations from bottom to top. Its sheer scale of the dataset (~900 hours of video) and its diversity in terms of scenarios and question types are decent.  The use of Ego4D ensures real-world relevance.\n2. The paper includes a thorough quantitative evaluation on LongViTU and several benchmark datasets, demonstrating the effectiveness of the dataset and highlighting the challenges it presents.  The use of GPT-4 for scoring is a reasonable approach given the open-ended nature of the QA pairs.  Qualitative examples further illustrate the dataset's capabilities. The availability of the dataset, fine-tuned models, and code is a valuable contribution to the community."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The reliance on LLMs (GPT-4) throughout the pipeline raises concerns about potential biases inherited from the pre-training data of these models.  Moreover, a hierarchical pipeline may cause error cumulation, making the bias even worse. A thorough analysis of potential biases in the generated QA pairs is missing. \n2. While self-revision is employed, a more robust human evaluation of the dataset quality would strengthen the paper's claims.  The current human evaluation seems limited to Appendix B.\n3. Experiments need improvements. The number of models evaluated in the benchmark is too limited, and some of the current long video large language models, such as LongVA, LongVILA, have not been included in the evaluation. The model performance used to validate the training dataset's effectiveness is too weak (for instance, LLama-VID performs below random chance on VideoMME), and the improvements achieved after fine-tuning are relatively minor."}},"nonreaders":[],"tmdate":1731428925998,"tcdate":1730625058143,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission7604/Reviewer_rfFt"],"signatures":["ICLR.cc/2025/Conference/Submission7604/Reviewer_rfFt"],"forum":"4j9plQoOH1","number":3,"license":"CC BY 4.0","cdate":1730625058143,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission7604/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428925998,"domain":"ICLR.cc/2025/Conference","replyto":"4j9plQoOH1","id":"jtgauTat7F","forumContent":{"TLDR":{"value":"We propose a large-scale instruction-tuning dataset for long-form video understanding."},"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["vision language models","instruction-tuning","long-form video understanding"]},"primary_area":{"value":"datasets and benchmarks"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"This paper presents LongViTU, a large-scale (~121k QA pairs, ~900h videos), automatically generated dataset for long-form video understanding. Our key idea is inspired by the success of Large Language Models (LLMs) and Multimodal Language Models (MLMs) that are fueled by machine-generated instruction-following data (*e.g.*, InstructGPT, LLaVA). We developed a *systematic* approach to produce massive question-answeringing pairs tailored to virtually unbounded long videos by organizing them into a ***hierarchical tree***, incorporating ***self-revision*** mechanisms to guarantee high quality. We curate LongViTU for each QA pair: 1) involves a long context (average *certificate length* of 4.6 minutes); 2) requires rich knowledge and condensed reasoning (commonsense, causality, planning, *etc.*); 3) explicit labels the timestamps of relevant events throughout the entire video. Furthermore, LongViTU provides a benchmark to facilitate future research in instruction-following for long-form videos. Our experiments first reveal the performance gap between open-source video MLMs and their commercial counterparts (*e.g.*, Gemini-1.5-Pro) on this benchmark. Supervised Fine-Tuning (SFT) on open-source models led to Video-LLaVA achieving the best performance, with a GPT-4 score of $50.7$, closely following $52.3$ by the leading closed-source model Gemini-1.5-Pro, underscoring the substantial challenge posed by our benchmark. Further SFT on LongViTU with Video-LLaVA resulted in improvements of $30.7$% on the In-Distribution (ID) benchmark EgoSchema; $12.9$% and $0.6$% on the Out-of-Distribution (OOD) benchmarks WorldQA and VideoMME, respectively. These outcomes demonstrate the effectiveness and robust OOD generalizability of our proposed instruction-tuning scheme for long-form video understanding. The dataset, SFT models, and code are publicly available on the anonymous page [LongViTU](https://longvitu.github.io)."},"_bibtex":{"value":"@misc{\nwu2024longvitu,\ntitle={LongVi{TU}: Instruction Tuning for Long-Form Video Understanding},\nauthor={Rujie Wu and Xiaojian Ma and Hai Ci and Yue Fan and Yuxuan Wang and Haozhe Zhao and Qing Li and Yizhou Wang},\nyear={2024},\nurl={https://openreview.net/forum?id=4j9plQoOH1}\n}"},"title":{"value":"LongViTU: Instruction Tuning for Long-Form Video Understanding"},"pdf":{"value":"/pdf/e663a2eb9e041444826a666f95acc8764c6e736b.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"wu|longvitu_instruction_tuning_for_longform_video_understanding"},"authorids":{"value":["~Rujie_Wu2","~Xiaojian_Ma1","~Hai_Ci1","~Yue_Fan2","~Yuxuan_Wang4","~Haozhe_Zhao1","~Qing_Li1","~Yizhou_Wang1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Rujie Wu","Xiaojian Ma","Hai Ci","Yue Fan","Yuxuan Wang","Haozhe Zhao","Qing Li","Yizhou Wang"]}},"version":2},{"content":{"summary":{"value":"This paper proposes CSVC (Causal Steering for Video Counterfactuals), a prompt-engineering framework that combines structural causal models (SCMs) with latent diffusion models (LDMs) for generating video counterfactuals. \nThe approach uses LLMs to generate counterfactual text prompts given a predefined causal graph, LDM-based video editors to generate corresponding videos, and a VLM-based textual loss (optimized via TextGrad) to refine prompts.\nThe framework is evaluated on 67 text-video pairs from CelebV-Text.\n\n**Disclaimer**: I don't have much experience in video editing, and therefore I will focus on evaluating the causal foundation and rigor of their methodology and might miss understanding of comparisons with existing literature."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"The current citation style makes it hard to distinguish normal text and citations."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":1},"strengths":{"value":"1. Applying causal reasoning to video generation is valuable and underexplored.\n2. The framework in Figure 2 is practical and easy to implement. It might be useful to help improve video-editing models (though I have doubts about whether it would really work as will be explained below).\n3. Despite some structural issues (why put 4.2 and 4.3 in methodology section?), the paper is overall easy to follow."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"**Causality and positioning** \n1. The causal framework is overclaimed. The text is saturated with SCM/causality-related language. However, the causal graph has a minimal role as far as I understand. It is essentially used as an input to the LLM to help generate the counterfactual prompt, which has minimal interaction with the LDM.\n2. Following this concern, I wonder what would happen if we did not include the causal graph and just let the LLM edit the prompt, keeping everything else the same.\n3. The framework assumes access to a causal graph, which limits its applicability.\n4. Line 97 – the authors claim that \"this stands in contrast to approaches based on attention engineering, which offer suboptimal solutions.\" I don't understand why those methods offer suboptimal solutions compared to the proposed method. The authors fail to justify this either from a theoretical perspective or an empirical perspective.\n\n**Other technical concerns**\n1. The authors use a VLM to compute \"textual loss\" to update the prompt. The authors also use VLMs for final evaluations. Is the same VLM used in both cases? If so, I think the comparison is not fair.\n2. I don't understand why the proposed framework would help with temporal coherence. Essentially, the counterfactual video is generated using an edited prompt. I am not sure why this helps with temporal consistency.\n3. Minimality is proposed for visual counterfactuals. However, the authors measured it in the text domain.\n\n**Empirical study**\n1. Could the authors explain why they exclude comparison against Video-P2P and FateZero. They claim that those methods require identical source and edited prompt structures, but I don't understand why that is a problem.\n2. The dataset seems small-scale and limited. Why don't the authors compare on the datasets the baselines used (DAVIS)?"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762933922935,"tcdate":1761965948571,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission20489/Reviewer_vXJD"],"signatures":["ICLR.cc/2026/Conference/Submission20489/Reviewer_vXJD"],"forum":"2nBjmUB5PM","number":4,"license":"CC BY 4.0","cdate":1761965948571,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission20489/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762933922935,"domain":"ICLR.cc/2026/Conference","replyto":"2nBjmUB5PM","id":"kueMXwFs1D","forumContent":{"TLDR":{"value":"A causal framework for counterfactual video generation, guided by a vision-language model (VLM)"},"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["counterfactual generation","causality","generative AI","diffusion models","VLMs"]},"supplementary_material":{"value":"/attachment/794ce3d2bc7ff1f831b2233210147c12c8f2d760.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Adapting text-to-image (T2I) latent diffusion models (LDMs) to video editing has shown strong visual fidelity and controllability, but challenges remain in maintaining causal relationships inherent to the video data generating process. In this work, we propose CSVC, a framework for counterfactual video generation grounded in structural causal models (SCMs) and formulated as an out-of-distribution (OOD) prediction task. CSVC builds on black-box counterfactual functions, which approximate SCM mechanisms without explicit structural equations. In our framework, large language models (LLMs) generate counterfactual prompts that are consistent with a predefined causal graph, while LDM-based video editors produce the corresponding video counterfactuals. To ensure faithful interventions, we introduce a vision–language model (VLM)-based textual loss that refines prompts to enforce counterfactual conditioning, steering the LDM latent space toward causally meaningful OOD variations without internal model access or fine-tuning. Experiments on real-world facial videos show that CSVC achieves state-of-the-art causal effectiveness while preserving temporal consistency and visual quality. By combining SCM reasoning with black-box generative models, CSVC enables realistic “what if” hypothetical video scenarios with applications in digital media and healthcare."},"_bibtex":{"value":"@misc{\nspyrou2025causally,\ntitle={Causally Steered Diffusion for Video Counterfactual Generation},\nauthor={Nikos Spyrou and Athanasios Vlontzos and Paraskevas Pegios and Thomas Melistas and Nefeli Gkouti and Yannis Panagakis and Giorgos Papanastasiou and Sotirios A. Tsaftaris},\nyear={2025},\nurl={https://openreview.net/forum?id=2nBjmUB5PM}\n}"},"title":{"value":"Causally Steered Diffusion for Video Counterfactual Generation"},"pdf":{"value":"/pdf/e68e9a649754408dc83b3556197f3d7edb923ce0.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"spyrou|causally_steered_diffusion_for_video_counterfactual_generation"},"authorids":{"value":["~Nikos_Spyrou1","~Athanasios_Vlontzos1","~Paraskevas_Pegios1","~Thomas_Melistas1","~Nefeli_Gkouti1","~Yannis_Panagakis1","~Giorgos_Papanastasiou2","~Sotirios_A._Tsaftaris1"]},"authors":{"value":["Nikos Spyrou","Athanasios Vlontzos","Paraskevas Pegios","Thomas Melistas","Nefeli Gkouti","Yannis Panagakis","Giorgos Papanastasiou","Sotirios A. Tsaftaris"]}},"version":2},{"content":{"venue":{"value":"CoRR 2025"},"pdf":{"value":"https://arxiv.org/pdf/2502.19737v2"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"he|xcomps_a_multilingual_benchmark_of_conceptual_minimal_pairs"},"html":{"value":"https://doi.org/10.48550/arXiv.2502.19737"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2502-19737,\n  publtype={informal},\n  author={Linyang He and Ercong Nie and Sukru Samet Dindar and Arsalan Firoozi and Adrian Florea and Van Nguyen and Corentin Puffay and Riki Shimizu and Haotian Ye and Jonathan Brennan and Helmut Schmid and Hinrich Schütze and Nima Mesgarani},\n  title={XCOMPS: A Multilingual Benchmark of Conceptual Minimal Pairs},\n  year={2025},\n  month={February},\n  cdate={1738368000000},\n  journal={CoRR},\n  volume={abs/2502.19737},\n  url={https://doi.org/10.48550/arXiv.2502.19737}\n}\n"},"abstract":{"value":"We introduce XCOMPS in this work, a multilingual conceptual minimal pair dataset covering 17 languages. Using this dataset, we evaluate LLMs' multilingual conceptual understanding through metalinguistic prompting, direct probability measurement, and neurolinguistic probing. By comparing base, instruction-tuned, and knowledge-distilled models, we find that: 1) LLMs exhibit weaker conceptual understanding for low-resource languages, and accuracy varies across languages despite being tested on the same concept sets. 2) LLMs excel at distinguishing concept-property pairs that are visibly different but exhibit a marked performance drop when negative pairs share subtle semantic similarities. 3) Instruction tuning improves performance in concept understanding but does not enhance internal competence; knowledge distillation can enhance internal competence in conceptual understanding for low-resource languages with limited gains in explicit task performance. 4) More morphologically complex languages yield lower concept understanding scores and require deeper layers for conceptual reasoning."},"title":{"value":"XCOMPS: A Multilingual Benchmark of Conceptual Minimal Pairs"},"authors":{"value":[{"fullname":"Linyang He","username":""},{"fullname":"Ercong Nie","username":""},{"fullname":"Sukru Samet Dindar","username":"~Sukru_Samet_Dindar1"},{"fullname":"Arsalan Firoozi","username":""},{"fullname":"Adrian Florea","username":""},{"fullname":"Van Nguyen","username":""},{"fullname":"Corentin Puffay","username":""},{"fullname":"Riki Shimizu","username":""},{"fullname":"Haotian Ye","username":""},{"fullname":"Jonathan Brennan","username":""},{"fullname":"Helmut Schmid","username":""},{"fullname":"Hinrich Schütze","username":""},{"fullname":"Nima Mesgarani","username":""}]}},"tmdate":1785774648894,"pdate":1767139200000,"externalIds":["dblp:journals/corr/abs-2502-19737"],"tcdate":1785774642912,"writers":["~","OpenReview.net/Public_Article/DBLP.org","OpenReview.net/Support"],"signatures":["~Sukru_Samet_Dindar1"],"forum":"lZfD59lWOo","license":"CC BY-SA 4.0","number":114185,"cdate":1738368000000,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/DBLP.org/-/Record","OpenReview.net/Public_Article/-/Edit"],"mdate":1785774648894,"domain":"OpenReview.net/Public_Article","id":"lZfD59lWOo","version":2},{"content":{"summary":{"value":"This paper proposes InfoAug, a contrastive learning framework that employs a novel way of selecting positive pairs. Besides positive pairs that are generated by augmentations, InfoAug also selects the \"twin patch\" which maximizes the mutual information of the original patch as another positive patch. Experiments show that when trained on a video dataset and evaluated on CIFAR-10, STL-10, and CIFAR-100, the proposed InfoAug brings 1-2% improvement over existing contrastive learning methods."},"presentation":{"value":"2 fair"},"contribution":{"value":"1 poor"},"soundness":{"value":"1 poor"},"strengths":{"value":"This paper proposes a novel way of selecting positive pairs in contrastive learning, which is to select a twin patch that maximizes mutual information between the two pairs. Such a design proposes a new direction in utilizing video datasets in contrastive learning."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Overall, the major problem of this paper is the weak experimental results. The experiment results are weak in two aspects:\n1. The accuracy of baselines on evaluated datasets is low. For example, when trained and tested on CIFAR-10, SimCLR can achieve >90% accuracy. However, in this paper, when trained on DAVIS20+GMOT40 and tested on CIFAR-10, the performance is only 60%-70%. I would suggest including CIFAR-10 as a pseudo-video dataset in the pre-training stage so that the paper can make a fair comparison with current methods on these datasets.\n2. The improvement of the proposed method is marginal. For example, when the accuracy on CIFAR-10 is 60%-70%, the standard deviation can be large, but the proposed method only improves the performance by 1-2% on each dataset. This makes the improvement shown in the paper not convincing enough.\n3. The paper does not include an analysis of computation overhead over baseline methods. Estimating the mutual information between two patches can introduce some computation overheads, making the framework slower than baseline methods.\n\nThe presentation of the paper is also not very clear and needs further polishing. There are multiple typos and grammar errors, and the font in the figures is too small."},"confidence":{"value":"4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work."},"questions":{"value":"See weaknesses"},"rating":{"value":"3: reject, not good enough"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636465957,"tcdate":1699117438669,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission4826/Reviewer_LYQ4"],"signatures":["ICLR.cc/2024/Conference/Submission4826/Reviewer_LYQ4"],"forum":"7GkdjhupsV","number":4,"license":"CC BY 4.0","cdate":1699117438669,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission4826/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636465957,"domain":"ICLR.cc/2024/Conference","replyto":"7GkdjhupsV","id":"YEKRoC1Htk","forumContent":{"venue":{"value":"Submitted to ICLR 2024"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["representation learning","mutual information","data augmentation"]},"primary_area":{"value":"unsupervised, self-supervised, semi-supervised, and supervised representation learning"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Representation learning methods utilizing the InfoNCE loss have demonstrated considerable capacity in reducing human annotation effort by training invariant neural feature extractors. Although different variants of the training objective adhere to the information maximization principle between the data and learned features, data selection and augmentation still rely on human hypotheses or engineering, which may be suboptimal. For instance, data augmentation in contrastive learning primarily focuses on color jittering, aiming to emulate real-world illumination changes. In this work, we investigate the potential of selecting training data based on their mutual information computed from real-world distributions, which, in principle, should endow the learned features with better generalization when applied in open environments. Specifically, we consider patches attached to scenes that exhibit high mutual information under natural perturbations, such as color changes and motion, as positive samples for learning with contrastive loss. We evaluate the proposed mutual-information-informed data augmentation method on several benchmarks across multiple state-of-the-art representation learning frameworks, demonstrating its effectiveness and establishing it as a promising direction for future research. The data and code will be available for further investigation."},"_bibtex":{"value":"@misc{\nchen2024infoaug,\ntitle={InfoAug: Mutual Information Informed Augmentation for Representation Learning},\nauthor={Hanyang Chen and Qingyuan Zheng and YANG ZONGRU and Yanchao Yang},\nyear={2024},\nurl={https://openreview.net/forum?id=7GkdjhupsV}\n}"},"title":{"value":"InfoAug: Mutual Information Informed Augmentation for Representation Learning"},"pdf":{"value":"/pdf/9efd9ec88b4702f88131fd1ecb2dfd6e9b5c509e.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Rejected_Submission"},"paperhash":{"value":"chen|infoaug_mutual_information_informed_augmentation_for_representation_learning"},"authorids":{"value":["~Hanyang_Chen2","~Qingyuan_Zheng1","~YANG_ZONGRU2","~Yanchao_Yang1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Hanyang Chen","Qingyuan Zheng","YANG ZONGRU","Yanchao Yang"]}},"version":2},{"content":{"summary":{"value":"This paper introduces Evidence-Gated Suppression (EGS), a lightweight, plug-in regularizer designed to combat shortcut learning in deep models without requiring group labels. EGS operates inside the network during training by tracking a class-conditional, confidence-weighted \"evidence energy\" for each neuron to identify which neurons contribute most strongly to the model's predictions. It then applies a percentile-based multiplicative decay to the weights of these extreme contributors, selectively suppressing overconfident shortcut pathways while leaving more robust features relatively influential. The authors demonstrate across several spurious correlation benchmarks (such as Waterbirds and CelebA) that EGS improves worst-group accuracy and calibration, achieving competitive performance with state-of-the-art methods while maintaining strong average accuracy and adding minimal training overhead."},"soundness":{"value":4},"confidence":{"value":3},"questions":{"value":"See weaknesses."},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":4},"strengths":{"value":"While many methods address spurious correlations by reweighting data (e.g., JTT, LfF) or modifying the loss function (e.g., GroupDRO), EGS introduces a fundamentally different locus of intervention: the internal evidence pathways of the network itself. The concept of a dynamic, training-time, *neuron-level* regularizer that is both class-conditional and group-agnostic is original. The \"evidence energy\" metric, which combines model confidence (`pk(x)`) with feature-weight alignment (`Wjk * φj(x)`), provides a simple yet powerful signal for identifying and suppressing over-reliant pathways."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The central heuristic of EGS is that the most negative (i.e., highest-confidence, class-aligned) evidence corresponds to spurious features. While this holds true in many shortcut-learning scenarios, it is not guaranteed. A genuinely robust and highly discriminative feature could also consistently produce very strong evidence and be inadvertently suppressed by the percentile gate. The paper would be strengthened by an analysis that more directly validates this core assumption. For example, the authors could run EGS on a dataset known to have minimal spurious correlations (e.g., a balanced version of a dataset or even standard CIFAR-10) to demonstrate that the method does not harm performance when strong shortcuts are absent."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762917553623,"tcdate":1761969473689,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission4753/Reviewer_2uU2"],"signatures":["ICLR.cc/2026/Conference/Submission4753/Reviewer_2uU2"],"forum":"L2L1hi0FGj","number":4,"license":"CC BY 4.0","cdate":1761969473689,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission4753/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762917553623,"domain":"ICLR.cc/2026/Conference","replyto":"L2L1hi0FGj","id":"zSBZSV9dBr","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Fairness","Regularization","Bias Free"]},"primary_area":{"value":"alignment, fairness, safety, privacy, and societal considerations"},"abstract":{"value":"Deep models often exploit spurious correlations (e.g., backgrounds or dataset artifacts), hurting worst-group performance. We propose \\textbf{Alignment-Gated Suppression (AGS)}, a lightweight, plug-in regularizer that intervenes inside the network during training. AGS tracks a class-conditional, confidence-weighted contribution for each neuron (more negative $\\Leftrightarrow$ stronger support) and applies a percentile-based, multiplicative decay to the most extreme contributors, reducing overconfident shortcut pathways while leaving other features relatively more influential. AGS integrates with standard ERM, requires no group labels, and adds $<5\\%$ training overhead. We provide analysis linking AGS to minority-margin gains, path-norm-like capacity control, and stability benefits via EMA-smoothed gating. Empirically, AGS improves worst-group accuracy and calibration vs.\\ ERM and is competitive with state-of-the-art methods across spurious-correlation benchmarks (e.g., Waterbirds, CelebA, BAR, COCO), while maintaining strong average accuracy. These results suggest that regulating internal alignment flow is a simple and scalable route to robustness without group labels."},"_bibtex":{"value":"@inproceedings{\ndwivedi2026regulating,\ntitle={Regulating Internal Alignment Flows for Robust Learning Under Spurious Correlations},\nauthor={Rajeev Ranjan Dwivedi and Mohammedkaif Mohammedrafiq Kalagond and Niramay M.Patel and Vinod K. Kurmi},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=L2L1hi0FGj}\n}"},"title":{"value":"Regulating Internal Alignment Flows for Robust Learning Under Spurious Correlations"},"pdf":{"value":"/pdf/b3c72ec1d354f283ae48dd0695ce72595743f706.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"dwivedi|regulating_internal_alignment_flows_for_robust_learning_under_spurious_correlations"},"authorids":{"value":["~Rajeev_Ranjan_Dwivedi1","~Mohammedkaif_Mohammedrafiq_Kalagond1","~Niramay_M.Patel1","~Vinod_K._Kurmi1"]},"authors":{"value":["Rajeev Ranjan Dwivedi","Mohammedkaif Mohammedrafiq Kalagond","Niramay M.Patel","Vinod K. Kurmi"]}},"version":2},{"content":{"summary":{"value":"The paper introduces Visionary-R1, a reinforcement learning (RL) framework for training visual-language models (VLMs) on visual reasoning tasks without any chain-of-thought (CoT) supervision. Inspired by DeepSeek-R1 and GRPO, the authors identify that direct RL fine-tuning on question–answer pairs causes shortcut learning, where the model overfits to easy samples by ignoring visual grounding. To mitigate this, Visionary-R1 enforces a caption–reason–answer output format, requiring the model to first describe the image (caption), then reason, then answer. The method introduces a caption reward (from AI feedback) to ensure informative visual grounding and applies cosine-annealed KL regularization to stabilize training."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"See Weaknesses"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1.The paper clearly identifies a practical failure mode — shortcut learning in visual RL — and provides strong empirical evidence of this phenomenon.\n2. The caption–reason–answer structure and caption reward are conceptually simple but yield measurable generalization improvements.\n3. The methodology is well-detailed, including architecture, rewards, and prompt templates. The authors commit to releasing code and models."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The caption reward relies on the model’s own LLM component to verify if the caption enables correct answering — this could lead to reward leakage or self-confirmation bias. There’s no analysis of how often the reward misfires.\n2. While outperforming on MathVista/MMBench, the improvement on some datasets (e.g., MathVision, MMStar) is modest, suggesting the gains may stem from formatting or stylistic changes rather than deeper reasoning.\n3. Longer outputs correlate with accuracy, but this metric may simply reflect verbosity rather than actual interpretive reasoning. There’s no human evaluation or visual attention analysis to confirm genuine grounding.\n4. The paper claims to “mitigate shortcut learning,” but does not quantify shortcut severity reduction.\n5. While the paper includes a few ablations (e.g., adding caption and caption reward), it does not disentangle all design contributions. Important components like the cosine-annealed KL penalty, format reward, or individual reward weights are not independently analyzed. Moreover, the ablations are limited to small subsets, making the conclusions less generalizable. The study also lacks statistical variance or multiple runs, so the reported gains may not be robust.\n6. Missing zero-shot caption–reason baseline: The paper claims that enforcing a caption–reason–answer structure via RL mitigates shortcut learning, but it does not include a simple zero-shot or instruction-tuned baseline where the base model is merely prompted to follow the same format without RL fine-tuning. Such a baseline would clarify whether the improvement actually comes from the reinforcement learning signal or simply from the structured prompting itself. Without this comparison, the central claim—that Visionary-R1 develops reasoning rather than format bias—remains unproven.\n7. Lack of quantitative evidence for grounding: The paper qualitatively argues that Visionary-R1 promotes visual grounding, but there is no metric measuring this (e.g., visual attention analysis, faithfulness, or caption relevance). The evaluation is purely based on accuracy, leaving it unclear whether the gains reflect genuine reasoning or just longer, well-structured responses."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762925261146,"tcdate":1761997223339,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission14918/Reviewer_i5cF"],"signatures":["ICLR.cc/2026/Conference/Submission14918/Reviewer_i5cF"],"forum":"bya3KOdLeS","number":4,"license":"CC BY 4.0","cdate":1761997223339,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission14918/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762925261146,"domain":"ICLR.cc/2026/Conference","replyto":"bya3KOdLeS","id":"Koo2RJFTEP","forumContent":{"TLDR":{"value":"We addressed a shortcut problem in applying reinforcement learning to VLMs with Visionary-R1, trained on 273K CoT-free visual question-answer pairs  using only reinforcement learning."},"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Large Vision Language Model","Visual Reasoning","Reinforcement Learning"]},"primary_area":{"value":"foundation or frontier models, including LLMs"},"abstract":{"value":"Learning general-purpose reasoning capabilities has long been a challenging problem in AI. Recent research in large language models (LLMs), such as DeepSeek-R1, has shown that reinforcement learning techniques like GRPO can enable pre-trained LLMs to develop reasoning capabilities using simple question-answer pairs. In this paper, we aim to train visual language models (VLMs) to perform reasoning on image data through reinforcement learning and visual question-answer pairs, without any explicit chain-of-thought (CoT) supervision. Our findings indicate that simply applying reinforcement learning to a VLM---by prompting the model to produce a reasoning chain before providing an answer---can lead the model to develop shortcuts from easy questions, thereby reducing its ability to generalize across unseen data distributions. We argue that the key to mitigating shortcut learning is to encourage the model to interpret images prior to reasoning. Therefore, we train the model to adhere to a caption-reason-answer output format: initially generating a detailed caption for an image, followed by constructing an extensive reasoning chain. When trained on 273K CoT-free visual question-answer pairs and using only reinforcement learning, our model, named Visionary-R1, outperforms strong multimodal models, such as GPT-4o, Claude3.5-Sonnet, and Gemini-1.5-Pro, on multiple visual reasoning benchmarks. Code and models will be publicly released."},"_bibtex":{"value":"@misc{\nxia2025visionaryr,\ntitle={Visionary-R1: Mitigating Shortcuts in Visual Reasoning with Reinforcement Learning},\nauthor={Jiaer Xia and Yuhang Zang and Peng Gao and Yixuan Li and Kaiyang Zhou},\nyear={2025},\nurl={https://openreview.net/forum?id=bya3KOdLeS}\n}"},"title":{"value":"Visionary-R1: Mitigating Shortcuts in Visual Reasoning with Reinforcement Learning"},"pdf":{"value":"/pdf/3dd28e833f99b895e121c749459d21c527822646.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"xia|visionaryr1_mitigating_shortcuts_in_visual_reasoning_with_reinforcement_learning"},"authorids":{"value":["~Jiaer_Xia1","~Yuhang_Zang1","~Peng_Gao3","~Yixuan_Li1","~Kaiyang_Zhou1"]},"authors":{"value":["Jiaer Xia","Yuhang Zang","Peng Gao","Yixuan Li","Kaiyang Zhou"]}},"version":2},{"content":{"summary":{"value":"This paper introduces an approach, Dual Temporal Adjacent Maps (DTAM), for enhancing video moment retrieval. DTAM separates visual appearance and semantic information, addressing issues in current methods that struggle to distinguish similar-looking moments with different meanings. DTAM uses two branches to encode visual and semantic features, with the appearance branch feeding signals to the semantic branch to improve differentiation. Additionally, a moment-aware mechanism is developed to optimize the model’s attention to relevant moments. Experiments on three video retrieval benchmarks demonstrate DTAM’s superior performance, highlighting its effectiveness in capturing complex temporal and semantic dependencies."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"Given DTAM’s dual-branch structure and additional moment-aware mechanism, which add complexity, what specific optimizations or architectural choices contribute to its minimal increase in inference time compared to simpler models like 2D-TAN?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. DTAM’s approach of separating visual appearance and semantic information addresses a key challenge in video retrieval, allowing the model to distinguish moments that look similar but have different meanings.\n2. The moment-aware mechanism in DTAM enhances the model’s sensitivity to specific moments by dynamically focusing on relevant segments.\n3. The paper presents extensive experiments on three challenging benchmarks, where DTAM consistently outperforms state-of-the-art methods."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"While DTAM achieves impressive retrieval accuracy, its dual-branch structure and moment-aware mechanism may increase computational demands. This complexity could limit its scalability and efficiency when applied to large-scale video datasets."}},"nonreaders":[],"tmdate":1731428021340,"tcdate":1730166487188,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission4047/Reviewer_HBE6"],"signatures":["ICLR.cc/2025/Conference/Submission4047/Reviewer_HBE6"],"forum":"l3CSCOnGPB","number":1,"license":"CC BY 4.0","cdate":1730166487188,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission4047/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428021340,"domain":"ICLR.cc/2025/Conference","replyto":"l3CSCOnGPB","id":"vrKm3u4ciK","forumContent":{"TLDR":{"value":"a novel semantic-enhanced dual temporal adjacent maps (DTAM) for effective video moment retrieval, which models temporal dependencies between moments in an appearance-semantic decoupled fashion."},"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Computer Vision","Muylti-modal Understanding","Video Grounding"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Retrieving a specific moment from an untrimmed video via a text description is a central problem in vision-language learning. It is a challenging task due to the sophisticated temporal dependency among moments. Existing methods fail to deal with this issue well since they establish temporal relations of moments in a way that visual content and semantics are coupled. This paper studies temporal dependence schemes that decouple content and semantic information, establishing semantic-enhanced Dual Temporal Adjacent Maps for video moment retrieval, conferred as DTAM. Specifically, DTAM designs two branches to encode visual appearance and semantic knowledge from video clips respectively, where knowledge from the appearance branch is distilled into the semantic branch to help DTAM distinguish features with the same visual content but different semantics with a well-designed semantic-aware contrastive loss. Besides, we also develop a moment-aware mechanism to assist temporal adjacent maps' learning for better video grounding. Finally, extensive experimental results and analysis demonstrate the superiority of the proposed DTAM over existing state-of-the-art approaches on three challenging video moment retrieval benchmarks, i.e., TACoS, Charades-STA, and ActivityNet Captions."},"_bibtex":{"value":"@misc{\nwang2025learning,\ntitle={Learning Semantic-Enhanced Dual Temporal Adjacent Maps for Video Moment Retrieval},\nauthor={Yu Wang and Shengjie Zhao and Shiwei Chen},\nyear={2025},\nurl={https://openreview.net/forum?id=l3CSCOnGPB}\n}"},"title":{"value":"Learning Semantic-Enhanced Dual Temporal Adjacent Maps for Video Moment Retrieval"},"pdf":{"value":"/pdf/83e3b8d8b0191f0c4f665fc641a4a682d0f67e99.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"wang|learning_semanticenhanced_dual_temporal_adjacent_maps_for_video_moment_retrieval"},"authorids":{"value":["~Yu_Wang32","~Shengjie_Zhao1","~Shiwei_Chen3"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Yu Wang","Shengjie Zhao","Shiwei Chen"]}},"version":2},{"content":{"summary":{"value":"This paper proposes an efficient LMM with minimal vision tokens. 1/ To achieve a high compression ratio of vision tokens while preserving visual information, the authors first analyze how LMMs understand vision tokens and find that most vision tokens only play a crucial role in the early layers, where they fuse visual information into text tokens. 2/ LLaVA-Mini introduces a novel modality pre-fusion to fuse visual information into text tokens in advance before feeding into LLM and a compression module to reduce #vision tokens into minimal ones. 3/ Experiments across 11 image-based and 7 video-based benchmarks demonstrate that LLaVA-Mini outperforms LLaVA-v1.5 with just 1 vision token instead of 576. Efficiency analyses reveal that LLaVA-Mini can reduce FLOPs by 77%, deliver low-latency responses within 40 milliseconds, and process over 10,000 frames of video on GPU hardware with 24GB of memory."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"Kindly refer to the weakness sec above"},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":4},"strengths":{"value":"1/ The problem of how to improve efficiency of MLLM is important. \n\n2/ The authors conducted a very interesting and detailed analysis of how LMMs understand vision tokens and find that most vision tokens only play a crucial role in the early layers, where they fuse visual information into text tokens. \n\n3/ LLaVA-Mini introduces a modality pre-fusion module to fuse visual information into text tokens in advance before feeding into LLM and a compression module to reduce #vision tokens into minimal ones. \n\n4/ Experimental results are strong and convincing."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1/ According to Sec 3, if I understand correctly, for each downstream task, this paper trains a new, independent compression and modality pre-fusion modules. In other words, for another task, we need to train another model. I wonder, how does this approach differ from task-specific distillation? It seems the reason why the efficiency can be improved significantly is because of removing unnecessary tokens for the target downstream task. Instead, would it be possible to train generic compression and modality pre-fusion modules in stage 1 that can be generalised to various downstream tasks?\n\n2/ For the proposed approach, a new modality pre-fusion module has been introduced. I wonder if this is necessary. Instead, can we re-use the first few layers, which have high activation for vision tokens, as the fusion module. To further include the compression module, in the first few layers of LLM, the part digesting vision tokens can have two pathways, one pathway is to output the compressed tokens while the other pathway is to be gradually fused together with the text tokens.  \n\n3/ When it comes to video, there are some recent streaming video LLM works such as \"VideoLLM-online: Online Video Large Language Model for Streaming Video. CVPR 2024\" which also advocates the idea of representing each frame with minimal vision tokens to improve efficiency. It might be worthwhile to compare or just discuss."}},"nonreaders":[],"tmdate":1731428799900,"tcdate":1729329142593,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission13752/Reviewer_CwER"],"signatures":["ICLR.cc/2025/Conference/Submission13752/Reviewer_CwER"],"forum":"UQJ7CDW8nb","number":1,"license":"CC BY 4.0","cdate":1729329142593,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission13752/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428799900,"domain":"ICLR.cc/2025/Conference","replyto":"UQJ7CDW8nb","id":"5W5uAaTJQs","forumContent":{"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Large Multimodal Models","Large Language Models"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"The advent of real-time large multimodal models (LMMs) like GPT-4o has sparked considerable interest in efficient LMMs. LMM frameworks typically encode visual inputs into vision tokens (continuous representations) and integrate them and textual instructions into the context of large language models (LLMs), where large-scale parameters and numerous context tokens (predominantly vision tokens) result in substantial computational overhead. Previous efforts towards efficient LMMs always focus on replacing the LLM backbone with smaller models, while neglecting the crucial issue of token quantity. In this paper, we introduce LLaVA-Mini, an efficient LMM with minimal vision tokens. To achieve a high compression ratio of vision tokens while preserving visual information, we first analyze how LMMs understand vision tokens and find that most vision tokens only play a crucial role in the early layers of LLM backbone, where they mainly fuse visual information into text tokens. Building on this finding, LLaVA-Mini introduces modality pre-fusion to fuse visual information into text tokens in advance, thereby facilitating the extreme compression of vision tokens fed to LLM backbone into one token. LLaVA-Mini is a unified large multimodal model that can support the understanding of images, high-resolution images, and videos in an efficient manner. Experiments across 11 image-based and 7 video-based benchmarks demonstrate that LLaVA-Mini outperforms LLaVA-v1.5 with just 1 vision token instead of 576. Efficiency analyses reveal that LLaVA-Mini can reduce FLOPs by 77%, deliver low-latency responses within 40 milliseconds, and process over 10,000 frames of video on the GPU hardware with 24GB of memory."},"_bibtex":{"value":"@inproceedings{\nzhang2025llavamini,\ntitle={{LL}a{VA}-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token},\nauthor={Shaolei Zhang and Qingkai Fang and Zhe Yang and Yang Feng},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=UQJ7CDW8nb}\n}"},"title":{"value":"LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token"},"pdf":{"value":"/pdf/efd2169a71f1800808f58038f0bf1023ce051103.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"zhang|llavamini_efficient_image_and_video_large_multimodal_models_with_one_vision_token"},"authorids":{"value":["~Shaolei_Zhang1","~Qingkai_Fang1","~Zhe_Yang7","~Yang_Feng4"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Shaolei Zhang","Qingkai Fang","Zhe Yang","Yang Feng"]}},"version":2},{"content":{"summary":{"value":"The paper proposes a novel method to perform instruction tuning of Large Multi-modal Models (LMMs) on labeled video datasets, where the label format is arbitrary and different from that needed for instruction tuning. The main idea is to use the original LMM itself to generate question-answer pairs for instruction tuning, while using the original dataset labels as a \"verifier\" to filter out low-quality generations. The authors perform multiple cycles of dataset generation and model training to obtain the final model Video-STAR. Video-STAR shows improved performance on TempCompass and zero-shot video QA datasets."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"* In Tables 6 and 7, the last row \"- Generation\" means that neither generation nor rationalization are used, or only that generation is not used while rationalization is used? Either way, it is important to see both experiments to better understand the contribution of each method.\n* When doing self-training in multiple cycles, every time the model is initialized from the same non-instruction-finetuned checkpoint. Why not reuse the model from the previous cycle of self-training as initialization? \n* How do the authors deal with numerical labels such as temporal localization, bounding box, performance score? Do they expect the LVM to predict those directly and compute the L1 distance in the verifier? If so, how accurate is the prediction of the model in this case? Is Answer Generation capable of generating good question-answer pairs or Label Rationalization is more helpful here?"},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"* The idea of turning existing labeled video datasets into instruction tuning dataset and use them for self-training is elegant and effective.\n* The paper is well written and is easy to follow\n* I appreciate the ablation study"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"* The main weakness of the paper is the Video-LLaVA baselines, i.e. Video-LLaVA+ and Video-LLaVA-gemini, which are the main points of comparison in the paper. Video-LLaVA-gemini has been fine-tuned only on several thousand sample pairs, which is incomparable to hundreds of thousands of video samples used for Video-STAR. As for Video-LLaVA+, the number of training samples is comparable, yet there is no information about how those samples were obtained. The way the samples in Video-LLaVA+ are constructed from the original video labels would affect the final model performance a lot, so it is important to explain this in details.\n* Additionally, one important baseline for dataset creation is missing. It looks like the SoTA method for creating instruction datasets from video is VideoInstruct [A]. I would expect the authors to compare to this method of dataset creation, both in answer generation and label rationalization settings, to understand how much the video stream actually helps.\n* It is somewhat unclear from the paper if Label Rationalization can be used to generate the entire dataset, without using Answer Generation. Somewhere in the paper the authors suggest that it leads to hallucinations, yet in Tables 6 and 7 they show the improvement.  \n\nThis is a preliminary rating and I will revise my score once the weaknesses and questions are addressed.\n\n[A] -  Maaz et al. \"Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models\""}},"nonreaders":[],"tmdate":1732209989883,"tcdate":1730593037886,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission826/Reviewer_ke7n"],"signatures":["ICLR.cc/2025/Conference/Submission826/Reviewer_ke7n"],"forum":"JYV2hrtFSv","number":1,"license":"CC BY 4.0","cdate":1730593037886,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission826/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732209989883,"domain":"ICLR.cc/2025/Conference","replyto":"JYV2hrtFSv","id":"7nuR8mR21Z","forumContent":{"venue":{"value":"ICLR 2025 Poster"},"TLDR":{"value":"Introducing the first self-training method for video-LMMs"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video Understanding","Visual Instruction Tuning","Self-Training","Chain-of-thought reasoning"]},"supplementary_material":{"value":"/attachment/9c72d32e9bce4627af3727c58ce7afee71870795.pdf"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"The performance and reasoning capabilities of Large Multi-modal Models (LMMs) is dependent on the size and quality of their training datasets. However, collecting datasets that support chain-of-thought instruction tuning is highly challenging. Existing video instruction tuning datasets are often derived by prompting large language models with video captions to generate question-answer pairs, which makes them predominantly descriptive rather than reasoning-focused. \nMeanwhile, many labeled video datasets with diverse labels and supervision exist -- however, we find that their integration into LMMs is non-trivial. \nHerein, we present $\\underline{\\text{Video}}$ $\\underline{\\text{S}}\\text{elf}$-$\\underline{\\text{T}}\\text{raining}$ $\\text{with}$ $\\underline{\\text{a}}\\text{ugmented}$ $\\underline{\\text{R}}\\text{easoning}$ (Video-STaR), the first self-training approach for video instruction tuning. \nVideo-STaR allows the utilization of *any* labeled video dataset for video instruction tuning.\nIn Video-STaR, an LMM cycles between instruction generation and finetuning, which we show (I) improves general video understanding and (II) adapts LMMs to novel downstream tasks with existing supervision. \nDuring instruction generation, an LMM is prompted to propose an answer.  The answers are then filtered only to those that contain the original video labels, and the LMM is then re-trained on the generated dataset. \nBy training exclusively on generated answers containing the correct video labels, Video-STaR leverages these existing labels as weak supervision for video instruction tuning.\nOur results demonstrate that Video-STaR-augmented LMMs achieve notable improvements in (I) general Video QA, where TempCompass performance improved by 6.1%, *and* (II) downstream tasks, with a 9.9% increase in Kinetics700-QA accuracy and a 4.0% improvement in action quality assessment on FineDiving, while also exhibiting better interpretability."},"_bibtex":{"value":"@inproceedings{\nzohar2025videostar,\ntitle={Video-{ST}aR: Self-Training Enables Video Instruction Tuning with Any Supervision},\nauthor={Orr Zohar and Xiaohan Wang and Yonatan Bitton and Idan Szpektor and Serena Yeung-Levy},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=JYV2hrtFSv}\n}"},"title":{"value":"Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision"},"pdf":{"value":"/pdf/2dbd229d9e63576bc3553bb16d276747ace91914.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"zohar|videostar_selftraining_enables_video_instruction_tuning_with_any_supervision"},"authorids":{"value":["~Orr_Zohar1","~Xiaohan_Wang2","~Yonatan_Bitton1","~Idan_Szpektor1","~Serena_Yeung-Levy1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Orr Zohar","Xiaohan Wang","Yonatan Bitton","Idan Szpektor","Serena Yeung-Levy"]}},"version":2},{"content":{"summary":{"value":"This paper introduces the first large-scale video action differencing dataset, presenting a novel task of identifying differences between videos depicting the same action. The authors compile over 500 video pairs from existing datasets across five categories: Fitness, Ball Sports, Diving, Music, and Surgery. These videos are then assigned to annotators along with 147 distinct descriptions. Annotators must indicate which video (A or B) most closely aligns with each description. For example, given two videos of different actors performing a squat, a description might read \"deeper squat,\" and the annotator would select A or B based on which video demonstrates the deeper squat. To ensure dataset quality, 25% of the initial annotations undergo re-annotation, revealing a very low discrepancy rate. The dataset also includes action localization (pinpointing where the action occurs in the video) and specific key points for each action (e.g., when knees start to bend).\n\nThe authors also develop an agentic model called VidDiff to address the action differencing challenge. VidDiff employs several Large Language Models (LLMs) and Vision Language Models (VLMs) as agents to solve specific aspects of the problem: proposing potential differences based on the action description, localizing frames where such actions might occur, and finally specifying which video (A or B) corresponds to the observed difference. VidDiff outperforms other zero-shot VLMs in this task.\n\nLastly, the authors provide ablation experiments that highlight the challenges presented by their new benchmark."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"See weaknesses, plus the following:\n\n- Which LLMs/VLMs are used for the *Difference Proposer* and *Action Differencer*?\n- How does the benchmark handle cases of inverse correlation? For example, would *lower squat in video A* be equivalent to *higher squat in video B*?\n- Since the videos are not curated, factors such as different camera angles, varying FPS, or differences in the actor's height could introduce biases in the annotations and results. How do the authors address these potential biases?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":3},"strengths":{"value":"### Originality\n- **A novel task**: This paper introduces the new task of video action differencing with natural language. While related tasks, such as difference captioning, have been explored to provide a coarse comparison between videos, no prior work has tackled video action differencing in the same way—focusing on fine-grained differences described in natural language.\n- **A challenging benchmark**: The proposed benchmark, VidDiffBench, is comprehensive, covering five categories of instructional videos. It has proven to be highly challenging, even for top-performing closed-source vision-language models (VLMs).\n- **An agent-based system**: The paper presents an agent-based system that decomposes the task, achieving better performance than existing VLMs.\n### Clarity\nThe flow of ideas is straightforward, making the paper easy to follow and understand.\n\n### Significance\nThe paper convincingly demonstrates the importance of video action differencing, and the introduction of the new benchmark is likely to inspire further research in this area."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"### Unproven claims\n- In the introduction, the authors claim they will address the challenges of *precise temporal alignment and the need for fine-grained understanding of action dynamics*. However, it remains unclear how they specifically solve the issue of temporal alignment. Could you elaborate on how you solve this issue or point us to the location where it is addressed?\n\n### Benchmark and results\n- Similar datasets are presented in the related work section; however, since this work is primarily a benchmark paper, more comparisons with existing benchmarks would be make the differences clearer (e.g., similar to Table 1 but with other datasets in the first column). Consider adding what is unique about each dataset and how the current dataset differs.\n- As a benchmark paper, we would expect more results from other open-source VLMs (especially those addressing video data such as LLaVA-video) to better understand their limitations and make it easier for other researchers to work with this benchmark. \n\n### Clarity\n- 557 or 656 video pairs? In the abstract, the authors state that the dataset contains *557 video pairs...  4,719 fine-grained action differences* (line 013-014), but on line 260, they mention *656 video pairs, 5,580 annotated differences*. Clarification needed on which is correct. \n- Figure 1: The distinction between the first and second row is unclear, yet the caption claims these represent two different challenges. These two challenges are not discussed elsewhere in the paper and don't seem to be related to the dataset splits. Please clarify this."}},"nonreaders":[],"tmdate":1732602783813,"tcdate":1729222996269,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission14085/Reviewer_8a8a"],"signatures":["ICLR.cc/2025/Conference/Submission14085/Reviewer_8a8a"],"forum":"3bcN6xlO6f","number":1,"license":"CC BY 4.0","cdate":1729222996269,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission14085/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732602783813,"domain":"ICLR.cc/2025/Conference","replyto":"3bcN6xlO6f","id":"Qqzm3PvjdP","forumContent":{"TLDR":{"value":"A new task and benchmark for comparing how an action is performed between two videos, with a zero-shot method"},"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video","Actions","Differencing","Zero-shot","benchmark","multimodal","lmm","llm"]},"primary_area":{"value":"datasets and benchmarks"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"How do two individuals differ when performing the same action? In this work, we introduce Video Action Differencing (VidDiff), the novel task of identifying subtle differences between videos of the same action, which has numerous applications, such as coaching and skill learning. To enable development on this new task, we first create VidDiffBench, a benchmark dataset containing 549 video pairs, with human annotations of 4,469 fine-grained action differences and 2,075 timestamps indicating where these differences occur. Our experiments demonstrate that VidDiffBench poses a significant challenge for state-of-the-art large multimodal models (LMMs), such as GPT-4o and Qwen2-VL. By analyzing the failure cases of LMMs on VidDiffBench, we highlight two key challenges for this task: localizing relevant sub-actions over two videos and fine-grained frame comparison. To overcome these, we propose the VidDiff method, an agentic workflow that breaks the task into three stages: action difference proposal, keyframe localization, and frame differencing, each stage utilizing specialized foundation models. To encourage future research in this new task, we release the benchmark and code."},"_bibtex":{"value":"@inproceedings{\nburgess2025video,\ntitle={Video Action Differencing},\nauthor={James Burgess and Xiaohan Wang and Yuhui Zhang and Anita Rau and Alejandro Lozano and Lisa Dunlap and Trevor Darrell and Serena Yeung-Levy},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=3bcN6xlO6f}\n}"},"title":{"value":"Video Action Differencing"},"pdf":{"value":"/pdf/102482b5babaacddfd916de17bda7c15b2020db5.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"burgess|video_action_differencing"},"authorids":{"value":["~James_Burgess2","~Xiaohan_Wang2","~Yuhui_Zhang3","~Anita_Rau1","~Alejandro_Lozano1","~Lisa_Dunlap1","~Trevor_Darrell2","~Serena_Yeung-Levy1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["James Burgess","Xiaohan Wang","Yuhui Zhang","Anita Rau","Alejandro Lozano","Lisa Dunlap","Trevor Darrell","Serena Yeung-Levy"]}},"version":2},{"content":{"summary":{"value":"This work introduces a novel approach for video editing that includes both Video-to-Paragraph (V2P) and Paragraph-to-Video (P2V) methodologies. Initially, the V2P method converts video content into a descriptive paragraph. Subsequently, users can edit this textual representation to facilitate video object addition, removal, and modification. Experimental results demonstrate that our approach outperforms existing state-of-the-art methods in the field."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"See Weakness."},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- The proposed Video-to-Paragraph (V2P) method serves as an effective video captioner, capable of capturing multi-granular spatiotemporal features.\n- This framework achieves state-of-the-art performance in video editing tasks.\n- A novel dataset is introduced, which can be used as a benchmark for video editing, which contains 7.2K high-quality detailed video paragraphs and 5.5K object-level detailed caption-mask pairs."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The proposed editing approach heavily relies on the accuracy of the Video-to-Paragraph (V2P) method, leading to potentially unnatural modifications. \n   - (1) For video object removal and modification, if the paragraph does not mention a specific object (e.g., in the description of Figure 4, the absence of the female character’s earrings), it raises the question of how to effectively remove or modify that object.\n   - (2) While adding objects does not depend on the paragraph's accuracy, it is unclear how to control the timing of additions through text. The mechanism for locating the spatiotemporal position of the mask for object addition is lacking in detail. Is the layout determined using a large language model (LLM), or is there another implementation?\n   - (3) The textual descriptions lack a temporal dimension, making it difficult to achieve fine-grained control over video content, such as adding objects at specific time intervals or altering object attributes over time.\n   - (4) There is no comparison of the Video-to-Paragraph method with more advanced video LLM techniques. In the single object prediction task, there is a lack of comparison with state-of-the-art methods like Video-Chat2 and Video-LLaVA.\n   - (5) While the editing pipeline boasts high accuracy, it is not user-friendly. Unlike methods such as Instruct Pix2Pix, which allow for straightforward edits, this approach requires modifications to the existing paragraph.\n\n\n2. The multi-granular pooling strategy lacks ablation studies, which are necessary to evaluate its effectiveness.\n\nAdditonal discussion:\nFor image editing, the workflow of converting images to text and then modifying through text works well; however, video editing inherently requires temporal control, and the current approach lacks mechanisms for time-based edits and fine-grained temporal adjustments—an essential aspect of video manipulation.\n\nThe pipeline could benefit from further streamlining. Ideally, users would simply input a high-level instruction, and the remainder of the editing process would be handled by the LLM, removing the need for users to read and modify the paragraph directly."}},"nonreaders":[],"tmdate":1731427999663,"tcdate":1730633755674,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission5042/Reviewer_h1ZE"],"signatures":["ICLR.cc/2025/Conference/Submission5042/Reviewer_h1ZE"],"forum":"qnGir4dyu9","number":3,"license":"CC BY 4.0","cdate":1730633755674,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission5042/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427999663,"domain":"ICLR.cc/2025/Conference","replyto":"qnGir4dyu9","id":"HHRYa1huvd","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video Inpainting","Video Editing","Video Caption","Multimodal Large Language Model"]},"supplementary_material":{"value":"/attachment/9c47b8f3ae2d08423fd5f0d8fb34f9be2dd9a9e1.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Recent video generative models primarily rely on carefully written text prompts for specific tasks, like inpainting or style editing. They require labor-intensive textual descriptions for input videos, hindering their flexibility to adapt personal/raw videos to user specifications. This paper proposes RACCooN, a versatile and user-friendly video-to-paragraph-to-video generative framework that supports multiple video editing capabilities, such as removal, addition, and modification, through a unified pipeline. RACCooN consists of two principal stages: Video-to-Paragraph (V2P) and Paragraph-to-Video (P2V). In the V2P stage, we automatically describe video scenes in well-structured natural language, capturing both the holistic context and focused object details. Subsequently, in the P2V stage, users can optionally refine these descriptions to guide the video diffusion model, enabling various modifications to the input video, such as removing, changing subjects, and/or adding new objects. The proposed approach stands out from other methods through several significant contributions: (1) RACCooN suggests a multi-granular spatiotemporal pooling strategy to generate well-structured video descriptions, capturing both the broad context and object details without requiring complex human annotations, simplifying precise video content editing based on text for users. (2) Our video generative model incorporates auto-generated narratives or instructions to enhance the quality and accuracy of the generated content. (3) RACCooN also plans to imagine new objects in a given video, so users simply prompt the model to receive a detailed video editing plan for complex video editing. The proposed framework demonstrates impressive versatile capabilities in video-to-paragraph generation (up to 9.4% absolute improvement in human evaluations against the baseline), video content editing (relative 49.7% in FVD), and can be incorporated into other SoTA video generative models for further enhancement."},"_bibtex":{"value":"@misc{\nyoon2025raccoon,\ntitle={{RACC}ooN: A Versatile Instructional Video Editing Framework with Auto-Generated Narratives},\nauthor={Jaehong Yoon and Shoubin Yu and Mohit Bansal},\nyear={2025},\nurl={https://openreview.net/forum?id=qnGir4dyu9}\n}"},"title":{"value":"RACCooN: A Versatile Instructional Video Editing Framework with Auto-Generated Narratives"},"pdf":{"value":"/pdf/92793d9a9c07981b6886e1172d1568eebdd6d952.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"yoon|raccoon_a_versatile_instructional_video_editing_framework_with_autogenerated_narratives"},"authorids":{"value":["~Jaehong_Yoon1","~Shoubin_Yu1","~Mohit_Bansal2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Jaehong Yoon","Shoubin Yu","Mohit Bansal"]}},"version":2},{"content":{"summary":{"value":"This paper introduces a large-scale image-text dataset sourced from 1.3B online videos. From this collection, 300M frame-caption pairs and 12M video-level text summaries are extracted for training vision-language models. By training CLIP models on the newly introduced OVID dataset, the authors demonstrate superior performance on both text-to-video and video-to-text retrieval."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"Please see the weaknesses above."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":1},"strengths":{"value":"1. The OVID dataset collects image-text pairs from large-scale online videos, which is significantly different from previous image-text datasets.\n\n2. Compared to existing open-ended video-language datasets, OVID contains a much larger number of publicly available videos."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. I am confused by the statement of contributions in L100–106. It appears that all the contributions focus solely on the release of a large-scale image-text (or frame-text) dataset. In other words, I believe the contributions of this paper are limited, as it primarily presents a dataset.\n\n2. Although OVID is significantly larger than previous datasets, it seems to sacrifice many details in its captions or summaries. As shown in Table 5, CLIP models trained on OVID perform almost the worst on ImageNet-1k, ImageNet-R, ImageNet-Sketch, and ImageNet-V2. The very simple prompt (“Provide a very coarse single line of caption”) used in OVID’s captioning pipeline likely omits too many details from video frames, resulting in unsatisfactory performance on zero-shot classification tasks.\n\n3. In my opinion, the related work section includes many unrelated topics, such as multimodal LLMs and large multimodal models. What is the purpose of connecting these MLLMs to the OVID dataset? Are there any notable similarities or differences between them? The single sentence introducing large multimodal models in  L194–196 seems particularly out of place.\n\n4. I suggest that this dataset paper be resubmitted to a dataset or benchmark track. Meanwhile, the authors may consider how to improve the overall quality of the captions while maintaining the scale of the dataset."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762933737648,"tcdate":1761535569314,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission20250/Reviewer_8Gxi"],"signatures":["ICLR.cc/2026/Conference/Submission20250/Reviewer_8Gxi"],"forum":"etFOgs8vIb","number":1,"license":"CC BY 4.0","cdate":1761535569314,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission20250/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762933737648,"domain":"ICLR.cc/2026/Conference","replyto":"etFOgs8vIb","id":"m1xvoRyAO7","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["video dataset","recaptioning","clip","open foundation models","open datasets"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"We present OVid, a large open video dataset comprising _10 million hours_ of diverse content collected from CommonCrawl. To complement the raw data, we generate image captions for scene-changing frames and video-level captions for a 300M frame–caption subset. Using this subset, we train CLIP models at multiple scales and benchmark them against reference CLIP models trained on DataComp, Re-LAION and DataComp recaptioned with the same captioning pipeline. Observed scaling trends for classification and retrieval show evidence that OVid can be another valuable and scalable source of image-text data, in addition to image-text pairs from public webpages. OVid marks a significant step towards democratizing access to large-scale video data and fostering the development of open multimodal foundation models. To this end, all the data will be freely available to research institutions."},"_bibtex":{"value":"@misc{\nhochlehnert2026ovid,\ntitle={{OV}id: Open Large-Scale Video Dataset as a Novel Source for Image-Text Data},\nauthor={Andreas Hochlehnert and Marianna Nezhurina and Thadd{\\\"a}us Wiedemer and Christoph Schuhmann and Mehdi Cherti and Romain Beaumont and Andrii Matiuk and Andrej Radonjic and Bernhard Sch{\\\"o}lkopf and Wieland Brendel and A. Sophia Koepke and Jenia Jitsev and Matthias Bethge},\nyear={2026},\nurl={https://openreview.net/forum?id=etFOgs8vIb}\n}"},"title":{"value":"OVid: Open Large-Scale Video Dataset as a Novel Source for Image-Text Data"},"pdf":{"value":"/pdf/7a5482f7f60923f65052b07107b7cc4fc627e148.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"hochlehnert|ovid_open_largescale_video_dataset_as_a_novel_source_for_imagetext_data"},"authorids":{"value":["~Andreas_Hochlehnert1","~Marianna_Nezhurina1","~Thaddäus_Wiedemer1","~Christoph_Schuhmann1","~Mehdi_Cherti2","~Romain_Beaumont1","~Andrii_Matiuk2","~Andrej_Radonjic1","~Bernhard_Schölkopf1","~Wieland_Brendel1","~A._Sophia_Koepke1","~Jenia_Jitsev1","~Matthias_Bethge1"]},"authors":{"value":["Andreas Hochlehnert","Marianna Nezhurina","Thaddäus Wiedemer","Christoph Schuhmann","Mehdi Cherti","Romain Beaumont","Andrii Matiuk","Andrej Radonjic","Bernhard Schölkopf","Wieland Brendel","A. Sophia Koepke","Jenia Jitsev","Matthias Bethge"]}},"version":2},{"content":{"summary":{"value":"This paper investigates into relaxing the faithfulness assumption in causal discovery and proposes a relaxed faithfulness, called minimal dependence faithfulness. Minimal dependence of a variable is a set of variables which individually independent to the variable but becomes dependent to the variable given the rest of the set. After introducing further restricted classes of minimal dependence (strong, symmetric), the paper presents the equivalence class of a DAG under minimal dependence (with additional assumptions to further orient edges), and devises a modified PC algorithm to detect minimal dependence and orient edges involving minimal dependencies."},"soundness":{"value":4},"confidence":{"value":4},"questions":{"value":"- Why is the minimal dependence, “minimal”? Is this because of “no proper subset…” part or something else with respect to faithfulness/dependence itself not about the set.\n- Can we have examples for Definition 2 showing the difference between Y^o and Y clearly? \n- Def. 2 U “\\in” should be “\\subseteq”\n- Line 251 uniform → universal?\n- Line 460 tripe → triple\n- Consider replacing ‘links’ to ‘edges’\n- Citation for FCI algorithm should be of Spirtes not a survey paper. (Line 046)\n\nRegarding \"Contributions\", there is a question about \"Are the results valuable to share with the *broader* ICLR community?\" I guess this paper would be a perfect fit for UAI or AAAI."},"rating":{"value":5},"details_of_ethics_concerns":{"value":"."},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"- The problem is well motivated: relaxing faithfulness is one of the most important topics in causal discovery.  \n- The concept of minimal dependence is easy to follow. Its relationships to previous notion of (weak) types of faithfulness are well explained.\n- The paper not only investigate into the implied properties of minimal dependence, but also covers equivalence class and a (theoretical) causal discovery algorithm."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- This paper employs causal sufficiency which is not explict as “assumption” but hidden in the definition of SCM in line 083. \n- How much can we expect that a given distribution exhibits ‘minimal dependence’ (isn’t it Lesbegue measure zero?) BTW, I am very thankful for the authors for coming up with examples other than XORs.\n- Despite of MD-PC’s theoretical theoretical guarantee, causal discovery almost always runs on a finite sample. It is crucial to empirically investigate into how MD-PC actually detects minimal dependencies. In some cases where no minimal dependence exists, MD-PC’s performance might (certainly?) be worse than traditional PC. (I “hopefully” expect that the false positive rate of minimal dependence is very low given that the sample size is often small or CI test is not so powerful. The issue will be true positive rate, unless we artificially create strong conditional dependence for minimal dependencies through carefully constructed synthetic datasets.)\n- The paper may require more visual elements although visualizing some of them may not be very obvious (e.g., Dep^o is more about functional/probabilistic aspects than (purely) graphical porperty.)\n- The organization of the paper makes readers somewhat exhausted with definitions and assumptions throughout the paper. One issue is that the authors present weak, strong, and then symmetric version where the the weak version is later dropped (i.e., unused). Further, MD-PC and equivalence only applies to the strong and symmetric version of minimal dependence. While it is the authors’ choice to build up notations slowly, sometimes it would be better to present what the authors feel “matter”."}},"nonreaders":[],"tmdate":1731427518240,"tcdate":1730718827251,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission1968/Reviewer_2fnM"],"signatures":["ICLR.cc/2025/Conference/Submission1968/Reviewer_2fnM"],"forum":"or8wkKoBP4","number":4,"license":"CC BY 4.0","cdate":1730718827251,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission1968/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427518240,"domain":"ICLR.cc/2025/Conference","replyto":"or8wkKoBP4","id":"x5wPIOmbvX","forumContent":{"TLDR":{"value":"Relaxing faithfulness in causal discovery, we introduce minimal dependence faithfulness to handle unfaithful structures like XOR, modifying the PC algorithm to detect these and output candidate DAGs."},"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Bayesian networks","Causality","Faithfulness","Structure learning"]},"primary_area":{"value":"causal reasoning"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Causality detection is to identify the ``true'' directed acyclic graph (DAG) of a causal model from the joint probability distribution of the observed variables.\nAlgorithms such as PC and its modified versions perform this task under the restrictive faithfulness assumption, that is the DAG encodes all conditional independencies imposed by the distribution. \nHowever, all existing algorithms fail to detect the simple structure where a variable is the XOR of several Bernoulli variables, violating faithfulness. We generalize this type of unfaithfulness that appears in other, non-XOR, examples and define the \\emph{minimal dependence} of a given variable $X$ as the set of variables, such that $X$ is independent of each variable in the set but depends on at least one of them, the \\emph{dependent member} if conditioned on the remainder of the set.\nMinimal dependencies of size at least two violate faithfulness. Consequently, we relax faithfulness to \\emph{minimal dependence faithfulness}, restricting the neighbors of a node to its dependent members, and impose \\emph{minimal orientation faithfulness} that generalizes the orientation rules under faithfulness.\nWe then determine the structure of the dependent members of a node $X$ in the true DAG and show that they are connected to $X$ either directly or indirectly by a collider. \nFinally, we provide a sound and complete modification of the PC algorithm to detect this kind of unfaithfulness and output all possible candidates for the true DAG."},"_bibtex":{"value":"@misc{\nramazi2025structure,\ntitle={Structure Learning for Unfaithful Distributions: The Minimal Dependence Faithfulness},\nauthor={Pouria Ramazi and Hamid Kalantari},\nyear={2025},\nurl={https://openreview.net/forum?id=or8wkKoBP4}\n}"},"title":{"value":"Structure Learning for Unfaithful Distributions: The Minimal Dependence Faithfulness"},"pdf":{"value":"/pdf/1680621371ff7f502a853684bb85826485542371.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"ramazi|structure_learning_for_unfaithful_distributions_the_minimal_dependence_faithfulness"},"authorids":{"value":["~Pouria_Ramazi1","~Hamid_Kalantari1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Pouria Ramazi","Hamid Kalantari"]}},"version":2},{"content":{"TLDR":{"value":"SFT builds a shortcut from the text in VideoLLMs that hides behind accuracy gains; VETO keeps the LoRA update out of the text-only subspace, suppressing the shortcut while preserving the gain of SFT and gaining more than SFT under shift."},"venue":{"value":"ICLR 2027 Conference Submission"},"keywords":{"value":["video large language models","supervised fine-tuning","shortcut","distribution shift","video question answering"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Supervised fine-tuning (SFT) is the stage in which Video Large Language Models (VideoLLMs) advance their capability to reason over video. However, the VideoLLM also learns to answer from textual regularities of its training data, a shortcut that bypasses the video. The regularities behind the shortcut are hard to define at the data level, and the shortcut forms regardless of training epoch, data size, or backbone. The shortcut is hidden behind a rise in accuracy within the training distribution, yet it is not required for the rise and even costs SFT the gain the video offers under distribution shift. We therefore propose Video-Evidence Training via Orthogonal projection (VETO), which intervenes on the fine-tuning update where the shortcut is written rather than on the data. VETO calibrates the subspace that the text alone occupies in the activations of the base model and projects the update onto its orthogonal complement. The VideoLLM is thus left to ground its answer in the full input, with the shortcut suppressed as it forms and never defined. Across seven VideoLLMs VETO preserves the gain of SFT on the training distribution and suppresses the shortcut. On six benchmarks under shift it gains more than SFT and answers from the full input rather than the shortcut."},"_bibtex":{"value":"@inproceedings{\nanonymous2026suppressing,\ntitle={Suppressing the Shortcut Hidden Behind the Gains of Video{LLM} Fine-Tuning},\nauthor={Anonymous},\nbooktitle={Submitted to The Fifteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=PPII4JrkqX},\nnote={under review}\n}"},"title":{"value":"Suppressing the Shortcut Hidden Behind the Gains of VideoLLM Fine-Tuning"},"pdf":{"value":"/pdf/e3e7665ab868f0e01cc76911246502363cf8162e.pdf"},"venueid":{"value":"ICLR.cc/2027/Conference/Submission"}},"tmdate":1791221112022,"tcdate":1788003664958,"writers":["ICLR.cc/2027/Conference","ICLR.cc/2027/Conference/Submission6121/Authors"],"signatures":["ICLR.cc/2027/Conference/Submission6121/Authors"],"forum":"PPII4JrkqX","license":"CC BY 4.0","number":6121,"cdate":1788003664958,"readers":["everyone"],"invitations":["ICLR.cc/2027/Conference/-/Submission","ICLR.cc/2027/Conference/-/Edit","ICLR.cc/2027/Conference/Submission6121/-/Full_Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Bidding","ICLR.cc/2027/Conference/-/Submission_Change_Before_Reviewing"],"mdate":1791221112022,"odate":1791031328286,"domain":"ICLR.cc/2027/Conference","id":"PPII4JrkqX","version":2},{"content":{"summary":{"value":"In this paper, the authors propose using video steganography techniques to protect video content shared on public platforms. Specifically, they first perform face swapping on the original video to create a cover video. Then, using a reversible neural network, the original video is embedded into the cover video. Since the cover video is altered, it helps protect the sensitive video information that users wish to safeguard. Through the reversible neural network, the original video can be seamlessly decoded from the cover video, ensuring secure transmission of the video content."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"Please see weaknesses."},"rating":{"value":5},"details_of_ethics_concerns":{"value":"The method could potentially allow attackers to safely disseminate harmful content on public platforms through video steganography."},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- Protecting the original video using video steganography techniques.\n\n- Using symmetric encryption to encode and decode the original video ensures the preservation of information during transmission."},"flag_for_ethics_review":{"value":["Yes, Privacy, security and safety","Yes, Responsible research practice (e.g., human subjects, data release)"]},"weaknesses":{"value":"- Application. The proposed method is positioned as a way to protect user privacy on public platforms. However, this presents an inherent contradiction: if users genuinely wish to protect their privacy, they wouldn’t need to upload videos to a public platform. The only scenario where this might make sense is if users intentionally want to share hidden information through public platforms, which raises potential societal concerns.\n\n- Security. While the use of a reversible neural network ensures video embedding and decoding with minimal loss of quality, the symmetric encryption method itself lacks strong security guarantees.\n\n- Experiments: The paper lacks metrics evaluating video smoothness and realism. The authors are encouraged to use metrics such Fréchet Inception Distance to provide a more detailed assessment of their method’s performance."}},"nonreaders":[],"tmdate":1731427783396,"tcdate":1729323534376,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission5154/Reviewer_Yq2Z"],"signatures":["ICLR.cc/2025/Conference/Submission5154/Reviewer_Yq2Z"],"forum":"waHmD2i1dv","number":1,"license":"CC BY 4.0","cdate":1729323534376,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission5154/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427783396,"domain":"ICLR.cc/2025/Conference","replyto":"waHmD2i1dv","id":"2PHVFdCRWZ","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Bioprivacy","Diffusion model","Face swapping","Video Prediction","Reversible neural networks","Video Hiding"]},"primary_area":{"value":"alignment, fairness, safety, privacy, and societal considerations"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Advanced facial recognition technologies and recommender systems with inadequate privacy technologies and policies for facial interactions increase concerns about bioprivacy violations. With the proliferation of video and live-streaming websites, public-face video distribution and interactions pose greater privacy risks. Existing techniques typically address the risk of sensitive biometric information leakage through various privacy enhancement methods but pose a higher security risk by corrupting the information to be conveyed by the interaction data, or by leaving certain biometric features intact that allow an attacker to infer sensitive biometric information from them. To address these shortcomings, in this paper, we propose a neural network framework, CausalVE. We obtain cover images by adopting a diffusion model to achieve face swapping with face guidance and use the speech sequence features and spatiotemporal sequence features of the secret video for dynamic video inference and prediction to obtain a cover video with the same number of frames as the secret video. In addition, we hide the secret video by using reversible neural networks for video hiding so that the video can also disseminate secret data. Numerous experiments prove that our CausalVE has good security in public video dissemination and outperforms state-of-the-art methods from a qualitative, quantitative, and visual point of view."},"_bibtex":{"value":"@misc{\nhuang2024causalve,\ntitle={Causal{VE}: Face Video Privacy Encryption via Causal Video Prediction},\nauthor={Yubo Huang and Wenhao Feng and Xin Lai and Zixi Wang and Jingzehua Xu and Shuai Zhang and Hongjie He and Fan Chen},\nyear={2024},\nurl={https://openreview.net/forum?id=waHmD2i1dv}\n}"},"title":{"value":"CausalVE: Face Video Privacy Encryption via Causal Video Prediction"},"pdf":{"value":"/pdf/8f95871091b3fc7498487a989666fbe91c5968e8.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"huang|causalve_face_video_privacy_encryption_via_causal_video_prediction"},"authorids":{"value":["~Yubo_Huang2","~Wenhao_Feng2","~Xin_Lai4","~Zixi_Wang2","~Jingzehua_Xu1","~Shuai_Zhang6","~Hongjie_He1","~Fan_Chen10"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Yubo Huang","Wenhao Feng","Xin Lai","Zixi Wang","Jingzehua Xu","Shuai Zhang","Hongjie He","Fan Chen"]}},"version":2},{"content":{"summary":{"value":"This paper addresses \"Shortcut Alignment\" in large reasoning models (LRMs), a phenomenon where models learn to issue templated refusals (e.g., \"I'm sorry...\") to harmful inputs without actually grounding the decision in their internal chain-of-thought (CoT) reasoning. This superficial alignment not only makes refusals uninformative but also causes models to become overly cautious, leading to excessive false refusals on benign, safe queries. To fix this, the authors propose Deep Instruct Fine-tuning (DIFT), a new method that uses a \"CMI-Loss\" function. This CMI-Loss specifically penalizes the direct input-to-refusal \"shortcut\" on harmful examples, forcing the model to rely on its CoT reasoning to make a safety decision. The results demonstrate that this method successfully alleviates erroneous refusals on benign inputs while preserving the model's safety and reasoning capabilities."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1. Line 192 can remove the \"(A)\" since there is no other paragraph.\n2. Strata-Sword is mentioned in line 308 but not used in the experiment."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. The paper introduces \"Shortcut Alignment\", an observation that current safety methods (SFT) can lead to models appearing safe by using superficial cues to generate templated refusals, while decoupling this refusal from their internal reasoning (CoT).\n2. The paper introduces Deep Instruct Fine-tuning (DIFT), which uses a CMI-Loss to penalize the x→y (input-to-answer) shortcut.\n3. The authors validate their method by showing it reduces over-refusal on benchmarks while maintaining safety on benchmarks and preserving general reasoning abilities. They further perform probe-based analyses that confirm the intended mechanism that refusals become more reliant on the CoT rather than generic cues."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. There is no description of the evaluation method and its justification. How is the refusal accuracy and non-refusal rate calculated? How reliable is the evaluation?\n2. The improvement in over-refusal is not significant. The p-value shows \"most likely\" that the proposed method improves over-refusal over the baseline, not indicating the scale of improvement.\n3. The abstract mentions one of the problems for \"Shortcut Alignment\" is \"refusals without reasoning carry no informative value\". The rest of the paper does not mention this point, and the proposed method does not improve on this problem.\n4. Why and whether \"Shortcut Alignment\" will lead to over-refusal needs more intuitive and quantitative justification. \n5. Presentation issues: For instance, in Figure 2, the reason for using normalized harmful dependence, and what it indicates, are not explained. Others see Questions."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762920832668,"tcdate":1760866610710,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission9146/Reviewer_TqNZ"],"signatures":["ICLR.cc/2026/Conference/Submission9146/Reviewer_TqNZ"],"forum":"3qHILWiEob","number":1,"license":"CC BY 4.0","cdate":1760866610710,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission9146/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762920832668,"domain":"ICLR.cc/2026/Conference","replyto":"3qHILWiEob","id":"4KI7XlNdRh","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["safety alignment","large reasoning model","over refusal"]},"supplementary_material":{"value":"/attachment/28cdf7b1b5c255402cc6319fab3039bdc650b07a.zip"},"primary_area":{"value":"alignment, fairness, safety, privacy, and societal considerations"},"abstract":{"value":"Large Reasoning models (LRMs) with reasoning capabilities have demonstrated remarkable performance on complex tasks, yet achieving robust safety alignment remains a significant challenge. Supervised fine-tuning (SFT) with safety data is a widely-used approach to improve the models' safety, however, we identify that current safety alignment methods with SFT often induce a phenomenon we term $\\textbf{Shortcut Alignment}$. \nIn this case, the model learns to recognize the patterns in harmful inputs and emit templated refusals (e.g., \"I'm sorry...\") while decoupling the final response from its internal chain-of-thought (CoT) reasoning. This superficiality leads to two critical problems: (i) refusals without reasoning carry no informative value, and (ii) models become overly cautious, leading to excessive false refusals on benign queries and thereby degrading their general helpfulness.\nTo understand this behavior, we formalize it through the lens of conditional mutual information (CMI), hypothesizing that when the information gain from CoT is low, such shortcuts become low-resistance solutions that reduce training loss with little cost. We empirically verify this hypothesis via probe experiments that estimate the gap between predictions with and without CoT on harmful versus benign data. \nMotivated by these insights, we propose Deep Instruct Fine-tuning  (DIFT), which uses $\\textbf{CMI-Loss}$, explicitly penalizing shortcut predictions while preserving original instruct-tuning on benign examples. Through theoretical analysis and empirical evidence, we show that our method offers a better solution. It alleviates erroneous refusals while preserving safety. Our work bridges theory and practice, offering the first fine-grained alignment method that explicitly targets shortcut alignment in LRMs."},"_bibtex":{"value":"@misc{\nliu2026beyond,\ntitle={Beyond Refusals: Fine-grained Safety Alignment for Reasoning {LLM}s},\nauthor={Zhendong Liu and Baihui Zheng and Hongqiong Zhong and Boren Zheng and Yingshui Tan and Xiaoyong Zhu and Bo Zheng},\nyear={2026},\nurl={https://openreview.net/forum?id=3qHILWiEob}\n}"},"title":{"value":"Beyond Refusals: Fine-grained Safety Alignment for Reasoning LLMs"},"pdf":{"value":"/pdf/10085969ef4991399b9fa16e5b35be31c94e5e8b.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"liu|beyond_refusals_finegrained_safety_alignment_for_reasoning_llms"},"authorids":{"value":["~Zhendong_Liu6","~Baihui_Zheng2","~Hongqiong_Zhong3","~Boren_Zheng1","~Yingshui_Tan1","~Xiaoyong_Zhu1","~Bo_Zheng5"]},"authors":{"value":["Zhendong Liu","Baihui Zheng","Hongqiong Zhong","Boren Zheng","Yingshui Tan","Xiaoyong Zhu","Bo Zheng"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a method for affordance-aware human video generation by conditioning a pretrained text-to-video model on scene images. The authors claim that the model implicitly learns affordance from human-scene interaction signals, and analyze cross-attention maps to support this claim."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"-The affordance perception claim seems largely based on attention maps. How do you justify that attention indicates true affordance understanding rather than text-token correlation?\n\n-How well does the approach generalize to unseen object categories or actions not present in the curated dataset? Any failure case analysis?\n\n-The human-removal (inpainting) pipeline appears to erase or distort key affordance-related objects in the scene. For instance, in Fig. 1 (top row, right example), the bicycle seat is already missing in the input scene after human removal, fundamentally altering the bike's action possibilities. How common are such affordance-distorting artifacts in the dataset? Could this lead the model to learn incorrect affordance priors or reduce interaction plausibility?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"-The paper tackles an interesting and relevant problem of affordance-aware human-scene video generation.\n\n-The paper provides a straightforward extension to condition a pretrained text-to-video model on scene images.\n\n-Qualitative examples are visually appealing and demonstrate some degree of human-scene interaction."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"-The paper mainly fine-tunes an existing model (e.g., MovieGen) with additional scene conditioning. The architectural modifications (latent concatenation + text-image fusion) appear incremental. \n\n-The affordance aspect, which the paper claims as an essential contribution, seems more like a re-interpretation of what attention maps already provide, rather than a fundamentally new capability.\n\n-The use of cross-attention heatmaps as evidence of affordance perception is weak. These maps do not prove that the model understands object functionality or the feasibility of object interactions. \n\n– The human-removal pipeline heavily depends on segmentation and inpainting models, which may introduce artifacts or incorrect affordance cues.\n\n-Most baselines are general video editing or text-to-video systems that are not designed for human-scene interaction, which makes the comparison less meaningful. \n\n– The dataset is not publicly available, limiting reproducibility and fair comparison."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762921462112,"tcdate":1762070747877,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission10069/Reviewer_J7Sg"],"signatures":["ICLR.cc/2026/Conference/Submission10069/Reviewer_J7Sg"],"forum":"hg0lpcHdWk","number":4,"license":"CC BY 4.0","cdate":1762070747877,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission10069/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762921462112,"domain":"ICLR.cc/2026/Conference","replyto":"hg0lpcHdWk","id":"o6FUFMMZEW","forumContent":{"TLDR":{"value":"We explore the affordance perception potential of text-to-video models by teaching them to predict human-environment interaction."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Video Generation","Human-Centric Generation","Affordance","Representation Visualization"]},"supplementary_material":{"value":"/attachment/f9f26569c797273b951a8f22e9a8d8d2b6e7c252.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Can a video generation model be repurposed as an interactive world simulator? We explore the affordance perception potential of text-to-video models by teaching them to predict human-environment interaction. Given a scene image and a prompt describing human actions, we fine-tune the model to insert a person into the scene, while ensuring coherent behavior, appearance, harmonization, and scene affordance. Unlike prior work, we infer human affordance for video generation (i.e., where to insert a person and how they should behave) from a single scene image, without explicit conditions like bounding boxes or body poses. An in-depth study of cross-attention heatmaps demonstrates that we can uncover the inherent affordance perception of a pre-trained video model without labeled affordance datasets."},"_bibtex":{"value":"@misc{\nshan2026populateascene,\ntitle={Populate-A-Scene: Affordance-Aware Human Video Generation},\nauthor={Mengyi Shan and Zecheng He and Haoyu Ma and Felix Juefei-Xu and Peizhao Zhang and Tingbo Hou and Ching-Yao Chuang},\nyear={2026},\nurl={https://openreview.net/forum?id=hg0lpcHdWk}\n}"},"title":{"value":"Populate-A-Scene: Affordance-Aware Human Video Generation"},"pdf":{"value":"/pdf/3859d5d6093b33616d7842b25ca3366cff213a88.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"shan|populateascene_affordanceaware_human_video_generation"},"authorids":{"value":["~Mengyi_Shan1","~Zecheng_He1","~Haoyu_Ma1","~Felix_Juefei-Xu1","~Peizhao_Zhang1","~Tingbo_Hou2","~Ching-Yao_Chuang1"]},"authors":{"value":["Mengyi Shan","Zecheng He","Haoyu Ma","Felix Juefei-Xu","Peizhao Zhang","Tingbo Hou","Ching-Yao Chuang"]}},"version":2},{"content":{"summary":{"value":"This paper introduces explicit timestamp tokens into MLLMs to improve time-aware video understanding and event localization. Meanwhile, it employs a two-stage semantic-guided key frame selection. Experiments are conducted on general video understanding benchmarks (VideoMME, LongVideoBench, and LVBench) and their event-aware tasks."},"soundness":{"value":2},"confidence":{"value":5},"questions":{"value":"- In Table 3, does the reported runtime of TASS include the DeepSeek-V3 recaptioning step? If not, the latency comparison against AKS may not be fair.\n\n- Why are general video understanding benchmarks used as primary evaluation instead of timestamp-aware grounding datasets such as ETBench or Charades-STA?\n\n- If reinforcement learning (e.g., GRPO) is added to the base model, do timestamp embeddings still yield improvements?\n\n- Does the method require additional training data or annotations, or is it completely training-free?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- The paper is well written and easy to follow.\n- The idea of injecting timestamp tokens into a video-language transformer is simple and intuitive.\n- The proposed approach is plug-and-play — it can be applied to existing multimodal LLMs without dense temporal grounding annotations or extensive retraining, making it appealing for low-resource or training-free settings."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. **Misalignment between problem definition and evaluation.**\n\nThe paper positions itself as addressing timestamp-aware video understanding and temporal/event grounding, yet the primary results are reported on general video understanding benchmarks (VideoMME, LongVideoBench, LVBench). Although some event-aware subsets are included, these benchmarks do not directly reflect timestamp reasoning or grounding capabilities.\nTo support the claimed contributions, evaluations on specialized temporal grounding datasets for MLLMs are necessary, such as:\n\n- E.T. Bench (NeurIPS 2024) — open-ended event-level temporal grounding for video LLMs\n- Charades-STA — standard video temporal localization benchmark\n\nAt present, the experiments do not sufficiently demonstrate that explicit timestamp modeling improves grounding over existing methods.\n\n**2. Limited novelty; similar ideas have been explored.**\n\nThe main technical idea—inserting explicit timestamp tokens into the multimodal input stream—has been extensively explored in temporal grounding literature and even appears in open-sourced MLLMs (e.g., Qwen3-VL). As a result, the contribution feels incremental. Several recent works already introduce timestamp or temporal tokens into video-LLMs or temporal grounding pipelines:\n\nNumber It: Temporal Grounding Videos like Flipping Manga, CVPR 2025\nGenS: Generative Frame Sampler for Long Video Understanding, ACL Findings 2025\n\nThus, timestamp token injection appears more like a commonly used trick rather than a novel methodological contribution.\n\n**3. Missing discussion on reinforcement learning (RL) effects.**\nRecent GRPO-based work (e.g., TimeR1: Qwen2.5-VL + GRPO) shows that reinforcement learning can significantly improve grounding capability even without explicit temporal embedding.\nTherefore, it remains unclear whether the observed gains originate from:\n\n- the timestamp embedding mechanism itself, or \n- insufficient grounding training of the base model.\n\nAblations such as timestamp embedding vs. RL-enhanced grounding would clarify this."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762923456604,"tcdate":1762351653872,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission12611/Reviewer_f994"],"signatures":["ICLR.cc/2026/Conference/Submission12611/Reviewer_f994"],"forum":"ABRR4fwXF2","number":3,"license":"CC BY 4.0","cdate":1762351653872,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission12611/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762923456604,"domain":"ICLR.cc/2026/Conference","replyto":"ABRR4fwXF2","id":"yiB336BFXe","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["video understanding"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Long video understanding remains a fundamental challenge for multimodal large language models (MLLMs), particularly in tasks requiring precise temporal reasoning and event localization. Existing approaches typically adopt uniform frame sampling and rely on implicit position encodings to model temporal order. However, these methods struggle with long-range dependencies, leading to critical information loss and degraded temporal comprehension. In this paper, we propose Dynamic Absolute Time Enhancement (DATE) that enhances temporal awareness in MLLMs through the Timestamp Injection Mechanism (TIM) and a semantically guided Temporal-Aware Similarity Sampling (TASS) strategy. Specifically, we interleave video frame embeddings with textual timestamp tokens to construct a continuous temporal reference system. We further reformulate the video sampling problem as a vision-language retrieval task and introduce a two-stage algorithm to ensure both semantic relevance and temporal coverage: enriching each query into a descriptive caption to better align with the vision feature, and sampling key event with a similarity-driven temporally regularized greedy strategy. Our method achieves remarkable improvements w.r.t. absolute time understanding and key event localization, resulting in state-of-the-art performance among 7B and 72B models on hour-long video benchmarks. Particularly, our 7B model even exceeds many 72B models on some benchmarks."},"_bibtex":{"value":"@misc{\nyuan2026date,\ntitle={{DATE}: Dynamic Absolute Time Enhancement for Long Video Understanding},\nauthor={Chao Yuan and Yang Yang and Yehui Yang and Zecheng Lin},\nyear={2026},\nurl={https://openreview.net/forum?id=ABRR4fwXF2}\n}"},"title":{"value":"DATE: Dynamic Absolute Time Enhancement for Long Video Understanding"},"pdf":{"value":"/pdf/be690f7da70dcda4269ad2363d4b7e8b9114fd12.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"yuan|date_dynamic_absolute_time_enhancement_for_long_video_understanding"},"authorids":{"value":["~Chao_Yuan3","~Yang_Yang1","~Yehui_Yang1","~Zecheng_Lin3"]},"authors":{"value":["Chao Yuan","Yang Yang","Yehui Yang","Zecheng Lin"]}},"version":2},{"content":{"summary":{"value":"This paper proposes Trajectory Consistent Flows (TCF), which jointly learns a flow model, which learns the flow velocity, and a trajectory model, which translates samples from one timepoint to another timepoint along the flow probability-flow ODE (PF-ODE) with a single function evaluation. Flow and trajectory models are combined into a single neural net, TCF, whose function value returns the jump along the flow PF-ODE, and time-derivative value returns the flow velocity. TCF is optimized with a combination of self-consistency loss (similar to the loss for shortcut models) which distills PF-ODE trajectories, and velocity consistency loss which matches TCF time-derivatives to flow velocities. TCF demonstrates competitive performance on the task of CIFAR10 and ImageNet-64 generation."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"- **[Q1] The precise training setup for TCFs in Tables 2 and 3 are unclear.** Are the models initialized from a pre-trained diffusion / flow model, or are they trained from scratch? Also, in lines 402-404, the authors write that they increase the batch size from 512 to 2048, but in  Appendix D.1, the authors write they use batch size 512. With experimental details scattered across the paper, it is difficult to fairly compare TCF against the baselines. Please describe the complete training setup (e.g., network initialization, number of training iterations, batch size, etc.) for the models in Tables 2 and 3."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"- **[S1] This paper is significant in the aspect that it provides several design choices which improve the generative performance of shortcut models.** Specifically, the authors demonstrate that by using a Taylor-expansion based parametrization Eq. (9) along with careful design choices such as time distributions, time conditioning, weighting coefficients for losses (Section 4.3), one can improve FID scores for shortcut models (Table 1)."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- **[W1] TCF lacks conceptual novelty, in the sense that it is very similar to shortcut models.** Specifically, TCF is also trained with a combination of flow matching and shortcut learning losses. The only difference between TCF and shortcut model is the choice of parametrization: TCF directly models the jump along PF-ODE trajectories whereas shortcut model learns the direction from one PF-ODE state to another.\n\n- **[W2] TCF lacks practical significance, as there are fast flow methods with similar or cheaper training costs but better generative performance.** For instance, the authors claim in lines 427-429 that baselines such as truncated consistency models (TCMs) or sCMs come at the cost of increased training or inference complexity. However, I would like to argue the contrary: TCM requires similar per-iteration cost, whereas sCM has cheaper per-iteration cost compared to TCF during training.\n\n  Specifically, the results in Tables 2 and 3 use Alg-C for TCF, which requires three forward passes and one backward pass of the network per iteration. TCM also requires three forward passes and one backward pass of the network per training iteration. While sCM requires JVP computation they use forward-mode automatic differentiation, which is significantly cheaper than backpropgation. Hence, as written in Discussions and Limitations of [1], the cost of sCM training per iteration is similar to two forward passes and one backward pass, so it is cheaper than TCF.\n\n[1] Simplifying, Stabilizing and Scaling Continuous-Time Consistency Models, ICLR, 2025"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762942899686,"tcdate":1761832948822,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission24017/Reviewer_Js8Q"],"signatures":["ICLR.cc/2026/Conference/Submission24017/Reviewer_Js8Q"],"forum":"VFYGBTSNWK","number":2,"license":"CC BY 4.0","cdate":1761832948822,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission24017/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762942899686,"domain":"ICLR.cc/2026/Conference","replyto":"VFYGBTSNWK","id":"aoCAQi2S4i","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"TLDR":{"value":"A fast sampling generative model."},"keywords":{"value":["flow matching","generative model"]},"primary_area":{"value":"generative models"},"abstract":{"value":"Diffusion and flow matching models have recently achieved remarkable generative performance, but their reliance on iterative ODE or SDE solvers results in slow and computationally expensive sampling.\nIn this work, we introduce Trajectory-Consistent Flows (TCF), a framework that unifies efficient training and accelerated sampling through a Taylor-expansion-based formulation. TCF jointly optimizes a flow matching model $p_{\\theta}$ and a fast-sampling surrogate $q_{\\theta}$ via a unified objective. We construct $q_{\\theta}$ using a second-order Taylor expansion as a trajectory-consistent approximation of $p_{\\theta}$'s ODE flow, enabling high-fidelity generation with as few as 5 sampling steps. We further extend this idea to a third-order expansion, achieving additional performance gains without increasing computational cost.  With further architectural and training enhancements, TCF achieves significantly improved sampling quality while retaining fast and stable training, making it particularly suitable for real-time generative applications."},"_bibtex":{"value":"@misc{\nmiao2025trajectoryconsistent,\ntitle={Trajectory-Consistent Flows: Enabling Fast Sampling for Flow Matching Models},\nauthor={Chenfeng Miao and Qingying Zhu and Minchuan Chen and Shaojun Wang and Jing Xiao},\nyear={2025},\nurl={https://openreview.net/forum?id=VFYGBTSNWK}\n}"},"title":{"value":"Trajectory-Consistent Flows: Enabling Fast Sampling for Flow Matching Models"},"pdf":{"value":"/pdf/86cfd2b2b95f833700bf790c4fe4c30dc0ba83a0.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"miao|trajectoryconsistent_flows_enabling_fast_sampling_for_flow_matching_models"},"authorids":{"value":["~Chenfeng_Miao1","~Qingying_Zhu1","~Minchuan_Chen1","~Shaojun_Wang1","~Jing_Xiao3"]},"authors":{"value":["Chenfeng Miao","Qingying Zhu","Minchuan Chen","Shaojun Wang","Jing Xiao"]}},"version":2},{"content":{"summary":{"value":"This paper introduces OmniVideoBench, a large-scale benchmark designed to evaluate the collaborative audiovisual reasoning capabilities of multimodal large language models. The benchmark comprises 1,000 manually annotated high-quality question-answer pairs (QA) across 628 videos, each featuring explicit step-by-step reasoning chains indicating modalities and evidence. OmniVideoBench spans 8 primary video genres, 68 subcategories, and 13 distinct task types (e.g., temporal, spatial, causal reasoning), structured to comprehensively evaluate modal complementarity and logical consistency. The paper evaluates both open-source and proprietary MLLMs on OmniVideoBench, revealing that model performance lags significantly behind human capabilities, particularly on tasks requiring genuine multimodal integration."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. Can the authors provide quantitative statistics on the efficacy of the automated filtering steps in Section 2.4? Specifically, what is the rejection or retention rate at each filtering stage, and what percentage of QA pairs end up truly requiring both modalities?\n2. Please clarify how \"semantic units\" ($S_i$) are operationalized for the semantic distance metric. Is this manual phrase decomposition, or is some NLP toolchain applied? This is crucial to evaluating distractor design reproducibility.\n3. Will human benchmark results (e.g., accuracy, agreement rates) on OmniVideoBench be reported?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. The paper ensures the richness and coverage of the dataset, comprising 1,000 distinct QA pairs that span a wide range of real-world scenarios, video durations, and audio types. It also contains 13 different task types covering diverse reasoning skills.\n2. The annotation protocol ensures that all questions require true audio-visual integration and stepwise reasoning, with multi-layered filtering to weed out unimodal or bias-prone items."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The paper does not provide statistical results comparing it with other relevant datasets, which fails to highlight the contribution of the proposed dataset.\n2 From Table 1 and Figure 3, it is evident that the vast majority of QA pairs relate to Speech (76.2%) versus Sound (14.7%) and Music (9.1%), creating considerable class imbalance.\n3 The benchmark's positioning relative to several directly analogous or recently proposed audio-visual (AV) benchmarks is incomplete. Recent works, such as AVHBench (Kim et al., 2025) and DAVE (Radevski et al., 2024), are not cited or discussed, nor are specialized audio-visual QA datasets, including AVQA (Yang et al., 2022) and MusicAVQA (Li et al., 2022).\n4. While human annotation is used in construction, the paper does not report human baseline accuracy or response variability for the main test set.\n5. While Figure 1 gives some specific sample breakdowns, the paper lacks a deeper set of qualitative analyses of successful versus failure cases, especially for (a) long video cases and (b) music understanding tasks."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762930897618,"tcdate":1761880415556,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission18908/Reviewer_rRon"],"signatures":["ICLR.cc/2026/Conference/Submission18908/Reviewer_rRon"],"forum":"ItRYEe8E61","number":3,"license":"CC BY 4.0","cdate":1761880415556,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission18908/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762930897618,"domain":"ICLR.cc/2026/Conference","replyto":"ItRYEe8E61","id":"B8gD35LbTb","forumContent":{"TLDR":{"value":"OmniVideoBench is a benchmark for comprehensive evaluation of synergistic audio–visual understanding in Omni MLLMs."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Multimodal Reasoning","MLLM Benchmark","Text","Audio","Video"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Recent advances in multimodal large language models (MLLMs) have demonstrated substantial potential in video understanding. However, existing benchmarks fail to comprehensively evaluate synergistic reasoning capabilities across audio and visual modalities, often neglecting either one of the modalities or integrating them in a logically inconsistent manner. To bridge this gap, we introduce OmniVideoBench, a large-scale and rigorously designed benchmark dedicated to assessing synergistic audio-visual understanding, with a strong emphasis on modality complementarity and logical consistency. Specifically, OmniVideoBench comprises 1000 high-quality question-answer(QA) pairs, each annotated with step-by-step reasoning traces, derived from 628 diverse videos ranging from several seconds to 30 minutes, and manually verified to guarantee complete correctness and uniqueness. Moreover, OmniVideoBench encompasses 13 carefully designed question types, covering temporal reasoning, spatial localization, counting, causal inference, summarization, and beyond, thereby capturing the essential challenges of video understanding. Evaluation of multiple MLLMs on OmniVideoBench reveals a pronounced gap between model performance and human reasoning, with open-source models lagging significantly behind their closed-source counterparts, underscoring the inherent difficulty of genuine audio-visual reasoning. We will release OmniVideoBench to foster the development of MLLMs with stronger and more generalizable reasoning capabilities."},"_bibtex":{"value":"@inproceedings{\nli2026omnivideobench,\ntitle={OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni {MLLM}s},\nauthor={Caorui Li and Yu Chen and Yiyan Ji and Jin Xu and Zhenyu Cui and Shihao Li and Yuanxing Zhang and Zhenghao Song and Dingling Zhang and Heying and Haoxiang Liu and Yuxuan Wang and Qiufeng Wang and Jiafu Tang and Zhenhe Wu and Jiehui Luo and Zhiyu Pan and Weihao Xie and Chenchen Zhang and Zhaohui Wang and Jiayi Tian and Yanghai Wang and Zhe Cao and Minxin Dai and ke wang and Runzhe Wen and Yinghao Ma and Yaning Pan and Sungkyun Chang and Termeh Taheri and Haiwen Xia and Christos Plachouras and Emmanouil Benetos and Yizhi LI and Ge Zhang and Jian Yang and Tianhao Peng and Zili Wang and Minghao Liu and Junran Peng and Zhaoxiang Zhang and Jiaheng Liu},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=ItRYEe8E61}\n}"},"title":{"value":"OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs"},"pdf":{"value":"/pdf/5966bcfa7745731e4a1c1d8d8f245167ce7c3991.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"li|omnivideobench_towards_audiovisual_understanding_evaluation_for_omni_mllms"},"authorids":{"value":["~Caorui_Li1","~Yu_Chen49","~Yiyan_Ji1","~Jin_Xu5","~Zhenyu_Cui4","~Shihao_Li5","~Yuanxing_Zhang3","~Zhenghao_Song1","~Dingling_Zhang1","~Heying2","~Haoxiang_Liu2","~Yuxuan_Wang4","~Qiufeng_Wang3","~Jiafu_Tang2","~Zhenhe_Wu1","~Jiehui_Luo2","~Zhiyu_Pan4","~Weihao_Xie1","~Chenchen_Zhang3","~Zhaohui_Wang10","~Jiayi_Tian1","~Yanghai_Wang1","~Zhe_Cao6","~Minxin_Dai1","~ke_wang40","~Runzhe_Wen1","~Yinghao_Ma1","~Yaning_Pan1","~Sungkyun_Chang1","~Termeh_Taheri1","~Haiwen_Xia1","~Christos_Plachouras1","~Emmanouil_Benetos1","~Yizhi_LI1","~Ge_Zhang5","~Jian_Yang10","~Tianhao_Peng1","~Zili_Wang1","~Minghao_Liu11","~Junran_Peng1","~Zhaoxiang_Zhang3","~Jiaheng_Liu1"]},"authors":{"value":["Caorui Li","Yu Chen","Yiyan Ji","Jin Xu","Zhenyu Cui","Shihao Li","Yuanxing Zhang","Zhenghao Song","Dingling Zhang","Heying","Haoxiang Liu","Yuxuan Wang","Qiufeng Wang","Jiafu Tang","Zhenhe Wu","Jiehui Luo","Zhiyu Pan","Weihao Xie","Chenchen Zhang","Zhaohui Wang","Jiayi Tian","Yanghai Wang","Zhe Cao","Minxin Dai","ke wang","Runzhe Wen","Yinghao Ma","Yaning Pan","Sungkyun Chang","Termeh Taheri","Haiwen Xia","Christos Plachouras","Emmanouil Benetos","Yizhi LI","Ge Zhang","Jian Yang","Tianhao Peng","Zili Wang","Minghao Liu","Junran Peng","Zhaoxiang Zhang","Jiaheng Liu"]}},"version":2},{"content":{"summary":{"value":"This paper presents VIDES, a framework for ultra-fast text-guided video editing using one-step diffusion models. Conventional video editing based on multi-step diffusion suffers from extreme latency (hours for a few minutes of video). VIDES introduces three key innovations to make one-step editing feasible and high-quality:\n1. A learnable inversion encoder that predicts the initial noise for each frame in one forward pass, eliminating costly multi-step inversion.\n2. A Structure-Aware Editing (SAE) loss trained on structurally aligned image pairs generated by prompt perturbation, ensuring geometry preservation during edits.\n3. A Unified-Frame Editing (UFE) mechanism that concatenates frame latents for joint processing, leveraging cross-frame attention for temporal consistency, with a sliding-window and anchor-frame strategy for long videos. Extensive experiments demonstrate a ~155× speedup over prior diffusion-based video editing methods while maintaining comparable or superior visual quality."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. How does VIDES perform on open-domain videos (e.g., YouTube, handheld footage) with uncontrolled motion?\n2. Can the SAE loss generalize beyond synthetic pairs to real edit pairs or user-provided source–target pairs?\n3. Does concatenating frame latents ever cause spatial artifacts due to large receptive fields near boundaries?\n4. How does the learned encoder generalize across different one-step diffusion backbones (e.g., SDXL-Lightning vs. DMD2)?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. Unified-Frame Editing elegantly leverages global attention for cross-frame consistency.\n2. Extensive experiments: both short and long videos, multiple baselines, detailed ablations (SAE loss, UFE, sliding window, anchor frame).\n3. Practical scalability — runs on a single GPU and aligns well with industry use cases for real-time editing."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Limited theoretical justification.\nThe encoder’s success is empirically shown but not theoretically characterized. There is no analysis of how its learned latent space aligns with that of the one-step generator.\n\n2. Dependence on dataset synthesis.\nThe “prompt perturbation” method for generating structure-aligned pairs is clever but synthetic, potentially limiting generalization to real-world video data.\n\n3. Possible overfitting to static structure.\nSince the model is trained on image pairs, it may not fully capture dynamic motion cues or 3D consistency, especially in non-rigid or fast-moving videos.\n\n4. Lack of generalization tests.\nEvaluation is limited to short clips (≤90 frames). There is no discussion of performance on videos with complex motion or strong occlusions.\n\n5. Missing resource analysis.\nWhile the paper claims 155× acceleration, a more detailed runtime breakdown (encoder vs. editing, memory footprint) would enhance reproducibility and credibility.\n\n6. Marginal novelty in components.\nEach component (encoder inversion, structure-aware training, latent concatenation) builds on existing paradigms, though their joint effectiveness is commendable."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762928447988,"tcdate":1761316116659,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission18735/Reviewer_6uqy"],"signatures":["ICLR.cc/2026/Conference/Submission18735/Reviewer_6uqy"],"forum":"DscflMFynS","number":1,"license":"CC BY 4.0","cdate":1761316116659,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission18735/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762928447988,"domain":"ICLR.cc/2026/Conference","replyto":"DscflMFynS","id":"i6mbPrqRB4","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Video Editing","Diffusion Models","Generative Models","Text editing"]},"supplementary_material":{"value":"/attachment/4105fea3522ed19959e01303b8da3ec2411c7f04.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Text-guided video editing with diffusion models is prohibitively slow, hindered by costly multi-step sampling and inversion. We present VIDES, the first framework to successfully adapt one-step text-to-image (T2I) models for high-quality video editing, addressing the core challenges of inversion, editability, and temporal consistency. To bypass slow iterative inversion, we train a learnable encoder that predicts the initial noise for each frame in a single forward pass. This encoder is trained with a novel Structure-Aware Editing (SAE) loss on a curated dataset of structurally-aligned image pairs, teaching it to preserve the source video's geometry during edits. For temporal coherence, we introduce Unified-Frame Editing (UFE), a technique that concatenates frame latents to facilitate cross-frame attention in a single generation step; for long videos, a sliding-window strategy with an anchor frame maintains global consistency. Our extensive experiments demonstrate that VIDES achieves editing quality comparable or superior to state-of-the-art multi-step methods, while operating approximately 155 times faster. This breakthrough paves the way for practical, real-time video editing applications."},"_bibtex":{"value":"@misc{\nlim2025vides,\ntitle={{VIDES}: {VIDEO} {EDITING} {IN} {SECONDS} {WITH} {ONE}-{STEP} {DIFFUSION} {MODELS}},\nauthor={Habin Lim and Gyeong-Moon Park},\nyear={2025},\nurl={https://openreview.net/forum?id=DscflMFynS}\n}"},"title":{"value":"VIDES: VIDEO EDITING IN SECONDS WITH ONE-STEP DIFFUSION MODELS"},"pdf":{"value":"/pdf/b776c3f6cdd85e19a19651574abda9a0b60de506.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"lim|vides_video_editing_in_seconds_with_onestep_diffusion_models"},"authorids":{"value":["~Habin_Lim2","~Gyeong-Moon_Park1"]},"authors":{"value":["Habin Lim","Gyeong-Moon Park"]}},"version":2},{"content":{"research_area_keywords":{"value":"sign languages, multi-channel structure, nonmanuals"},"keywords":{"value":["American Sign Language","minimal pair analysis","sign language translation","linguistic evaluations"]},"A1_potential_risks":{"value":"N/A"},"languages_studied":{"value":"American Sign Language, English"},"B3_data_contains_personally_identifying_info":{"value":"No"},"checklist_separator":{"readers":["everyone"]},"B4_data_contains_offensive_content":{"value":"N/A"},"venue":{"value":"ACL ARR 2026 August Submission"},"_bibtex":{"value":"@inproceedings{\nanonymous2026targeted,\ntitle={Targeted Linguistic Analysis of Sign Language Models with Minimal Translation Pairs},\nauthor={Anonymous},\nbooktitle={Submitted to ACL Rolling Review - August 2026},\nyear={2026},\nurl={https://openreview.net/forum?id=Ags6ER6rnm},\nnote={under review}\n}"},"title":{"value":"Targeted Linguistic Analysis of Sign Language Models with Minimal Translation Pairs"},"contribution_types":{"value":["Model analysis & interpretability","Approaches to low-resource settings","Data resources","Data analysis"]},"abstract":{"value":"Models of sign language have historically lagged behind those for spoken language (text and speech). Recent work has greatly improved their performance on tasks like sign language translation and isolated sign recognition. However, it remains unclear to what extent existing models capture various linguistic phenomena of sign language, and how well they use cues from the multiple articulators used in sign language (hands, upper body, face). We introduce a new benchmark dataset for American Sign Language, ASL Minimal Translation Pairs (ASL-MTP), divided into multiple types of sign language phenomena and corresponding minimal pairs of translations, for performing such linguistic analyses. As a case study, we use ASL-MTP to analyze a state-of-the-art ASL-to-English translation model. We conduct a targeted analysis of the model by ablating various input cues during training and inference and evaluating on the phenomena in ASL-MTP. Our results show that, while the model performs above chance level on most of the phenomena, it relies strongly on manual cues while often missing crucial non-manual cues."},"paper_type":{"value":"Long"},"pdf":{"value":"/pdf/b79bf99f018ed869f84ae3c5a92efba3eff9d064.pdf"},"research_area":{"value":"Linguistic theories, Cognitive Modeling and Psycholinguistics"},"venueid":{"value":"aclweb.org/ACL/ARR/2026/August/Submission"},"D5_annotator_population":{"value":"N/A"}},"tmdate":1787842582699,"tcdate":1785790504670,"writers":["aclweb.org/ACL/ARR/2026/August","aclweb.org/ACL/ARR/2026/August/Submission2115/Authors"],"signatures":["aclweb.org/ACL/ARR/2026/August/Submission2115/Authors"],"forum":"Ags6ER6rnm","license":"CC BY 4.0","number":2115,"cdate":1785790504670,"readers":["everyone"],"invitations":["aclweb.org/ACL/ARR/2026/August/-/Submission","aclweb.org/ACL/ARR/2026/August/-/Edit","aclweb.org/ACL/ARR/2026/August/-/Preprint_Post_Submission"],"mdate":1787842582699,"odate":1785846543432,"domain":"aclweb.org/ACL/ARR/2026/August","id":"Ags6ER6rnm","version":2},{"content":{"venue":{"value":"CoRR 2026"},"pdf":{"value":"https://arxiv.org/pdf/2604.27232v3"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"karabüklü|targeted_linguistic_analysis_of_sign_language_models_with_minimal_translation_pairs"},"html":{"value":"https://doi.org/10.48550/arXiv.2604.27232"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2604-27232,\n  publtype={informal},\n  author={Serpil Karabüklü and Kanishka Misra and Shester Gueuwou and Diane Brentari and Greg Shakhnarovich and Karen Livescu},\n  title={Targeted Linguistic Analysis of Sign Language Models with Minimal Translation Pairs},\n  year={2026},\n  month={April},\n  cdate={1775001600000},\n  journal={CoRR},\n  volume={abs/2604.27232},\n  url={https://doi.org/10.48550/arXiv.2604.27232}\n}\n"},"abstract":{"value":"Models of sign language have historically lagged behind those for spoken language (text and speech). Recent work has greatly improved their performance on tasks like sign language translation and isolated sign recognition. However, it remains unclear to what extent existing models capture various linguistic phenomena of sign language, and how well they use cues from the multiple articulators used in sign language (hands, upper body, face). We introduce a new benchmark dataset for American Sign Language, ASL Minimal Translation Pairs (ASL-MTP), divided into multiple types of sign language phenomena and corresponding minimal pairs of translations, for performing such linguistic analyses. As a case study, we use ASL-MTP to analyze a state-of-the-art ASL-to-English translation model. We conduct a targeted analysis of the model by ablating various input cues during training and inference and evaluating on the phenomena in ASL-MTP. Our results show that, while the model performs above chance level on most of the phenomena, it relies strongly on manual cues while often missing crucial non-manual cues."},"title":{"value":"Targeted Linguistic Analysis of Sign Language Models with Minimal Translation Pairs"},"authors":{"value":[{"fullname":"Serpil Karabüklü","username":"~Serpil_Karabüklü1"},{"fullname":"Kanishka Misra","username":""},{"fullname":"Shester Gueuwou","username":""},{"fullname":"Diane Brentari","username":""},{"fullname":"Greg Shakhnarovich","username":""},{"fullname":"Karen Livescu","username":""}]}},"tmdate":1785858145996,"pdate":1798675200000,"externalIds":["dblp:journals/corr/abs-2604-27232"],"tcdate":1785858141983,"writers":["~","OpenReview.net/Public_Article/DBLP.org","OpenReview.net/Support"],"signatures":["~Serpil_Karabuklu1"],"forum":"8pI6boanz4","license":"CC BY-SA 4.0","number":117370,"cdate":1775001600000,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/DBLP.org/-/Record","OpenReview.net/Public_Article/-/Edit"],"mdate":1785858145996,"domain":"OpenReview.net/Public_Article","id":"8pI6boanz4","version":2},{"content":{"summary":{"value":"4DNeX is a feed-forward framework that generates dynamic 4D scene representations (RGB + XYZ sequences) directly from a single image. It reformulates 4D generation as 6D video diffusion, fine-tuning a 14B Wan2.1 video diffusion model via a lightweight width-wise fusion and modality-aware normalization strategy. It introduces 4DNeX-10M, a large-scale pseudo-annotated dataset curated from DUSt3R / MonST3R / MegaSaM pipelines, and performs minimal post-optimization to recover camera poses. The method achieves competitive or superior motion quality to Free4D / 4Real / Animate124 while being significantly faster (~15 min vs 1+ hr), positioning itself as a scalable path toward single-image 4D world modeling."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. How to guarantee that the generated XYZ is geometrically accurate? Please add quantitative evaluations of 3D error or global pose alignment.\n\n2. Is your method actually SE(3)-consistent, or does the 4D break when camera motion is large? Please show failure cases with fast curved trajectories, not only forward panning.\n\n3. The sloped plane (Eq. 6) is a very strong inductive bias. Do you ablate training without any depth init? Does the model still converge?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. “RGB + XYZ as 6D video” is conceptually elegant — unifies appearance & geometry without NeRF volume rendering or Gaussian splats during training.\n\n2. Width-wise fusion is empirically validated and actually justified with token interaction distance。\n\n3. LoRA-only tuning on 14B Wan2.1 while preserving pretrained RGB distributions is a strong practical insight, compared to Cat4D / Free4D which lose RGB appearance fidelity when optimizing all parameters.\n\n4. Post-optimization for camera recovery is lightweight & well-justified."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Core novelty is incremental, effectively “fine-tuning Wan2.1 to predict XYZ instead of RGB” with latent concatenation and normalization.\nThere is no fundamentally new 4D generative architecture. This feels closer to a careful repurposing of a large pretrained model rather than a new generative paradigm.\n\n2. No true camera control or 3D consistency is learned during generation itself. The “feed-forward” claim is slightly misleading — XYZ is predicted in image-plane coordinates, not explicitly SE(3)-aware. \n\n3. The point cloud visualization in the video shows that the geometry of the generated scenes is unnatural."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762916687768,"tcdate":1761438146076,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission3354/Reviewer_sm7K"],"signatures":["ICLR.cc/2026/Conference/Submission3354/Reviewer_sm7K"],"forum":"guUrm5IRQS","number":1,"license":"CC BY 4.0","cdate":1761438146076,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission3354/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762916687768,"domain":"ICLR.cc/2026/Conference","replyto":"guUrm5IRQS","id":"0dDpZM0vJF","forumContent":{"TLDR":{"value":"4DNeX is a feed-forward framework that generates dynamic 4D scene representations from a single image, built upon our curated large-scale dataset 4DNeX-10M."},"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Image-to-4D Modeling","Generative 4D World Models","4D Dataset"]},"supplementary_material":{"value":"/attachment/935d2c9c9517f18b32872a6cb2536b6483a4271d.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"We present 4DNeX, the first feed-forward framework for generating 4D (i.e., dynamic 3D) scene representations from a single image. In contrast to existing methods that rely on computationally intensive optimization or require multi-frame video inputs, 4DNeX enables efficient, end-to-end image-to-4D generation by fine-tuning a pretrained video diffusion model. Specifically, 1) to alleviate the scarcity of 4D data, we construct 4DNeX-10M, a large-scale dataset with high-quality 4D annotations generated using advanced reconstruction approaches. 2) we introduce a unified 6D video representation that jointly models RGB and XYZ sequences, facilitating structured learning of both appearance and geometry. 3) we propose a set of simple yet effective adaptation strategies to repurpose pretrained video diffusion models for 4D modeling. 4DNeX produces high-quality dynamic point clouds that enable novel-view video synthesis. Extensive experiments demonstrate that 4DNeX outperforms existing 4D generation methods in efficiency and generalizability, offering a scalable solution for image-to-4D modeling and laying the foundation for generative 4D world models that simulate dynamic scene evolution."},"_bibtex":{"value":"@misc{\nchen2025dnex,\ntitle={4{DN}eX: Feed-Forward 4D Generative Modeling Made Easy},\nauthor={Zhaoxi Chen and Tianqi Liu and Long Zhuo and Jiawei Ren and Zeng Tao and He Zhu and Fangzhou Hong and Liang Pan and Ziwei Liu},\nyear={2025},\nurl={https://openreview.net/forum?id=guUrm5IRQS}\n}"},"title":{"value":"4DNeX: Feed-Forward 4D Generative Modeling Made Easy"},"pdf":{"value":"/pdf/894ee98bf5ad3c73854d49d115076037d37b7e47.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"chen|4dnex_feedforward_4d_generative_modeling_made_easy"},"authorids":{"value":["~Zhaoxi_Chen1","~Tianqi_Liu3","~Long_Zhuo1","~Jiawei_Ren1","~Zeng_Tao1","~He_Zhu7","~Fangzhou_Hong1","~Liang_Pan2","~Ziwei_Liu1"]},"authors":{"value":["Zhaoxi Chen","Tianqi Liu","Long Zhuo","Jiawei Ren","Zeng Tao","He Zhu","Fangzhou Hong","Liang Pan","Ziwei Liu"]}},"version":2},{"content":{"summary":{"value":"This paper tackles the controllable video relighting task in a training-free manner. For this, LightCtrl builds on a pre-trained single-image relighting model, i.e., IC-Light, and tailors it to video relighting with several designs. First, it proposes to conduct frame-wise relighting on the video diffusion model's paired VAE decoder. Besides, it proposes to progressively fuse the source video and relit video during the diffusion process to maintain temporal coherence. Further, a user-defined light trajectory injection module is applied. Finally, they use off-the-shelf normal estimation to provide the video's normal maps to enable a geometry-aware relighting module. Experiments demonstrate the effectiveness of the proposed approach."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"See \"weakness\""},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- originality-wise: the idea of using a single-image relighting model to enable video relighting is interesting.\n- quality-wise: the qualitative and quantitative results are promising.\n- clarity-wise: the presentation is good.\n- significance-wise: the video relighting is important for a lot of downstream tasks, e.g., content creation."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"I feel the ablations in the paper are quite inadequate. Even though there are some ablations in the appendix that are good, they are not the core part of the model design.\n\nFor example, can authors provide both **quantitative and qualitative** results to show the gradual improvement from the pre-trained IC-Light? From my understanding, there are several enhancements, but I have no clue which one contributes. Feel free to add things that I miss.\n- decoded frames + residual instead of raw frames (L216)\n- progressive fusion in Eq. (2)\n- geometry-aware feature in latent space\n- geometry-aware in frequency space.\n\nI am actually confused why not directly use the raw frames in the source video? Since the authors add back the difference between the raw videos and the decoded videos (L216), isn't this the same as just using the original video?"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762923009172,"tcdate":1761957723461,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission12029/Reviewer_bAHR"],"signatures":["ICLR.cc/2026/Conference/Submission12029/Reviewer_bAHR"],"forum":"5ft8vd9rwc","number":4,"license":"CC BY 4.0","cdate":1761957723461,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission12029/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762923009172,"domain":"ICLR.cc/2026/Conference","replyto":"5ft8vd9rwc","id":"DWM6rgF1OM","forumContent":{"TLDR":{"value":"LightCtrl relight a given video via a user-specific light trajectory in a zero-shot manner."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["video relighting; controllable video editing"]},"supplementary_material":{"value":"/attachment/9620cf406218d8104168386c0036f3d45fbf2f7a.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Recent diffusion models have achieved remarkable success in image relighting, and this success has quickly been reproduced in video relighting. Although these methods can relight videos under various conditions, their ability to explicitly control the illumination in the relighted video remains limited. Therefore, we present \\name, the first controllable video relighting method that offers explicit control over the video illumination through a user-supplied light trajectory in a training-free manner. This is essentially achieved by leveraging a hybrid approach that combines pre-trained diffusion models: a pre-trained image relighting diffusion model is used to relight each frame individually, followed by a video diffusion prior that enhances the temporal consistency of the relighted sequence. In particular, to enable explicit control over dynamically varying lighting in the relighted video, we introduce two key components. \nFirst, the Light Map Injection module samples light trajectory-specific noise and injects it into the latent representation of the source video, significantly enhancing illumination coherence with respect to the conditional light trajectory. \nSecond, the Geometry-Aware Relighting module dynamically combines RGB and normal map latents in the frequency domain to suppress the influence of the original lighting in the input video, thereby further improving the relighted video's adherence to the input light trajectory. \nOur experiments demonstrate that \\name can generate high-quality video results with diverse illumination changes closely following the light trajectory condition, indicating improved controllability over baseline methods. The code will be released at: https://github.com/GVCLab/LightCtrl."},"_bibtex":{"value":"@inproceedings{\npeng2026lightctrl,\ntitle={LightCtrl: Training-free Controllable Video Relighting},\nauthor={Yizuo Peng and Xuelin Chen and Kai Zhang and Xiaodong Cun},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=5ft8vd9rwc}\n}"},"title":{"value":"LightCtrl: Training-free Controllable Video Relighting"},"pdf":{"value":"/pdf/ebecc07688f72c477de1cca270c7a0498aa35be7.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"peng|lightctrl_trainingfree_controllable_video_relighting"},"authorids":{"value":["~Yizuo_Peng1","~Xuelin_Chen1","~Kai_Zhang16","~Xiaodong_Cun1"]},"authors":{"value":["Yizuo Peng","Xuelin Chen","Kai Zhang","Xiaodong Cun"]}},"version":2},{"content":{"summary":{"value":"The paper targets the problem that MLU models pick up spurious shortcuts and fail under distribution shifts. It proposes CaMIB: first apply an Information Bottleneck to each modality to filter noise, then fuse features and use a learnable mask to split them into causal and shortcut subrepresentations, guided by an instrumental variable built with self-attention to capture global inter-modal and token-level dependencies; finally, use a backdoor adjustment by randomly recombining causal and shortcut features so the model relies on causal information. Experiments on sentiment, humor, and sarcasm (including OOD tests) show better robustness and accuracy than prior methods. The work contributes an SCM for MLU that treats redundancy-induced spurious correlations as confounders, and a simple end-to-end framework that improves interpretability and OOD generalization."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. The current ablation conflates “having IV” with “having any global cross-modal/token-level dependency modeling.” To know what truly drives the gains, you need to isolate the effect of dependencies from the IV alignment itself under capacity-matched settings.\n(1) Build V via cross-modal, token-level attention but do not use the IV alignment constraint.\n(2) Keep the IV alignment constraint, but replace the attention-built V with a capacity-matched non-attention IV (e.g., SimpleSum / MeanPool, Max/Avg pooling).\n2. How do you ensure that V does not contain environment or shortcut information that could bias the mask towards spurious correlations?\n3. Please add a Pure-causal (no recombination) ablation: use only the causal representation for prediction and do not perform any random recombination with shortcut features; keep all other modules and settings the same."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"The paper proposes a task-agnostic Structural Causal Model for MLU that treats redundancy-induced spurious correlations as confounders, rather than limiting attention to predefined bias types. Building on this causal view, it introduces an attention-derived instrumental variable to anchor causal signals, a learnable mask that splits fused multimodal features into causal vs. shortcut parts, and a backdoor “remix” intervention at the representation level to break residual shortcut reliance. This combination is a fresh, creative, and lightweight approach to multimodal debiasing, and is broadly applicable across MLU tasks.\n\nRobustness to distribution shift is central in multimodal learning. By operationalizing causal principles with simple modules compatible with modern encoders, the approach is practical."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The “with IV” model includes strong cross-modal global attention; if “w/o IV” lacks a capacity-matched global dependency module, the ablation conflates “having IV” with “having any global dependency modeling,” making the comparison unfair.\n2. The alignment loss ||Zc−V||^2 encourages Zc to align with V; however, if V contains environment or shortcut information, the mask might be \"misled,\" causing the model to treat spurious correlations as causal relationships. This could undermine the validity of the causal separation.\n3. The current “w/o INTV” ablation only sets the intervention weight to zero, which conflates removing a loss term with removing the intervention mechanism itself. It does not test a setting where the model relies solely on the causal path without any random recombination with shortcut features, so it’s unclear whether the robustness gain comes from the intervention operation or simply from using the causal representation."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762931394667,"tcdate":1761902133157,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission19490/Reviewer_C3c7"],"signatures":["ICLR.cc/2026/Conference/Submission19490/Reviewer_C3c7"],"forum":"CULACouTam","number":1,"license":"CC BY 4.0","cdate":1761902133157,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission19490/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762931394667,"domain":"ICLR.cc/2026/Conference","replyto":"CULACouTam","id":"ForhkHL4BG","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Human Multimodal Language Understanding","Causal Representation","Information Bottleneck","Out-of-Distribution generalization"]},"supplementary_material":{"value":"/attachment/109a3cfc07cb724f9c248be94313d40ad9d2af2d.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Human Multimodal Language Understanding (MLU) aims to infer human intentions by integrating related cues from heterogeneous modalities. Existing works predominantly follow a ``learning to attend\" paradigm, which maximizes mutual information between data and labels to enhance predictive performance. However, such methods are vulnerable to unintended dataset biases, causing models to conflate statistical shortcuts with genuine causal features and resulting in degraded out-of-distribution (OOD) generalization. To alleviate this issue, we introduce a Causal Multimodal Information Bottleneck (CaMIB) model that leverages causal principles rather than traditional likelihood. Concretely, we first applies the information bottleneck to filter unimodal inputs, removing task-irrelevant noise. A parameterized mask generator then disentangles the fused multimodal representation into causal and shortcut subrepresentations. To ensure global consistency of causal features, we incorporate an instrumental variable constraint, and further adopt backdoor adjustment by randomly recombining causal and shortcut features to stabilize causal estimation. Extensive experiments on multimodal sentiment analysis, humor detection, and sarcasm detection, along with OOD test sets, demonstrate the effectiveness of CaMIB. Theoretical and empirical analyses further highlight its interpretability and soundness."},"_bibtex":{"value":"@misc{\njiang2026towards,\ntitle={Towards Minimal Causal Representations for Human Multimodal Language Understanding},\nauthor={Menghua Jiang and Yuncheng Jiang and Haifeng Hu and Sijie Mai},\nyear={2026},\nurl={https://openreview.net/forum?id=CULACouTam}\n}"},"title":{"value":"Towards Minimal Causal Representations for Human Multimodal Language Understanding"},"pdf":{"value":"/pdf/3686917b41ef8328da878a997d6605b745904cd8.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"jiang|towards_minimal_causal_representations_for_human_multimodal_language_understanding"},"authorids":{"value":["~Menghua_Jiang1","~Yuncheng_Jiang1","~Haifeng_Hu3","~Sijie_Mai1"]},"authors":{"value":["Menghua Jiang","Yuncheng Jiang","Haifeng Hu","Sijie Mai"]}},"version":2},{"content":{"summary":{"value":"In this paper, the authors introduce SVBench, a novel benchmark designed to evaluate the capabilities of LVLMs in long-context streaming video understanding. SVBench features temporal multi-turn question-answering chains comprising 49,979 QA pairs across 1,353 streaming videos. The study reveals that while the closed-source GPT-4o model outperforms others, most open-source LVLMs struggle with long-context streaming video understanding. \n\nThe authors also develop StreamingChat, a model that outperforms existing open-source LVLMs on SVBench and achieves comparable performance on diverse vision-language benchmarks."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Q1: The videos of the benchmark come from some publicly available datasets; how do you ensure that the models used for evaluation have not encountered this data during their training?\n\nQ2: How do QA chains address the issue of information sparsity in videos? For lengthy videos with few key events, how does SVBench construct meaningful QA chains? Is there a mechanism to handle or generate question chains that effectively deal with scenarios where information is sparse or low-density?\n\nQ3: In your benchmark, do the videos retain their audio tracks? Multimodal information significantly aids in understanding videos. If audio is included, does it genuinely assist in answering the questions?"},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":3},"strengths":{"value":"S1: The paper introduces SVBench, a benchmark explicitly designed for evaluating LVLMs in long-context streaming video understanding. I believe this fills a notable gap in existing benchmarks, which typically focus on isolated text inputs rather than sustained temporal reasoning across video streams. The comparison between the benchmarks also shows the advantages of SVBench.\n\nS2: The QA pairs in this work are annotated semi-automatically, making it a large-scale dataset with high-quality annotations.\n\nS3: The experiments and evaluation provide insights into open-source models' struggles with long-context video understanding.\n\nS4: The paper is well-organized, with clear explanations of SVBench, the QA chain structure, and the temporal linkages created for sustained reasoning tasks."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"W1: The paper does not include a comparison with human performance. Incorporating such a comparison would provide valuable insights into the gap between current models and human capabilities in long-context streaming video understanding.\n\nW2: The paper does not analyze the impact of language model size on performance. Considering that models like InternVL2 have versions with 1B, 2B, 4B, 8B, 26B, 40B, and 72B parameters, and Video-LLaMA2 also have 72B versions, expanding experiments to include these variations and providing more detailed analysis would enhance the work. Additionally, exploring the number of frames the model can process would offer valuable insights. In addition to model size and the input length, analyzing the amount and diversity of training data used for each model would provide a more comprehensive understanding of their performance in long-context streaming video understanding. Training data quality and quantity are crucial factors influencing model capabilities.\n\nW3: The evaluation overlooks several important models capable of long-video understanding, such as LLaVA-OneVision [1], Qwen2-VL [2], LongVILA [3], Long-LLaVA [4], and Oryx [5]. Including these models would enhance the comparison, providing a more comprehensive view of current capabilities in this area.\n\n[1] https://github.com/LLaVA-VL/LLaVA-NeXT\n\n[2] https://github.com/QwenLM/Qwen2-VL\n\n[3] https://github.com/NVlabs/VILA/blob/main/LongVILA.md\n\n[4] https://github.com/FreedomIntelligence/LongLLaVA\n\n[5] https://github.com/Oryx-mllm/Oryx"}},"nonreaders":[],"tmdate":1731428906113,"tcdate":1730669056374,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission7487/Reviewer_k44W"],"signatures":["ICLR.cc/2025/Conference/Submission7487/Reviewer_k44W"],"forum":"Hz4BYVY8YM","number":3,"license":"CC BY 4.0","cdate":1730669056374,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission7487/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428906113,"domain":"ICLR.cc/2025/Conference","replyto":"Hz4BYVY8YM","id":"oECM4PRYKA","forumContent":{"venue":{"value":"ICLR 2025 Spotlight"},"TLDR":{"value":"A benchmark with temporal multi-turn dialogues specifically designed to thoroughly assess the capabilities of long-context streaming video understanding of current LVLMs."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Multimodal large language model","Streaming video analysis","Video understanding"]},"primary_area":{"value":"datasets and benchmarks"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Despite the significant advancements of Large Vision-Language Models (LVLMs) on established benchmarks, there remains a notable gap in suitable evaluation regarding their applicability in the emerging domain of long-context streaming video understanding. Current benchmarks for video understanding typically emphasize isolated single-instance text inputs and fail to evaluate the capacity to sustain temporal reasoning throughout the entire duration of video streams. To address these limitations, we introduce SVBench, a pioneering benchmark with temporal multi-turn question-answering chains specifically designed to thoroughly assess the capabilities of streaming video understanding of current LVLMs. We design a semi-automated annotation pipeline to obtain 49,979 Question-Answer (QA) pairs of 1,353 streaming videos, which includes generating QA chains that represent a series of consecutive multi-turn dialogues over video segments and constructing temporal linkages between successive QA chains. Our experimental results, obtained from 14 models in dialogue and streaming evaluations, reveal that while the closed-source GPT-4o outperforms others, most open-source LVLMs struggle with long-context streaming video understanding. We also construct a StreamingChat model, which significantly outperforms open-source LVLMs on our SVBench and achieves comparable performance on diverse vision-language benchmarks. We expect SVBench to advance the research of streaming video understanding by providing a comprehensive and in-depth analysis of current LVLMs. Our benchmark and model can be accessed at https://yzy-bupt.github.io/SVBench."},"_bibtex":{"value":"@inproceedings{\nyang2025svbench,\ntitle={{SVB}ench: A Benchmark with Temporal Multi-Turn Dialogues for Streaming Video Understanding},\nauthor={Zhenyu Yang and Yuhang Hu and Zemin Du and Dizhan Xue and Shengsheng Qian and Jiahong Wu and Fan Yang and Weiming Dong and Changsheng Xu},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=Hz4BYVY8YM}\n}"},"title":{"value":"SVBench: A Benchmark with Temporal Multi-Turn Dialogues for Streaming Video Understanding"},"pdf":{"value":"/pdf/570cf446ad6d2a9682beb96af5d08dfb6b98d95a.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"yang|svbench_a_benchmark_with_temporal_multiturn_dialogues_for_streaming_video_understanding"},"authorids":{"value":["~Zhenyu_Yang6","~Yuhang_Hu3","~Zemin_Du2","~Dizhan_Xue1","~Shengsheng_Qian1","~Jiahong_Wu2","~Fan_Yang30","~Weiming_Dong1","~Changsheng_Xu1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Zhenyu Yang","Yuhang Hu","Zemin Du","Dizhan Xue","Shengsheng Qian","Jiahong Wu","Fan Yang","Weiming Dong","Changsheng Xu"]}},"version":2},{"content":{"summary":{"value":"This paper tackles long horizon decision making in robotics where classic task and motion planning relies on hand specified skills and often yields long satisficing plans. The authors propose Shortcut Learning for Abstract Planning, or SLAP, which automatically learns new low level options that act as shortcuts between abstract states in a planning graph induced by the existing skill library. The system builds a two level abstract planning graph, searches for shortest executions at the low level, then augments the graph with learned options trained through model free RL in self contained shortcut MDPs with step penalties and goal termination. A simple but effective pruning strategy screens shortcut candidates using random rollouts before launching PPO training. At test time, SLAP plans with both given skills and learned options, prunes failing edges online, and selects shorter plans via Dijkstra on the ground graph. To generalize beyond the training object set, SLAP projects observations onto relevant objects determined from add and delete atom sets and applies object substitution to reuse learned policies when object identities differ. This preserves the relational inductive bias of TAMP while enabling dynamic behaviors beyond fingertip grasp and place such as slap, wiggle, and wipe."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"- How many shortcut policies were ultimately trained and kept per environment during the runs reported in Table 1, and what was the total wall clock training time per environment including pruning and PPO updates?\n- What termination condition is used for each learned option at test time beyond the abstract state check and step limit. Do you implement safety guards such as maximum end effector velocity or minimum clearance during dynamic actions like a slap or wipe?\n- You evaluate separate PPO policies per shortcut pair. How do results compare to a single goal conditioned policy trained across shortcut goals when both are tuned equally, and how does data efficiency change as the number of abstract states grows?\n- The object substitution criterion checks add and delete inclusions. How is the mapping computed in practice when multiple candidates exist, and what is its computational cost during planning on the larger PyBullet graphs?\n- For hierarchical RL, did you try curriculum learning, intrinsic motivation, or hindsight relabeling at the high level to mitigate sparse rewards before concluding failure on the three PyBullet domains?"},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"Framing and practical value\n- Clear problem statement that existing TAMP systems rely on hand designed skills which limits efficiency and expressivity. SLAP focuses squarely on reducing execution time without discarding the benefits of abstraction and search.\n- Elegant algorithmic design that keeps planning and learning modular. The abstract planning graph yields a well defined search problem, and shortcuts are learned in parallel MDPs with simple goal conditions and dense step penalties.\n- Relational generalization via relevant object projection and object substitution is well motivated and effective, letting the same shortcut policy apply across object sets and counts.\n\nEmpirical evidence\n- Consistent improvements in plan length across four diverse domains with long horizons and sparse rewards. Reductions are large in magnitude, for example about 69 percent in Obstacle Tower and over 50 percent in the PyBullet tasks.\n- Clear wins over strong baselines representing three regimes: pure planning, pure RL with PPO and SAC plus HER, and hierarchical RL with access to the same predefined skills.\n- Training steps analysis and shortcut discovery counts show monotonic gains as more shortcuts are learned, connecting learning progress to planning improvements.\n- Generalization experiments demonstrate robustness to different numbers of obstacles and to additional distractor objects. The learned dynamic behaviors like slap and wipe act on multiple objects and keep plan lengths stable."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Assumptions and scope\n- Relies on a known transition function and fully observable deterministic settings for graph construction, although variants relax these assumptions. Real systems often face sensing delays, latency, and controller noise which may require tighter integration of failure recovery and uncertainty aware planning.\n- Assumes the provided option set enables task completion. When this is not true, completeness can be lost. Appendix results discuss such cases but a stronger treatment of failure detection and fallback would help.\n\nScalability and compute\n- The number of candidate shortcuts scales with the square of the number of abstract states. The pruning heuristic is simple and effective but a learned prioritizer or search over the shortcut space could further reduce training cost. Wall clock budgets and GPU hours per environment would make the compute footprint transparent.\n- Planning time modestly increases compared to pure planning in two domains before being outweighed by shorter execution. A more thorough analysis of search heuristics and graph size versus latency would help practitioners choose configurations."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762942071616,"tcdate":1761987561968,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission22118/Reviewer_AFvJ"],"signatures":["ICLR.cc/2026/Conference/Submission22118/Reviewer_AFvJ"],"forum":"enprG5H9aD","number":4,"license":"CC BY 4.0","cdate":1761987561968,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission22118/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762942071616,"domain":"ICLR.cc/2026/Conference","replyto":"enprG5H9aD","id":"ScDvSP9YeO","forumContent":{"TLDR":{"value":"We use RL to learn shortcuts in the abstract planning graph induced by predefined options."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Robot Planning","Reinforcement Learning"]},"supplementary_material":{"value":"/attachment/963c7ff7b5a8b143b0ebb6b42b4266d8ba9dbeab.zip"},"primary_area":{"value":"applications to robotics, autonomy, planning"},"abstract":{"value":"Long-horizon decision-making with sparse rewards and continuous states and actions remains a fundamental challenge in AI and robotics. Task and motion planning (TAMP) is a model-based framework that addresses this challenge by planning hierarchically with abstract actions (options). These options are manually defined, limiting the agent to behaviors that we as human engineers know how to program (pick, place, move). In this work, we propose Shortcut Learning for Abstract Planning (SLAP), a method that leverages existing TAMP options to automatically discover new ones. Our key idea is to use model-free reinforcement learning (RL) to learn *shortcuts* in the abstract planning graph induced by the existing options in TAMP. Without any additional assumptions or inputs, shortcut learning leads to shorter solutions than pure planning, and higher task success rates than flat and hierarchical RL. Qualitatively, SLAP discovers dynamic physical improvisations (e.g., slap, wiggle, wipe) that differ significantly from the manually-defined ones. In experiments in four simulated robotic environments, we show that SLAP solves and generalizes to a wide range of tasks, reducing overall plan lengths by over 50\\% and consistently outperforming planning and RL baselines."},"_bibtex":{"value":"@inproceedings{\nliu2026slap,\ntitle={{SLAP}: Shortcut Learning for Abstract Planning},\nauthor={Y. Isabel Liu and Bowen Li and Benjamin Eysenbach and Tom Silver},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=enprG5H9aD}\n}"},"title":{"value":"SLAP: Shortcut Learning for Abstract Planning"},"pdf":{"value":"/pdf/4f54aa8a88c1a1d21a3c9878134690e1432728d8.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"liu|slap_shortcut_learning_for_abstract_planning"},"authorids":{"value":["~Y._Isabel_Liu1","~Bowen_Li7","~Benjamin_Eysenbach1","~Tom_Silver1"]},"authors":{"value":["Y. Isabel Liu","Bowen Li","Benjamin Eysenbach","Tom Silver"]}},"version":2},{"content":{"summary":{"value":"This paper proposes Stable Video-Driven Portraits, a video-driven portrait generation framework based on the Diffusion Transformer (DiT) architecture. The model takes a single source image and a driving video as input and synthesizes high-fidelity dynamic videos that reenact the target person’s motion and expressions.\nThe main contributions claimed are as follows:\n1. Adoption of a cross-identity training scheme combined with facial-region masking (eyes, nose, mouth) to prevent identity leakage;\n2. Minimal-parameter adaptation on top of Stable Diffusion 3.5, achieving faster convergence and improved generalization;\n3. Introduction of a full spatio-temporal attention mechanism to capture fine-grained motion and enhance temporal coherence."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1. Since LivePortrait-generated paired data are used for training, how does the proposed model outperform LivePortrait itself?\n2. Given that only masked facial regions (eyes, nose, mouth) are used as motion input, why do the generated results still include head pose movements?\n3. In classifier-free guidance, all control signals—including the source image—are discarded in one configuration. Does this impact identity preservation?\n4. Do the authors plan to release the internal 20k video dataset? If not, please describe its distribution, diversity, and sampling details.\n5. Could the proposed framework be extended to audio-driven or emotion-conditioned portrait generation?\n6. Have the authors evaluated longer sequences (e.g., >16 frames) or longer reference clips (>3 frames) for temporal consistency analysis?\n7. Does the full-video attention scale quadratically with frame length, and how is inference efficiency maintained in such cases?"},"rating":{"value":2},"details_of_ethics_concerns":{"value":"None"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. The Supplementary Material includes detailed technical explanations and training pipeline descriptions, which improve readability and aid understanding.\n2. The experiments are comprehensive, covering self-reenactment, cross-reenactment, and stylized portrait animation.\n3. The evaluation metrics are diverse, including PSNR, FVD, CSIM, and MAE, allowing a thorough quantitative evaluation.\n4. The method effectively leverages the pretrained SD3.5 model with minimal parameter adjustments, demonstrating practical adaptation of large diffusion transformers.\n5. The cross-identity training strategy effectively prevents appearance leakage and improves generalization.\n6. The full spatio-temporal attention mechanism significantly enhances temporal smoothness and consistency in generated videos.\n7. The ablation studies verify the rationale of the training curriculum and attention structure.\n8. The paper includes comparisons with non-Diffusion Transformer architectures, validating the advantages of the proposed DiT-based design."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Writing and organizational quality are significantly below ICLR standards. The paper suffers from vague phrasing, redundant sections, missing implementation details, and low-quality figure presentation. For example, the Broader Impact and Societal Impact sections in the supplementary material repeat nearly identical content, showing a lack of editorial care. The overall narrative feels fragmented and lacks the polished academic tone typically found in ICLR-accepted papers.\n2. The abstract and introduction are overly generic and imprecise. Keywords such as “minimal new parameters” and “superior temporal consistency” are not supported with specific numbers or quantitative evidence. For instance, it is unclear exactly how many additional parameters were introduced or what percentage of the base SD3.5 model this represents. These missing details weaken the clarity and perceived magnitude of the contribution.\n3. Limited novelty. The core technical design heavily depends on Stable Diffusion 3.5 with only minor adaptations. The use of masked facial regions (eyes, nose, mouth) as input features is already a well-established practice in talking-head generation literature. Consequently, the originality of the approach appears limited and largely incremental.\n4. Method–result mismatch. The paper states that only masked facial regions are used for motion input, which theoretically should exclude head pose changes. However, the supplementary video results clearly show head movements, suggesting inconsistencies between the described methodology and the observed behavior.\n5. Excessive reliance on LivePortrait. The training pairs are generated using LivePortrait, which likely introduces error accumulation. Moreover, the data generation process using LivePortrait is not well-documented. It is also unclear why a model trained on LivePortrait-generated data produces results superior to LivePortrait itself, and no ablation or justification is provided for this discrepancy.\n6. Insufficient comparison with other methods. The experiments compare against only three baselines, which is inadequate for a top-tier conference. Stronger and more recent competitors such as VividPortraits, DiffPortrait, and AniFace should be included to strengthen the argument.\n7. Implementation details are lacking. Critical aspects such as token dimensionality, dropout scheduling, temporal encoding design, and optimizer settings are omitted, making reproducibility difficult.\n8. Ablation studies are incomplete. The paper only provides ablations on attention factorization and training curriculum. Important factors such as region-wise mask contributions (eyes vs. mouth vs. nose), sensitivity to fusion weight λ, and additional parameter count are missing. Including these would significantly improve the paper’s technical credibility.\n9. Internal dataset not released. The claimed 20k internal dataset is neither released nor described in sufficient detail (e.g., domain diversity, subject identity distribution, video length), preventing reproducibility and verification by the research community."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762923946430,"tcdate":1761300037727,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission13269/Reviewer_rTAL"],"signatures":["ICLR.cc/2026/Conference/Submission13269/Reviewer_rTAL"],"forum":"dpVQPM6P3q","number":1,"license":"CC BY 4.0","cdate":1761300037727,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission13269/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762923946430,"domain":"ICLR.cc/2026/Conference","replyto":"dpVQPM6P3q","id":"tpVCUJrplr","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Diffusion model","face","reenactment"]},"supplementary_material":{"value":"/attachment/67dc3f56d8f55ac9b10fcbd2e3e8c5dbc0ed51e1.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Portrait animation aims to generate photo-realistic videos from a single source image by reenacting the expression and pose from a driving video. While early methods relied on 3D morphable models or feature warping techniques, they often suffered from limited expressivity, temporal inconsistency, and poor generalization to unseen identities or large pose variations. Recent advances using diffusion models have demonstrated improved quality but remain constrained by weak control signals and architectural limitations. In this work, we propose a novel diffusion-based framework that leverages masked facial regions—specifically the eyes, nose, and mouth—from the driving video as strong motion control cues. To enable robust training without appearance leakage, we adopt cross-identity supervision. To leverage the strong prior from the pre-trained diffusion model, our novel architecture introduces minimal new parameters that converge faster and help in better generalization. We introduce spatial-temporal attention mechanisms that allow inter-frame and intra-frame interactions, effectively capturing subtle motions and reducing temporal artifacts. Our model uses history frames to ensure continuity across segments. At inference, we propose a novel signal fusion strategy that balances motion fidelity with identity preservation. Our approach achieves superior temporal consistency and accurate expression control, enabling high-quality, controllable portrait animation suitable for real-world applications."},"_bibtex":{"value":"@misc{\nreddy2026stable,\ntitle={Stable Video-Driven Portraits},\nauthor={Mallikarjun Byrasandra Ramalinga Reddy and Fei Yin and Vikram Voleti and Nikita Drobyshev and Maksim Lapin and Aaryaman Vasishta and Varun Jampani},\nyear={2026},\nurl={https://openreview.net/forum?id=dpVQPM6P3q}\n}"},"title":{"value":"Stable Video-Driven Portraits"},"pdf":{"value":"/pdf/a16ccb4e94a4164853b85295ae2cbb7acd5e9c39.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"reddy|stable_videodriven_portraits"},"authorids":{"value":["~Mallikarjun_Byrasandra_Ramalinga_Reddy2","~Fei_Yin3","~Vikram_Voleti1","~Nikita_Drobyshev1","~Maksim_Lapin1","~Aaryaman_Vasishta1","~Varun_Jampani2"]},"authors":{"value":["Mallikarjun Byrasandra Ramalinga Reddy","Fei Yin","Vikram Voleti","Nikita Drobyshev","Maksim Lapin","Aaryaman Vasishta","Varun Jampani"]}},"version":2},{"content":{"venue":{"value":"ACM Trans. Intell. Syst. Technol. 2024"},"venueid":{"value":"dblp.org/journals/TIST/2024"},"paperhash":{"value":"li|perceiving_actions_via_temporal_video_frame_pairs"},"authorids":{"value":["~Rongchang_Li1","~Tianyang_Xu1","https://dblp.org/search/pid/api?q=author:Xiaojun_Wu_0001:","~Zhongwei_Shen3","https://dblp.org/search/pid/api?q=author:Josef_Kittler:"]},"html":{"value":"https://doi.org/10.1145/3652611"},"_bibtex":{"value":"@article{DBLP:journals/tist/LiXWSK24,\n  author={Rongchang Li and Tianyang Xu and Xiaojun Wu and Zhongwei Shen and Josef Kittler},\n  title={Perceiving Actions via Temporal Video Frame Pairs},\n  year={2024},\n  month={June},\n  cdate={1717200000000},\n  journal={ACM Trans. Intell. Syst. Technol.},\n  volume={15},\n  number={3},\n  pages={58:1-58:20},\n  url={https://doi.org/10.1145/3652611}\n}\n"},"abstract":{"value":"Video action recognition aims at classifying the action category in given videos. In general, semantic-relevant video frame pairs reflect significant action patterns such as object appearance variation and abstract temporal concepts like speed, rhythm, and so on. However, existing action recognition approaches tend to holistically extract spatiotemporal features. Though effective, there is still a risk of neglecting the crucial action features occurring across frames with a long-term temporal span. Motivated by this, in this article, we propose to perceive actions via frame pairs directly and devise a novel Nest Structure with frame pairs as basic units. Specifically, we decompose a video sequence into all possible frame pairs and hierarchically organize them according to temporal frequency and order, thus transforming the original video sequence into a Nest Structure. Through naturally decomposing actions, the proposed structure can flexibly adapt to diverse action variations such as speed or rhythm changes. Next, we devise a Temporal Pair Analysis module (TPA) to extract discriminative action patterns based on the proposed Nest Structure. The designed TPA module consists of a pair calculation part to calculate the pair features and a pair fusion part to hierarchically fuse the pair features for recognizing actions. The proposed TPA can be flexibly integrated into existing backbones, serving as a side branch to capture various action patterns from multi-level features. Extensive experiments show that the proposed TPA module can achieve consistent improvements over several typical backbones, reaching or updating CNN-based SOTA results on several challenging action recognition benchmarks."},"title":{"value":"Perceiving Actions via Temporal Video Frame Pairs"},"authors":{"value":["Rongchang Li","Tianyang Xu","Xiaojun Wu","Zhongwei Shen","Josef Kittler"]}},"tmdate":1784434160831,"pdate":1704067200000,"tcdate":1731465989639,"writers":["~"],"signatures":["~Tianyang_Xu1"],"forum":"Tt5t7i0yEC","license":"CC BY-SA 4.0","number":191649,"cdate":1717200000000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1784434160831,"domain":"DBLP.org","id":"Tt5t7i0yEC","version":2},{"content":{"summary":{"value":"The paper presents a diffusion-based framework for portrait animation that generates realistic talking-head videos from a single source image and a driving video. Unlike earlier warping or landmark-based methods that struggle with temporal stability and subtle motion, the proposed approach leverages the powerful spatial-temporal reasoning of Diffusion Transformers (DiT) to achieve stable, identity-preserving reenactment. The model reuses the pretrained DiT backbone with minimal new parameters, maintaining pretrained priors while allowing conditional control. A full-video attention mechanism captures inter- and intra-frame dependencies, and history frames ensure continuity across temporal chunks. During inference, a multi-configuration CFG-style fusion balances motion fidelity with identity consistency."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"Please check the weakness part."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"This paper presents a clean and stable implementation of video-driven portrait generation built on the powerful SD3.5, demonstrating that large pretrained T2I diffusion backbones can be effectively adapted for temporal control with minimal modification, showing clear improvements in visual fidelity, temporal consistency, and expression synchronization compared to prior U-Net–based methods such as X-Portrait and LivePortrait."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The method is largely an incremental upgrade of X-Portrait—essentially porting the cross-id driving signal and masked-region conditioning to the SD3.5 Diffusion Transformer backbone.\n2. The method section is notably under-detailed and heavily relies on a single overview figure to convey the overall architecture.\n3. The paper’s cross-identity supervision fully depends on LivePortrait-generated videos to form training pairs. However, if LivePortrait fails, or is insensitive to certain extreme or rare motions, this will prevente the DiT from learning to handle such challenging dynamics.\n4. The experimental comparison is unbalanced and somewhat outdated. Most baselines (LivePortrait, AniPortrait, X-Portrait) are based on U-Net or landmark-driven architectures, while the proposed method uses a much stronger SD3.5 DiT backbone."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762923943221,"tcdate":1761978087800,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission13269/Reviewer_Nrdd"],"signatures":["ICLR.cc/2026/Conference/Submission13269/Reviewer_Nrdd"],"forum":"dpVQPM6P3q","number":4,"license":"CC BY 4.0","cdate":1761978087800,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission13269/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762923943221,"domain":"ICLR.cc/2026/Conference","replyto":"dpVQPM6P3q","id":"WOa1VAWaFR","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Diffusion model","face","reenactment"]},"supplementary_material":{"value":"/attachment/67dc3f56d8f55ac9b10fcbd2e3e8c5dbc0ed51e1.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Portrait animation aims to generate photo-realistic videos from a single source image by reenacting the expression and pose from a driving video. While early methods relied on 3D morphable models or feature warping techniques, they often suffered from limited expressivity, temporal inconsistency, and poor generalization to unseen identities or large pose variations. Recent advances using diffusion models have demonstrated improved quality but remain constrained by weak control signals and architectural limitations. In this work, we propose a novel diffusion-based framework that leverages masked facial regions—specifically the eyes, nose, and mouth—from the driving video as strong motion control cues. To enable robust training without appearance leakage, we adopt cross-identity supervision. To leverage the strong prior from the pre-trained diffusion model, our novel architecture introduces minimal new parameters that converge faster and help in better generalization. We introduce spatial-temporal attention mechanisms that allow inter-frame and intra-frame interactions, effectively capturing subtle motions and reducing temporal artifacts. Our model uses history frames to ensure continuity across segments. At inference, we propose a novel signal fusion strategy that balances motion fidelity with identity preservation. Our approach achieves superior temporal consistency and accurate expression control, enabling high-quality, controllable portrait animation suitable for real-world applications."},"_bibtex":{"value":"@misc{\nreddy2026stable,\ntitle={Stable Video-Driven Portraits},\nauthor={Mallikarjun Byrasandra Ramalinga Reddy and Fei Yin and Vikram Voleti and Nikita Drobyshev and Maksim Lapin and Aaryaman Vasishta and Varun Jampani},\nyear={2026},\nurl={https://openreview.net/forum?id=dpVQPM6P3q}\n}"},"title":{"value":"Stable Video-Driven Portraits"},"pdf":{"value":"/pdf/a16ccb4e94a4164853b85295ae2cbb7acd5e9c39.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"reddy|stable_videodriven_portraits"},"authorids":{"value":["~Mallikarjun_Byrasandra_Ramalinga_Reddy2","~Fei_Yin3","~Vikram_Voleti1","~Nikita_Drobyshev1","~Maksim_Lapin1","~Aaryaman_Vasishta1","~Varun_Jampani2"]},"authors":{"value":["Mallikarjun Byrasandra Ramalinga Reddy","Fei Yin","Vikram Voleti","Nikita Drobyshev","Maksim Lapin","Aaryaman Vasishta","Varun Jampani"]}},"version":2},{"content":{"TLDR":{"value":"We build a multi-discipline multi-faceted multimodal video understanding benchmark towards world model evaluation."},"venue":{"value":"Video-Langauge Models Poster"},"pdf":{"value":"/pdf/7c026d60dc74d947ac65c9e87abb525ba43d887c.pdf"},"keywords":{"value":["Video Understanding","Benchmark"]},"venueid":{"value":"NeurIPS.cc/2024/Workshop/Video-Langauge_Models"},"paperhash":{"value":"he|mmworld_towards_multidiscipline_multifaceted_world_model_evaluation_in_videos"},"authorids":{"value":["~Xuehai_He1","~Weixi_Feng2","~Kaizhi_Zheng1","~Yujie_Lu1","~Wanrong_Zhu1","~Jiachen_Li6","~Yue_Fan3","~Jianfeng_Wang4","~Linjie_Li1","~Zhengyuan_Yang1","~Kevin_Lin3","~William_Yang_Wang2","~Lijuan_Wang1","~Xin_Eric_Wang2"]},"abstract":{"value":"Multimodal Language Language Models (MLLMs) demonstrate the emerging abilities of \"world models\"---interpreting and reasoning about complex real-world dynamics. To assess these abilities, we posit videos are the ideal medium, as they encapsulate rich representations of real-world dynamics and causalities. To this end, we introduce MMWorld, a new benchmark for multi-discipline, multi-faceted multimodal video understanding. MMWorld distinguishes itself from previous video understanding benchmarks with two unique advantages: (1) multi-discipline, covering various disciplines that often require domain expertise for comprehensive understanding; (2) multi-faceted reasoning, including explanation, counterfactual thinking, future prediction, etc. MMWorld consists of a human-annotated dataset to evaluate MLLMs with questions about the whole videos and a synthetic dataset to analyze MLLMs within a single modality of perception. Together, MMWorld encompasses 1,910 videos across seven broad disciplines and 69 subdisciplines, complete with 6,627 question-answer pairs and associated captions. The evaluation includes 2 proprietary and 10 open-source MLLMs, which struggle on MMWorld (e.g., GPT-4V performs the best with only 52.3% accuracy), showing large room for improvement. Further ablation studies reveal other interesting findings such as models' different skill sets from humans. We hope MMWorld can serve as an essential step towards world model evaluation in videos."},"_bibtex":{"value":"@inproceedings{\nhe2025mmworld,\ntitle={{MMW}orld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos},\nauthor={Xuehai He and Weixi Feng and Kaizhi Zheng and Yujie Lu and Wanrong Zhu and Jiachen Li and Yue Fan and Jianfeng Wang and Linjie Li and Zhengyuan Yang and Kevin Lin and William Yang Wang and Lijuan Wang and Xin Eric Wang},\nbooktitle={Workshop on Video-Language Models @ NeurIPS 2024},\nyear={2025},\nurl={https://openreview.net/forum?id=rsjkujjBts}\n}"},"title":{"value":"MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos"},"track":{"value":"Short Paper Track (up to 3 pages)"},"authors":{"value":["Xuehai He","Weixi Feng","Kaizhi Zheng","Yujie Lu","Wanrong Zhu","Jiachen Li","Yue Fan","Jianfeng Wang","Linjie Li","Zhengyuan Yang","Kevin Lin","William Yang Wang","Lijuan Wang","Xin Eric Wang"]}},"tmdate":1736861080409,"pdate":1730081752582,"tcdate":1725917886708,"writers":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission25/Authors"],"signatures":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission25/Authors"],"forum":"rsjkujjBts","license":"CC BY 4.0","number":25,"cdate":1725917886708,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Post_Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Edit"],"mdate":1736861080409,"odate":1736861080395,"domain":"NeurIPS.cc/2024/Workshop/Video-Langauge_Models","id":"rsjkujjBts","version":2},{"content":{"venue":{"value":"CoRR 2026"},"pdf":{"value":"https://arxiv.org/pdf/2603.16195v2"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"yan|svam_shortcut_videoaction_model_by_selfdistilling_geometric_and_semantic_foresight"},"html":{"value":"https://doi.org/10.48550/arXiv.2603.16195"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2603-16195,\n  publtype={informal},\n  author={Haodong Yan and Zhide Zhong and Jiaguan Zhu and Junjie He and Weilin Yuan and Wenxuan Song and Xin Gong and Yingjie Cai and Guanyi Zhao and Xu Yan and Bingbing Liu and Ying-Cong Chen and Haoang Li},\n  title={S-VAM: Shortcut Video-Action Model by Self-Distilling Geometric and Semantic Foresight},\n  year={2026},\n  month={March},\n  cdate={1772323200000},\n  journal={CoRR},\n  volume={abs/2603.16195},\n  url={https://doi.org/10.48550/arXiv.2603.16195}\n}\n"},"abstract":{"value":"Video action models (VAMs) have emerged as a promising paradigm for robot learning, owing to their powerful visual foresight for complex manipulation tasks. However, current VAMs, typically relying on either slow multi-step video generation or noisy one-step feature extraction, cannot simultaneously guarantee real-time inference and high-fidelity foresight. To address this limitation, we propose S-VAM, a shortcut video-action model that foresees coherent geometric and semantic representations via a single forward pass. Serving as a stable blueprint, these foreseen representations significantly simplify the action prediction. To enable this efficient shortcut, we introduce a novel self-distillation strategy that condenses structured generative priors of multi-step denoising into one-step inference. Specifically, vision foundation model (VFM) representations extracted from the diffusion model's own multi-step generated videos provide teacher targets. Lightweight decouplers, as students, learn to directly map noisy one-step features to these targets. Extensive experiments in simulation and the real world demonstrate that our S-VAM outperforms state-of-the-art methods, enabling efficient and precise manipulation in complex environments. Our project page is https://haodong-yan.github.io/S-VAM/"},"title":{"value":"S-VAM: Shortcut Video-Action Model by Self-Distilling Geometric and Semantic Foresight"},"authors":{"value":[{"fullname":"Haodong Yan","username":""},{"fullname":"Zhide Zhong","username":"~Zhide_Zhong1"},{"fullname":"Jiaguan Zhu","username":""},{"fullname":"Junjie He","username":""},{"fullname":"Weilin Yuan","username":""},{"fullname":"Wenxuan Song","username":""},{"fullname":"Xin Gong","username":""},{"fullname":"Yingjie Cai","username":""},{"fullname":"Guanyi Zhao","username":""},{"fullname":"Xu Yan","username":""},{"fullname":"Bingbing Liu","username":""},{"fullname":"Ying-Cong Chen","username":""},{"fullname":"Haoang Li","username":""}]}},"tmdate":1785321645655,"pdate":1798675200000,"externalIds":["dblp:journals/corr/abs-2603-16195"],"tcdate":1785321639389,"writers":["~","OpenReview.net/Public_Article/DBLP.org","OpenReview.net/Support"],"signatures":["~Zhide_Zhong1"],"forum":"cPK0u6Jw4d","license":"CC BY-SA 4.0","number":111360,"cdate":1772323200000,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/DBLP.org/-/Record","OpenReview.net/Public_Article/-/Edit"],"mdate":1785321645655,"domain":"OpenReview.net/Public_Article","id":"cPK0u6Jw4d","version":2},{"content":{"summary":{"value":"This paper presents LIFT (Logging Improvement via Fine-Tuned Trajectories), a framework for data augmentation in offline RL applied to active positioning tasks, such as optical or robotic alignment. The key idea is to exploit the geometric structure of positioning tasks to derive “shortcut augmentations”, i.e., trajectory-level perturbations that preserve the underlying task geometry while improving sample diversity and policy support. The authors propose two complementary modes of augmentation: 1) Static trajectory augmentation, which uses structure-aware perturbations (shortcuts) derived from the geometry of the transition dynamics and value functions. 2) Policy-time augmentation, which injects optimistic off-policy actions into the logging process, guided by a Q-function trained on augmented data. They derive theoretical guarantees for when these shortcuts improve expected return, under assumptions of Lipschitz continuity of the value function, linear placement error (LPE) in the transition function, and f-contraction of the policy. Empirically, the method is validated on synthetic and semi-realistic active positioning environments with various movement distortions and observation models (e.g., 2D/5D positional data, optical image generators). Results show that LIFT and its variant LIFT-SC (with shortcut-augmented CQL training) improve data quality and final policy performance over CQL, IORL, and warm-start SAC baselines."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"Q1: The theoretical results presume accurate $V_{\\pi}$. In practice, when $V_{\\pi}$ is estimated from noisy data or with function approximation, how robust is Algorithm 1 to error propagation in shortcut detection?\n\nQ2: The augmentor is trained using synthetic shortcuts to guide real data collection. Given the policy–Q coupling, does this training exhibit instability akin to standard off-policy actor–critic divergence?\n\nQ3: Many domains (e.g., autonomous driving, manipulation) exhibit structured dynamics and suboptimal experts. Could the proposed framework generalize beyond additive action spaces and static contexts?"},"rating":{"value":6},"details_of_ethics_concerns":{"value":"N/A"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"The authors establish a connection between trajectory perturbations, geometry of value landscapes, and movement dynamics, grounding shortcut augmentations in formal RL theory. The derived theorems provide interpretable conditions under which shortcuts are guaranteed to improve performance."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1) The theoretical results rely on Lipschitz continuity, f-contraction, and linear placement errors (LPE), assumptions that may not hold in more complex, discontinuous real-world systems (e.g., frictional or hysteretic actuators). The paper acknowledges this but does not propose methods for verifying or relaxing these assumptions.\n\n2) Experiments are conducted in semi-simulated optical systems. While these are realistic, validation on a physical robotic or optical alignment platform would greatly strengthen the paper’s impact and credibility.\n\n3) Shortcut sampling (Algorithm 1) requires O($n^2$) pairwise comparisons within trajectories. Although feasible for short logs, this may become expensive in longer sequences or higher-frequency data."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762927125207,"tcdate":1761940496742,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission17132/Reviewer_edFm"],"signatures":["ICLR.cc/2026/Conference/Submission17132/Reviewer_edFm"],"forum":"AsVH1FQGuR","number":3,"license":"CC BY 4.0","cdate":1761940496742,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission17132/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762927125207,"domain":"ICLR.cc/2026/Conference","replyto":"AsVH1FQGuR","id":"aF6r3Z5i85","forumContent":{"TLDR":{"value":"We introduce a trajectory-based data augmentation method that improves offline reinforcement learning in active positioning tasks by leveraging task structure and geometric properties of rewards, values, and logging policies."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Offline Reinforcement Learning","Reinforcement Learning","Active Position","Off-Policy Learning","Value Function Geometry"]},"supplementary_material":{"value":"/attachment/bc599ad161c10f4d3c9c0b2faba2a5e8b9f311e9.zip"},"primary_area":{"value":"reinforcement learning"},"abstract":{"value":"We propose a method for data augmentation in offline reinforcement learning applied to active positioning problems. \nThe approach enables the training of off-policy models from a limited number of trajectories generated by a suboptimal logging policy.  \nOur method is a trajectory-based augmentation technique that exploits task structure and quantify the effect of admissible perturbations on the data using the geometric\ninterplay of properties of the reward, the value function, and the logging policy.\nMoreover, we show that by training an off-policy model with our augmentation while collecting data, the suboptimal logging policy can be supported \nduring collection, leading to higher data quality and improved offline reinforcement learning performance.\nWe provide theoretical justification for these strategies and validate them empirically across positioning tasks of varying dimensionality and under partial observability."},"_bibtex":{"value":"@misc{\nschmahling2026augmentations,\ntitle={Augmentations in Offline Reinforcement Learning for Active Positioning},\nauthor={Tobias Schm{\\\"a}hling and Matthias Burkhardt and Tobias Windisch},\nyear={2026},\nurl={https://openreview.net/forum?id=AsVH1FQGuR}\n}"},"title":{"value":"Augmentations in Offline Reinforcement Learning for Active Positioning"},"pdf":{"value":"/pdf/457214e7d63837bd6841b8548c5854336912ba8c.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"schmähling|augmentations_in_offline_reinforcement_learning_for_active_positioning"},"authorids":{"value":["~Tobias_Schmähling1","~Matthias_Burkhardt1","~Tobias_Windisch1"]},"authors":{"value":["Tobias Schmähling","Matthias Burkhardt","Tobias Windisch"]}},"version":2},{"content":{"research_area_keywords":{"value":"In-Context Learning, Demonstration Retrieval, Example Selection, Prompting, Large Language Models"},"keywords":{"value":["In-Context Learning","Example Retrieval","Demonstration Selection","Large Language Models","Retrieval-Augmented Prompting"]},"B_use_or_create_scientific_artifacts":{"value":"Yes"},"languages_studied":{"value":"English"},"D4_ethics_review_board_approval":{"value":"N/A"},"C4_parameters_for_packages":{"value":"Yes"},"B6_elaboration":{"value":"Section 4 and Appendix."},"B1_cite_creators_of_artifacts":{"value":"Yes"},"B2_discuss_the_license_for_artifacts":{"value":"No"},"B4_elaboration":{"value":"The work uses standard publicly available NLP benchmarks and does not involve collection of personally identifying information."},"B5_elaboration":{"value":"Section 4 and Appendix."},"D_human_subjects_including_annotators":{"value":"No"},"A2_elaboration":{"value":"Section 7 (Limitations)."},"C3_elaboration":{"value":"We report single-run results due to computational constraints and leave multi-run statistical analysis for future work."},"B2_elaboration":{"value":"We use publicly available datasets and pretrained models commonly used in NLP research. Detailed license discussion is omitted due to space constraints."},"B3_elaboration":{"value":"Section 4."},"C2_elaboration":{"value":"Section 4 and Appendix B."},"C4_elaboration":{"value":"Section 4 and Appendix B."},"E_ai_assistants_in_research_or_writing":{"value":"Yes"},"C1_elaboration":{"value":"Appendix B."},"B1_elaboration":{"value":"Section 4 and References."},"C2_experimental_setup_and_hyperparameters":{"value":"Yes"},"E1_elaboration":{"value":"AI assistants were used for language polishing and writing refinement. All technical content, experiments, and conclusions were verified and revised by the authors."},"C1_model_size_and_budget":{"value":"Yes"},"venue":{"value":"ACL ARR 2026 May Submission"},"D3_data_consent":{"value":"N/A"},"_bibtex":{"value":"@inproceedings{\nanonymous2026beyond,\ntitle={Beyond Semantic Similarity: Task-Aware Anti-Shortcut Example Retrieval for In-Context Learning},\nauthor={Anonymous},\nbooktitle={Submitted to ACL Rolling Review - May 2026},\nyear={2026},\nurl={https://openreview.net/forum?id=u9zAR0T7jg},\nnote={under review}\n}"},"title":{"value":"Beyond Semantic Similarity: Task-Aware Anti-Shortcut Example Retrieval for In-Context Learning"},"C3_descriptive_statistics":{"value":"No"},"contribution_types":{"value":["Model analysis & interpretability","NLP engineering experiment","Approaches to low-compute settings (efficiency)"]},"A1_limitations_section":{"value":"This paper has a limitations section."},"B6_statistics_for_data":{"value":"Yes"},"B4_data_contains_personally_identifying_info_or_offensive_content":{"value":"No"},"B3_artifact_use_consistent_with_intended_use":{"value":"Yes"},"A2_potential_risks":{"value":"Yes"},"E1_information_about_use_of_ai_assistants":{"value":"Yes"},"abstract":{"value":"In-context learning (ICL) is highly sensitive to the choice of demonstrations. Most existing retrieval methods select demonstrations mainly by semantic similarity, but this strategy has two important weaknesses: it does not explicitly encode task information, and it often prefers examples with strong lexical overlap, which can encourage shortcut-based predictions. We propose \\textbf{T}ask-aware \\textbf{A}nti-shortcut \\textbf{E}xample \\textbf{R}etrieval (TAER), a simple two-stage retrieval framework for ICL. TAER first prepends a short task description to the test input and candidate examples, and retrieves task-relevant examples according to task-aware similarity. It then reranks these candidates using shallow-layer surface-feature similarity and keeps the least superficially similar examples, reducing the chance that the model relies on shallow correlations. Experiments on 15 natural language understanding and generation datasets, using LLMs from multiple families and scales, show that TAER consistently outperforms strong retrieval baselines. Additional ablations and robustness analyses further support the value of combining task-aware retrieval with anti-shortcut filtering."},"D1_instructions_given_to_participants":{"value":"N/A"},"paper_type":{"value":"Long"},"C_computational_experiments":{"value":"Yes"},"pdf":{"value":"/pdf/2aa2af12152443ee52371654df849c57f73e4076.pdf"},"research_area":{"value":"Machine Learning for NLP"},"B5_documentation_of_artifacts":{"value":"Yes"},"EMNLP_2026_AI_Reviewing_Experiment":{"value":"no"},"D2_recruitment_and_payment":{"value":"N/A"},"venueid":{"value":"aclweb.org/ACL/ARR/2026/May/Submission"}},"tmdate":1790697017833,"tcdate":1779784120510,"writers":["aclweb.org/ACL/ARR/2026/May","aclweb.org/ACL/ARR/2026/May/Submission13290/Authors"],"signatures":["aclweb.org/ACL/ARR/2026/May/Submission13290/Authors"],"forum":"u9zAR0T7jg","license":"CC BY 4.0","number":13290,"cdate":1779784120510,"readers":["everyone"],"invitations":["aclweb.org/ACL/ARR/2026/May/-/Submission","aclweb.org/ACL/ARR/2026/May/-/Edit","aclweb.org/ACL/ARR/2026/May/-/Post_Submission","aclweb.org/ACL/ARR/2026/May/Submission13290/-/Blind_Submission_License_Agreement","aclweb.org/ACL/ARR/2026/May/-/Preprint_Post_Submission"],"mdate":1790697017833,"odate":1780385598983,"domain":"aclweb.org/ACL/ARR/2026/May","id":"u9zAR0T7jg","version":2},{"content":{"summary":{"value":"This paper proposes SLAP (Shortcut Learning for Abstract Planning), a method that integrates model-free reinforcement learning (RL) with task and motion planning (TAMP) to automatically discover shortcut policies between abstract states. The main idea is to augment a symbolic planning graph with additional edges learned via RL. Each shortcut policy is trained to transition directly between two abstract states which effectively bypasses multiple intermediate actions and reduces plan length."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"In addition to the remarks made for the Weaknesses section, here are some additional questions:\n\n- In the pruning phase, a shortcut is only retained for RL training if random rollouts reach the terminal abstract state s_{\\text{term}} in at least K_{\\text{rollout}} / N_{\\text{rollout}} (which is \\approx 5% with the mentioned parameters) of attempts. Given that this requires non-negligible random success (especially for a 3-D environment) between abstract states, how are such reachable abstract-state pairs obtained? Is there any mechanism during the construction of the abstract planning graphs or the definition of abstract states that biases the graph toward physically close or feasible state pairs, making this 5 % success rate attainable? Because, even for the simplest 3-D tasks, having more than a 5% success rate by taking random actions is highly unlikely. \n\n- My understanding is that Figure 4 shows how, as training progresses, more shortcut policies become reliable and therefore less likely to be excluded from use during inference. In other words, the number of shortcuts that remain usable in the evaluation graph increases as RL training improves their success rates. Is this interpretation correct? The current wording (“increasing the number of training steps leads to more shortcuts being successfully learned and incorporated as graph edges”) is somewhat unclear because the number of learned shortcuts does not change with training.\n\n- “To create an initial state distribution, we do not assume that we can sample directly from sinit; instead, we sample from the states encountered in the abstract planning graphs for the training tasks.” I suppose the graph creation process is deterministic. So, doesn’t this limit the diversity of the initial state distribution? How robust is SLAP to out-of-distribution physical configurations at test time, and could additional randomization during graph construction improve generalization?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":1},"strengths":{"value":"- The intuition of the paper (discovering shortcut connections in an abstract planning graph) is conceptually sound and easy to grasp. It is a simple yet effective idea that directly addresses the inefficiency of long hierarchical plans in TAMP frameworks.\n\n- The approach can produce genuinely new high-level actions that are not part of the manually defined option set (e.g., slap). This demonstrates the potential of the method to extend the agent’s action set beyond what is explicitly encoded by human designers.\n\n- The paper is clearly written and well-structured, making the overall framework and experiments easy to follow.\n\n- Long horizon is one of the reasons why RL does not work effectively. Solving small RL problems with short horizons makes sense."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"A primary weakness of this work lies in its conceptual alignment with the TAMP paradigm and the choice of experimental baselines.\n\n1. Conceptual Dissonance with TAMP: The core philosophy of TAMP is to find plans that are valid with respect to a given symbolic model, ensuring that high-level action sequences are grounded and physically feasible according to predefined rules. The proposed method, SLAP, learns \"shortcut\" policies (e.g., \"slapping\" a tower) that achieve a goal state by operating outside of this symbolic structure. While this can yield more efficient plans, it fundamentally reframes the problem from one of constrained, symbolic planning to one of unconstrained trajectory optimization. This raises the question of whether SLAP is improving upon a TAMP solution or solving a different, less constrained problem altogether. The potential for emergent, \"destructive,\" or unpredictable behaviors is at odds with the safety and predictability that motivates the use of TAMP in the first place.\n\n2. Inadequate Baselines: The comparison to model-free RL baselines is not particularly insightful. It is well-established that vanilla RL algorithms struggle with the long-horizon, sparse-reward problems that TAMP is designed to solve. Since SLAP heavily leverages the strong structural priors from the TAMP framework (the abstract graph and initial options), outperforming these baselines is an expected outcome and does not sufficiently isolate the contribution of the shortcut-learning mechanism itself.\n\nIn summary, the paper positions itself as a hybrid TAMP-RL method, yet it diverges from the core principles of TAMP and compares against RL methods on a task setup where they are known to be inefficient. A clearer articulation of its contribution, perhaps as a method for discovering new symbolic operators rather than just bypassing them, and a comparison against more relevant baselines would significantly strengthen the paper.\n\nBelow are some further remarks:\n- The methodology depends on finding two abstract states where there is at least a success rate of K_{rollout}/N_{rollout} transitioning from one to another state by taking random actions. I find this very restrictive and dependent on the abstract planning graph generation.\n\n- The paper lacks a clear and fair baseline. It does not make sense to use pure RL algorithms as baselines for such complex, long-horizon tasks that fundamentally require high-level task planning. Even in simpler motion-planning problems, RL methods typically demand more training steps than the 500k used here to achieve meaningful performance. Since SLAP is built on a TAMP framework with strong structural priors, comparisons to model-free RL is not ideal. That’s why it cannot be claimed that the proposed method increases the performance. The paper already includes a comparison with the no-shortcut case. But this is more of an ablation study rather than a baseline. Finally, reporting success rates over only 10 trials per task is insufficient for statistical reliability; larger-scale evaluations or confidence intervals are needed to support the claimed improvements.\n\n- The core idea of using RL to learn shortcut policies between abstract states is not a significant conceptual contribution. The method essentially applies standard model-free RL within a predefined TAMP structure to learn transitions that skip multiple existing options. While this integration is practically useful, it does not introduce new algorithmic insights or theoretical advances in either reinforcement learning or task and motion planning. The contribution is therefore more of an engineering combination of known components than a novel methodological development.\n\n- The notation in Section 4.4 is not intuitive. add(â) and del(â) denote sets of atoms, whereas rel(â) denotes a set of objects involved in those atoms. The subsequent use of expressions such as add(â_train)[σ(rel(â_train))] ⊆ add(â_eval) is confusing (the square-bracket substitution notation is unconventional and not clearly defined). I can understand what is meant but a more rigorous and transparent mathematical formulation (or pseudocode) of the object substitution process would significantly improve clarity.\n\n- The paper claims that learned shortcut policies can be reused on new objects through an object-substitution mapping σ, which aligns the add/del atom sets after substitution. However, the method for finding σ is not clearly described. When multiple objects are involved, determining consistent multi-object mappings becomes nontrivial, yet the paper provides no explanation of how σ is computed, whether it must be one-to-one, or how ambiguities are resolved. This lack of detail makes it difficult to evaluate the reliability and scalability of the object substitution mechanism proposed in Section 4.4.\n\n- Figure 4 is hard to read. I would create two figures: one with graphs, one showing a learned shortcut (and improve this image).\n\n- The analysis for Q5 (“Which RL design decisions are important for learning shortcuts?”) is limited in scope. The section “Shortcut Policy Learning Analysis” examines only one factor (whether shortcut policies are trained independently or as a shared universal policy). Other key RL design aspects (e.g., algorithm choice, reward shaping, exploration strategy) are not explored. As a result, the section provides a partial answer to Q5 and does not fully justify the general phrasing of the question.\n\n- The claimed generalization ability of SLAP is limited. Section 4.2 describes “generalization over objects,” but the mechanism is a simple object-substitution procedure based on symbolic equivalence of add/del atoms. This allows policy reuse only when new tasks are structurally isomorphic to training ones. The method does not learn to generalize over varying object geometries, dynamics, or unseen relational structures (its transfer is purely symbolic). Consequently, the generalization claim overstates the scope of what the approach can handle.\n\n- The abstract claims that SLAP “consistently outperforms planning and RL baselines”. As Table 1 demonstrates, SLAP matches the Pure Planning baseline in success rate (100%) and only improves efficiency by reducing plan length. The improvement therefore lies in shorter trajectories, not higher task success."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762942072200,"tcdate":1761916804896,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission22118/Reviewer_aDVn"],"signatures":["ICLR.cc/2026/Conference/Submission22118/Reviewer_aDVn"],"forum":"enprG5H9aD","number":2,"license":"CC BY 4.0","cdate":1761916804896,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission22118/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762942072200,"domain":"ICLR.cc/2026/Conference","replyto":"enprG5H9aD","id":"8A7YGdls86","forumContent":{"TLDR":{"value":"We use RL to learn shortcuts in the abstract planning graph induced by predefined options."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Robot Planning","Reinforcement Learning"]},"supplementary_material":{"value":"/attachment/963c7ff7b5a8b143b0ebb6b42b4266d8ba9dbeab.zip"},"primary_area":{"value":"applications to robotics, autonomy, planning"},"abstract":{"value":"Long-horizon decision-making with sparse rewards and continuous states and actions remains a fundamental challenge in AI and robotics. Task and motion planning (TAMP) is a model-based framework that addresses this challenge by planning hierarchically with abstract actions (options). These options are manually defined, limiting the agent to behaviors that we as human engineers know how to program (pick, place, move). In this work, we propose Shortcut Learning for Abstract Planning (SLAP), a method that leverages existing TAMP options to automatically discover new ones. Our key idea is to use model-free reinforcement learning (RL) to learn *shortcuts* in the abstract planning graph induced by the existing options in TAMP. Without any additional assumptions or inputs, shortcut learning leads to shorter solutions than pure planning, and higher task success rates than flat and hierarchical RL. Qualitatively, SLAP discovers dynamic physical improvisations (e.g., slap, wiggle, wipe) that differ significantly from the manually-defined ones. In experiments in four simulated robotic environments, we show that SLAP solves and generalizes to a wide range of tasks, reducing overall plan lengths by over 50\\% and consistently outperforming planning and RL baselines."},"_bibtex":{"value":"@inproceedings{\nliu2026slap,\ntitle={{SLAP}: Shortcut Learning for Abstract Planning},\nauthor={Y. Isabel Liu and Bowen Li and Benjamin Eysenbach and Tom Silver},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=enprG5H9aD}\n}"},"title":{"value":"SLAP: Shortcut Learning for Abstract Planning"},"pdf":{"value":"/pdf/4f54aa8a88c1a1d21a3c9878134690e1432728d8.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"liu|slap_shortcut_learning_for_abstract_planning"},"authorids":{"value":["~Y._Isabel_Liu1","~Bowen_Li7","~Benjamin_Eysenbach1","~Tom_Silver1"]},"authors":{"value":["Y. Isabel Liu","Bowen Li","Benjamin Eysenbach","Tom Silver"]}},"version":2},{"content":{"summary":{"value":"This paper present a novel concept-based video editing framework. The framework includes a Concept-Augmented Textual Inversion (CATI) to extract the target concept from given video, and then use a Dual Prior Supervision (DPS) to manipulate the attention matrices to ensure the stability of output videos. For the CATI part, optimization is performed on both textual embeddings and LoRA modules added to the cross-attention layers. And for the DPS part, loss on both irrelevant parts and edited subject are used to prevent unwanted changes.\n\nResults are conducted over several with or without concept video pairs and was compared with several methods. The overall results are good."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. The pipeline for editing without concept video is not very clear for me, and in the code repository provided, a pretrained lora/concept model is required for all the cases including car to porsche. Can authors provide a more detailed explanation?\n2. It would be interesting to know the range of concepts. More examples of diverse concepts or attributes could clarify the extent to which CATI handles generalization across different object classes and scenarios."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The motivation for introducing the concept of video editing is clear and good. \n2. The results are impressive, the concepts injected are consistent, and the backgrounds are generally consistent with the target videos.\n3. The paper reads well and is easy to follow."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. How authors obtain the results from comparing methods is not explained in the paper. Some of those don’t accept concept video/images input, and if not given the concept video, then direct comparison is a bit unfair. \n2. The quantitative comparison table seems to be evaluated on both with or without concept videos. And the ablation for using or not using concept videos is missing. I suggest the author having separate dataset for using or not using concept videos, and this can further strengthen the contribution.\n3. Missing references and some for comparisons:\n    \n    *[1] AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks\n    [2] GenVideo: One-shot target-image and shape aware video editing using T2I diffusion models*\n    \n    *[3] Video Editing via Factorized Diffusion Distillation*\n    \n    *[4] VidToMe: Video Token Merging for Zero-Shot Video Editing*\n    \n    *[5] Pix2video: Video editing using image diffusion*\n    \n    *[6] Slicedit: Zero-Shot Video Editing With Text-to-Image Diffusion Models Using Spatio-Temporal Slices*\n    \n    *[7] TokenFlow: Consistent Diffusion Features for Consistent Video Editing*\n    \n    *[8] Video-P2P: Video Editing with Cross-attention Control*"}},"nonreaders":[],"tmdate":1731427306244,"tcdate":1730711676157,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission716/Reviewer_HLfd"],"signatures":["ICLR.cc/2025/Conference/Submission716/Reviewer_HLfd"],"forum":"CxS8mlkOH7","number":3,"license":"CC BY 4.0","cdate":1730711676157,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission716/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427306244,"domain":"ICLR.cc/2025/Conference","replyto":"CxS8mlkOH7","id":"pvGNEPP2YQ","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Text-Guided Video Editing"]},"primary_area":{"value":"generative models"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Text-driven video editing utilizing generative diffusion models has garnered significant attention due to their potential applications. However, existing approaches are constrained by the limited word embeddings provided in pre-training, which hinders nuanced editing targeting open concepts with specific attributes. Directly altering the keywords in target prompts often results in unintended disruptions to the attention mechanisms. To achieve more flexible editing easily, this work proposes an improved concept-augmented video editing approach that generates diverse and stable target videos flexibly by devising abstract conceptual pairs. Specifically, the framework involves concept-augmented textual inversion and a dual prior supervision mechanism. The former enables plug-and-play guidance of stable diffusion for video editing, effectively capturing target attributes for more stylized results. The dual prior supervision mechanism significantly enhances video stability and fidelity. Comprehensive evaluations demonstrate that our approach generates more stable and lifelike videos, outperforming state-of-the-art methods. The anonymous code is available at \\href{https://anonymous.4open.science/w/STIVE-PAGE-B4D4/}{https://anonymous.4open.science/w/STIVE-PAGE-B4D4/}."},"_bibtex":{"value":"@misc{\nguo2025shaping,\ntitle={Shaping a Stabilized Video by Mitigating Unintended Changes for Concept-Augmented Video Editing},\nauthor={Mingce Guo and Jingxuan He and Shengeng Tang and Zhangye Wang and Lechao Cheng},\nyear={2025},\nurl={https://openreview.net/forum?id=CxS8mlkOH7}\n}"},"title":{"value":"Shaping a Stabilized Video by Mitigating Unintended Changes for Concept-Augmented Video Editing"},"pdf":{"value":"/pdf/6ee805d6f2ff14368b91a3b977f139f02d6d57a4.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"guo|shaping_a_stabilized_video_by_mitigating_unintended_changes_for_conceptaugmented_video_editing"},"authorids":{"value":["~Mingce_Guo1","~Jingxuan_He2","~Shengeng_Tang1","~Zhangye_Wang1","~Lechao_Cheng2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Mingce Guo","Jingxuan He","Shengeng Tang","Zhangye Wang","Lechao Cheng"]}},"version":2},{"content":{"venue":{"value":"AAAI 2025"},"pdf":{"value":"https://ojs.aaai.org/index.php/AAAI/article/download/32276/34431"},"venueid":{"value":"dblp.org/conf/AAAI/2025"},"paperhash":{"value":"deng|boundaryaware_temporal_dynamic_pseudosupervision_pairs_generation_for_zeroshot_natural_language_video_localization"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Xiongwen_Deng:","~Haoyu_Tang1","~Han_Jiang10","~Qinghai_Zheng1","https://dblp.org/search/pid/api?q=author:Jihua_Zhu:"]},"html":{"value":"https://doi.org/10.1609/aaai.v39i3.32276"},"_bibtex":{"value":"@inproceedings{DBLP:conf/aaai/DengTJZZ25,\n  author={Xiongwen Deng and Haoyu Tang and Han Jiang and Qinghai Zheng and Jihua Zhu},\n  title={Boundary-Aware Temporal Dynamic Pseudo-Supervision Pairs Generation for Zero-Shot Natural Language Video Localization},\n  year={2025},\n  cdate={1735689600000},\n  pages={2717-2725},\n  url={https://doi.org/10.1609/aaai.v39i3.32276},\n  booktitle={AAAI},\n  crossref={conf/aaai/2025}\n}\n"},"abstract":{"value":"Zero-shot Natural Language Video Localization (NLVL) aims to automatically generate moments and corresponding pseudo queries from raw videos for the training of the localization model without any manual annotations. Existing approaches typically produce pseudo queries as simple words, which overlook the complexity of queries in real-world scenarios. Considering the powerful text modeling capabilities of large language models (LLMs), leveraging LLMs to generate complete queries that are closer to human descriptions is a potential solution. However, directly integrating LLMs into existing approaches introduces several issues, including insensitivity, isolation, and lack of regulation, which prevent the full exploitation of LLMs to enhance zero-shot NLVL performance. To address these issues, we propose BTDP, an innovative framework for Boundary-aware Temporal Dynamic Pseudo-supervision pairs generation. Our method contains two crucial operations: 1) Boundary Segmentation that identifies both visual boundaries and semantic boundaries to generate the atomic segments and activity descriptions, tackling the issue of insensitivity. 2) Context Aggregation that employs the LLMs with a self-evaluation process to aggregate and summarize global video information for optimized pseudo moment-query pairs, tackling the issue of isolation and lack of regulation. Comprehensive experimental results on the Charades-STA and ActivityNet Captions datasets demonstrate the effectiveness of our BTDP method."},"title":{"value":"Boundary-Aware Temporal Dynamic Pseudo-Supervision Pairs Generation for Zero-Shot Natural Language Video Localization"},"authors":{"value":["Xiongwen Deng","Haoyu Tang","Han Jiang","Qinghai Zheng","Jihua Zhu"]}},"tmdate":1784451417825,"pdate":1735689600000,"tcdate":1753024406827,"writers":["~"],"signatures":["~Haoyu_Tang1"],"forum":"3Vgk8pfgpz","license":"CC BY-SA 4.0","number":577414,"cdate":1735689600000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1784451417825,"domain":"DBLP.org","id":"3Vgk8pfgpz","version":2},{"content":{"summary":{"value":"The paper proposes Video-EM, a training-free framework for long-form video understanding inspired by human episodic memory. Instead of treating keyframes as isolated tokens to VLLMS, Video-EM groups them into temporally ordered key events, expands events to recover missing context, and builds rich episodic memory representations that capture when, where, what, and which objects. It then uses Chain-of-Thought reasoning to iteratively select a minimal but informative subset of episodic memories before passing them to a VLLM"},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"* How do you capture object-level semantics $q_0$ and scene-level context $q_s$?\n* Why do you need an adaptive event expansion mechanism based on the spatio-temporal difference metric? What does such complex method bring to the table? Couldn't simpler methods yield similar results?"},"rating":{"value":4},"details_of_ethics_concerns":{"value":"I do not identify any significant ethical issues in this paper. The method operates on publicly available video datasets commonly used in the community, and there is no indication of privacy violations, harmful content generation, or misuse potential beyond standard concerns in video understanding research. Therefore, I do not see any ethical concerns requiring further attention."},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"* Clear motivation.\n* Training-free approach that can equip state-of-the-art VLLMs with improved performance.\n* The paper appears to be aware of the related work.\n* The key event selection is sound and they expand each event to recover query-relevant context that similarity-based approaches may miss. This sounds novel and important as pure semantic retrieval can yield a sparse set of disjoint frames. \n* Video-EM reduces frames while improving accuracy.\n* The paper provides ablation studies."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"* The adaptive event expansion module feels somewhat heavy for a training-free method; simpler alternatives (e.g., adjacent-frame motion thresholds) could be discussed or compared.\n* Heavy reliance on object detectors and captioners where failure cases of these modules may propagate.\n* A notable limitation is that the method introduces several hyperparameters across multiple stages (e.g., similarity thresholds, expansion limits, CoT confidence and depth, temporal gap $\\Delta t$). While the authors provide ablations showing relative robustness, the number of hyperparameters is still large, and tuning them in practice may be non-trivial.\n* As a video agent Video-EM requires too many different models which can lead to efficiency problems and lack of end-to-end practicality. \n* The baselines differ from dataset to dataset. While this is ok it is a bit difficult to assess Video-EM's capabilities.\n* You should test Video-EM with LLMs other than the Qwen family (such as VideoLLaMA3, InterVL3...).\n* Results are not state-of-the-art, however, improvements over backbone models are achieved.\n\n\nMinor comments:\n* Authors use \\citet instead of \\citep.\n* Use Gemini 2.5 pro instead of 1.5 version."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762920279524,"tcdate":1760711682853,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission8370/Reviewer_EH6e"],"signatures":["ICLR.cc/2026/Conference/Submission8370/Reviewer_EH6e"],"forum":"aLQsPnVNnk","number":1,"license":"CC BY 4.0","cdate":1760711682853,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission8370/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762920279524,"domain":"ICLR.cc/2026/Conference","replyto":"aLQsPnVNnk","id":"v0fW1roTVK","forumContent":{"TLDR":{"value":"We introduce Video-EM, a training‑free framework that treats long video question answering as an episodic memory retrieval-and‑reasoning problem inspired by human cognitive psychology."},"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Multi-modal Vision","Video Understanding"]},"primary_area":{"value":"foundation or frontier models, including LLMs"},"abstract":{"value":"Video Large Language Models (Video-LLMs) excel at general video understanding but struggle with long-form videos due to context-window limits. Consequently, recent approaches focus on keyframe retrieval, condensing lengthy videos into a small set of informative frames. Despite their practicality, these methods simplify the problem to static text-image matching, overlooking spatio-temporal relationships crucial for capturing scene transitions and contextual continuity, and may yield redundant keyframes with limited information, diluting salient cues essential for accurate video question answering. To address these limitations, we introduce Video-EM, a training-free framework inspired by the principles of human episodic memory, designed to facilitate robust and contextually grounded reasoning. Rather than treating keyframes as isolated visual entities, Video-EM explicitly models them as temporally ordered episodic events, capturing both spatial relationships and temporal dynamics necessary for accurately reconstructing the underlying narrative. Furthermore, the framework leverages chain-of-thought (CoT) thinking with LLMs to iteratively identify a minimal yet highly informative subset of episodic memories, enabling efficient and accurate question answering by Video-LLMs. Extensive evaluations on multiple mainstream long-video benchmarks demonstrate the superiority of Video-EM, which achieves highly competitive results while using fewer frames."},"_bibtex":{"value":"@misc{\nwang2026episodic,\ntitle={Episodic Memory Representation for Long Video Understanding},\nauthor={Yun wang and Long Zhang and Jingren Liu and Jiaqi Yan and Zhanjie Zhang and Jiahao Zheng and Ao Ma and Xun Yang and Dapeng Wu and Xiangyu Chen and Xuelong Li},\nyear={2026},\nurl={https://openreview.net/forum?id=aLQsPnVNnk}\n}"},"title":{"value":"Episodic Memory Representation for Long Video Understanding"},"pdf":{"value":"/pdf/f020aad44e91a19ec0c3c30fd66c852f9da5374f.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"wang|episodic_memory_representation_for_long_video_understanding"},"authorids":{"value":["~Yun_wang13","~Long_Zhang4","~Jingren_Liu1","~Jiaqi_Yan3","~Zhanjie_Zhang2","~Jiahao_Zheng1","~Ao_Ma2","~Xun_Yang1","~Dapeng_Wu1","~Xiangyu_Chen5","~Xuelong_Li2"]},"authors":{"value":["Yun wang","Long Zhang","Jingren Liu","Jiaqi Yan","Zhanjie Zhang","Jiahao Zheng","Ao Ma","Xun Yang","Dapeng Wu","Xiangyu Chen","Xuelong Li"]}},"version":2},{"content":{"summary":{"value":"The paper addresses the problem of sampling from a target density when we only have access to the unnormalized density and no samples from it. The authors propose to adapt shortcut models to this setting by transporting samples between marginals of a linear interpolation of the densities in log-space to enable faster sampling, and by introducing an improved estimator for $\\partial_t \\log Z_t$ that uses the learned velocity to move samples between marginals and HMC to mix samples within these marginals. The method is evaluated on toy systems (DW-4, LJ13, GMM-40, and MW-32) and includes several ablations on model architectures and regularization for shortcut learning."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. In Figure 1, what exactly does “SMC” denote? Is this a sequential Monte Carlo pipeline that uses $\\nabla \\log p_t(x)$ inside HMC to move particles between intermediate marginals, rather than using the learned velocity field?\n\n\n2. Do the authors have an explanation for why “Velocity + SMC” underperforms at 10 steps in Figure 1? \n\n3. It would help to include an ablation comparing (a) Velocity + SMC + learned $\\partial_t \\log Z_t$ vs. (b) Velocity + SMC + Stein control variates. Right now the paper shows that the Stein-based estimator reduces variance during training, but it is not clear how much of that translates to improved inference-time performance compared to simply learning $\\partial_t \\log Z_t$. This would also clarify the difference from NETS, which I believe those authors estimate $\\partial_t \\log Z_t$ via the reweighted expectation of $\\nabla \\cdot v_t(x) + v_t(x) \\cdot \\nabla \\log p_t(x) + \\partial_t \\log \\hat{p}_t(x)$, rather than using a separate trained predictor [3].\n\n\n4. Could the authors describe in more detail the augmentation strategy that samples proportional to the residual error, and provide an ablation on how this affects training stability? It would be useful to have the exact procedure (either as equations or short pseudocode) in the appendix.\n\n\n5. How does the method perform without resampling during the transport? This would help isolate the contribution of the learned velocity $+$ HMC move from the resampling step.\n\n\n6. What are the ESS values for the proposed model (and for the baselines)?\n\n\n7. Have the authors tested the method on a larger Lennard–Jones system such as LJ-55? This would illustrate how well the approach scales beyond the toy/N-body setups shown.\n\n\n**Typos**\n\n1. Lines 218–219: there seems to be a missing “=” in the expression involving $\\partial_t \\log Z_t , \\mathbb{E}[\\cdot]$.\n\n[3] Albergo, Michael S., and Eric Vanden-Eijnden. \"Nets: A non-equilibrium transport sampler.\" arXiv preprint arXiv:2410.02711 (2024)."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":1},"contribution":{"value":2},"strengths":{"value":"1. Adapts the shortcut-modeling idea to a flow/SMC-based sampling setup, showing it can be used to reduce the number of transport steps.\n\n2. Achieves consistently strong results on the reported toy and small N-body benchmarks (DW-4, LJ-13, GMM-40, MW-32).\n\n3. Empirically shows that the proposed estimator yields lower-variance estimates of $\\partial_t \\log Z_t(x)$ during training, which is important for stabilizing PINN-style objectives."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The paper is hard to follow and poorly structured. It reads like a mix of several ideas and lacks a cohesive structure. The title focuses on shortcut models, but the first thing introduced is variance reduction for $\\partial_t \\log Z_t​$, rather than the shortcut method, which seems to be the main topic according to the title. Reordering the sections would help where the method should be introduced first and then talk about improving stability by having a lower variance estimator of $\\partial_t \\log Z_t(x)$ as an additional benefit. \n\n2. While the paper shows strong empirical performance on the toy targets, the theoretical contribution is comparatively modest, since it mainly adapts shortcut modeling to this particular SMC/flow-based sampling setup rather than introducing something fundamentally new. That said, showing that shortcut models can be made to work in this setting is still interesting\n\n3. he experiments are restricted to synthetic or small N-body systems (DW-4, LJ-13, GMM-40, MW-32), so it is hard to assess whether the method remains effective on practical sampling problems such as Boltzmann sampling of peptide conformations, either in dihedral space as in [1] or in Cartesian space as in [2]. A demonstration on at least one peptide-like system (e.g., alanine dipeptide with torsions, or a small capped peptide in Cartesian coordinates) would make the empirical claims more convincing.\n\n4. Below are comments about the presentation of results and figures:\n\n    a. Figure 2: add spacing between the caption and the following text. \n\n    b. Figures 5 and 6: currently poorly visualized, I would suggest removing the bars with the means; Table 1 should highlight better performance (bolding or coloring).\n\n    c. Table 1: metrics are inconsistent (e.g., E-TV missing for GMM-40, E-W2 missing for MW-32). It would be better to make the metrics uniform across tasks and to bold best results.\n\n    d. The architecture ablations are better placed in the appendix, since the method is architecture-agnostic and mainly benefits from using the best-performing architecture (which appears to be a transformer). The main text should focus more on shortcut ablations and on the importance of accurate $\\partial_t \\log Z_t$ estimates.\n\n    e. The results are missing ESS metrics.\n\n\n[1] Midgley, Laurence Illing, et al. \"Flow annealed importance sampling bootstrap.\" arXiv preprint arXiv:2208.01893 (2022).\n\n[2] Akhound-Sadegh, Tara, et al. \"Progressive Inference-Time Annealing of Diffusion Models for Sampling from Boltzmann Densities.\" arXiv preprint arXiv:2506.16471 (2025)."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762923220278,"tcdate":1761862674480,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission12288/Reviewer_3Rm5"],"signatures":["ICLR.cc/2026/Conference/Submission12288/Reviewer_3Rm5"],"forum":"hHfUwjl3hF","number":3,"license":"CC BY 4.0","cdate":1761862674480,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission12288/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762923220278,"domain":"ICLR.cc/2026/Conference","replyto":"hHfUwjl3hF","id":"WHAHuwu2FH","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["generative models"]},"supplementary_material":{"value":"/attachment/d3f073fd3c20a63d9ef676d79183c3c325fb563a.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Sampling from unnormalized densities presents a fundamental challenge with wide-ranging applications, from posterior inference to molecular dynamics simulations. Continuous flow-based neural samplers offer a promising approach, learning a velocity field that satisfies key principles of marginal density evolution (e.g., the continuity equation) to generate samples. However, this learning procedure requires accurate estimation of intractable terms linked to the computationally challenging partition function, for which existing estimators often suffer from high variance or low accuracy. To overcome this, we introduce an improved estimator for these challenging quantities, employing a velocity-driven Sequential Monte Carlo method enhanced with control variates. Furthermore, we introduce a shortcut consistency model to boost the runtime efficiency of the flow-based neural sampler by minimizing its required sampling steps. Our proposed Neural Flow Shortcut Sampler empirically outperforms existing flow-based neural samplers on both synthetic datasets and complex n-body system targets"},"_bibtex":{"value":"@misc{\nchen2025neural,\ntitle={Neural Flow Samplers with Shortcut Models},\nauthor={Wuhao Chen and Zijing Ou and Yingzhen Li},\nyear={2025},\nurl={https://openreview.net/forum?id=hHfUwjl3hF}\n}"},"title":{"value":"Neural Flow Samplers with Shortcut Models"},"pdf":{"value":"/pdf/6bb36d45cc2323b345570fb49768c9ab441cf3df.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"chen|neural_flow_samplers_with_shortcut_models"},"authorids":{"value":["~Wuhao_Chen1","~Zijing_Ou1","~Yingzhen_Li1"]},"authors":{"value":["Wuhao Chen","Zijing Ou","Yingzhen Li"]}},"version":2},{"content":{"summary":{"value":"This paper presents a systematic study on the trade-off between predictivity and availability in deep learning, proposing a theoretical foundation for shortcut learning in neural networks. The key contributions are:\n1.The paper introduces quantitative notions of predictivity and availability using a generative model to synthesize classification datasets where two latent features vary in terms of these attributes. A measure of \"shortcut bias\" is proposed to quantify a model's over-reliance on more available but less predictive features.\n2.Through controlled experiments, it is shown that nonlinear models exhibit greater shortcut bias compared to linear models, and model depth amplifies this effect.\n3.A theoretical analysis based on Neural Tangent Kernels proves the inevitability of availability bias for nonlinear architectures like ReLU networks, while linear networks are unbiased.\n4.Experiments on natural image datasets demonstrate that widely used vision models are sensitive to availability, not just predictivity, of non-core features. Explicit availability manipulations of images are shown to alter models' reliance on different features.\nTaken together, the empirical and theoretical findings reveal shortcut learning as an inherent characteristic of nonlinear deep models that needs to be studied systematically. The framework presented lays the groundwork for further investigating architectural choices, dataset factors and methods to mitigate shortcut bias.\n5.Overall, this is a thorough, well-executed study that makes important theoretical and practical contributions towards understanding shortcut learning in deep neural networks. The paper is clearly written and the empirical methodology is sound. The theorems and proofs rigorously analyze the interplay between predictive value and feature availability. This work provides key insights into model failures related to over-reliance on spuriously predictive features, and tools to improve model robustness."},"presentation":{"value":"2 fair"},"contribution":{"value":"3 good"},"soundness":{"value":"3 good"},"strengths":{"value":"1.The proposed generative framework for synthesizing datasets with controllable levels of predictivity and availability is creative and enables controlled experimentation.\n2.The empirical methodology is thorough and sound. The datasets, models, and evaluation metrics are carefully designed. Results are reported over multiple random seeds to ensure significance.\n3.Shortcut learning is a pivotal concept in deep learning with implications for generalization and fairness. This work significantly advances our understanding of why models fail in this manner.\nIn summary, this is a paper of exceptional quality and scientific merit that offers significant theoretical and practical value to the field. The novel concepts, thorough empirics, and rigorous theory set a new standard and open up many promising research directions."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"This is an good paper overall, but a few minor weaknesses could be addressed to further improve it:\n\n- The theoretical analysis makes approximations to obtain tractability (small covariance, quadratic approximation of ReLU kernel). It would strengthen the analysis to discuss the impact of these approximations. Are the core insights still valid without them?\n\n- The proposed notion of availability is intuitively reasonable but remains a hypothesis. Additional ablation studies that directly validate the choice of manipulations affecting availability could make this more concrete.\n\n- The measures of reliance and shortcut bias, though well-motivated, are indirectly quantified through alignment with idealized classifiers. More analysis could be provided to justify these metrics over alternatives.\n\n- There is scope for further investigating architectural manipulations that could mitigate shortcut bias, beyond just model depth and activation function. This could lead to practical guidelines.\n\nOverall these are minor limitations that do not diminish the quality of the work. The paper thoroughly delivers on its core promises. Addressing the above points where possible would make it even stronger. The work clearly advances our understanding of an important problem and provides a solid foundation for reducing shortcut reliance in deep learning models."},"confidence":{"value":"4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work."},"questions":{"value":"Here are some questions and suggestions to further improve the paper:\n\n- The quadratic approximation for the ReLU kernel greatly simplified the theoretical analysis. Could you provide some empirical verification that this does not alter the core findings? For example, compare the availability bias for the true ReLU kernel versus the quadratic approximation.\n\n- Have you experimented with any architectural manipulations beyond depth and activation functions that could potentially reduce shortcut bias? Things like skip connections, normalization layers etc. Exploring this could provide practical guidelines. \n\n- For the image experiments, it would be interesting to also show the impact of availability manipulations on a model pretrained on Imagenet. Does pretraining affect sensitivity to shortcuts?\n\n- The measures of reliance and bias involve probing a model in latent space. For vision experiments, could you provide some visualizations of model decisions before/after availability manipulations to build intuition?\n\n- The notion of availability is central but still somewhat conceptual. Are there any additional experiments you could do to further validate the manipulations proposed to affect availability?\n\nI hope these suggestions help further refine the work. The key results seem solid and demonstrate the importance of availability bias. Some additional experiments and discussion along the lines above could make it even more convincing and applicable across domains. I look forward to the authors' response."},"rating":{"value":"8: accept, good paper"},"details_of_ethics_concerns":{"value":"none."},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636702538,"tcdate":1698672421079,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission6363/Reviewer_G1nt"],"signatures":["ICLR.cc/2024/Conference/Submission6363/Reviewer_G1nt"],"forum":"Tj3xLVuE9f","number":2,"license":"CC BY 4.0","cdate":1698672421079,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission6363/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636702538,"domain":"ICLR.cc/2024/Conference","replyto":"Tj3xLVuE9f","id":"J711UT4PUu","forumContent":{"venue":{"value":"ICLR 2024 spotlight"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["shortcut learning","spurious correlations","architectural inductive bias"]},"primary_area":{"value":"visualization or interpretation of learned representations"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Deep-learning models can extract a rich assortment of features from data. Which features a model uses depends not only on *predictivity*---how reliably a feature indicates training-set labels---but also on *availability*---how easily the feature can be extracted from inputs. The literature on shortcut learning has noted examples in which models privilege one feature over another, for example texture over shape and image backgrounds over foreground objects. Here, we test hypotheses about which input properties are more available to a model, and systematically study how predictivity and availability interact to shape models' feature use. We construct a minimal, explicit generative framework for synthesizing classification datasets with two latent features that vary in predictivity and in factors we hypothesize to relate to availability, and we quantify a model's shortcut bias---its over-reliance on the shortcut (more available, less predictive) feature at the expense of the core (less available, more predictive) feature. We find that linear models are relatively unbiased, but introducing a single hidden layer with ReLU or Tanh units yields a bias. Our empirical findings are consistent with a theoretical account based on Neural Tangent Kernels. Finally, we study how models used in practice trade off predictivity and availability in naturalistic datasets, discovering availability manipulations which increase models' degree of shortcut bias. Taken together, these findings suggest that the propensity to learn shortcut features is a fundamental characteristic of deep nonlinear architectures warranting systematic study given its role in shaping how models solve tasks."},"_bibtex":{"value":"@inproceedings{\nhermann2024on,\ntitle={On the Foundations of Shortcut Learning},\nauthor={Katherine Hermann and Hossein Mobahi and Thomas FEL and Michael Curtis Mozer},\nbooktitle={The Twelfth International Conference on Learning Representations},\nyear={2024},\nurl={https://openreview.net/forum?id=Tj3xLVuE9f}\n}"},"title":{"value":"On the Foundations of Shortcut Learning"},"pdf":{"value":"/pdf/3f47b29f0e35691e7047d9fbfa0e4c47ea966e49.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference"},"paperhash":{"value":"hermann|on_the_foundations_of_shortcut_learning"},"authorids":{"value":["~Katherine_Hermann1","~Hossein_Mobahi2","~Thomas_FEL1","~Michael_Curtis_Mozer1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Katherine Hermann","Hossein Mobahi","Thomas FEL","Michael Curtis Mozer"]}},"version":2},{"content":{"data_release":{"value":"We authorize the release of our submission and author names to the public in the event of acceptance."},"venue":{"value":"ICML 2026 GlobalSouthML"},"email_sharing":{"value":"We authorize the sharing of all author emails with Program Chairs."},"pdf":{"value":"/pdf/0201fe45201bfc8997af302c7b7652901eded312.pdf"},"keywords":{"value":["3d Vision","Robotics","Robotic perception"]},"venueid":{"value":"ICML.cc/2026/Workshop/GlobalSouthML"},"paperhash":{"value":"bhosale|moreg_minimaloverlap_3d_point_cloud_registration_via_visibilityaware_matching"},"authorids":{"value":["~Rushikesh_Bhosale1","~Sundaram_Muthu2"]},"abstract":{"value":"3D point cloud registration is widely used in ap-plications such as autonomous navigation, 3D re-construction, mapping, and scene understanding ,where partially overlapping observations must be aligned. Registration under minimal over-lap (10–30%) remains challenging, as limited shared regions often lead to unreliable correspondences. Most existing methods rely on sufficient overlap, which affects performance in such set-tings. We introduce MOREG, a framework for minimal-overlap registration that uses visibility-aware matching and overlap-guided feature inter-action to focus matching on shared regions. We evaluate MOREG on ModelNet40 under varying overlap conditions and validate it on a real-world dataset collected using a UR7e robot across different overlap pairs. Our method consistently improves performance over PREDATOR in minimal-overlap settings"},"_bibtex":{"value":"@inproceedings{\nbhosale2026moreg,\ntitle={MoReg: Minimal-Overlap 3D Point Cloud Registration via Visibility-Aware Matching},\nauthor={Rushikesh Bhosale and Sundaram Muthu},\nbooktitle={ICML 2026 GlobalSouthML},\nyear={2026},\nurl={https://openreview.net/forum?id=8Nj6VnPmHk}\n}"},"title":{"value":"MoReg: Minimal-Overlap 3D Point Cloud Registration via Visibility-Aware Matching"},"authors":{"value":["Rushikesh Bhosale","Sundaram Muthu"]}},"tmdate":1786857549331,"pdate":1780623000721,"tcdate":1778067643308,"writers":["ICML.cc/2026/Workshop/GlobalSouthML","ICML.cc/2026/Workshop/GlobalSouthML/Submission139/Authors"],"signatures":["ICML.cc/2026/Workshop/GlobalSouthML/Submission139/Authors"],"forum":"8Nj6VnPmHk","license":"CC BY 4.0","number":139,"cdate":1778067643308,"readers":["everyone"],"invitations":["ICML.cc/2026/Workshop/GlobalSouthML/-/Submission","ICML.cc/2026/Workshop/GlobalSouthML/-/Submission_Change_Before_Bidding","ICML.cc/2026/Workshop/GlobalSouthML/-/Submission_Change_Before_Reviewing","ICML.cc/2026/Workshop/GlobalSouthML/-/Submission_Release","ICML.cc/2026/Workshop/GlobalSouthML/Submission139/-/Camera_Ready_Revision","ICML.cc/2026/Workshop/GlobalSouthML/-/Edit"],"mdate":1786857549331,"odate":1780623000721,"domain":"ICML.cc/2026/Workshop/GlobalSouthML","id":"8Nj6VnPmHk","version":2},{"content":{"summary":{"value":"In this paper, a new task called video action differencing is proposed which aims for models to be able to understand fine-grained differences between multiple videos of people performing the same action. A new benchmark dataset is collected, named VidDiffBench, which includes 5 categories from 4 different existing datasets. Annotations are collected from pairs of videos with statements given per pair of video based on the action (for example video A includes someone jumping higher than Video B for a lay-up shot). There are two main evaluation protocols for this task, a closed set setting, in which the model must predict A or B for each possible description, and a closed set setting in which the method must generate the description. A new method which combines stages named VidDiff is proposed which outperforms standard LMMs on the dataset."},"soundness":{"value":2},"confidence":{"value":5},"questions":{"value":"1. Does the sampling of pairs of videos mean that these might not contain the same action (see above)? Or the actions are sampled first to ensure a wide range of comparison difficulty over actions before video pairs are sampled within action?\n2. Were the annotators skilled/experts/knowledgeable in the actions that they were annotating? Or was this found to not be that important given the annotation task? Additionally, how many total annotators were used and were they renumerated for their time?\n3. Has the potential bias of the video pairs in the closed task been checked to ensure that naive performance should be 50% instead of video A (or B) occurring as the answer more than 50% of the time. Additionally, I would be interested to know if candidates which could be categorised as C (for insignificant differences) can be understood by the model as this would also increase the naive difficulty to 33% before taking into account a non-uniform distribution.\n4. The evaluation protocol for the open-set task seems like it could include errors/inconsistencies depending on the LLM output. Has there been any investigation into this and how much it differs per run and how much it aligns with a human? Currently, the prompts are also given in the appendix with little to no discussion as to why these prompts were chosen, if they were developed over multiple iterations to find the best performing prompt, etc.\n5. Did the easy/medium/hard classifications align with experts' opinions for each of the actions? It would be good to know the types of actions that are classed as easy/medium/hard as these are not present within the paper as far as I could tell. It's not clear why an LLM was chosen to do this task.\n6. Could more qualitative results and statistics be provided about the dataset? For example, there is very little in the paper regarding the retrieval task of localising the differences: How much of the video does the method need to localise? Are there any temporal biases regarding the timestamps from the videos? Additionally, under the closed set task, more statistics over the number of As, Bs, and Cs that have been annotated and included for each action would be interesting to see. Other statistics that feel like they are missing are the average length of each video (potentially broken down per category) as well as the total number of hours within the dataset.\n7. As a thought, has an experiment where the same video is duplicated and provided into the methods, would the output predictions (esp. for the closed set task) give a 50% response rate? Ideally, this is where a method could predict the difference is negligible also."},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":1},"contribution":{"value":4},"strengths":{"value":"* The new task of Video Action Differencing is an interesting new task for video understanding, forcing models to recognise and understand fine-grained differences between two very similar videos.\n* The collected dataset combines four datasets with 5 different categories of video, providing a varied test bed for this new task.\n* The proposed method performs well on the dataset, outperforming off the shelf LMMs on the task yet still showcase that there is a lot still to work on in this area for future work."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"# Weaknesses\n\n* There are some missing references for skill determination within the related work [a, b, c, d] as another example of fine-grained differences between videos containing the same action.\n* Line 196: It is mentioned here in the text that *\"Video pairs are randomly sampled within each dataset to ensure a wide range of comparison difficulty, from simple actions to more advanced tasks requiring fine-grained understanding\"* This implies that videos of differing actions are compared against one another. \n* Section 3.3.2: There are some missing information about the annotators, regarding skill level, total number, renumeration etc.\n* For the closed set, a binary classification setup was used as all candidate difference statements which is mentioned to be unbiased on Line 298. However, has this been checked? If videos are not randomly swapped at inference/training time there could have been a bias towards one video or another.\n* The open set evaluation seems like it could be prone to some errors/inconsistencies depending on the LLM chosen and how much it could hallucinate/not understand the task and doesn't represent a potentially sound evaluation protocol.\n* It is not clear within the paper as to why an LLM was used to choose the easy/medium/hard splits for each of the actions.\n* This paper did not feel like an easy read, whilst the grammar/sentence clarity was good. There was a lot of information that is split across the main paper and the appendix which necessitates jumping between them. The structure of the paper could also be improved, the method details occur within the experiments yet are given as a main contribution within the introduction with only a small amount of space given to explain the model. Another major factor for this is that details of the dataset are given before the task is formally defined, which given this is a new task, makes it harder to read than it should be.\n\n# Additional Comments\nLine 158 is referring to the wrong table, this should be Table 1\nLine 1040 (in supp.) vary -> very\nSection D.1 in the appendix is empty, maybe D.2 is meant to be a subheading of D.1?\nFor results tables, it would be good to include a random performance row.\n\n# References\n[a] Doughty, Hazel, Dima Damen, and Walterio Mayol-Cuevas. \"Who's better? who's best? pairwise deep ranking for skill determination.\" Proceedings of the IEEE conference on computer vision and pattern recognition. 2018.\n\n[b] Doughty, Hazel, Walterio Mayol-Cuevas, and Dima Damen. \"The pros and cons: Rank-aware temporal attention for skill determination in long videos.\" Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2019.\n\n[c] Pan, Jia-Hui, Jibin Gao, and Wei-Shi Zheng. \"Adaptive action assessment.\" IEEE Transactions on Pattern Analysis and Machine Intelligence 44.12 (2021): 8779-8795.\n\n[d] Zhang, Shao-Jie, et al. \"Adaptive stage-aware assessment skill transfer for skill determination.\" IEEE Transactions on Multimedia (2023)"}},"nonreaders":[],"tmdate":1732519954308,"tcdate":1730450431936,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission14085/Reviewer_dA39"],"signatures":["ICLR.cc/2025/Conference/Submission14085/Reviewer_dA39"],"forum":"3bcN6xlO6f","number":2,"license":"CC BY 4.0","cdate":1730450431936,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission14085/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732519954308,"domain":"ICLR.cc/2025/Conference","replyto":"3bcN6xlO6f","id":"HqXyDMc7C2","forumContent":{"TLDR":{"value":"A new task and benchmark for comparing how an action is performed between two videos, with a zero-shot method"},"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video","Actions","Differencing","Zero-shot","benchmark","multimodal","lmm","llm"]},"primary_area":{"value":"datasets and benchmarks"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"How do two individuals differ when performing the same action? In this work, we introduce Video Action Differencing (VidDiff), the novel task of identifying subtle differences between videos of the same action, which has numerous applications, such as coaching and skill learning. To enable development on this new task, we first create VidDiffBench, a benchmark dataset containing 549 video pairs, with human annotations of 4,469 fine-grained action differences and 2,075 timestamps indicating where these differences occur. Our experiments demonstrate that VidDiffBench poses a significant challenge for state-of-the-art large multimodal models (LMMs), such as GPT-4o and Qwen2-VL. By analyzing the failure cases of LMMs on VidDiffBench, we highlight two key challenges for this task: localizing relevant sub-actions over two videos and fine-grained frame comparison. To overcome these, we propose the VidDiff method, an agentic workflow that breaks the task into three stages: action difference proposal, keyframe localization, and frame differencing, each stage utilizing specialized foundation models. To encourage future research in this new task, we release the benchmark and code."},"_bibtex":{"value":"@inproceedings{\nburgess2025video,\ntitle={Video Action Differencing},\nauthor={James Burgess and Xiaohan Wang and Yuhui Zhang and Anita Rau and Alejandro Lozano and Lisa Dunlap and Trevor Darrell and Serena Yeung-Levy},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=3bcN6xlO6f}\n}"},"title":{"value":"Video Action Differencing"},"pdf":{"value":"/pdf/102482b5babaacddfd916de17bda7c15b2020db5.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"burgess|video_action_differencing"},"authorids":{"value":["~James_Burgess2","~Xiaohan_Wang2","~Yuhui_Zhang3","~Anita_Rau1","~Alejandro_Lozano1","~Lisa_Dunlap1","~Trevor_Darrell2","~Serena_Yeung-Levy1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["James Burgess","Xiaohan Wang","Yuhui Zhang","Anita Rau","Alejandro Lozano","Lisa Dunlap","Trevor Darrell","Serena Yeung-Levy"]}},"version":2},{"content":{"summary":{"value":"The paper proposes denoising generative model called shortcut model. These models are trained end-to-end and can generate samples in both single-step and few-step setting. The core idea is to condition the model on step size in addition to time step, which is typically done in conventional diffusion models. The hypothesis is that the additional conditioning on step size allows shortcut models to account for the future curvature. The model is trained via a loss objective that combines the optimal-transport flow matching loss with a self-consistency loss. The self-consistency loss is enforced by ensuring that the prediction of model at a given skip size matches the target obtained by two sequential smaller skips (for the same skip distance) with the shortcut model. The high computational cost as well as variance during training is dealt with by computing self-consistency loss on a small subset of samples in the batch. The proposed method has been tested on CelebAHQ-256 and ImageNet-256, and has good single-step image generation quality."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. For the results of Progressive Distillation in Table 1, a reasonable 128-step model would be the obtained after taking the initial (say, T=1024) step model and then distilling it progressively until we get 128-step model. Was this model used to report metrics in the table or, did the authors use the final 1-step model to generate samples with 128 steps? \n2. Appendix B.1 - could the authors specify the amount of TPUv3 nodes needed to run these experiments in 1-2 days? \n3. For consistency training, is there a reason why the authors chose to deviate from the proposed discretization schedule by Song et al. (2023) during training? From my experience with these models, the choice of discretization schedule matters a lot, and deviations affect the performance. Further, does this paper implement consistency training (which includes perceptual loss)[1] or the follow up paper [2] which uses pseudo-Huber loss and proposes several improvements over their previous paper.\n4. Could the authors include architectural details/positional encoding details of both of the conditioning (time/noise level and step size)?\n5. Could the authors include some qualitative results of one-step generation of some other methods they have trained on in the appendix? FID score doesn’t always correlate with human perception of images. \n6. Minor: Related work: for consistency models, the line needs revision - “shortcut models are practically simpler, as many consistency model tricks - e.g. using a strict discretization schedule, using perceptual loss rather than L1 loss - can be bypassed entirely. ” In their latest paper, Improved Consistency Training [2], Song et al. already bypass the need of perceptual loss; they use only pseudo-Huber loss.\n7. Minor: some typos: competetive -> competitive (Section 5, line 2), qualatative -> qualitative (Figure 4 caption)\n\n[1] Song, Yang, et al. \"Consistency models.\" arXiv preprint arXiv:2303.01469 (2023).\n[2] Song, Yang, and Prafulla Dhariwal. \"Improved techniques for training consistency models.\" arXiv preprint arXiv:2310.14189 (2023)."},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. _Simplicity_: The proposed method is simple, easy to understand and simplifies some challenges previously encountered in other methods to obtain one-step or few-step models. For eg. It gets rid of two-stage training required for diffusion distillation as well as complex training scheduling required for consistency models. The loss objective is intuitive and builds upon flow matching loss and introduces an additional self-consistency condition. \n2. _Writing_: The paper is well-written. The core method has been explained well. The paper talks about challenges that can be encountered while training shortcut models and ways to address these challenges. \n3. _Computational efficiency_: The method seems computationally efficient as with an additional (marginal) cost of training diffusion models, shortcut models can be trained to generate samples in single-step as well as few steps. Further, even on larger datasets such as ImageNet-256 and CelebAHQ-256, these models can be trained in 1-2 days (though the exact number of TPUv3 nodes is unclear).  \n4. _Applications_: The authors demonstrate advantages of this method on both image-generation as well as on robotic control tasks showing ways to generate an iterative sequence of actions in a single denoising step per action."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Note: The template of this paper seems to be different and doesn’t have line numbers present in the standard template for papers under review. \n\n1. The proposed method for shortcut sampling is general and should work for any gaussian probability paths ($x_t = \\alpha_t x_0 + \\sigma_t \\epsilon$) however this hasn’t been explored or described in the paper. The paper considers only one choice of optimal transport probability path ($x_t = (1-t) x_0 + t \\epsilon$). However, nothing constraints the method to be restricted to this particular choice. (One motivation for this choice of parametrization could be that OT paths are said to have less curvature but in this case, these are conditional OT paths and are not exactly straight lines paths between $x_1$ and $x_0$.) It is worth exploring ablations on how shortcut models perform on other parametrization of gaussian probability paths/sampling paths. In addition, the text should also be updated to have a more general parametrization.\n2. The paper should also include quantitative performance on more commonly used datasets such as ImageNet-64 and CIFAR-10. While reporting results on larger datasets such as CelebAHQ-256 and ImageNet-256 is impressive, reporting metrics on these smaller datasets is also important. Most of the prior papers report metrics on these datasets which makes it easier to compare the performance of shortcut models to a larger number of previously proposed methods/models which are not currently reported in the paper (Note: I understand that it is impossible to replicate all the previously proposed methods due to the vastness of literature in the space).\n3. Minor: I personally feel that images such as in Figure 1 are misleading as the original flow-matching objective was not specifically meant for single step generation, and the performance is expected to deteriorate because these models don’t learn an ODE integrator. There are some advantages of the original flow-matching models as well over shortcut models. These models are continuous time models which can be evaluated at any time between [0, 1] whereas shortcut models are discrete time and can only be evaluated at specific discrete time steps."}},"nonreaders":[],"tmdate":1732483968688,"tcdate":1730613571600,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission5115/Reviewer_PitZ"],"signatures":["ICLR.cc/2025/Conference/Submission5115/Reviewer_PitZ"],"forum":"OlzB6LnXcS","number":2,"license":"CC BY 4.0","cdate":1730613571600,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission5115/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732483968688,"domain":"ICLR.cc/2025/Conference","replyto":"OlzB6LnXcS","id":"OTSwXtQ5wo","forumContent":{"venue":{"value":"ICLR 2025 Oral"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["diffusion","flow-matching","fast inference","distillation"]},"primary_area":{"value":"generative models"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Diffusion models and flow matching models have enabled generating diverse and realistic images by learning to transfer noise to data. However, sampling from these models involves iterative denoising over many neural network passes, making generation slow and expensive. Previous approaches for speeding up sampling require complex training regimes, such as multiple training phases, multiple networks, or fragile scheduling. We introduce Shortcut Models, a family of generative models that use a single network and training phase to produce high-quality samples in a single or multiple sampling steps. Shortcut models condition the network not only on the current noise level but also on the desired step size, allowing the model to skip ahead in the generation process. Across a wide range of sampling step budgets, shortcut models consistently produce higher quality samples than previous approaches, such as consistency models and reflow. Compared to distillation, shortcut models reduce complexity to a single network and training phase and additionally allow varying step budgets at inference time."},"_bibtex":{"value":"@inproceedings{\nfrans2025one,\ntitle={One Step Diffusion via Shortcut Models},\nauthor={Kevin Frans and Danijar Hafner and Sergey Levine and Pieter Abbeel},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=OlzB6LnXcS}\n}"},"title":{"value":"One Step Diffusion via Shortcut Models"},"pdf":{"value":"/pdf/834505d749dfe267a983986c87c26443a58835cb.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"frans|one_step_diffusion_via_shortcut_models"},"authorids":{"value":["~Kevin_Frans1","~Danijar_Hafner1","~Sergey_Levine1","~Pieter_Abbeel2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Kevin Frans","Danijar Hafner","Sergey Levine","Pieter Abbeel"]}},"version":2},{"content":{"summary":{"value":"This paper presents I4VGEN, a novel video diffusion inference pipeline that enhances pre-trained text-to-video models without additional training. It tackles the complexities of spatio-temporal modeling by utilizing advanced image techniques in two stages: first, synthesizing anchor images with a strategy to ensure visual realism and semantic accuracy; second, augmenting these images to generate videos through a noise-invariant video score distillation sampling (NI-VSDS) method. This process also includes a video regeneration step for refinement. Experiments show that I4VGEN significantly improves the visual quality and textual accuracy of generated videos and can be easily integrated into existing models, boosting overall video quality."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"1. Concerns about inference time. The proposed method consists of two stages: anchor image synthesis and anchor image-augmented video synthesis. The reviewer wants to know whether the time in the table includes the time used in the first stage.\n2. Low-quality face video in Fig.6. In 1st row of Fig6, SparseCtrl produces a video in extremely low quality, could the authors explain reasons?"},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. This paper introduces a training-free pipeline called I4VGen to improve the performance of text-to-video diffusion models throught image reference information.\n2. A simple yet effective generation-selection strategy is proposed to obtain high-quality-images, while a noise-invariant video score distillation sampling is introduced for image animation.\n3. Extensive experiments show that the proposed method comsiderably outperforms the performance of video diffusion baselines in terms of video quliaty."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The technical contributions of the paper are somewhat limited. The proposed noise-invariant video score distillation only modifies some hyper-parameters of the original SDS techinque.\n2. Compared to the baseline results, the video actions enhanced using the proposed method in this paper are minimal or essentially stationary. The metrics in Table 1 also show that the proposed method heavily harm the dynamic degree of generated videos.\n3. AnimateDiff relies on high-quality LoRAs to improve the quality and consistency of generated videos. Please provide generated videos of AnimateDiff with high-quality LoRAs for a fair comparison."}},"nonreaders":[],"tmdate":1731428849761,"tcdate":1730538046552,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission7261/Reviewer_hdjr"],"signatures":["ICLR.cc/2025/Conference/Submission7261/Reviewer_hdjr"],"forum":"PKAZzhcIrP","number":2,"license":"CC BY 4.0","cdate":1730538046552,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission7261/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428849761,"domain":"ICLR.cc/2025/Conference","replyto":"PKAZzhcIrP","id":"HKFp5DUGXq","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Text-to-Video","Video Diffusion Models","Video Synthesis"]},"supplementary_material":{"value":"/attachment/e17ad8465374e66c804b8c9e47235e379c9b2e31.zip"},"primary_area":{"value":"generative models"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Text-to-video generation has trailed behind text-to-image generation in terms of quality and diversity, primarily due to the inherent complexities of spatio-temporal modeling and the limited availability of video-text datasets. Recent text-to-video diffusion models employ the image as an intermediate step, significantly enhancing overall performance but incurring high training costs. In this paper, we present I4VGen, a novel video diffusion inference pipeline to leverage advanced image techniques to enhance pre-trained text-to-video diffusion models, which requires no additional training. Instead of the vanilla text-to-video inference pipeline, I4VGen consists of two stages: anchor image synthesis and anchor image-augmented text-to-video synthesis. Correspondingly, a simple yet effective generation-selection strategy is employed to achieve visually-realistic and semantically-faithful anchor image, and an innovative noise-invariant video score distillation sampling (NI-VSDS) is developed to animate the image to a dynamic video by distilling motion knowledge from video diffusion models, followed by a video regeneration process to refine the video. Extensive experiments show that the proposed method produces videos with higher visual realism and textual fidelity. Furthermore, I4VGen also supports being seamlessly integrated into existing image-to-video diffusion models, thereby improving overall video quality."},"_bibtex":{"value":"@misc{\nguo2025ivgen,\ntitle={I4{VG}en: Image as Free Stepping Stone for Text-to-Video Generation},\nauthor={Xiefan Guo and Jinlin Liu and Miaomiao Cui and Liefeng Bo and Di Huang},\nyear={2025},\nurl={https://openreview.net/forum?id=PKAZzhcIrP}\n}"},"title":{"value":"I4VGen: Image as Free Stepping Stone for Text-to-Video Generation"},"pdf":{"value":"/pdf/41a342ae5a16dce21f9e8d2374dbcd67681e2bd1.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"guo|i4vgen_image_as_free_stepping_stone_for_texttovideo_generation"},"authorids":{"value":["~Xiefan_Guo1","~Jinlin_Liu1","~Miaomiao_Cui2","~Liefeng_Bo1","~Di_Huang4"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Xiefan Guo","Jinlin Liu","Miaomiao Cui","Liefeng Bo","Di Huang"]}},"version":2},{"content":{"summary":{"value":"In this paper, the authors present a text-to-video model consisting of multiple components including text-2-video, temporal interpolation and super-resolution components. The text-2-video model is built on the pre-trained text-2-image model. Additionally, they introduce Vimeo25M dataset to enhance the quality of text-to-video generation. The method is straightforward and the paper showcases the versatility of the model in long video generation and personalized video generation."},"presentation":{"value":"3 good"},"contribution":{"value":"2 fair"},"soundness":{"value":"3 good"},"strengths":{"value":"- The paper is well-written and easy to follow\n\n- The resulting model is capable of handling both T2I and T2V tasks.\n\n- The paper provides a human evaluation of video generation quality. \n\n- The paper introduces Vimeo25M dataset which is a collection of 25 million text-video pairs. The dataset aids in boosting the performance of model in terms of quality and diversity.\n\n- Joint image-video training of model is interesting and seems a reasonable approach for training video generation models."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1- The technical novelty of the proposed method is limited as it is a combination of existing techniques, including text-to-image generation, frame interpolation, and super-resolution. \n\n2- Dandi et al.'s work on “Jointly Trained Image and Video Generation using Residual Vectors\" is relevant to this paper's exploration of joint image-video training. However, this paper is neither cited nor discussed.\n\n3- There is no analysis of the contribution of individual components or techniques on the video generation performance. The paper mentions the temporal module, joint image-video training, and usage of Vimeo25M dataset for training. However, we don’t see a comprehensive analysis on the impact of each of them on the performance. It’s hard to understand which component has a higher impact on the final model. The only analysis we see is in Fig 10 which is very limited by showing only three images.\n\n4- The role of the temporal self-attention module (SA-T) during image-only training phases is ambiguous. It's unclear from Figure 3 whether SA-T is frozen or entirely excluded from the process.\n\n5- It’s not clear what is the benefit of the technique explained on page 4: \"our approach differs from conventional video frame interpolation methods, as each frame generated through interpolation replaces the corresponding input frame. In other words, every frame in the output is newly synthesized\"\n\n6- How did authors make sure there is enough correlation between text and video segment? Details of text-video pair selection are missing.\n\n7- In Fig. 9 (b) statistics on resolution are not clear. what “99.9%” means?\n\n8- There is no comprehensive evaluation of diversity (since the paper claimed to improve it) besides Fig. 4. Did the authors consider evaluating diversity with human evaluation?"},"confidence":{"value":"4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work."},"questions":{"value":"see the weakness section."},"rating":{"value":"6: marginally above the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636218269,"tcdate":1699030736492,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission2758/Reviewer_m4rL"],"signatures":["ICLR.cc/2024/Conference/Submission2758/Reviewer_m4rL"],"forum":"p09XyFxZkc","number":4,"license":"CC BY 4.0","cdate":1699030736492,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission2758/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636218269,"domain":"ICLR.cc/2024/Conference","replyto":"p09XyFxZkc","id":"DsLCp9L75h","forumContent":{"venue":{"value":"Submitted to ICLR 2024"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["text-to-video generation","diffusion models"]},"supplementary_material":{"value":"/attachment/9b7be39ae37699c99be3e7c7ae7b8d3b62214416.zip"},"primary_area":{"value":"generative models"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"This work aims to learn a high-quality text-to-video (T2V) generative model by leveraging a pre-trained text-to-image (T2I) model as a basis. It is a highly desirable yet challenging task to simultaneously a) accomplish the synthesis of visually realistic and temporally coherent videos while b) preserving the strong creative generation nature of the pre-trained T2I model. To this end, we propose LaVie, an integrated video generation framework that operates on cascaded video latent diffusion models, comprising a base T2V model, a temporal interpolation model, and a video super-resolution model. Our key insights are two-fold: 1) We reveal that the incorporation of simple temporal self-attentions, coupled with relative positional encoding, adequately captures the temporal correlations inherent in video data. 2) Additionally, we validate that the process of joint image-video fine-tuning plays a pivotal role in producing high-quality and creative outcomes. To enhance the performance of LaVie, we contribute a comprehensive and diverse video dataset named Vimeo25M, consisting of 25 million text-video pairs that prioritize quality, diversity, and aesthetic appeal. Extensive experiments demonstrate that LaVie achieves state-of-the-art performance both quantitatively and qualitatively. Furthermore, we showcase the versatility of pre-trained LaVie models in various long video generation and personalized video synthesis applications."},"_bibtex":{"value":"@misc{\nwang2024lavie,\ntitle={LaVie: High-Quality Video Generation with Cascaded Latent Diffusion Models},\nauthor={Yaohui Wang and Xinyuan Chen and Xin Ma and Shangchen Zhou and Ziqi Huang and Yi Wang and Ceyuan Yang and Yinan He and Jiashuo Yu and Peiqing Yang and Yuwei Guo and Tianxing Wu and Chenyang Si and Yuming Jiang and Cunjian Chen and Chen Change Loy and Bo Dai and Dahua Lin and Yu Qiao and Ziwei Liu},\nyear={2024},\nurl={https://openreview.net/forum?id=p09XyFxZkc}\n}"},"title":{"value":"LaVie: High-Quality Video Generation with Cascaded Latent Diffusion Models"},"pdf":{"value":"/pdf/f40f1ea778821a8640f919df133835970623a510.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Rejected_Submission"},"paperhash":{"value":"wang|lavie_highquality_video_generation_with_cascaded_latent_diffusion_models"},"authorids":{"value":["~Yaohui_Wang1","~Xinyuan_Chen1","~Xin_Ma3","~Shangchen_Zhou1","~Ziqi_Huang2","~Yi_Wang19","~Ceyuan_Yang2","~Yinan_He1","~Jiashuo_Yu1","~Peiqing_Yang1","~Yuwei_Guo1","~Tianxing_Wu2","~Chenyang_Si2","~Yuming_Jiang1","~Cunjian_Chen2","~Chen_Change_Loy2","~Bo_Dai2","~Dahua_Lin1","~Yu_Qiao1","~Ziwei_Liu1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Yaohui Wang","Xinyuan Chen","Xin Ma","Shangchen Zhou","Ziqi Huang","Yi Wang","Ceyuan Yang","Yinan He","Jiashuo Yu","Peiqing Yang","Yuwei Guo","Tianxing Wu","Chenyang Si","Yuming Jiang","Cunjian Chen","Chen Change Loy","Bo Dai","Dahua Lin","Yu Qiao","Ziwei Liu"]}},"version":2},{"content":{"venue":{"value":"ICLR 2027 Conference Submission"},"keywords":{"value":["dataset watermarking","shortcut learning","causal reasoning","distributionally robust optimization","adversarial attacks","robustness"]},"primary_area":{"value":"causal reasoning"},"abstract":{"value":"Spurious correlations (SC) are correlated features in datasets that allow for better benchmark predictions, but are not robust to domain change and so might harm generalisation when relied upon.\nSC can be unintended, such as background bias in object classification and diagnosis markings in medical images, but they have also been used intentionally, such as in dataset protection.\nTraditional methods for addressing SC often require expert knowledge and task-specific operations, leading to recent research to detect and correct SC without putting shortcut labels in the training objective.\nIn this paper, we follow this line of research, both when SC is unintended and when it is deliberately injected.\nWe propose a method that does not use shortcut labels during training; instead, we create label-conflicting counterexample mixtures so shortcut-following predictions become costly.\nWe study this as a clean-counterexample supervision regime: biased training data and a small clean counterexample pool are available, but shortcut labels are not used in the training objective.\nIn a controlled hidden-watermark setting, the method suppresses watermark-only prediction to chance, validating the intended mechanism.\nOn augmented Waterbirds, where the shortcut is a replaceable background cue, we compare against previous methods, attaining higher worst-group accuracy without domain labels in the objective.\nAdditional CelebA results test a visually entangled attribute shortcut and show weaker, more variable gains.\nThese results suggest that clean counterexamples can provide useful weak supervision for shortcut attenuation, with effectiveness governed by shortcut separability and counterexample quality."},"_bibtex":{"value":"@inproceedings{\nanonymous2026counterexampleguided,\ntitle={Counterexample-Guided Shortcut Suppression without Shortcut Labels},\nauthor={Anonymous},\nbooktitle={Submitted to The Fifteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=Xzo6qIcdSW},\nnote={under review}\n}"},"title":{"value":"Counterexample-Guided Shortcut Suppression without Shortcut Labels"},"pdf":{"value":"/pdf/8150d537acd6ff076d490e9d865444165c27043c.pdf"},"venueid":{"value":"ICLR.cc/2027/Conference/Submission"}},"tmdate":1791231863370,"tcdate":1789682317309,"writers":["ICLR.cc/2027/Conference","ICLR.cc/2027/Conference/Submission34754/Authors"],"signatures":["ICLR.cc/2027/Conference/Submission34754/Authors"],"forum":"Xzo6qIcdSW","license":"CC BY 4.0","number":34754,"cdate":1789682317309,"readers":["everyone"],"invitations":["ICLR.cc/2027/Conference/-/Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Bidding","ICLR.cc/2027/Conference/Submission34754/-/Full_Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Reviewing","ICLR.cc/2027/Conference/-/Edit"],"mdate":1791231863370,"odate":1791031328286,"domain":"ICLR.cc/2027/Conference","id":"Xzo6qIcdSW","version":2},{"content":{"TLDR":{"value":"We propose hierarchical keyframe selector to avoid redundant information for long-form video question answering"},"venue":{"value":"Video-Langauge Models Poster"},"keywords":{"value":["video question answering (VQA)","VLM","LLM"]},"supplementary_material":{"value":"/attachment/fb88e1dbafa3b0918d498e59aa3a058959bfe115.pdf"},"abstract":{"value":"Long-form videos that span across wide temporal intervals are highly information redundant and contain multiple distinct events or entities that are often loosely related. Therefore, when performing long-form video question answering (LVQA), all information necessary to generate a correct response can often be contained within a small subset of frames. Recent literature explore the use of large language models (LLMs) in LVQA benchmarks, achieving exceptional performance, while relying on vision language models (VLMs) to convert all visual content within videos into natural language. Such VLMs often independently caption a large number of frames uniformly sampled from long videos, which is not efficient and can mostly be redundant. Questioning these decision choices, we explore optimal strategies for key-frame selection that can significantly reduce these redundancies, namely Hierarchical Keyframe Selector. Our proposed framework, LVNet, achieves state-of-the-art performance at a comparable caption scale across three benchmark LVQA datasets: EgoSchema, NExT-QA, IntentQA. The code can be found at https://github.com/jongwoopark7978/LVNet"},"_bibtex":{"value":"@inproceedings{\npark2025too,\ntitle={Too Many Frames, Not All Useful: Efficient Strategies for Long-Form Video {QA}},\nauthor={Jongwoo Park and Kanchana Ranasinghe and Kumara Kahatapitiya and Wonjeong Ryu and Donghyun Kim and Michael S Ryoo},\nbooktitle={Workshop on Video-Language Models @ NeurIPS 2024},\nyear={2025},\nurl={https://openreview.net/forum?id=K3xNKR3yOf}\n}"},"title":{"value":"Too Many Frames, Not All Useful: Efficient Strategies for Long-Form Video QA"},"pdf":{"value":"/pdf/36cf8b6beaeb9d9132c908a24aa669f08533fdce.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Workshop/Video-Langauge_Models"},"paperhash":{"value":"park|too_many_frames_not_all_useful_efficient_strategies_for_longform_video_qa"},"authorids":{"value":["~Jongwoo_Park1","~Kanchana_Ranasinghe1","~Kumara_Kahatapitiya1","~Wonjeong_Ryu1","~Donghyun_Kim2","~Michael_S_Ryoo1"]},"track":{"value":"Long Paper Track (up to 9 pages)"},"authors":{"value":["Jongwoo Park","Kanchana Ranasinghe","Kumara Kahatapitiya","Wonjeong Ryu","Donghyun Kim","Michael S Ryoo"]}},"tmdate":1736861079999,"pdate":1730081752047,"tcdate":1725487184846,"writers":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission11/Authors"],"signatures":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission11/Authors"],"forum":"K3xNKR3yOf","license":"CC BY 4.0","number":11,"cdate":1725487184846,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Post_Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission11/-/Full_Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Edit","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission11/-/Camera-Ready_Revision"],"mdate":1736861079999,"odate":1736861079971,"domain":"NeurIPS.cc/2024/Workshop/Video-Langauge_Models","id":"K3xNKR3yOf","version":2},{"content":{"summary":{"value":"The authors tackle the problem of partial observability video prediction, which deals with video prediction problems in which the camera is also moving, in which case the video is influenced by both the scene dynamics and the camera's motion -- common in autonomous vehicles and robot manipulators. \n\nThis work explicitly models camera motion dynamics by extending the observed image state (existing settings) by introducing three models build upon prior works. Two models are based upon SVG-lp that learn image-action priors (LeAP) -- \n- (i) vg-leap (imagen-action pairs generated using a single stochastic process), and \n- (ii) causal-leap (causal relationship between image and action), and \n- (iii) RAFI -- that augments the image-action state pair of an existing flow-matching model RIVER. \n\nFor learning the latent action prior, two variational approaches are presented -- (i) combined image-action prior derived from both the observed image and action, with the assumption that image and action states are conditionally independent, and (ii) two separate posterior priors learnt for the image and action latent variables assuming causal interlinking. \n\nThe experimental evaluations are performed on the RoAM dataset which consists of synchronized image-action pairs recorded with a Turtlebot robot using a stereo camera setup. The dataset has 45 training videos and 5 testing sequences (300k videos sequences of 25 frames each used for training). The models are evaluated on the task of video generation for generating 10 frames conditioned of randomly sampled 5 previous consecutive frames. Perceptual and semantic quantitative metrics are used for evaluation."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"- How does the model perform on long-horizon video prediction problems. \n- How is the performance in different robot settings? (for instance, robot manipulators, drones, etc.)\n- Were any other latent conditioning methods evaluated?"},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"The problem of partially observable video prediction is quite interesting and has applications in autonomous vehicles that have an onboard camera (such as autonomous cars/taxis), drones (with an onboard camera), and robot manipulators (that have wrist-mounted cameras). \n\nTackling video prediction under partially observable settings (where the acting agent is not visible on the camera) can be benefit robot applications (for instance pedestrian intent detection and prediction could influence autonomous driver decisions, video prediction networks as world models for robot manipulators etc.)\n\nThis work tackles the interesting idea of incorporating (robot) actions (as a learned latent) into the video prediction/generation task that have been traditionally conditioned on just image frames. \n\nThe problem statement is interesting the authors motivate it well. The paper is also well-presented and articulated except for a few spelling errors (for instance lossses pg 6, divergance pg 6)."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Although sound, I find the contributions (incorporating actions) to be minimal additions to the existing frameworks. For instance, in VG-leap, the extended image-action state pair is used to condition the SVG-lp model instead of just the images, and the latent posterior approximated with recurrent modules. \nSimilarly, in Causal-leap, two stochastic posteriors are learned -- one each for image and action, learnt using recurrent modules. \nFor RAFI, the image latent is concatenated with the action state along the image latent's channel dimension of an existing RIVER model. Given the extensive literature on latent conditioning methods, I would have liked to see comparisons/discussions with different latent conditioning methods for incorporating actions. \n\nI find the experimental evaluations quite weak. The experiments are performed just on the RoAM dataset which has just 45 video sequences for training and 5 for testing/inference I would have liked to see comprehensive comparisons -- for instance on the robot manipulation settings in which case the camera is mounted on the robot's wrist. The Causal-leap model should be evaluated on such a setting -- which has a larger action space (7dof instead of the 2-dimensional in RoAM) that was used to motivate the problem. A similar comparison could have been done on drones for instance. \n\nEvaluations on short sequences. Predicting just 10 frames or evaluating on 25 frame length long video sequences does not quantify as long-term video prediction (which was used to motivate the manuscript). In such short horizon prediction problems, it is harder to quantify the effect of the moving camera. I would have liked to see examples of video predictions that are influenced by external factors (for instance the car turning right due to an obstacle on the left -- collision avoidance being one of the motivations). The quantitative metrics used in the paper evaluate semantic/perceptual quality of the generated videos -- fvd, lpips, vgg-16 etc., and don't necessarily motivate incorporating actions into the video prediction problem. TLDR; how obstacles or other hinderances influence video predictions is quite unclear as the metrics primarily evaluate video quality and not decisions. This would also strengthen the applications of such systems."}},"nonreaders":[],"tmdate":1731428773402,"tcdate":1730829728374,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission13808/Reviewer_u7qA"],"signatures":["ICLR.cc/2025/Conference/Submission13808/Reviewer_u7qA"],"forum":"VAvZ4oinpa","number":4,"license":"CC BY 4.0","cdate":1730829728374,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission13808/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428773402,"domain":"ICLR.cc/2025/Conference","replyto":"VAvZ4oinpa","id":"QUfos4V8b9","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"TLDR":{"value":"We propose variational models for learning action priors for video generation tasks in situations where the camera is also moving like in autonomous cars or robots."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Stochastic Video Generation","Variational Inference"]},"supplementary_material":{"value":"/attachment/ba85e8fd60e0de14ef7300be056e5cbf5f02a58a.pdf"},"primary_area":{"value":"probabilistic methods (Bayesian methods, variational inference, sampling, UQ, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Long-term stochastic video generation remains challenging, especially with moving cameras. This scenario introduces complex interactions between camera movement and observed pixels, resulting in intricate spatio-temporal dynamics and partial observability issues. Current approaches often focus on pixel-level image reconstruction, neglecting explicit modeling of camera motion dynamics. Our proposed solution incorporates camera motion or action as an extended part of the observed image state, employing a multi-modal learning framework to simultaneously model both image and action. We introduce three models: (i) Video Generation with Learning Action Prior (VG-LeAP) that treats the image-action pair as an augmented state generated from a single latent stochastic process and uses variational inference to learn the image-action latent prior; (ii) Causal-LeAP, which establishes a causal relationship between action and the observed image frame, and learns a seperate action prior, conditioned on the observed image states along with the image prior; and (iii) RAFI, which integrates the augmented image-action state concept with a conditional flow matching framework, demonstrating that this action-conditioned image generation concept can be extended to other transformer-based architectures. Through comprehensive empirical studies on robotic video dataset, RoAM, we highlight the importance of multi-modal training in addressing partially observable video generation problems."},"_bibtex":{"value":"@misc{\nsarkar2025video,\ntitle={Video Generation with Learned Action Prior},\nauthor={Meenakshi Sarkar and Devansh Bhardwaj and Debasish Ghose},\nyear={2025},\nurl={https://openreview.net/forum?id=VAvZ4oinpa}\n}"},"title":{"value":"Video Generation with Learned Action Prior"},"pdf":{"value":"/pdf/7e0cbd10c17c4802fbb1a9110ab3deec2d204710.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"sarkar|video_generation_with_learned_action_prior"},"authorids":{"value":["~Meenakshi_Sarkar1","~Devansh_Bhardwaj1","~Debasish_Ghose1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Meenakshi Sarkar","Devansh Bhardwaj","Debasish Ghose"]}},"version":2},{"content":{"summary":{"value":"This paper discusses whether current LLMs truly understand the content of short videos by constructing a temporal counterfactual evaluation benchmark, Vinoground. The benchmark includes 1000 video-caption pairs and assesses the shortcomings of existing LLMs through text, video, and group scores. The evaluation results indicate that dense temporal reasoning within short videos remains a capability that LLMs need to improve and focus on."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"please refer to \"Weaknesses\"."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"- The three evaluation metrics proposed in the paper—text, video, and group score—effectively reflect the capabilities of LLM in Vinoground and can expose the model's random guessing behavior.\n\n- This work established a human baseline, providing a good reference for evaluating both open-source and closed-source LLMs.\n\n- This paper reports the performance of GPT-4o under 0-shot sampling as a control for testing bias in these models. This is a highlight.\n\n- The paper presents some new findings, such as the 64-frame variant of GPT-4o performs 5% worse than its 32-frame variant across all three metrics."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The author should further elaborate on the principles and methods of generating counterfactual captions, such as how GPT4 is guided to generate counterfactual captions through what kind of prompts. This part of the process is very confusing.\n\n- The author should explain why the generated captions can be found corresponding to videos in the VATEX dataset. I have some doubts about the process of benchmark construction. Shouldn't we first obtain the video and then generate the counterfactual captions?\n\n- The author should provide more detailed classification criteria, such as whether human classification is used, to determine if there is overlap or similarity between action and viewpoint in video-caption pairs.\n\n- Why do some pairs belong to a multitude of these minor categories, while some do not belong to any of them? The minor categories in the benchmark include 350 videos, and I have doubts about the data volume. Due to the imbalance of categories and the reduced number of videos, is the benchmark capable of reasonably evaluating the capabilities of LLMs?\n\n- The author should improve the readability of the article, for example: \"For generative LMMs, we can only provide inputs (e.g., 2 captions and 1 video) to the model and ask it to output an answer A or B.\" Without reading the subsequent experimental setups, the reference to A and B here may lead to confusion.\n\n- The author should provide the video length distribution for each category of the benchmark to illustrate the generalizability of the method where \"we concatenate the positive and negative videos into a single video with a 2 second black screen in between.\" If there are too many short videos, it may result in a significant loss of evaluation data.\n\n- While the article establishes a strong human baseline, the author should provide information on the distribution of factors such as age, gender, educational background, etc., for the human participants. This information would impact the reliability of the human baseline.\n\n- Figures 3, 6, 7, and 8 have low resolution, affecting readability.\n\nI think that there are many unclear points in this paper, especially in the crucial sections on dataset construction and model evaluation, which can lead to misunderstandings. \n\nIf my concerns can be addressed, I will raise my rating."}},"nonreaders":[],"tmdate":1732582762315,"tcdate":1729519059405,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission1636/Reviewer_nPZz"],"signatures":["ICLR.cc/2025/Conference/Submission1636/Reviewer_nPZz"],"forum":"a1P5kh2oo8","number":1,"license":"CC BY 4.0","cdate":1729519059405,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission1636/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732582762315,"domain":"ICLR.cc/2025/Conference","replyto":"a1P5kh2oo8","id":"916HrJcN9w","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"TLDR":{"value":"Modern SoTA LMMs still demonstrates subpar performance at temporal reasoning with our temporal counterfactual benchmark composed of natural videos."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["temporal reasoning; counterfactual reasoning; short video comprehension"]},"supplementary_material":{"value":"/attachment/fd093e5ac3c36911b68e297a2932dcf6ce1ed019.zip"},"primary_area":{"value":"datasets and benchmarks"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"There has been growing sentiment recently that modern large multimodal models (LMMs) have addressed most of the key challenges related to short video comprehension. As a result, both academia and industry are gradually shifting their attention towards the more complex challenges posed by understanding long-form videos. \nHowever, is this really the case?  Our studies indicate that LMMs still lack many fundamental reasoning capabilities even when dealing with short videos.  We introduce Vinoground, a temporal counterfactual LMM evaluation benchmark encompassing 1000 short and natural video-caption pairs. We demonstrate that existing LMMs severely struggle to distinguish temporal differences between different actions and object transformations.  For example, the best model GPT-4o only obtains $\\sim$50\\% on our text and video scores, showing a large gap compared to the human baseline of $\\sim$90\\%. All open-source multimodal models and CLIP-based models perform much worse, producing mostly random chance performance. Through this work, we shed light onto the fact that temporal reasoning in short videos is a problem yet to be fully solved. We will make our benchmark publicly available."},"_bibtex":{"value":"@misc{\nzhang2025vinoground,\ntitle={Vinoground: Scrutinizing {LMM}s over Dense Temporal Reasoning with Short Videos},\nauthor={Jianrui Zhang and Mu Cai and Yong Jae Lee},\nyear={2025},\nurl={https://openreview.net/forum?id=a1P5kh2oo8}\n}"},"title":{"value":"Vinoground: Scrutinizing LMMs over Dense Temporal Reasoning with Short Videos"},"pdf":{"value":"/pdf/2e0dd590d0c11d11d7071fa71d7eaea83ded91ac.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"zhang|vinoground_scrutinizing_lmms_over_dense_temporal_reasoning_with_short_videos"},"authorids":{"value":["~Jianrui_Zhang1","~Mu_Cai1","~Yong_Jae_Lee2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Jianrui Zhang","Mu Cai","Yong Jae Lee"]}},"version":2},{"content":{"summary":{"value":"This paper proposes DrivingWorld, a self-driving world model based on the GPT architecture for generating high-fidelity and long-term video sequence predictions. The model improves performance through three key innovations: Temporal-Aware Tokenization, Hybrid Token Prediction, and Long-time Controllable Strategies. Experiments show that the model can generate more than 100 seconds of high-quality video at a frequency of 5Hz. Compared with the traditional GPT structure, this method significantly reduces the computational cost by decoupling spatiotemporal processing while maintaining better temporal consistency and structural integrit"},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"The paper contains inconsistent  descriptions regarding the video generation frequency, stating 6Hz in the method section ( line 161 \"capable of extending predictions beyond 30 seconds at a frequency of 6Hz. \" ) but 5Hz in the experiment section (line 416 \"our model can generate up to 640 future frames at 5 Hz, resulting in 128-second videos with strong temporal consistency.\")"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The paper demonstrates clarity in its presentation and organization. The writing flows logically from motivation to implementation, while the figures illustrate key concepts and results. \n2. The paper's core innovation, applying temporal-aware GPT architecture to autonomous driving world modeling, represents an advancement in the field. \n3. The introduction of Dropout for Drifting-free Autoregression is a good solution to generating long-driving video sequences. Experiments show this technique addresses the common problem of quality degradation in long-term predictions."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The biggest concern lies in the model's controllability claims. While the paper demonstrates models of \"controllability\" with vehicle trajectory control, this is primarily restricted to lane changes on straight roads with minimal viewpoint variations. The surrounding environment, including other vehicles and road structures, remains purely autoregressive without direct control. \nBesides, the absence of more challenging scenarios like turning left/right in intersections raises questions about the model's true generalization capabilities and control flexibility.\n2. The generated videos still exhibit noticeable visual artifacts and physical inconsistencies. For instance, the black car in Figure 5's \"300s\" subfigure and the project webpage's Example 3 demonstrate unrealistic vehicle collision scenarios at the 14-second mark. \nThese issues highlight a considerable challenge: the model's inability to handle physically complex scenarios where vehicles are in close proximity or potential collision situations, which are crucial for autonomous driving applications.\n3. The method's long video generation is excellent, but the model's FID (16.4) and FVD (174.4) scores, while reasonable, do not lead the benchmark comparisons in Table 1. This becomes more apparent when considering recent SOTA works like DiVE and Vista, which are absent from the comparison."}},"nonreaders":[],"tmdate":1731427498224,"tcdate":1730695425269,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission1834/Reviewer_ceym"],"signatures":["ICLR.cc/2025/Conference/Submission1834/Reviewer_ceym"],"forum":"xJtWqVBZya","number":3,"license":"CC BY 4.0","cdate":1730695425269,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission1834/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427498224,"domain":"ICLR.cc/2025/Conference","replyto":"xJtWqVBZya","id":"HQ73Uay110","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["world model","video generation"]},"primary_area":{"value":"generative models"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Recent successes in autoregressive (AR) generation models, such as the GPT series in natural language processing, have motivated efforts to replicate this success in visual tasks. By leveraging the next-token prediction strategy, GPT-style models can forecast future events from past data. Some research aims to extend this approach to autonomous driving by building video-based world models capable of generating realistic future video sequences and predicting the ego state. However, the prior works tend to produce unsatisfactory results, since the classic GPT framework is designed to handle 1D contextual information, such as text, and lacks the inherent capability to model the spatial and temporal dynamics necessary for video generation. In this paper, we present DrivingWorld, a video-based world model for autonomous driving via a new GPT structure with spatial-temporal design. The key idea is to disentangle temporal and spatial information in the generation. Specifically, we first propose next-frame-prediction strategy to model temporal coherence between consecutive frames and then apply next-token-prediction strategy to capture spatial information within a frame. With the hybrid design, our work is capable of producing high-fidelity and consistent video clips with long-time duration. Experiments show that compared to the prior works, our method presents better quality of visual effects and more accurate controllable future video generation."},"_bibtex":{"value":"@misc{\nhu2024drivingworld,\ntitle={DrivingWorld: Constructing World Model for Autonomous Driving via Video {GPT}},\nauthor={Xiaotao Hu and Wei Yin and Mingkai Jia and Junyuan Deng and Xiaoyang Guo and Qian Zhang and Xiaoxiao Long and Ping Tan},\nyear={2024},\nurl={https://openreview.net/forum?id=xJtWqVBZya}\n}"},"title":{"value":"DrivingWorld: Constructing World Model for Autonomous Driving via Video GPT"},"pdf":{"value":"/pdf/aad7fc36550d4db84152c09c15aae1687253abc5.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"hu|drivingworld_constructing_world_model_for_autonomous_driving_via_video_gpt"},"authorids":{"value":["~Xiaotao_Hu1","~Wei_Yin2","~Mingkai_Jia1","~Junyuan_Deng1","~Xiaoyang_Guo1","~Qian_Zhang7","~Xiaoxiao_Long2","~Ping_Tan2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Xiaotao Hu","Wei Yin","Mingkai Jia","Junyuan Deng","Xiaoyang Guo","Qian Zhang","Xiaoxiao Long","Ping Tan"]}},"version":2},{"content":{"summary":{"value":"The paper proposes a counterfactual video editing process using multiple pretrained foundation models to engineer causality-aware text prompts for video editing.\nThe core idea is to embed Pearl-style counterfactual ideas to create meta prompts that enable causality-aware edits to a video editing prompt.\nThe paper does this by iteratively refining an editing text prompt given the causal graph, some in-context-learning examples and some meta instructions.\nThe paper then demonstrates the method empirically both quantitatively and qualitatively on real human face videos."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"- Why do you remove all causal DAG variables from minimality? It seems you should only exclude intervention and downstream variables, not upstream variables. In fact, upstream variables should be preserved."},"rating":{"value":4},"details_of_ethics_concerns":{"value":"- This paper on counterfactual video editing definitely needs some ethical discussion that is missing from the current paper.\n- Possible issue could include impersonating people or editing them older or younger without their knowledge.\n- There could also be ethical issues with changing someone from an adult to a minor.\n- The ethical questions are more intense because videos can be more convincing than other forms of media like images.\n- Another concern is modifying gender of a factual person against their knowledge.\n- There seems to be many ways that this could be harmful or unethical so a strong discussion section seems critical."},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- Introduces a simple prompting method with causal-awareness to produce better counterfactual prompts for video generation (a causal-aware type of prompt engineering).\n- The paper is well-written and easy to understand.\n- The paper appropriately leverages pretrained models and does not add unnecessary complexity to the methods.\n- Introduce some simple VLM-based metrics for counterfactual video generation evaluation."},"flag_for_ethics_review":{"value":["Yes, Discrimination / bias / fairness concerns","Yes, Privacy, security and safety","Yes, Potentially harmful insights, methodologies and applications"]},"weaknesses":{"value":"- The paper is an application-focused paper since it composes multiple other models to accomplish it's objectives. From an application standpoint, this is a strong and well-explained approach that doesn't do anything unnecessary and uses the natural tools to accomplish the goal. From a fundamental side, it does not make a significant contribution since even the optimization is based on TextGrad. Thus, the methodological novelty is relatively low---though this is not necessarily a rejection-worthy concern given the application focus.\n\n- This method does require a causal graph to be known a priori. This should be noted as a key limitation for more general causal video editing. \n\n- One potential issue with the proposed method is that the LLMs or VLMs may be using implicit world knowledge rather than actually using the provided causal graph. Currently, you only use a \"realisitic\" causal graph that would align with the pretrained knowledge of foundation models. Thus, it is unclear whether the editing ability comes from the provided causal graph or simply world knowledge.\n   - To more carefully explore your method, I would be interested in another experiment with a non-realistic causal graph (e.g., one that is opposite of true causal factors). Would your method still obey the (incorrect) causal graph? This would be an interesting way to stress test whether it is actually using your causal graph or just using world knowledge.\n\n- Experimental baselines are limited\n   - Because this is primarily prompt engineering, I feel that only having a few baselines is a little odd. You have an initial prompt and LLM paraphrasing as the two main baselines. However, it seems there should be other baselines like a meta-prompt or thinking-based prompt enhancement or other types of prompt editing techniques. For example, first describe the video verbosely and then ask it to causally change a part of the prompt to create a complex causally modified prompt.\n   - Needless to say, more prompt engineering baselines seem appropriate. What is the best zero-shot prompt engineering technique that can leverage VLMs first to get textual respresentation of video and then the target modification.\n\n- (Minor/Typo) Please use \\citep and \\citet macros correctly. If taking out a citation leaves the sentence structure intact, you should use \\citep. If the citation is needed for the sentence, then you should use \\citet.\n\n(I would probably increase my score if you could demonstrate that your method actually respects the input causal graph, which could differ from reality or differ from pretrained knowledge (like a near inverse of the given causal graph). This would show it's not just using pretrained knowledge but actually making edits based on the given causal graph.)"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762933923833,"tcdate":1761924432468,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission20489/Reviewer_gSfz"],"signatures":["ICLR.cc/2026/Conference/Submission20489/Reviewer_gSfz"],"forum":"2nBjmUB5PM","number":2,"license":"CC BY 4.0","cdate":1761924432468,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission20489/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762933923833,"domain":"ICLR.cc/2026/Conference","replyto":"2nBjmUB5PM","id":"T1fZdzAE8m","forumContent":{"TLDR":{"value":"A causal framework for counterfactual video generation, guided by a vision-language model (VLM)"},"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["counterfactual generation","causality","generative AI","diffusion models","VLMs"]},"supplementary_material":{"value":"/attachment/794ce3d2bc7ff1f831b2233210147c12c8f2d760.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Adapting text-to-image (T2I) latent diffusion models (LDMs) to video editing has shown strong visual fidelity and controllability, but challenges remain in maintaining causal relationships inherent to the video data generating process. In this work, we propose CSVC, a framework for counterfactual video generation grounded in structural causal models (SCMs) and formulated as an out-of-distribution (OOD) prediction task. CSVC builds on black-box counterfactual functions, which approximate SCM mechanisms without explicit structural equations. In our framework, large language models (LLMs) generate counterfactual prompts that are consistent with a predefined causal graph, while LDM-based video editors produce the corresponding video counterfactuals. To ensure faithful interventions, we introduce a vision–language model (VLM)-based textual loss that refines prompts to enforce counterfactual conditioning, steering the LDM latent space toward causally meaningful OOD variations without internal model access or fine-tuning. Experiments on real-world facial videos show that CSVC achieves state-of-the-art causal effectiveness while preserving temporal consistency and visual quality. By combining SCM reasoning with black-box generative models, CSVC enables realistic “what if” hypothetical video scenarios with applications in digital media and healthcare."},"_bibtex":{"value":"@misc{\nspyrou2025causally,\ntitle={Causally Steered Diffusion for Video Counterfactual Generation},\nauthor={Nikos Spyrou and Athanasios Vlontzos and Paraskevas Pegios and Thomas Melistas and Nefeli Gkouti and Yannis Panagakis and Giorgos Papanastasiou and Sotirios A. Tsaftaris},\nyear={2025},\nurl={https://openreview.net/forum?id=2nBjmUB5PM}\n}"},"title":{"value":"Causally Steered Diffusion for Video Counterfactual Generation"},"pdf":{"value":"/pdf/e68e9a649754408dc83b3556197f3d7edb923ce0.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"spyrou|causally_steered_diffusion_for_video_counterfactual_generation"},"authorids":{"value":["~Nikos_Spyrou1","~Athanasios_Vlontzos1","~Paraskevas_Pegios1","~Thomas_Melistas1","~Nefeli_Gkouti1","~Yannis_Panagakis1","~Giorgos_Papanastasiou2","~Sotirios_A._Tsaftaris1"]},"authors":{"value":["Nikos Spyrou","Athanasios Vlontzos","Paraskevas Pegios","Thomas Melistas","Nefeli Gkouti","Yannis Panagakis","Giorgos Papanastasiou","Sotirios A. Tsaftaris"]}},"version":2},{"content":{"venue":{"value":"CoRR 2024"},"pdf":{"value":"https://arxiv.org/pdf/2403.17299v1"},"venueid":{"value":"dblp.org/journals/CORR/2024"},"paperhash":{"value":"he|decoding_probing_revealing_internal_linguistic_structures_in_neural_language_models_using_minimal_pairs"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Linyang_He:","https://dblp.org/search/pid/api?q=author:Peili_Chen:","https://dblp.org/search/pid/api?q=author:Ercong_Nie:","~Yuanning_Li2","https://dblp.org/search/pid/api?q=author:Jonathan_R._Brennan:"]},"html":{"value":"https://doi.org/10.48550/arXiv.2403.17299"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2403-17299,\n  publtype={informal},\n  author={Linyang He and Peili Chen and Ercong Nie and Yuanning Li and Jonathan R. Brennan},\n  title={Decoding Probing: Revealing Internal Linguistic Structures in Neural Language Models using Minimal Pairs},\n  year={2024},\n  cdate={1704067200000},\n  journal={CoRR},\n  volume={abs/2403.17299},\n  url={https://doi.org/10.48550/arXiv.2403.17299}\n}\n"},"abstract":{"value":"Inspired by cognitive neuroscience studies, we introduce a novel `decoding probing' method that uses minimal pairs benchmark (BLiMP) to probe internal linguistic characteristics in neural language models layer by layer. By treating the language model as the `brain' and its representations as `neural activations', we decode grammaticality labels of minimal pairs from the intermediate layers' representations. This approach reveals: 1) Self-supervised language models capture abstract linguistic structures in intermediate layers that GloVe and RNN language models cannot learn. 2) Information about syntactic grammaticality is robustly captured through the first third layers of GPT-2 and also distributed in later layers. As sentence complexity increases, more layers are required for learning grammatical capabilities. 3) Morphological and semantics/syntax interface-related features are harder to capture than syntax. 4) For Transformer-based models, both embeddings and attentions capture grammatical features but show distinct patterns. Different attention heads exhibit similar tendencies toward various linguistic phenomena, but with varied contributions."},"title":{"value":"Decoding Probing: Revealing Internal Linguistic Structures in Neural Language Models using Minimal Pairs"},"authors":{"value":["Linyang He","Peili Chen","Ercong Nie","Yuanning Li","Jonathan R. Brennan"]}},"tmdate":1769641067332,"pdate":1735603200000,"externalIds":["dblp:journals/corr/abs-2403-17299"],"tcdate":1769641064022,"writers":["~"],"signatures":["~Yuanning_Li2"],"forum":"7rMVGnVWYB","license":"CC BY-SA 4.0","number":813954,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1769641067332,"domain":"DBLP.org","id":"7rMVGnVWYB","version":2},{"content":{"venue":{"value":"LREC/COLING 2024"},"pdf":{"value":"https://aclanthology.org/2024.lrec-main.402.pdf"},"venueid":{"value":"dblp.org/conf/COLING/2024"},"paperhash":{"value":"he|decoding_probing_revealing_internal_linguistic_structures_in_neural_language_models_using_minimal_pairs"},"authorids":{"value":["~Linyang_He2","https://dblp.org/search/pid/api?q=author:Peili_Chen:","~Ercong_Nie1","~Yuanning_Li2","https://dblp.org/search/pid/api?q=author:Jonathan_R._Brennan:"]},"html":{"value":"https://aclanthology.org/2024.lrec-main.402"},"_bibtex":{"value":"@inproceedings{DBLP:conf/coling/HeCNLB24,\n  author={Linyang He and Peili Chen and Ercong Nie and Yuanning Li and Jonathan R. Brennan},\n  title={Decoding Probing: Revealing Internal Linguistic Structures in Neural Language Models Using Minimal Pairs},\n  year={2024},\n  cdate={1704067200000},\n  pages={4488-4497},\n  url={https://aclanthology.org/2024.lrec-main.402},\n  booktitle={LREC/COLING},\n  crossref={conf/coling/2024}\n}\n"},"abstract":{"value":"Inspired by cognitive neuroscience studies, we introduce a novel “decoding probing” method that uses minimal pairs benchmark (BLiMP) to probe internal linguistic characteristics in neural language models layer by layer. By treating the language model as the brain and its representations as “neural activations”, we decode grammaticality labels of minimal pairs from the intermediate layers’ representations. This approach reveals: 1) Self-supervised language models capture abstract linguistic structures in intermediate layers that GloVe and RNN language models cannot learn. 2) Information about syntactic grammaticality is robustly captured through the first third layers of GPT-2 and also distributed in later layers. As sentence complexity increases, more layers are required for learning grammatical capabilities. 3) Morphological and semantics/syntax interface-related features are harder to capture than syntax. 4) For Transformer-based models, both embeddings and attentions capture grammatical features but show distinct patterns. Different attention heads exhibit similar tendencies toward various linguistic phenomena, but with varied contributions."},"title":{"value":"Decoding Probing: Revealing Internal Linguistic Structures in Neural Language Models Using Minimal Pairs"},"authors":{"value":["Linyang He","Peili Chen","Ercong Nie","Yuanning Li","Jonathan R. Brennan"]}},"tmdate":1769641064190,"pdate":1704067200000,"tcdate":1729183238668,"writers":["~"],"signatures":["~Linyang_He2"],"forum":"3U0Bdf3fYp","license":"CC BY-SA 4.0","number":154196,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1769641064190,"domain":"DBLP.org","id":"3U0Bdf3fYp","version":2},{"content":{"venue":{"value":"ICLR 2027 Conference Submission"},"TLDR":{"value":"SceShortcut enables few-step and one-step full-sequence scene-aware motion synthesis via a global-to-local shortcut model with training augmented by supervision of the combined root transport velocity and regularization of the local root residual."},"keywords":{"value":["Scene-aware human motion generation","diffusion acceleration","controllable generation"]},"supplementary_material":{"value":"/attachment/47da2db6e683852145c26487ea38f9f624e71352.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Previous works on scene-aware human motion synthesis have mainly focused on motion-scene compatibility, with less attention to runtime efficiency.\nIn this work, we present SceShortcut, a shortcut-based framework that accelerates scene-aware motion synthesis through one-step or few-step sampling while producing plausible and scene-compatible motion.\nTo preserve fine-grained scene adaptation under a limited sampling budget, we introduce a global-to-local architecture with two stages trained jointly end-to-end.\nSpecifically, at each sampling step, conditioned on the text instruction and global geometric-semantic scene features, the global stage produces intermediate motion features and predicts a root transport velocity.\nThis velocity yields a coarse clean root motion estimate for querying fine-grained local scene geometry from a dense height field.\nThe local stage incorporates text conditioning and the queried geometry into these intermediate features to predict transport velocities for the remaining motion components and a residual correction to the global-stage root transport velocity.\nTo improve coarse root motion estimates and thereby enhance generation quality, we augment shortcut training with additional root-transport-velocity supervision and root-residual regularization.\nOn TRUMANS and HUMANISE, one-step sampling with SceShortcut achieves approximately $50\\sim100\\times$ speedups over recent advanced open-source baselines while producing plausible motion with comparable motion-scene alignment."},"_bibtex":{"value":"@inproceedings{\nanonymous2026sceshortcut,\ntitle={SceShortcut: Efficient Scene-Aware Motion Synthesis via Shortcut Modeling},\nauthor={Anonymous},\nbooktitle={Submitted to The Fifteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=DLOUGy0vS5},\nnote={under review}\n}"},"title":{"value":"SceShortcut: Efficient Scene-Aware Motion Synthesis via Shortcut Modeling"},"pdf":{"value":"/pdf/741aed70150e7acf5eee9cf1e99d0493078ef503.pdf"},"venueid":{"value":"ICLR.cc/2027/Conference/Submission"}},"tmdate":1791230227159,"tcdate":1789545285911,"writers":["ICLR.cc/2027/Conference","ICLR.cc/2027/Conference/Submission24026/Authors"],"signatures":["ICLR.cc/2027/Conference/Submission24026/Authors"],"forum":"DLOUGy0vS5","license":"CC BY 4.0","number":24026,"cdate":1789545285911,"readers":["everyone"],"invitations":["ICLR.cc/2027/Conference/-/Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Bidding","ICLR.cc/2027/Conference/Submission24026/-/Full_Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Reviewing","ICLR.cc/2027/Conference/-/Edit"],"mdate":1791230227159,"odate":1791031328286,"domain":"ICLR.cc/2027/Conference","id":"DLOUGy0vS5","version":2},{"content":{"summary":{"value":"The paper introduces VideoGPT+, a novel video understanding model that integrates image and video encoders to enhance video comprehension by leveraging detailed spatial information and temporal context. It presents VCG+ 112K, a new dataset developed through a semi-automatic annotation pipeline, which improves model performance. Additionally, the paper proposes VCGBench-Diverse, a benchmark covering 18 video categories, for a comprehensive evaluation of video LMMs."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"1. Are there any studies comparing the effects of using only an image encoder versus using both an image and a video encoder?\n\n2. Can you compare the performance effects between having information interaction and not having information interaction between them?"},"rating":{"value":5},"details_of_ethics_concerns":{"value":"NO"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The paper's strengths include its innovative dual-encoder design that effectively combines image and video encoders for richer spatiotemporal understanding of videos. \n2. The introduction of the VCG+ 112K dataset, developed through a semi-automatic annotation pipeline, enhances the model's performance by providing dense video captions and reasoning-based QA pairs. \n3. Furthermore, the proposal of the VCGBench-Diverse benchmark allows for a more comprehensive evaluation of video LMMs across diverse video types and dynamics."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The paper has the following weaknesses.\n1. The simultaneous use of both an image and a video encoder increases the number of parameters in the model, making it difficult to determine whether the effectiveness of the paper is accurate or if the performance improvement is simply due to the increased parameters. Are there any studies comparing the effects of using only an image encoder versus using both an image and a video encoder?\n\n2. The image encoder and the video encoder are independent of each other, with no interaction between them. Can you compare the performance effects between having information interaction and not having information interaction between them?\n\n3. The paper lacks some theoretical derivations."}},"nonreaders":[],"tmdate":1731428662442,"tcdate":1730535790631,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission11214/Reviewer_yRXZ"],"signatures":["ICLR.cc/2025/Conference/Submission11214/Reviewer_yRXZ"],"forum":"YGWxpOI6Y0","number":2,"license":"CC BY 4.0","cdate":1730535790631,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission11214/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428662442,"domain":"ICLR.cc/2025/Conference","replyto":"YGWxpOI6Y0","id":"HRv5UBSXTO","forumContent":{"TLDR":{"value":"VideoGPT+: Spatiotemporal Aware Video Conversation Model"},"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["video-conversation-model","large multi-modal model","multi-modal","video-conversation","image-and-video","phi-3-min","vision-language","video-chatbot"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Building on the advances of language models, Large Multimodal Models (LMMs) have contributed significant improvements in video understanding. While the current video LMMs utilize advanced Large Language Models (LLMs), they rely on either image or video encoders to process visual inputs, each of which has its own limitations. Image encoders excel at capturing rich spatial details from frame sequences but lack explicit temporal context, which can be important in videos with intricate action sequences. On the other hand, video encoders provide temporal context but are often limited by computational constraints that lead to processing only sparse frames at lower resolutions, resulting in reduced contextual and spatial understanding. To this end, we introduce our model, which combines the complementary benefits of the image encoder (for detailed spatial understanding) and the video encoder (for global temporal context modeling). The model processes videos by dividing them into smaller segments and applies an adaptive pooling strategy on features extracted by both image and video encoders. Our architecture showcases improved performance across multiple video benchmarks, including VCGBench, MVBench and Zero-shot question-answering. Further, we develop 112K video-instruction set using a novel semi-automatic annotation pipeline which further improves the model performance. Additionally, to comprehensively evaluate video LMMs, we present our bench, covering 18 broad video categories such as lifestyle, sports, science, gaming, and surveillance videos. This benchmark with 4,354 question-answer pairs evaluates the generalization of existing LMMs on dense video captioning, spatial and temporal understanding, and complex reasoning, ensuring comprehensive assessment across diverse video types and dynamics. Our code, dataset, and pre-trained models will be publicly released."},"_bibtex":{"value":"@misc{\nmaaz2024videogpt,\ntitle={Video{GPT}+: Integrating Image and Video Encoders for Enhanced Video Understanding},\nauthor={Muhammad Maaz and Hanoona Abdul Rasheed and Salman Khan and Fahad Shahbaz Khan},\nyear={2024},\nurl={https://openreview.net/forum?id=YGWxpOI6Y0}\n}"},"title":{"value":"VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding"},"pdf":{"value":"/pdf/ec0f125e42fed1a2714c46de829865aaa0e50324.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"maaz|videogpt_integrating_image_and_video_encoders_for_enhanced_video_understanding"},"authorids":{"value":["~Muhammad_Maaz1","~Hanoona_Abdul_Rasheed1","~Salman_Khan4","~Fahad_Shahbaz_Khan1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Muhammad Maaz","Hanoona Abdul Rasheed","Salman Khan","Fahad Shahbaz Khan"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a framework that investigates the tracking ability of video diffusion models.  It first trains a motion perception model to predict the dense trajectory.  After that, the proposed method identifies the motion-aware features with a new dataset, and injects these features into identically the same video diffusion models. This paper evaluates long-range 2D optical flow and dense 3D tracking tasks."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"Please refer to the weaknesses."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The idea of training a video diffusion model to predict dense trajectories is interesting.\n2. This paper designs a motion-labeled video dataset for the analysis of motion-aware features is insightful.\n3. The paper is well-written and easy to follow."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Despite the interesting design, the proposed method achieves worse results than the baselines, as shown in Table 1 and Table 2. This limits the application of the proposed method.\n\n2. This paper highlights the point tracking task, but lacks one important benchmark, TAP-Vid [1]. How about the tracking performance compared to non-diffusion methods such as CoTracker3 on TAP-Vid?\n\n[1] TAP-Vid: A Benchmark for Tracking Any Point in a Video\n\n3. This paper evaluates videos of no more than 48 frames. However, in real applications, videos could be much longer than that. How does the method track longer videos?\n\n4. According to line 89, this paper claims comparable accuracy in occlusion handling. However, the only result is OA reported in Table 2, which is lower than both baseline methods. This makes the claim not solid.\n\n5. Figure 5 and Figure 6 show some visualizations of the proposed method. However, these are limited to scenarios with objects. How about real-world applications, such as those involving human motions?\n\n6. How about performing a similar PCA analysis using features from the original video diffusion models in Figure 4? Will the trend of selecting layers be the same as in the motion perception model?"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762917270343,"tcdate":1762005645354,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission4276/Reviewer_yiqk"],"signatures":["ICLR.cc/2026/Conference/Submission4276/Reviewer_yiqk"],"forum":"E87rJRkmTk","number":4,"license":"CC BY 4.0","cdate":1762005645354,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission4276/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762917270343,"domain":"ICLR.cc/2026/Conference","replyto":"E87rJRkmTk","id":"AmoDNE2NNx","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"TLDR":{"value":"A framework that bridges generation and understanding in video diffusion models."},"keywords":{"value":["Video Diffusion Model; Point Tracking"]},"supplementary_material":{"value":"/attachment/49fc5b0cf9362e1836a0285ffcfe1e3efa33ea77.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Video diffusion models, trained on large-scale datasets, naturally capture correspondences of shared features across frames. \nRecent works have exploited this property for tasks such as optical flow prediction and tracking in a zero-shot setting.\nMotivated by these findings, we investigate whether supervised training can more fully harness the tracking capability of video diffusion models. To this end, we propose Moaw, a framework that unleashes motion awareness for video diffusion models and leverages it to facilitate motion transfer. Specifically, we train a diffusion model for motion perception, shifting its modality from image-to-video generation to video-to-dense-tracking. We then construct a motion-labeled dataset to identify features that encode the strongest motion information, and inject them into a structurally identical video generation model.  Owing to the homogeneity between the two networks, these features can be naturally adapted in a zero-shot manner, enabling motion transfer without additional adapters. Our work provides a new paradigm for bridging generative modeling and motion understanding, paving the way for more unified and controllable video learning frameworks."},"_bibtex":{"value":"@misc{\nzhang2026moaw,\ntitle={Moaw: Unleashing Motion Awareness for Video Diffusion Models},\nauthor={Tianqi Zhang and Ziyi Wang and Wenzhao Zheng and Weiliang Chen and Yuanhui Huang and Zhengyang Huang and Jie Zhou and Jiwen Lu},\nyear={2026},\nurl={https://openreview.net/forum?id=E87rJRkmTk}\n}"},"title":{"value":"Moaw: Unleashing Motion Awareness for Video Diffusion Models"},"pdf":{"value":"/pdf/945b8d2efaea8f066154744686782feab4f804db.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"zhang|moaw_unleashing_motion_awareness_for_video_diffusion_models"},"authorids":{"value":["~Tianqi_Zhang4","~Ziyi_Wang3","~Wenzhao_Zheng1","~Weiliang_Chen1","~Yuanhui_Huang1","~Zhengyang_Huang2","~Jie_Zhou3","~Jiwen_Lu1"]},"authors":{"value":["Tianqi Zhang","Ziyi Wang","Wenzhao Zheng","Weiliang Chen","Yuanhui Huang","Zhengyang Huang","Jie Zhou","Jiwen Lu"]}},"version":2},{"content":{"venue":{"value":"ICLR 2024 spotlight"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["shortcut learning","spurious correlations","architectural inductive bias"]},"primary_area":{"value":"visualization or interpretation of learned representations"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Deep-learning models can extract a rich assortment of features from data. Which features a model uses depends not only on *predictivity*---how reliably a feature indicates training-set labels---but also on *availability*---how easily the feature can be extracted from inputs. The literature on shortcut learning has noted examples in which models privilege one feature over another, for example texture over shape and image backgrounds over foreground objects. Here, we test hypotheses about which input properties are more available to a model, and systematically study how predictivity and availability interact to shape models' feature use. We construct a minimal, explicit generative framework for synthesizing classification datasets with two latent features that vary in predictivity and in factors we hypothesize to relate to availability, and we quantify a model's shortcut bias---its over-reliance on the shortcut (more available, less predictive) feature at the expense of the core (less available, more predictive) feature. We find that linear models are relatively unbiased, but introducing a single hidden layer with ReLU or Tanh units yields a bias. Our empirical findings are consistent with a theoretical account based on Neural Tangent Kernels. Finally, we study how models used in practice trade off predictivity and availability in naturalistic datasets, discovering availability manipulations which increase models' degree of shortcut bias. Taken together, these findings suggest that the propensity to learn shortcut features is a fundamental characteristic of deep nonlinear architectures warranting systematic study given its role in shaping how models solve tasks."},"_bibtex":{"value":"@inproceedings{\nhermann2024on,\ntitle={On the Foundations of Shortcut Learning},\nauthor={Katherine Hermann and Hossein Mobahi and Thomas FEL and Michael Curtis Mozer},\nbooktitle={The Twelfth International Conference on Learning Representations},\nyear={2024},\nurl={https://openreview.net/forum?id=Tj3xLVuE9f}\n}"},"title":{"value":"On the Foundations of Shortcut Learning"},"pdf":{"value":"/pdf/3f47b29f0e35691e7047d9fbfa0e4c47ea966e49.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference"},"paperhash":{"value":"hermann|on_the_foundations_of_shortcut_learning"},"authorids":{"value":["~Katherine_Hermann1","~Hossein_Mobahi2","~Thomas_FEL1","~Michael_Curtis_Mozer1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Katherine Hermann","Hossein Mobahi","Thomas FEL","Michael Curtis Mozer"]}},"tmdate":1710544298787,"pdate":1705410973678,"tcdate":1695418758228,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission6363/Authors"],"signatures":["ICLR.cc/2024/Conference/Submission6363/Authors"],"forum":"Tj3xLVuE9f","number":6363,"cdate":1695418758228,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/-/Submission","ICLR.cc/2024/Conference/-/Post_Submission","ICLR.cc/2024/Conference/Submission6363/-/Revision","ICLR.cc/2024/Conference/Submission6363/-/Rebuttal_Revision","ICLR.cc/2024/Conference/-/Edit","ICLR.cc/2024/Conference/Submission6363/-/Camera_Ready_Revision"],"mdate":1710544298787,"odate":1697213872796,"domain":"ICLR.cc/2024/Conference","id":"Tj3xLVuE9f","version":2},{"content":{"comment":{"value":"Reviewer 82h7: \"Better justify the practical importance of focusing on fully diacritized Arabic by discussing how frequently such text occurs in real-world corpora and evaluating performance on more realistic settings where Arabic text is predominantly undiacritized.\"\n\nReviewer PP8e: \"Is undiacritized Arabic more commonly used in the writing systems and electronic systems in the Arabic-culture world? ... Can (a) outperform (c) as well? Is this scenario a common case in Arabic processing?\"\n\nReviewer HvSw: \"...permute the order of the answer presentation to ensure that we can disentangle that the gains are attributed to the tokenizer (not its higher fertility, or answer prior).\"\n\nThese bear on the same question from different directions, so we answer them together and refer to this response from the individual threads. Because of the character limit on comments, the response is split into two parts. This is Part 1, covering the motivation and the experiment design. Part 2 follows immediately below and reports the scoring, results, and controls.\n\n- Our answer in short. We built a controlled test of 110 Arabic minimal pairs whose meaning is fixed only by the diacritics. On this test the morphology aware tokenizer lets the model select the correct reading, and the BPE tokenizer, trained on exactly the same diacritized text, does not. Averaged over eight paired checkpoints, NOHA reaches 0.525 accuracy [0.509, 0.543] where chance is exactly 0.500, while the BPE model trained on identical text reaches 0.494 [0.476, 0.510] and does not exceed chance. The paired per item contrast favours NOHA on both measures: +3.1 percentage points of accuracy [+0.7, +5.5] and +0.54 nats of displacement [+0.11, +1.00], the latter meaning the vowel marks shift NOHA's preference toward the correct continuation roughly 1.7 times more strongly than they shift BPE's. Both intervals exclude zero. The undiacritized models cannot answer these items at all, and not because they are weak: no entry in their vocabulary can encode a case vowel, so both readings of every item map to identical token identifiers and the outcome is exactly chance by construction. The systems differ only in tokenizer and in whether the marks are present during training, so the ordering is attributable to segmentation rather than to model capacity or training budget. The remainder of Part 1 and the whole of Part 2 explain the design and the controls behind those numbers.\n\n- Undiacritized text is the common case and we do not dispute it. Short vowel marks are omitted in most modern writing, and removing them is routine preprocessing in Arabic NLP. Our claim is not that diacritized text is frequent. It is that the omitted marks carry information the common case has discarded, and that this information is not recoverable from the consonantal string alone.\n\n- What the marks encode. In Arabic they can carry vocalic form, morphological structure, grammatical case, and syntactic role. Iʿrāb is the traditional analysis of how each word functions grammatically in a sentence, identifying roles such as subject, object, and predicate together with their inflectional case endings, and these endings are often written as vowel marks (Al-Sabahi et al., 2026, https://www.sciencedirect.com/science/article/pii/S2667096826000170; Mubarak et al., 2026, https://aclanthology.org/2026.eacl-long.296/). Arabic curricula teach grammar, morphology, and iʿrāb formally, and learners are taught to connect word form, case marking, grammatical function, and meaning.\n\n- Why MCQ accuracy does not settle this. AraLingBench finds that models can show strong surface level proficiency while struggling with deeper grammatical and syntactic reasoning (Zbib et al., 2026, https://aclanthology.org/2026.abjadnlp-1.45/), and Nahw reports substantial remaining gaps in Arabic grammar understanding (Mubarak et al., 2026, https://aclanthology.org/2026.eacl-long.296/). Broad accuracy does not establish that a model uses diacritic dependent information.\n\n- The experiment. We release the item set as TAAD, Teacher-Approved Arabic Diacritics (https://huggingface.co/datasets/ali-issa/TAAD). Each of the 110 items is one sentence written two ways, evaluated under both of its vocalized readings, giving 220 forced choices. The two prompts have the same consonantal string once diacritics are removed, so only the vocalization distinguishes the readings. For example أُحِبُّ الشِّعْرَ العَرَبِيَّ, I love Arabic poetry, and أُحِبُّ الشَّعْرَ الطَّوِيلَ, I love long hair, both reduce to أحب الشعر. Items are constrained so both continuations carry the same final vowel, which removes a shortcut in which a model succeeds by copying the case vowel of the preceding word. An earlier version of the design without this constraint was solved by that shortcut alone at every position, 238 of 238."},"title":{"value":"Arabic minimal pairs isolating the diacritic signal from fertility and answer prior - Part 1"}},"parentInvitations":"TMLR/-/Official_Comment","tmdate":1786825289036,"tcdate":1786825289036,"writers":["TMLR","TMLR/Paper9841/Authors"],"signatures":["TMLR/Paper9841/Authors"],"forum":"Vyc4Ormkqp","number":8,"license":"CC BY 4.0","cdate":1786825289036,"readers":["everyone"],"invitations":["TMLR/Paper9841/-/Official_Comment"],"mdate":1786825289036,"domain":"TMLR","replyto":"Vyc4Ormkqp","id":"vGbegOGSjV","forumContent":{"submission_length":{"value":"Long submission (more than 12 pages of main content)"},"venue":{"value":"Accepted by TMLR"},"abstract":{"value":"Arabic is a morphologically rich language whose complex phonological and morphological structure presents fundamental challenges for statistical tokenization methods used in large language model (LLM) pre-training. Although statistical methods such as Byte-Pair Encoding (BPE) are widely used, they are not optimally suited to morphologically rich languages such as Arabic, where frequency-based subword discovery ignores phonological and morphological boundaries and produces inconsistent segmentation of diacritized text. This paper presents NOHA, a fully deterministic diacritic-aware Arabic tokenizer that replaces statistical subword discovery with a rule-based chunking algorithm based on Arabic morphological and phonological structure. NOHA applies clitic preservation and diacritic-driven syllabic chunking to produce subword units that reflect the natural sound and grammatical structure of Arabic, without any statistical training. Despite producing approximately $1.49\\times$ more tokens per sequence than BPE (and thus observing fewer unique training documents under identical compute and hyperparameter conditions), a GPT-2-scale model trained with NOHA achieves higher average accuracy across 38 matched-compute evaluation checkpoints on all three targeted Arabic linguistic benchmarks: $+2.42\\%$ on GAT Reading Comprehension, $+2.65\\%$ on the three Arabic language subsets of ArabicMMLU, and $+2.54\\%$ on TAAG, a 171-question teacher-authored Arabic grammar diagnostic set written for this study by a certified Arabic language teacher. On general knowledge benchmarks, BPE leads by $0.26\\%$ on a weighted average across 18 ArabicMMLU subsets, indicating that NOHA preserves general knowledge capacity while improving linguistic understanding. BPE achieves lower bits-per-byte at all 38 matched-compute evaluation checkpoints, with a mean compression advantage of $+2.87\\%$. Both models perform near the $25\\%$ random baseline on the Nahw-MCQ Arabic grammar benchmark, the remaining GATLc subtasks (Verbal Analogy, Sentence Completion, Contextual Error, and Semantic Association), and most mathematical reasoning subtasks, confirming that model scale rather than tokenizer design is the limiting factor for these tasks. For mathematical reasoning subtasks, limited exposure to mathematical content in the training corpus is an additional contributing factor. These findings show that, at one model scale and one training seed, linguistically grounded tokenization yields measurable improvements on Arabic linguistic tasks at an identical computational budget, and that compression efficiency and linguistic understanding are dissociable properties of Arabic tokenizer design."},"_bibtex":{"value":"@article{\nissa2026noha,\ntitle={{NOHA}: A Diacritic-Aware Rule-Based Morphophonological Tokenizer for Arabic},\nauthor={Ali Issa and Tina Yaacoub and Patrick Abi Salloum},\njournal={Transactions on Machine Learning Research},\nissn={2835-8856},\nyear={2026},\nurl={https://openreview.net/forum?id=Vyc4Ormkqp},\nnote={}\n}"},"title":{"value":"NOHA: A Diacritic-Aware Rule-Based Morphophonological Tokenizer for Arabic"},"changes_since_last_submission":{"value":"Relative to the desk-rejected submission (9683), the abstract was reformatted into a single continuous block, addressing the Editors-in-Chief comment. This camera-ready version additionally implements the minor revisions requested by the Action Editor. The changes are wording only, with no changes to experiments or results. (1) The main conclusions are stated as established at one model scale and one training seed, in the abstract, the introduction, Section 6.1, the Model scale paragraph and a new Single training run limitation in Section 6.4, and the conclusion. (2) Section 6.2 no longer attributes the gains to a particular component of NOHA and presents its two properties as candidate explanations, since no component-level ablations were performed. (3) A new Tokenizer baselines limitation in Section 6.4 states that morphology-aware and Unigram-style baselines were not compared. (4) TAAG is described as a diagnostic set rather than a standalone benchmark contribution, its validation procedure is stated accurately in Section 3.2.3, and it is no longer listed as a separate contribution."},"pdf":{"value":"/pdf/fb33b3e7407a43b77fa6c1e4fddf208e7490aa0f.pdf"},"venueid":{"value":"TMLR"},"paperhash":{"value":"issa|noha_a_diacriticaware_rulebased_morphophonological_tokenizer_for_arabic"},"authorids":{"value":["~Ali_Issa2","~Tina_Yaacoub1","~Patrick_Abi_Salloum1"]},"previous_TMLR_submission_url":{"value":"https://openreview.net/forum?id=wq4UWPP6MB"},"assigned_action_editor":{"value":"~Sachin_Kumar1"},"authors":{"value":["Ali Issa","Tina Yaacoub","Patrick Abi Salloum"]}},"version":2},{"content":{"summary":{"value":"This paper focuses on the identity shortcut issue in multi-class unsupervised anomaly detection and proposes a feature-reconstruction method using two components: a low-rank noisy bottleneck (LRNB) and global perturbation attention (GPA). The method achieves strong results on four datasets."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"And Other Comments\n* Will sigmoid attention lead to issues such as gradient saturation and scale drift?\n* Could the authors provide visualizations of intermediate reconstruction before/after LRNB and GPA to validate the claimed mechanism? \n* Papers on this topic typically include anomaly score maps. Including such visualizations could partially address weakness 4.\n* Some figures are too small to read clearly.\n\nPlease note I did not consider this section in the rating. I was going to give 5 and was willing to increase to 6 at least if the authors could address some of my concerns. But since only options 4 and 6 are available, I will give 6 in advance."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"* This paper presents a clear motivation for addressing the identity shortcut issue. Two components are closely related to this motivation. The authors also provide a clear illustration of their specific method.\n* The promising results demonstrate the effectiveness of the proposed method. Additionally, the authors conducted a comprehensive ablation study to evaluate the contribution of each component and parameter."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"* Property 1 of LRNB is very similar to the idea of VAE. It only rules out global identity functions. It is not clear if there are local approximate identity mappings\n* It is not clear why \"noisy operations are often limited in effectively preventing shortcut learning in complex scenarios (lines 207-210).\" Perhaps the authors should consider comparing their approach with previous techniques for resolving identity shortcut issues, rather than merely presenting results or exclusively.\n* The reason for introducing noise in the LRNB decoder is not clarified.\n* All techniques, including bottleneck and noise in LRNB, sigmoid function and global mask in GAP, will compress/lose detailed information. Could this pose a problem when detecting very small anomalies, such as small defects? The paper does not analyze these side effects. It’s difficult to determine from the current results (both pixel-level and image-level), since all four datasets contain both large and small anomalies.\n* The connection between the two components is not discussed. For example, let's say \"the global constraint is derived from LRNB, while the local constraint is based on GPA\".\n* No discussion about computing resources and time."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762921922640,"tcdate":1762565446451,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission10673/Reviewer_muHh"],"signatures":["ICLR.cc/2026/Conference/Submission10673/Reviewer_muHh"],"forum":"HJIm7Ds3kF","number":4,"license":"CC BY 4.0","cdate":1762565446451,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission10673/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762921922640,"domain":"ICLR.cc/2026/Conference","replyto":"HJIm7Ds3kF","id":"qGoulW0psn","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["anomaly detection"]},"supplementary_material":{"value":"/attachment/51b3e322a35074936a627c206650c12f1a4aa51c.pdf"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Multi-class unsupervised anomaly detection (MUAD) has garnered growing research interest, as it seeks to develop a unified model for anomaly detection across multiple classes—eliminating the need to train separate models for distinct objects and thereby saving substantial computational resources.\n\tUnder the MUAD setting, while advanced Transformer-based architectures have brought significant performance improvements, identity shortcuts persist: they directly copy inputs to outputs, narrowing the gap in reconstruction errors between normal and abnormal cases, and thereby making the two harder to distinguish.\n\tTherefore, we propose ShortcutBreaker, a novel unified feature-reconstruction framework for MUAD tasks, featuring two key innovations to address the issue of shortcuts.\n\tFirst, drawing on matrix rank inequality, we design a low-rank noisy bottleneck (LRNB) to project high-dimensional features into a low-rank latent space, and theoretically demonstrate its capacity to prevent trivial identity reproduction.\n\tSecond, leveraging ViT’s global modeling capability instead of merely focusing on local features, we incorporate a global perturbation attention to prevent information shortcuts in the decoders.\n\tExtensive experiments are performed on four widely used anomaly detection benchmarks, including three industrial datasets (MVTec-AD, ViSA, and Real-IAD) and one medical dataset (Universal Medical). The proposed method achieves a remarkable image-level AUROC of 99.8%, 98.9%, 90.6%, and 87.8% on these four datasets, respectively, consistently outperforming previous MUAD methods across different scenarios."},"_bibtex":{"value":"@misc{\ntang2025shortcutbreaker,\ntitle={ShortcutBreaker: Low-Rank Noisy Bottleneck with Global Perturbation Attention for Multi-Class Unsupervised Anomaly Detection},\nauthor={Peng Tang and Xiaoxiao Yan and Xiaobin Hu and Yuning Cui and Donghao Luo and Jiangning Zhang and Pengcheng Xu and Jinlong Peng and Qingdong He and Feiyue Huang and Song Xue and Tobias Lasser},\nyear={2025},\nurl={https://openreview.net/forum?id=HJIm7Ds3kF}\n}"},"title":{"value":"ShortcutBreaker: Low-Rank Noisy Bottleneck with Global Perturbation Attention for Multi-Class Unsupervised Anomaly Detection"},"pdf":{"value":"/pdf/f7ab943e9cd1c121c17f46f68ee7d9dec7ed2813.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"tang|shortcutbreaker_lowrank_noisy_bottleneck_with_global_perturbation_attention_for_multiclass_unsupervised_anomaly_detection"},"authorids":{"value":["~Peng_Tang5","~Xiaoxiao_Yan1","~Xiaobin_Hu1","~Yuning_Cui1","~Donghao_Luo1","~Jiangning_Zhang1","~Pengcheng_Xu1","~Jinlong_Peng1","~Qingdong_He1","~Feiyue_Huang2","~Song_Xue6","~Tobias_Lasser1"]},"authors":{"value":["Peng Tang","Xiaoxiao Yan","Xiaobin Hu","Yuning Cui","Donghao Luo","Jiangning Zhang","Pengcheng Xu","Jinlong Peng","Qingdong He","Feiyue Huang","Song Xue","Tobias Lasser"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a controllable first-frame-guided video editing framework that adapts a pretrained image-to-video diffusion model using mask-aware LoRA. The central idea is to repurpose the I2V model’s spatiotemporal mask not just as an inference constraint but also as a learning signal that tells LoRA what to learn. The method (i) performs per-video LoRA tuning so the model learns the input video’s motion pattern while preserving unedited regions via a mask; and (ii) optionally performs a second, brief LoRA pass where edited reference frames at later timestamps teach the model the desired appearance evolution inside masked regions (e.g., a flower gradually becoming a red rose). Experiments on first-frame-guided and reference-guided editing show better results compared with AnyV2V, Go-with-the-Flow, and I2VEdit."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"See weaknesses"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"- **Simple and efficient idea**\nUsing a spatiotemporal mask as an explicit LoRA supervision signal is a clean and reusable concept that doesn’t require architectural changes. This is likely to inspire follow-up work on using conditioning signals to steer what LoRA learns. This strategy addresses two persistent failure modes in first-frame editing: (i) unwanted background drift/leakage and (ii) insufficient control of how the edited object looks as it rotates/deforms or reveals disocclusions.\n\n- **Low engineering overhead**\nPer-video LoRA is a common practice. Adding mask-aware conditioning with edited frames is easy to deploy and works across different I2V-based models.\n\n- Qualitative results look compelling on diverse manipulations (object add/replace/style, clothing/hair edits), with visibly better background preservation and temporal coherence than the chosen baselines."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- **Incremental Novelty**\nPer-video LoRA for motion propagation and mask-conditioned inpainting are known practices in video editing and inpainting. The novelty is mainly in using the mask to steer LoRA’s learning objective across space-time. This is elegant but not a large algorithmic leap.\n\n- **Mask acquisition and robustness are under-specified**\nIt is not clear how masks are obtained for training and inference. Is it automatically calculated from the first-frame edit difference, and then propagated with optical flow/segmentation? How to handle later edited references?\nRobustness to mask errors (boundary misalignment, sparsity, temporal jitter) is not evaluated. Since the method’s key promise is precise, region-specific control, an analysis with sensitivity to mask quality is critical.\n\nThe paper tackles a high-impact practical problem (controllable propagation of first-frame edits) with a minimal, effective mechanism. The idea is easy to adopt on top of modern I2V models, and qualitative outcomes are persuasive. While the novelty is incremental, the contribution is useful and timely for creators and researchers. I tend to give a positive score."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762918436773,"tcdate":1762002804404,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission6060/Reviewer_eyka"],"signatures":["ICLR.cc/2026/Conference/Submission6060/Reviewer_eyka"],"forum":"xkRMJ1Y7Um","number":2,"license":"CC BY 4.0","cdate":1762002804404,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission6060/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762918436773,"domain":"ICLR.cc/2026/Conference","replyto":"xkRMJ1Y7Um","id":"v2oPC2Y8cI","forumContent":{"TLDR":{"value":"The paper introduces a mask-based LoRA tuning method for highly flexible video editing using the pre-trained Image-to-Video model."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Video Editing"]},"supplementary_material":{"value":"/attachment/034635334583436f67f6ff70e86f40fbc1976fcf.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Video editing using diffusion models has achieved remarkable results in generating high-quality edits for videos. However, current methods often rely on large-scale pretraining, limiting flexibility for specific edits. First-frame-guided editing provides control over the first frame, but lacks fine-grained control over the edit's subsequent temporal evolution. To address this, we propose a mask-based LoRA (Low-Rank Adaptation) tuning method that adapts pretrained Image-to-Video models for flexible video editing.\nOur key innovation is using a spatiotemporal mask to strategically guide the LoRA fine-tuning process. This teaches the model two distinct skills: first, to interpret the mask as a command to either preserve content from the source video or generate new content in designated regions. Second, for these generated regions, LoRA learns to synthesize either temporally consistent motion inherited from the video or novel appearances guided by user-provided reference frames.\nThis dual-capability LoRA grants users control over the edit's entire temporal evolution, allowing complex transformations like an object rotating or a flower blooming. Experimental results show our method achieves superior video editing performance compared to baseline methods. The code and video results are available at our project website: https://cjeen.github.io/LoRAEdit."},"_bibtex":{"value":"@inproceedings{\ngao2026controllable,\ntitle={Controllable First-Frame-Guided Video Editing via Mask-Aware Lo{RA} Fine-Tuning},\nauthor={Chenjian Gao and Lihe Ding and Xin Cai and Zhanpeng Huang and Zibin Wang and Tianfan Xue},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=xkRMJ1Y7Um}\n}"},"title":{"value":"Controllable First-Frame-Guided Video Editing via Mask-Aware LoRA Fine-Tuning"},"pdf":{"value":"/pdf/a7fc2c198d3de6a267b714b5b442ae140e592d4c.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"gao|controllable_firstframeguided_video_editing_via_maskaware_lora_finetuning"},"authorids":{"value":["~Chenjian_Gao1","~Lihe_Ding1","~Xin_Cai2","~Zhanpeng_Huang1","~Zibin_Wang1","~Tianfan_Xue2"]},"authors":{"value":["Chenjian Gao","Lihe Ding","Xin Cai","Zhanpeng Huang","Zibin Wang","Tianfan Xue"]}},"version":2},{"content":{"summary":{"value":"This paper reassess the Knowledge Neuron Thesis in two ways, with syntatic minimal pairs and with generalizing to bijective relationships and synonyms. Theranalysis shows the limitation of current knowledge identification and editing, suggesting the need for more sophisticated understanding of inner mechanism of a language model."},"presentation":{"value":"3 good"},"contribution":{"value":"4 excellent"},"soundness":{"value":"3 good"},"strengths":{"value":"1. This paper introduces many new practices for the rigorous study of knowledge neuron thesis, including using minimal pairs and t-test.\n2. Broadening the definition of knowledge neural and connecting it to prior works in linguist phenomena.\n3. Through and diverse analysis."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. If I have to nitpick, section 4 felt a bit disjoint from the rest of the paper and is not fully fledged."},"confidence":{"value":"4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work."},"questions":{"value":"n/a"},"rating":{"value":"8: accept, good paper"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636143088,"tcdate":1698855403865,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission2106/Reviewer_c8FP"],"signatures":["ICLR.cc/2024/Conference/Submission2106/Reviewer_c8FP"],"forum":"2HJRwwbV3G","number":2,"license":"CC BY 4.0","cdate":1698855403865,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission2106/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636143088,"domain":"ICLR.cc/2024/Conference","replyto":"2HJRwwbV3G","id":"Lshau2x4wY","forumContent":{"venue":{"value":"ICLR 2024 spotlight"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["language model","knowledge neuron","model editing","formal and function competence","syntax","fact"]},"primary_area":{"value":"representation learning for computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"We reassess the Knowledge Neuron (KN) Thesis: an interpretation of the mechanism underlying the ability of large language models to recall facts from a training corpus. This nascent thesis proposes that facts are recalled from the training corpus through the MLP weights in a manner resembling key-value memory, implying in effect that \"knowledge\" is stored in the network. Furthermore, by modifying the MLP modules, one can control the language model's generation of factual information. The plausibility of the KN thesis has been demonstrated by the success of KN-inspired model editing methods (Dai et al., 2022; Meng et al., 2022).\n\nWe find that this thesis is, at best, an oversimplification. Not only have we found that we can edit the expression of certain linguistic phenomena using the same model editing methods but, through a more comprehensive evaluation, we have found that the KN thesis does not adequately explain the process of factual expression. While it is possible to argue that the MLP weights store complex patterns that are interpretable both syntactically and semantically, these patterns do not constitute \"knowledge.\" To gain a more comprehensive understanding of the knowledge representation process, we must look beyond the MLP weights and explore recent models' complex layer structures  and attention mechanisms."},"_bibtex":{"value":"@inproceedings{\nniu2024what,\ntitle={What does the Knowledge Neuron Thesis Have to do with Knowledge?},\nauthor={Jingcheng Niu and Andrew Liu and Zining Zhu and Gerald Penn},\nbooktitle={The Twelfth International Conference on Learning Representations},\nyear={2024},\nurl={https://openreview.net/forum?id=2HJRwwbV3G}\n}"},"title":{"value":"What does the Knowledge Neuron Thesis Have to do with Knowledge?"},"pdf":{"value":"/pdf/04389528ee4dfce215a52f1dd0987190ad2adc1d.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference"},"paperhash":{"value":"niu|what_does_the_knowledge_neuron_thesis_have_to_do_with_knowledge"},"authorids":{"value":["~Jingcheng_Niu1","~Andrew_Liu6","~Zining_Zhu1","~Gerald_Penn1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Jingcheng Niu","Andrew Liu","Zining Zhu","Gerald Penn"]}},"version":2},{"content":{"summary":{"value":"This paper studies multimodal retrieval with MLLM unified encoders and identifies a modality shortcut issue, where models over-rely on one modality instead of jointly leveraging both, leading to reduced robustness under distribution shifts. To mitigate this, the authors introduce a modality composition awareness framework, combining a preference loss to favor multimodal over unimodal representations and a composition regularization to align multimodal features with prototypes composed from unimodal parts. Experiments on several OOD retrieval benchmarks show consistent gains. Overall, the paper addresses a relevant problem with a simple and intuitive solution, but the experimental scope is limited, particularly in baseline selection, dataset diversity, and clarity of the shortcut analysis."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"1. Can you provide a clearer formal definition and more qualitative evidence (e.g., retrieval examples showing failure due to shortcut)? Could you compare the phenomenon with known shortcut learning literature?\n\n2. What is the sensitivity to α and β? Could you provide empirical results or theoretical reasoning for their chosen values?\n\n3. Why not compare against stronger contemporary retrieval backbones (e.g., BLIP-2/Flamingo-style encoders, EVA-CLIP, SigLip)? Do your findings still hold when the base MLLM has already undergone multimodal alignment training?\n\n4. Can the method apply to tasks beyond retrieval (e.g., VQA, caption-grounded generation)? Does the gain still exist on in-distribution benchmarks?\n\n5. Could you show ablation on unimodal vs multimodal dominance (e.g., mask image vs mask text cases)?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. Addresses modality shortcut behavior in unified-encoder MLLMs, a relevant issue as unified retrieval becomes more common.\n\n2. The proposed preference and composition regularization losses are easy to implement and computationally lightweight.\n\n3. Shows consistent OOD retrieval improvements and basic ablations confirming each loss contributes."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The concept is compelling but currently presented somewhat vaguely. The illustrative figure does not intuitively convey shortcut behavior; more concrete failure case analysis would strengthen the argument.\n\n2. Baselines are relatively old (mostly CLIP-style models). Missing comparisons with recent multimodal LLMs and retrieval-tuning frameworks would improve credibility.\n\n3. Manual selection without principled justification or tuning range exploration weakens the training claim. Lack of theoretical or empirical guidance.\n\n4. Although improvements are shown, the number of OOD scenarios and datasets is limited, weakening the robustness claim.\n\n5. Strengthening the narrative, why unified encoders for retrieval matter long-term, would increase appeal."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762922965122,"tcdate":1762094511164,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission11968/Reviewer_o4QY"],"signatures":["ICLR.cc/2026/Conference/Submission11968/Reviewer_o4QY"],"forum":"EzBRI0Llxk","number":3,"license":"CC BY 4.0","cdate":1762094511164,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission11968/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762922965122,"domain":"ICLR.cc/2026/Conference","replyto":"EzBRI0Llxk","id":"35aNMoitkw","forumContent":{"TLDR":{"value":"Mitigating the modality shortcut problem in MLLM-based multimodal retrieval with modal composition awareness."},"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Multimodal Retrieval","Modality Shortcut","MLLMs","Modality Composition"]},"primary_area":{"value":"unsupervised, self-supervised, semi-supervised, and supervised representation learning"},"abstract":{"value":"Multimodal retrieval, which seeks to retrieve relevant content across modalities such as text or image, supports applications from AI search to contents production. Despite the success of separate-encoder approaches like CLIP align modality-specific embeddings with contrastive learning, recent multimodal large language models (MLLMs) enable a unified encoder that directly processes composed inputs. While flexible and advanced, we identify that unified encoders trained with conventional contrastive learning are prone to learn modality shortcut, leading to poor robustness under distribution shifts. We propose a modality composition awareness framework to mitigate this issue. Concretely, a preference loss enforces multimodal embeddings to outperform their unimodal counterparts, while a composition regularization objective aligns multimodal embeddings with prototypes composed from its unimodal parts. These objectives explicitly model structural relationships between the composed representation and its unimodal counterparts. Experiments on various benchmarks show gains in out-of-distribution retrieval, highlighting modality composition awareness as a effective principle for robust composed multimodal retrieval when utilizing MLLMs as the unified encoder."},"_bibtex":{"value":"@misc{\nwu2025mca,\ntitle={{MCA}: Modality Composition Awareness for Robust Composed Multimodal Retrieval},\nauthor={Qiyu Wu and Shuyang Cui and Satoshi Hayakawa and Wei-Yao Wang and Hiromi Wakaki and Yuki Mitsufuji},\nyear={2025},\nurl={https://openreview.net/forum?id=EzBRI0Llxk}\n}"},"title":{"value":"MCA: Modality Composition Awareness for Robust Composed Multimodal Retrieval"},"pdf":{"value":"/pdf/2c99bf21eae5f0db226f555f11aba423303addf3.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"wu|mca_modality_composition_awareness_for_robust_composed_multimodal_retrieval"},"authorids":{"value":["~Qiyu_Wu2","~Shuyang_Cui1","~Satoshi_Hayakawa1","~Wei-Yao_Wang1","~Hiromi_Wakaki1","~Yuki_Mitsufuji1"]},"authors":{"value":["Qiyu Wu","Shuyang Cui","Satoshi Hayakawa","Wei-Yao Wang","Hiromi Wakaki","Yuki Mitsufuji"]}},"version":2},{"content":{"summary":{"value":"This paper presents LongVTG-R1, a framework applying Reinforcement Learning to enhance the long-video temporal grounding capabilities of Multimodal Large Language Models. The authors identify that standard Supervised Fine-Tuning can lead to catastrophic forgetting  and propose RL as a better alternative. The paper introduces three main technical contributions to make RL viable for this task: (1) Token-aware KL Regularization to balance exploration (for timestamps) and exploitation (for general language); (2) a Center Distance Reward to create a denser reward signal and alleviate the sparsity of the standard IoU metric ; and (3) a new dataset, SceneTG, constructed automatically using scene boundaries and a uniqueness filter to provide clear and unambiguous training signals. The authors show that their model achieves state-of-the-art results on several zero-shot VTG benchmarks and claim this skill also generalizes to improve general-purpose long-video Question-Answering."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. The paper claims \"consistent gains\" on general QA tasks , but the results in Table 3 appear to contradict this, showing that LongVTG-R1 (58.7) actually underperforms the baseline Qwen2.5VL (60.5) on LongVideoBench. Could the authors please address this discrepancy and clarify why the model fails to improve on this specific benchmark? Does this not suggest that the claimed generalization to direct QA is more limited than stated?\n\n2. The SceneTG dataset is presented as a core contribution, yet the paper lacks crucial statistics for reproducibility and assessment. Could the authors please provide the final size of the dataset (number of QA pairs and total video hours), the acceptance/rejection rate of their uniqueness filter , and clarify whether the \"40k\" training data mentioned  refers to the SceneTG dataset alone or the total mixture of all three datasets?\n\n3. The proposed Center Distance Reward (CenDist) addresses the issue of reward sparsity from IoU, which is a well-studied problem. Could the authors justify why they chose to design this new reward function  instead of adapting existing, established dense regression losses (like GIoU or DIoU), which were created to solve the exact same problem of zero gradients for non-overlapping predictions?\n\n4. For the Token-aware KL regularization, the author make the strong design choice of setting the penalty for time tokens, $\\beta_{time}$, to exactly zero. The ablation in Table 6 only compares this against the baseline, not against other small, non-zero values for $\\beta_{time}$. Could the authors provide justification for this specific choice, as setting the penalty to zero could theoretically risk divergence and seems to warrant a more thorough ablation?"},"rating":{"value":4},"details_of_ethics_concerns":{"value":"None."},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. Relevant Problem and Sound Motivation: The paper tackles the challenging and important problem of long-video temporal grounding. The motivation to use RL as an alternative to SFT to prevent catastrophic forgetting is sound and follows a logical trend from recent works on shorter video grounding.\n\n2. Clear Identification of Challenges: The authors correctly identify the key bottlenecks in applying RL to this domain: the exploration-exploitation trade-off , reward sparsity , and data ambiguity. The paper is well-structured around proposing a specific solution for each of these problems.\n\n3. Intuitive Technical Solutions: The proposed contributions are clever and intuitive. The Token-aware KL regularization, in particular, is a task-specific and sensible way to manage the KL penalty by partitioning the vocabulary. The \"propose-then-annotate\" paradigm for the SceneTG dataset is also a practical approach to generating higher-quality data.\n\n4. Strong Empirical Grounding Results: The paper demonstrates strong performance on its primary task: temporal grounding. As shown in Table 2, LongVTG-R1 achieves state-of-the-art (SOTA) results across a suite of six zero-shot benchmarks. Achieving this with a smaller dataset (40k) than some competitors highlights the data efficiency of the approach."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Overstated Generalization to QA: A major weakness is the overstatement of the paper's \"most surprising\" finding—generalization to direct long-video QA. This claim is not well-supported by the results in Table 3. The model actually performs worse than the Qwen2.5VL baseline on LongVideoBench (58.7 vs 60.5) and is only comparable on MLVU (69.1 vs 68.6). The claim of \"consistent gains\"  is therefore inaccurate. The strong improvements are only clearly demonstrated in the \"ground-then-answer\" pipeline (Table 4). This is a different and less significant finding, as it's expected that providing a QA model with the correct pre-localized clip would improve its performance. This discrepancy undermines one of the paper's central narratives.\n\n2. SceneTG Dataset is Underspecified: The SceneTG dataset is presented as a key contribution, but it is poorly described. The paper details the method of construction (Sec 3.4)  but provides no statistics on the resulting dataset. Critical details are missing, such as the final number of (video, QA) pairs, the total hours of video, the source of the 20k Prime Videos, or the rejection rate of the uniqueness filter. The \"40k\" data mentioned  is ambiguous. This lack of detail makes the contribution hard to evaluate and the work difficult to reproduce.\n\n3. Incremental Novelty and Insufficient Ablation:\n- The overall design of the proposed approach is simple and straightforward; however, no notably impressive performance improvement has been observed compared to previous methods.\n- The overall framework is a careful adaptation of existing methods (GRPO , and concepts from models like VideoChat-R1 ) to the long-video domain. The novelty is more incremental than foundational.\n- The technical contributions, while intuitive, lack thorough validation. For CenDist, the paper does not compare it against the large body of existing work on dense regression losses (e.g., GIoU, DIoU)  that were designed to solve the exact same problem (zero gradients for non-overlapping predictions).\n- For Token-aware KL, the choice of $\\beta_{time}=0$ is an extreme value. An ablation study comparing this to other small, non-zero values would have been necessary to fully justify this design choice."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762925945287,"tcdate":1761567411432,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission15694/Reviewer_oC5c"],"signatures":["ICLR.cc/2026/Conference/Submission15694/Reviewer_oC5c"],"forum":"8H1HmGH8ua","number":1,"license":"CC BY 4.0","cdate":1761567411432,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission15694/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762925945287,"domain":"ICLR.cc/2026/Conference","replyto":"8H1HmGH8ua","id":"GPynosR2Fg","forumContent":{"TLDR":{"value":"We present LongVTG-R1, the first RL-based framework for long-video temporal grounding, which leverages novel regularization, reward design, and dataset construction to achieve state-of-the-art performance and generalize to QA tasks."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["video understanding","temporal grounding","multimodal large language model (MLLM)"]},"supplementary_material":{"value":"/attachment/a7526f2447a3b01b5d091f2498a198a2aae0ee49.zip"},"primary_area":{"value":"foundation or frontier models, including LLMs"},"abstract":{"value":"We present the first Reinforcement Learning (RL)-based framework that equips Multimodal Large Language Models (MLLMs) with long-video temporal grounding skills, and demonstrate that this approach also generalizes to improve performance on general question-answering (QA) tasks. Unlike dominant supervised fine-tuning (SFT) methods, RL enables models to acquire temporal grounding abilities without risking catastrophic forgetting of their core understanding. However, adopting RL for long-video temporal grounding reveals a challenge in balancing exploitation of pre-trained knowledge with exploration of new localization skills. To address this, we propose Token-aware KL Regularization, which selectively relaxes the KL-divergence regularization on timestamp-related tokens to guide exploration. Moreover, effective optimization requires a learning signal that alleviates the sparsity of key events in long videos, for which we introduce a denser reward, the Center Distance Reward (CenDist). To further mitigate grounding ambiguity between language queries and visually similar content, and to facilitate effective RL training, we propose an automatic data construction method and construct a small but high-quality dataset, SceneTG. Our resulting model, QwenLongTG, delivers substantial improvements across three long-video temporal grounding datasets among efficiently fine-tuned MLLMs, and even approaches the performance of densely pre-trained or continually trained models. Beyond temporal grounding, we further verify its generalization to long-video QA: under a “Ground-then-Answer” strategy, QwenLongTG consistently enhances downstream QA performance, serving as an effective first-stage grounding module."},"_bibtex":{"value":"@misc{\nzhang2026longvtgr,\ntitle={Long{VTG}-R1: Reinforcement Learning for Robust Long-Video Temporal Grounding},\nauthor={Zheyu Aqa Zhang and Shixing Chen and Ziqi Pang and Xiang Hao and Kushan Thakkar and Yu-Xiong Wang},\nyear={2026},\nurl={https://openreview.net/forum?id=8H1HmGH8ua}\n}"},"title":{"value":"LongVTG-R1: Reinforcement Learning for Robust Long-Video Temporal Grounding"},"pdf":{"value":"/pdf/1a44d3a0d532c7e585afb40db0defe29d293c3cd.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"zhang|longvtgr1_reinforcement_learning_for_robust_longvideo_temporal_grounding"},"authorids":{"value":["~Zheyu_Aqa_Zhang1","~Shixing_Chen1","~Ziqi_Pang1","~Xiang_Hao1","~Kushan_Thakkar1","~Yu-Xiong_Wang1"]},"authors":{"value":["Zheyu Aqa Zhang","Shixing Chen","Ziqi Pang","Xiang Hao","Kushan Thakkar","Yu-Xiong Wang"]}},"version":2},{"content":{"venue":{"value":"ECCV (25) 2026"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-032-37461-5_9.pdf"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"huang|no_place_to_hide_benchmarking_video_hallucination_with_backgroundcontrolled_pairs"},"html":{"value":"https://doi.org/10.1007/978-3-032-37461-5_9"},"_bibtex":{"value":"@inproceedings{DBLP:conf/eccv/HuangCLDXCHLC26,\n  author={Haojian Huang and Harold Haodong Chen and Meng Luo and Junjia Du and Shanqing Xu and Ziheng Chen and Yanxiang Huang and Yinchuan Li and Ying-Cong Chen},\n  title={No Place to Hide: Benchmarking Video Hallucination with Background-Controlled Pairs},\n  year={2026},\n  cdate={1767225600000},\n  pages={155-175},\n  url={https://doi.org/10.1007/978-3-032-37461-5_9},\n  booktitle={ECCV (25)},\n  crossref={conf/eccv/2026-25}\n}\n"},"abstract":{"value":"We introduce VidPair-Halluc, a new benchmark for evaluating video hallucination in large video models (LVMs) under rigorous and controlled conditions. Unlike previous benchmarks that primarily rely on text-based perturbations or adversarial questions while neglecting the consistency of visual backgrounds, VidPair-Halluc features video pairs with highly similar backgrounds but distinctly different foreground semantics, enabling precise attribution of model errors to genuine hallucination rather than background variation. The benchmark is constructed through PairFlow, a pipeline that leverages recent advances in text-to-image and video generation to systematically compose stories, generate coherent video clips, and assemble them into adversarial pairs. Covering both spatial and temporal reasoning across ten semantic aspects, VidPair-Halluc comprises 1K high-quality adversarial video pairs and 11K spatio-temporal QA pairs with control over background and foreground variations. Evaluations on mainstream LVMs show persistent difficulty with robust fine-grained video understanding in adversarial settings, and code and data are available at the project page."},"title":{"value":"No Place to Hide: Benchmarking Video Hallucination with Background-Controlled Pairs"},"authors":{"value":[{"fullname":"Haojian Huang","username":""},{"fullname":"Harold Haodong Chen","username":"~Haodong_Chen2"},{"fullname":"Meng Luo","username":""},{"fullname":"Junjia Du","username":""},{"fullname":"Shanqing Xu","username":""},{"fullname":"Ziheng Chen","username":""},{"fullname":"Yanxiang Huang","username":""},{"fullname":"Yinchuan Li","username":""},{"fullname":"Ying-Cong Chen","username":""}]}},"tmdate":1791703935100,"pdate":1798675200000,"externalIds":["dblp:conf/eccv/HuangCLDXCHLC26"],"tcdate":1791703932036,"writers":["~","OpenReview.net/Public_Article/DBLP.org","OpenReview.net/Support"],"signatures":["~Haodong_Chen2"],"forum":"16kEKXyjIF","license":"CC BY-SA 4.0","number":170145,"cdate":1767225600000,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/DBLP.org/-/Record","OpenReview.net/Public_Article/-/Edit"],"mdate":1791703935100,"domain":"OpenReview.net/Public_Article","id":"16kEKXyjIF","version":2},{"content":{"venue":{"value":"CoRR 2026"},"pdf":{"value":"https://arxiv.org/pdf/2606.31933v1"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"huang|no_place_to_hide_benchmarking_video_hallucination_with_backgroundcontrolled_pairs"},"html":{"value":"https://doi.org/10.48550/arXiv.2606.31933"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2606-31933,\n  publtype={informal},\n  author={Haojian Huang and Harold Haodong Chen and Meng Luo and Junjia Du and Shanqing Xu and Ziheng Chen and Yanxiang Huang and Yinchuan Li and Ying-Cong Chen},\n  title={No Place to Hide: Benchmarking Video Hallucination with Background-Controlled Pairs},\n  year={2026},\n  month={June},\n  cdate={1780272000000},\n  journal={CoRR},\n  volume={abs/2606.31933},\n  url={https://doi.org/10.48550/arXiv.2606.31933}\n}\n"},"abstract":{"value":"We introduce VidPair-Halluc, a new benchmark for evaluating video hallucination in large video models (LVMs) under rigorous and controlled conditions. Unlike previous benchmarks that primarily rely on text-based perturbations or adversarial questions while neglecting the consistency of visual backgrounds, VidPair-Halluc features video pairs with highly similar backgrounds but distinctly different foreground semantics, enabling precise attribution of model errors to genuine hallucination rather than background variation. The benchmark is constructed through PairFlow, a pipeline that leverages recent advances in text-to-image and video generation to systematically compose stories, generate coherent video clips, and assemble them into adversarial pairs. Covering both spatial and temporal reasoning across ten semantic aspects, VidPair-Halluc comprises 1K high-quality adversarial video pairs and 11K spatio-temporal QA pairs with control over background and foreground variations. Evaluations on mainstream LVMs show persistent difficulty with robust fine-grained video understanding in adversarial settings, and code and data are available at the https://jethrojames.github.io/VidPair-Halluc/."},"title":{"value":"No Place to Hide: Benchmarking Video Hallucination with Background-Controlled Pairs"},"authors":{"value":[{"fullname":"Haojian Huang","username":"~Haojian_Huang1"},{"fullname":"Harold Haodong Chen","username":""},{"fullname":"Meng Luo","username":""},{"fullname":"Junjia Du","username":""},{"fullname":"Shanqing Xu","username":""},{"fullname":"Ziheng Chen","username":""},{"fullname":"Yanxiang Huang","username":""},{"fullname":"Yinchuan Li","username":""},{"fullname":"Ying-Cong Chen","username":""}]}},"tmdate":1784163949000,"pdate":1798675200000,"externalIds":["dblp:journals/corr/abs-2606-31933"],"tcdate":1784136346365,"writers":["~","OpenReview.net/Public_Article/DBLP.org","OpenReview.net/Support"],"signatures":["~Haojian_Huang1"],"forum":"X2VcQjtNno","license":"CC BY-SA 4.0","number":57914,"cdate":1780272000000,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/DBLP.org/-/Record","OpenReview.net/Public_Article/-/Edit","OpenReview.net/Public_Article/-/Author_Removal","OpenReview.net/Public_Article/-/Authorship_Claim"],"mdate":1784163949000,"domain":"OpenReview.net/Public_Article","id":"X2VcQjtNno","version":2},{"content":{"summary":{"value":"This paper presents QueryStream, a framework for streaming video understanding that builds query-aware temporal representations to improve performance. The method dynamically fuses incoming frame features with past context, guided by query-conditioned similarity maps, allowing the model to emphasize temporally relevant segments as the video unfolds. The paper reports improvements on multiple benchmarks (streming + long video benchmarks) compared with existing streaming or offline VLMs."},"soundness":{"value":2},"confidence":{"value":5},"questions":{"value":"- Why was OpenCLIP chosen as the encoder instead of more recent options like SigLIP that are already aligned with VLMs?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":1},"strengths":{"value":"- The introduction is very well written. It clearly defines the gap between offline video-language models and real-time streaming setups, motivating the need for a query-aware temporal design.\n\n- The method is well organized, with a clear hierarchy between query-aware temporal modules, memory updates, and frame-level fusion."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The method is presented as a new query-aware temporal representation, but in practice it relies on several manually tuned parameters and heuristic weighting functions between query and frame embeddings. The current mothod design appears quite sensitive to these hyperparameters even though they have discussed in A.4. Despite the hierarchical structure, the approach still feels heuristic and lacks a clear principle.\n\n- The proposed framework mainly reorganizes existing components such as query-conditioned similarity and temporal fusion into a clean, well-structured pipeline. In the end, it looks more like a refined arrangement of standard similarity-based scoring rather than a fundamentally new approach, so the novelty feels limited.\n\n- The paper emphasizes low-latency and lightweight design, but there is no analysis of runtime, throughput, or resource usage. In streaming scenarios, maintaining query-aware similarity maps and recurrent updates for each frame can be computationally heavy. Without latency or efficiency analysis, it is unclear whether the method is actually suitable for real-time use."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762925778502,"tcdate":1761993647246,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission15489/Reviewer_6ZdJ"],"signatures":["ICLR.cc/2026/Conference/Submission15489/Reviewer_6ZdJ"],"forum":"738HjJEbml","number":4,"license":"CC BY 4.0","cdate":1761993647246,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission15489/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762925778502,"domain":"ICLR.cc/2026/Conference","replyto":"738HjJEbml","id":"7A5ecxnc9V","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Multimodal Large Language Model","Online Video Understanding","Streaming Video Understanding"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"The increasing demand for real-time interaction in online video scenarios necessitates a new class of efficient streaming video understanding models. However, existing approaches often rely on a query-agnostic ''change-is-important'' assumption, which conflates visual dynamics with semantic relevance, leading to computational redundancy and mistimed responses. To address this, we propose QueryStream, a novel framework that integrates query-awareness into the core of video processing and response scheduling. QueryStream features two synergistic components: (1) Query-Aware Differential Pruning (QDP), a policy that filters the token stream by jointly assessing semantic relevance to the query and temporal novelty against a dynamically smoothed history; and (2) Relevance-Triggered Active Response (RTAR), a dual-gated mechanism that schedules responses based on both high query relevance and significant information density. As a lightweight, training-free module, QueryStream achieves state-of-the-art performance on benchmarks such as StreamingBench and OVO-Bench under moderate pruning, and matches full-token baselines while pruning over 70\\% of visual tokens. Notably, our pruning mechanism generalizes to offline tasks, where it serves as a context-denoising module that benefits long-form video understanding. This work not only reveals the vast semantic redundancy in video streams relative to user intent but also establishes a promising, intent-driven direction for efficient and robust online video understanding. Code is available at: https://github.com/Zhangkr2003/QueryStream."},"_bibtex":{"value":"@inproceedings{\nzhang2026querystream,\ntitle={QueryStream: Advancing Streaming Video Understanding with Query-Aware Pruning and Proactive Response},\nauthor={Kairui Zhang and Zhenyu Yang and Bing Wang and Shengsheng Qian and Changsheng Xu},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=738HjJEbml}\n}"},"title":{"value":"QueryStream: Advancing Streaming Video Understanding with Query-Aware Pruning and Proactive Response"},"pdf":{"value":"/pdf/cd7eb509371a7f036f82fb9c86406725b54a1cbb.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"zhang|querystream_advancing_streaming_video_understanding_with_queryaware_pruning_and_proactive_response"},"authorids":{"value":["~Kairui_Zhang1","~Zhenyu_Yang6","~Bing_Wang20","~Shengsheng_Qian1","~Changsheng_Xu1"]},"authors":{"value":["Kairui Zhang","Zhenyu Yang","Bing Wang","Shengsheng Qian","Changsheng Xu"]}},"version":2},{"content":{"summary":{"value":"The authors present a novel annotated dataset of EEG-video pairs and an approach to reconstruct videos from EEG brain activity data. The dataset contains brain responses of 20 subjects watching 2-s videos from 40 general concepts. A total of 7 classification tasks are built based on the metadata available in the dataset (e.g. finegrained concept, coarse concept, color, etc.) and a video diffusion pipeline is trained to reconstruct videos. Classification results on the different tasks are presented, along with generated video frames and image quality metrics."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1. What is the architecture and hyperparameters for the EEG encoder (Section 4.2)?\n2. In Section 4.1: What is meant by “treating all channels equally”? Most deep learning encoders trained on EEG data have some kind of spatial processing layer, e.g. a convolutional layer, that learns to reweigh different channels end-to-end [1].\n3. The 40-class classification performance appears very low, however generations are qualitatively very good. Can you describe the process for selecting the generations shown in the paper?\n4. In Table 2, what does 40-way classification refer to when fewer than 40 classes are used?\n\n[1] Schirrmeister, Robin Tibor, et al. \"Deep learning with convolutional neural networks for EEG decoding and visualization.\" Human brain mapping 38.11 (2017): 5391-5420."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"* Originality: the study of video decoding from EEG has not been the focus of much attention yet, and the presentation of a new dataset is both original and useful to the community.\n* Clarity: the manuscript is overall clearly written and the general approach is well motivated.\n* Quality: interesting analysis of what information can be decoded (color, optical flow, object number, human face, human) through the different classification tasks described in Section 3.5.\n* Significance: the presented dataset and analysis set the stage for more generalizable results in visual decoding from EEG."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The dataset contains videos spanning 40 concepts which are seen in both training and test sets as described in Section B.2, i.e. there is \"categorical leakage\" between the two sets. This makes it very likely that the model learns to mostly (or solely) predict a concept, rather than predict the finer grained visual information contained in a video. Following this hypothesis, the EEG encoder, the seq2seq model and the semantic predictor could all be replaced by a single classifier that outputs the label of one of 40 concepts, followed by a lookup table that returns the corresponding concept-specific conditioning vector to be fed to the video diffusion model, and generation performance would remain similar. To test whether that is indeed the case, it would be interesting to train the model on e.g. 25 concepts, and test it on the remaining left out 5 concepts (also taking into account the next point about finetuning the video model).\n2. Moreover, as described in Section 4.2, line 244: “[...] all video-text pairs are used for fine-tuning the Stable Diffusion Model [...]”. If that is indeed the case, this means that the video diffusion model has already seen the specific videos it tries to predict later on, which makes the generation task significantly easier."},"limitations":{"value":"Yes."}},"nonreaders":[],"tmdate":1730879302191,"tcdate":1720840257022,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission9270/Reviewer_ACPC"],"signatures":["NeurIPS.cc/2024/Conference/Submission9270/Reviewer_ACPC"],"forum":"RfsfRn9OFd","number":2,"license":"CC BY 4.0","cdate":1720840257022,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission9270/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730879302191,"domain":"NeurIPS.cc/2024/Conference","replyto":"RfsfRn9OFd","id":"75PbXQiNW9","forumContent":{"TLDR":{"value":"We build an EEG-video dataset and propose a framework to generate videos from EEG signals, taking an important step towards decoding dynamic visual perception from EEG."},"venue":{"value":"NeurIPS 2024 poster"},"keywords":{"value":["EEG","video generation","diffusion model","brain-computer interface"]},"supplementary_material":{"value":"/attachment/44c231ba4e1dda99312f7354b6c439294f8a61c6.zip"},"primary_area":{"value":"neuroscience_and_cognitive_science"},"flagged_for_ethics_review":{"value":true},"abstract":{"value":"Our visual experience in daily life are dominated by dynamic change. Decoding such dynamic information from brain activity can enhance the understanding of the brain’s visual processing system. However, previous studies predominately focus on reconstructing static visual stimuli. In this paper, we explore to decode dynamic visual perception from electroencephalography (EEG), a neuroimaging technique able to record brain activity with high temporal resolution (1000 Hz) for capturing rapid changes in brains. Our contributions are threefold: Firstly, we develop a large dataset recording signals from 20 subjects while they were watching 1400 dynamic video clips of 40 concepts. This dataset fills the gap in the lack of EEG-video pairs. Secondly, we annotate each video clips to investigate the potential for decoding some specific meta information (e.g., color, dynamic, human or not) from EEG. Thirdly, we propose a novel baseline EEG2Video for video reconstruction from EEG signals that better aligns dynamic movements with high temporal resolution brain signals by Seq2Seq architecture. EEG2Video achieves a 2-way accuracy of 79.8% in semantic classification tasks and 0.256 in structural similarity index (SSIM). Overall, our works takes an important step towards decoding dynamic visual perception from EEG signals. Our dataset and code will be released soon."},"_bibtex":{"value":"@inproceedings{\nliu2024eegvideo,\ntitle={{EEG}2Video: Towards Decoding Dynamic Visual Perception from {EEG} Signals},\nauthor={Xuanhao Liu and Yan-Kai Liu and Yansen Wang and Kan Ren and Hanwen Shi and Zilong Wang and Dongsheng Li and Bao-liang Lu and Wei-Long Zheng},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=RfsfRn9OFd}\n}"},"title":{"value":"EEG2Video: Towards Decoding Dynamic Visual Perception from EEG Signals"},"pdf":{"value":"/pdf/a5ebcd48c768c5e7095143727c967cdbc60cc9f7.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"liu|eeg2video_towards_decoding_dynamic_visual_perception_from_eeg_signals"},"authorids":{"value":["~Xuanhao_Liu2","~Yan-Kai_Liu1","~Yansen_Wang2","~Kan_Ren1","~Hanwen_Shi1","~Zilong_Wang8","~Dongsheng_Li2","~Bao-liang_Lu1","~Wei-Long_Zheng1"]},"authors":{"value":["Xuanhao Liu","Yan-Kai Liu","Yansen Wang","Kan Ren","Hanwen Shi","Zilong Wang","Dongsheng Li","Bao-liang Lu","Wei-Long Zheng"]}},"version":2},{"content":{"summary":{"value":"The paper proposes TemporalBench, a new benchmark aimed at evaluating fine-grained temporal dynamics understanding in multimodal video models. The authors argue that most existing video QA and video understanding benchmarks can be solved from a single frame or even purely from text, due to coarse annotations that describe high-level actions (“cooking”, “playing guitar”) rather than temporal structure. TemporalBench is built from ~2K videos collected from multiple existing datasets (ActivityNet Captions, Charades, COIN, EgoExo4D, movie clips, Oops, FineGym), and augmented with dense, human-refined temporal captions that describe precise action sequences, frequencies, durations, and temporal relations. This produces ~15K QA pairs."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"Please refer to weaknesses."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. Clear and compelling motivation\n\nThe paper articulates the single-frame / text-only bias problem in existing video benchmarks convincingly, citing prior work showing that models can often solve them with a single frame or even without visual input. TemporalBench is explicitly designed to counter this by making temporal details the core of the task. The argument that coarse annotations lead to pseudo-video tasks that are effectively static-image recognition is well made and timely.\n\n2. Well-designed benchmark with fine-grained temporal focus\n\nThe dataset design targets temporal dynamics explicitly: action frequency (“two vs three slices”), sequence (“push glasses, then drink” vs “drink, then push glasses”), motion direction, and effector changes (right vs left hand), which cannot be reliably guessed from a single frame. The dense captions (≈42 words per clip on average, significantly more words per second than existing benchmarks) support more fine-grained tasks than typical datasets like MSRVTT or TGIF."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. My major concern is how the long videos (e.g. 20 mins) are annotated. Does it keep the temporal dynamic over a long span of the video? If it has, how are the long-time dependencies annotated?\n\n2. While ~2K videos and ~15K QA pairs are reasonable for a benchmark, this is still modest compared to many modern video datasets. It would be useful to more explicitly discuss how representative 2K videos are of the wider space of temporal phenomena.\n\n3. The manuscript seems do not provide quantitative measures of annotation quality, such as inter-annotator agreement, consistency checks on counts, or error rates on held-out verification sets. Given that many questions hinge on very fine distinctions (e.g., two vs three swings, direction “corner to center” vs “center to corner”), even small annotation inaccuracies could significantly affect scores."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762915558335,"tcdate":1762176275977,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission588/Reviewer_RXQd"],"signatures":["ICLR.cc/2026/Conference/Submission588/Reviewer_RXQd"],"forum":"XQfRnmOzY8","number":2,"license":"CC BY 4.0","cdate":1762176275977,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission588/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762915558335,"domain":"ICLR.cc/2026/Conference","replyto":"XQfRnmOzY8","id":"L0SYhqrfwk","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"TLDR":{"value":"TemporalBench: Evaluating Fine-Grained Temporal Dynamics Understanding for Multimodal Models"},"keywords":{"value":["Video; Benchmark; Temporal"]},"supplementary_material":{"value":"/attachment/9160fda348c6f1df228c11cd91238e3ac9d7b8df.zip"},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Understanding fine-grained temporal dynamics is crucial for multimodal video comprehension and generation. Due to the lack of fine-grained temporal annotations, existing video benchmarks mostly resemble static image benchmarks and are insufficient at evaluating models for temporal understanding.  In this paper, we introduce *TemporalBench*, a benchmark dedicated to evaluating **fine-grained temporal understanding** in videos. *TemporalBench* consists of $\\sim$15K video question-answer pairs, derived from $\\sim$2K high-quality human annotations detailing the temporal dynamics. As a result, our benchmark provides a unique testbed for evaluating various temporal understanding and reasoning abilities such as *action frequency, motion magnitude, event order*, etc.  Moreover, it enables evaluations on various tasks such as both short and long video understanding, as well as different models including multimodal embedding models and text generation models.  Furthermore, we notice a critical pitfall for multi-choice QA where LLMs can detect the subtle changes in negative captions and find a \"centralized” description as a cue for its prediction. To correct such bias, we propose **Multiple Binary Accuracy (MBA)**, a new metric for dense temporal understanding.  Results show that state-of-the-art models like GPT-4o achieve only **38.5%** short video QA, demonstrating a significant gap ($\\sim$30%) between humans and AI in temporal understanding.  We hope that *TemporalBench* can foster research on improving models' temporal reasoning capabilities. Both dataset and code will be available."},"_bibtex":{"value":"@misc{\ncai2025temporalbench,\ntitle={TemporalBench: Evaluating Fine-Grained Temporal Dynamics Understanding for Multimodal Models},\nauthor={Mu Cai and Reuben Tan and Jianrui Zhang and Bocheng Zou and Kai Zhang and Feng Yao and Fangrui Zhu and Jing Gu and Yiwu Zhong and Yuzhang Shang and Yao Dou and Jaden Park and Jianfeng Gao and Yong Jae Lee and Jianwei Yang},\nyear={2025},\nurl={https://openreview.net/forum?id=XQfRnmOzY8}\n}"},"title":{"value":"TemporalBench: Evaluating Fine-Grained Temporal Dynamics Understanding for Multimodal Models"},"pdf":{"value":"/pdf/99e03a75194d1ccd20c1d2ebaf5eb24601970f43.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"cai|temporalbench_evaluating_finegrained_temporal_dynamics_understanding_for_multimodal_models"},"authorids":{"value":["~Mu_Cai1","~Reuben_Tan1","~Jianrui_Zhang1","~Bocheng_Zou1","~Kai_Zhang10","~Feng_Yao1","~Fangrui_Zhu1","~Jing_Gu2","~Yiwu_Zhong1","~Yuzhang_Shang1","~Yao_Dou1","~Jaden_Park1","~Jianfeng_Gao1","~Yong_Jae_Lee2","~Jianwei_Yang1"]},"authors":{"value":["Mu Cai","Reuben Tan","Jianrui Zhang","Bocheng Zou","Kai Zhang","Feng Yao","Fangrui Zhu","Jing Gu","Yiwu Zhong","Yuzhang Shang","Yao Dou","Jaden Park","Jianfeng Gao","Yong Jae Lee","Jianwei Yang"]}},"version":2},{"content":{"summary":{"value":"This paper presents a training-free weight update framework for removing unsafe concepts from video diffusion models. The method leverages low-rank refusal vectors derived from pairs of safe/unsafe prompts, refined via contrastive PCA  to disentangle target concepts from unrelated semantics. By integrating these vectors into model weights through closed-form updates, the approach achieves permanent suppression of unsafe content without retraining or inference overhead. Experiments on OPEN-SORA and ZEROSCOPET2V across T2VSafetyBench and SafeSora benchmarks show average reductions in unsafe generations, while preserving video quality and prompt alignment."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Given that video diffusion models differ structurally from LLMs, have you validated that unsafe concepts in these models indeed form linear, separable directions in latent space? \n\nHow are \"neutral concepts\" operationalized? If a neutral concept is highly correlated with an unsafe concept, does subtracting Ce reduce the magnitude of the unsafe direction in Cr - αCe? Please provide experimental analysis."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The work introduces a training-free weight update framework for video unlearning, addressing a critical gap in scalable safety for video generative models. Unlike filtering methods or fine-tuning approaches, it enables irreversible, weight-level suppression with no computational overhead at inference.\n\n2. The method demonstrates strong empirical performance across diverse benchmarks and models, with significant reductions in unsafe content while maintaining visual quality and semantic alignment. The use of only five prompt pairs and closed-form updates enhances its practicality for real-world deployment.\n\n3. The integration of contrastive low-rank factorization to isolate unsafe concepts from neutral semantics is a thoughtful extension of prior concept-editing work to video domains, leveraging both text and image conditioning for improved precision."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Overreliance on linear representation assumptions: The core assumption—that unsafe concepts correspond to linear directions in latent space—is justified via references to LLM studies (e.g., Nanda et al., 2023), but video diffusion models have distinct architectures. No empirical validation such as visualization or concept disentanglement tests is provided to confirm this assumption holds for video models.\n\n2. Ambiguity in \"Neutral Concepts\" for cPCA: The paper does not clearly define how \"neutral concepts\" are selected for cPCA. If neutral concepts share semantic overlap with unsafe concepts (e.g., \"knife\" vs. \"violence\"), subtracting their covariance (Ce) could inadvertently weaken the unsafe concept direction, reducing suppression efficacy. This risk is not discussed or evaluated.\n\n3. While the method is tested on two models and benchmarks, it is unclear how it scales to multi-concept erasure (e.g., simultaneous removal of nudity and gore) or rare/unseen unsafe concepts. The mass erasure experiment (Section C.3) shows degraded performance for combined concepts, but the cause (e.g., overlapping directions) is not explored."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762920563101,"tcdate":1761218457462,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission8789/Reviewer_zvo7"],"signatures":["ICLR.cc/2026/Conference/Submission8789/Reviewer_zvo7"],"forum":"U1XBHtXl7Y","number":1,"license":"CC BY 4.0","cdate":1761218457462,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission8789/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762920563101,"domain":"ICLR.cc/2026/Conference","replyto":"U1XBHtXl7Y","id":"Jvu5nmvORL","forumContent":{"TLDR":{"value":"This work introduces the first training-free weight update framework for concept removal in video diffusion models, removing harmful concepts using just five safe/unsafe prompt pairs."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["video generation","machine unlearning"]},"primary_area":{"value":"alignment, fairness, safety, privacy, and societal considerations"},"abstract":{"value":"Video generative models achieve high-quality synthesis from natural-language prompts by leveraging large-scale web data. However, this training paradigm inherently exposes them to unsafe biases and harmful concepts, introducing the risk of generating undesirable or illicit content. To mitigate unsafe generations, existing machine unlearning approaches either rely on filtering, and can therefore be bypassed, or they update model weights, but with costly fine-tuning or training-free closed-form edits. We propose the first training-free weight update framework for concept removal in video diffusion models.\nFrom five paired safe/unsafe prompts, our method estimates a refusal vector and integrates it into the model weights as a closed-form update. A contrastive low-rank factorization further disentangles the target concept from unrelated semantics, it ensures a selective concept suppression and it does not harm generation quality. Our approach reduces unsafe generations on the Open-Sora and ZeroScopeT2V models across the T2VSafetyBench and SafeSora benchmarks, with average reductions of 36.3% and 58.2% respectively, while preserving prompt alignment and video quality. This establishes an efficient and scalable solution for safe video generation without retraining nor any inference overhead."},"_bibtex":{"value":"@inproceedings{\nfacchiano2026video,\ntitle={Video Unlearning via Low-Rank Refusal Vector},\nauthor={Simone Facchiano and Stefano Saravalle and Matteo Migliarini and Edoardo De Matteis and Alessio Sampieri and Andrea Pilzer and Emanuele Rodol{\\`a} and Indro Spinelli and Luca Franco and Fabio Galasso},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=U1XBHtXl7Y}\n}"},"title":{"value":"Video Unlearning via Low-Rank Refusal Vector"},"pdf":{"value":"/pdf/dce5198a701c8c8a7ebc30914ecc6fa75b673c48.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"facchiano|video_unlearning_via_lowrank_refusal_vector"},"authorids":{"value":["~Simone_Facchiano1","~Stefano_Saravalle1","~Matteo_Migliarini1","~Edoardo_De_Matteis1","~Alessio_Sampieri1","~Andrea_Pilzer1","~Emanuele_Rodolà1","~Indro_Spinelli1","~Luca_Franco1","~Fabio_Galasso1"]},"authors":{"value":["Simone Facchiano","Stefano Saravalle","Matteo Migliarini","Edoardo De Matteis","Alessio Sampieri","Andrea Pilzer","Emanuele Rodolà","Indro Spinelli","Luca Franco","Fabio Galasso"]}},"version":2},{"content":{"venue":{"value":"ICLR 2027 Conference Submission"},"keywords":{"value":["Multimodal Agent","Artist-Grounded Image Generation","Vague Intent Translation","Iterative Refinement","Style Authenticity Evaluation"]},"primary_area":{"value":"generative models"},"abstract":{"value":"Artist-grounded image generation requires more than appending an artist name to a prompt. Image models often respond to artist names through canonical shortcuts, such as recurring motifs, generic palettes, or overrepresented period signatures, rather than preserving the user's intended scene. We introduce Shortcut-Aware Artistic Control Planning (SACP), a formulation that translates underspecified artistic intent into an explicit control state separating scene anchors, preserve/transform decisions, style-regime hypotheses, role-bound artist evidence, and shortcut-avoidance constraints, and Atelier, its closed-loop runtime. Atelier grounds this state using artist-level knowledge and local patch references, compiles backend-aware generation plans, and iteratively refines candidates through global and local authenticity feedback. We further introduce ArtIntentBench, a benchmark covering Van Gogh and Qi Baishi across artwork re-rendering, period/style-controlled generation, historically unseen subjects, shortcut auditing, and human preference evaluation. Across open-weight and closed-source generators, Atelier improves artist-level style fidelity, preserves source structure more faithfully, and substantially reduces shortcut substitution compared with prompt-engineered, retrieval-augmented, and general-purpose agent baselines. These results suggest that artist-grounded generation is bottlenecked not only by image synthesis, but by the upstream inference of explicit, evidence-grounded artistic controls."},"_bibtex":{"value":"@inproceedings{\nanonymous2026beyond,\ntitle={Beyond Starry Night: Shortcut-Aware Control-State Planning for Artist-Grounded Text-to-Image Generation},\nauthor={Anonymous},\nbooktitle={Submitted to The Fifteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=iSDJKIKtok},\nnote={under review}\n}"},"title":{"value":"Beyond Starry Night: Shortcut-Aware Control-State Planning for Artist-Grounded Text-to-Image Generation"},"pdf":{"value":"/pdf/30b6b5801f472319285114b8eb2cef859c2c06e8.pdf"},"venueid":{"value":"ICLR.cc/2027/Conference/Submission"}},"tmdate":1791228912390,"tcdate":1789409091820,"writers":["ICLR.cc/2027/Conference","ICLR.cc/2027/Conference/Submission18828/Authors"],"signatures":["ICLR.cc/2027/Conference/Submission18828/Authors"],"forum":"iSDJKIKtok","license":"CC BY 4.0","number":18828,"cdate":1789409091820,"readers":["everyone"],"invitations":["ICLR.cc/2027/Conference/-/Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Bidding","ICLR.cc/2027/Conference/Submission18828/-/Full_Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Reviewing","ICLR.cc/2027/Conference/-/Edit"],"mdate":1791228912390,"odate":1791031328286,"domain":"ICLR.cc/2027/Conference","id":"iSDJKIKtok","version":2},{"content":{"venue":{"value":"ICLR 2026 Conference Desk Rejected Submission"},"keywords":{"value":["Video Hallucination","Large Multimodal Models"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"We introduce VidPair-Halluc, a new benchmark for evaluating video hallucination in large video models (LVMs) under rigorous and controlled conditions. Unlike previous benchmarks that primarily rely on text-based perturbations or adversarial questions while neglecting the consistency of visual backgrounds, VidPair-Halluc features video pairs with highly similar backgrounds but distinctly different foreground semantics, enabling precise attribution of model errors to genuine hallucination rather than background variation. The benchmark is constructed through PairFlow, a pipeline that leverages recent advances in text-to-image and video generation to systematically compose stories, generate coherent video clips, and assemble them into adversarial pairs. Covering both spatial and temporal reasoning across ten semantic aspects, VidPair-Halluc comprises $1$K high-quality adversarial video pairs and $11$K spatio-temporal QA pairs with control over background and foreground variations. We evaluate mainstream LVMs on VidPair-Halluc, and our results show that current models still struggle with robust and fine-grained video understanding in adversarial settings. Our code and data will be released."},"_bibtex":{"value":"@misc{\nanonymous2025no,\ntitle={No Place to Hide: Benchmarking Video Hallucination with Background-Controlled Pairs},\nauthor={Anonymous},\nyear={2025},\nurl={https://openreview.net/forum?id=11JBQuHDLb}\n}"},"title":{"value":"No Place to Hide: Benchmarking Video Hallucination with Background-Controlled Pairs"},"pdf":{"value":"/pdf/cfba3777ce45f510fcd82e64aa343c4bc44f85fb.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Desk_Rejected_Submission"},"paperhash":{"value":"huang|no_place_to_hide_benchmarking_video_hallucination_with_backgroundcontrolled_pairs"},"authorids":{"value":["~Haojian_Huang1","~Harold_Haodong_Chen1","~Meng_Luo2","~Junjia_Du1","~Shanqing_Xu1","~Ziheng_Chen7","~Yanxiang_Huang1","~Yinchuan_Li1","~Ying-Cong_Chen1"]},"authors":{"value":["Haojian Huang","Harold Haodong Chen","Meng Luo","Junjia Du","Shanqing Xu","Ziheng Chen","Yanxiang Huang","Yinchuan Li","Ying-Cong Chen"]}},"tmdate":1770907924931,"tcdate":1757153838598,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission2585/Authors"],"signatures":["ICLR.cc/2026/Conference/Submission2585/Authors"],"forum":"11JBQuHDLb","license":"CC BY 4.0","number":2585,"cdate":1757153838598,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/-/Submission","ICLR.cc/2026/Conference/-/Post_Submission","ICLR.cc/2026/Conference/Submission2585/-/Full_Submission","ICLR.cc/2026/Conference/-/Desk_Rejected_Submission","ICLR.cc/2026/Conference/-/Edit"],"mdate":1770907924931,"odate":1759896705795,"domain":"ICLR.cc/2026/Conference","id":"11JBQuHDLb","version":2},{"content":{"summary":{"value":"This paper presents a novel video benchmark aimed at evaluating the capabilities of large multi-modal models (LMMs) in understanding and assessing video quality. The benchmark stands out due to its inclusion of a wide range of videos sourced from diverse platforms, with a particular emphasis on AI-generated content (AIGC) and CG videos. The authors strategically categorize evaluation questions into three distinct groups to enable a thorough analysis of video quality. Notably, the inclusion of open-ended questions allows for the assessment of complex, nuanced scenarios beyond simple objective measures. Last but not least, the authors conducted extensive comparisons of current LLMs on this benchmark, showing their limited performance on evaluating video quality. Despite these contributions, there are concerns that need to be further addressed before this paper can be accepted."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"- How to generate questions for different categories? If you use LLM to generate the questions, please include the prompt or template in the paper for readers to study. If not, illustrate your approach with details."},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- The paper is well-presented, with adequate examples and figures to show the query designs in the benchmark.\n\n- The questions can cover extensive aspects of video quality, including technical, aesthetic, temporal, and AIGC distortions. \n\n- The video pairs comparison is a good task for evaluating LLM's capability, which can be considered as an advantage over existing video benchmarks."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Motivation of the Benchmark: The paper's motivation for building the benchmark is not sufficiently strong. While the introduction states the importance of video quality for viewer experience and high-quality video generation, what viewers primarily care about is the classification of video quality (e.g., low, medium, high). A well-trained classifier could fulfill this need more directly. Additionally, for assessing video generation quality, designing robust evaluation metrics is more crucial than simply building a benchmark to test LLMs.\n\n- Question Design Limitations: The paper’s three question categories are not comprehensive enough. It would be beneficial to include overarching questions such as \"What is the quality level of this video?\" with clear, concise answers like \"low/medium/high.\" This would better align with practical applications and user expectations.\n\n- Evaluation Metric Weaknesses: The evaluation metrics used to assess LLMs' QA performance lack robustness. The current scoring prompt only accounts for completeness, accuracy, and relevance, which are insufficient for open-ended questions. Recent studies (e.g., RAGchecker [NeurIPS 2024]) suggest including metrics like faithfulness (to detect hallucinations), truthfulness, correctness, etc. Additionally, prompts should be specifically designed with definitions and in-context examples to help LLMs understand each metric. Without refining the prompt and using comprehensive evaluation metrics, the reported performance scores are likely unreliable."}},"nonreaders":[],"tmdate":1731427490316,"tcdate":1730687668621,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission1786/Reviewer_BVGh"],"signatures":["ICLR.cc/2025/Conference/Submission1786/Reviewer_BVGh"],"forum":"VaUy5GZO3f","number":2,"license":"CC BY 4.0","cdate":1730687668621,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission1786/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427490316,"domain":"ICLR.cc/2025/Conference","replyto":"VaUy5GZO3f","id":"HQoMpi0e3S","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Large multi-modal model","benchmark","video quality assessment"]},"primary_area":{"value":"datasets and benchmarks"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"With the rising interest in research on Large Multi-modal Models (LMMs) for video understanding, many studies have emphasized general video comprehension capabilities, neglecting the **systematic exploration into video quality understanding**. To address this oversight, we introduce **Q-Bench-Video** in this paper, a new benchmark specifically designed to evaluate LMMs' proficiency in discerning video quality. **a)** To ensure the diversity of video sources, Q-Bench-Video encompasses videos from natural scenes, computer graphics (CG), and AI-generated content (AIGC). **b)** Building on the traditional multiple-choice questions format with the *Yes-or-No* and *What-How* categories, we include *Open-ended* questions to better evaluate complex scenarios. Additionally, we incorporate the **video pair quality comparison** question to enhance comprehensiveness. **c)** Beyond the traditional *Technical*, *Aesthetic*, and *Temporal* distortions, we have expanded our evaluation aspects to include the dimension of *AIGC* distortions, which addresses the increasing demand for video generation. Finally, we collect a total of 2,378 question-answer pairs and test them on 12 open-source & 5 proprietary LMMs. Our findings indicate that while LMMs have a foundational understanding of video quality, their performance remains incomplete and imprecise, with a notable discrepancy compared to human-level performance. Through **Q-Bench-Video**, we seek to catalyze community interest, stimulate further research, and unlock the untapped potential of LMMs to close the gap in video quality understanding."},"_bibtex":{"value":"@misc{\nzhang2024qbenchvideo,\ntitle={Q-Bench-Video: Benchmarking the Video Quality Understanding of {LMM}s},\nauthor={Zicheng Zhang and Ziheng Jia and Haoning Wu and Chunyi Li and Zijian Chen and Yingjie Zhou and Wei Sun and Xiaohong Liu and Xiongkuo Min and Weisi Lin and Guangtao Zhai},\nyear={2024},\nurl={https://openreview.net/forum?id=VaUy5GZO3f}\n}"},"title":{"value":"Q-Bench-Video: Benchmarking the Video Quality Understanding of LMMs"},"pdf":{"value":"/pdf/b6665a67a0e193ca939b3b34f488e5e0b380c738.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"zhang|qbenchvideo_benchmarking_the_video_quality_understanding_of_lmms"},"authorids":{"value":["~Zicheng_Zhang7","~Ziheng_Jia1","~Haoning_Wu1","~Chunyi_Li1","~Zijian_Chen1","~Yingjie_Zhou1","~Wei_Sun12","~Xiaohong_Liu2","~Xiongkuo_Min1","~Weisi_Lin1","~Guangtao_Zhai1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Zicheng Zhang","Ziheng Jia","Haoning Wu","Chunyi Li","Zijian Chen","Yingjie Zhou","Wei Sun","Xiaohong Liu","Xiongkuo Min","Weisi Lin","Guangtao Zhai"]}},"version":2},{"content":{"TLDR":{"value":"MLLMs control interfaces but struggle with dynamic GUIs. The new GUI-World dataset evaluates their understanding, revealing challenges and enhance their capabilities in handling dynamic content and complex tasks."},"venue":{"value":"Video-Langauge Models Poster"},"keywords":{"value":["GUI Agent","Video LLM","Multimodal Large Language Model","Benchmark","Dataset","Instruction Tuning"]},"supplementary_material":{"value":"/attachment/1e27054aa6cf4dea3f9c1a4b82c096a259e2d0c1.zip"},"abstract":{"value":"Recently, Multimodal Large Language Models (MLLMs) have been used as agents to control keyboard and mouse inputs by directly perceiving the Graphical User Interface (GUI) and generating corresponding code. However, current agents primarily exhibit excellent understanding capabilities in static environments and are predominantly applied in relatively simple domains, such as Web or mobile interfaces. We argue that a robust GUI agent should be capable of perceiving temporal information on the GUI, including dynamic Web content and multi-step tasks. Additionally, it should possess a comprehensive understanding across various GUI scenarios, including desktop software and multi-window interactions. To this end, this paper introduces a new dataset, termed GUI-World, which features meticulously crafted Human-MLLM annotations, extensively covering six GUI scenarios and eight types of GUI-orientated questions in three formats. We evaluate the capabilities of current state-of-the-art MLLMs, including ImageLLMs and VideoLLMs, in understanding various types of GUI content, especially dynamic and sequential content. Our findings reveal that ImageLLMs struggle with dynamic GUI content without manually annotated keyframes or operation history. On the other hand, VideoLLMs fall short in all GUI-orientated tasks given the sparse of GUI video dataset. Based on GUI-World, we take the initial step of leveraging a fine-tuned VideoLLM as a GUI agent, demonstrating an improved understanding of various GUI tasks. However, due to the limitations in the performance of base LLMs, we conclude that using VideoLLMs as GUI agents remains a significant challenge. We believe our work provides valuable insights for future research in dynamic GUI content understanding."},"_bibtex":{"value":"@inproceedings{\nchen2025guiworld,\ntitle={{GUI}-{WORLD}: A {GUI}-oriented Video Dataset for Multimodal {LLM}-based Agents},\nauthor={Dongping Chen and Yue Huang and Siyuan Wu and Jingyu Tang and Huichi Zhou and Qihui Zhang and Zhigang He and Yilin Bai and Chujie Gao and Liuyi Chen and Yiqiang Li and Chenlong Wang and Yue Yu and Tianshuo Zhou and Zhen Li and Yi Gui and Yao Wan and Pan Zhou and Jianfeng Gao and Lichao Sun},\nbooktitle={Workshop on Video-Language Models @ NeurIPS 2024},\nyear={2025},\nurl={https://openreview.net/forum?id=SkZRB75Q3H}\n}"},"title":{"value":"GUI-WORLD: A GUI-oriented Video Dataset for Multimodal LLM-based Agents"},"pdf":{"value":"/pdf/fb060096421505a738b39c87f74eee8851d0a94e.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Workshop/Video-Langauge_Models"},"paperhash":{"value":"chen|guiworld_a_guioriented_video_dataset_for_multimodal_llmbased_agents"},"authorids":{"value":["~Dongping_Chen1","~Yue_Huang9","~Siyuan_Wu6","~Jingyu_Tang1","~Huichi_Zhou2","~Qihui_Zhang1","~Zhigang_He3","~Yilin_Bai1","~Chujie_Gao1","~Liuyi_Chen1","~Yiqiang_Li3","~Chenlong_Wang3","~Yue_Yu9","~Tianshuo_Zhou2","~Zhen_Li24","~Yi_Gui3","~Yao_Wan2","~Pan_Zhou5","~Jianfeng_Gao1","~Lichao_Sun1"]},"track":{"value":"Long Paper Track (up to 9 pages)"},"authors":{"value":["Dongping Chen","Yue Huang","Siyuan Wu","Jingyu Tang","Huichi Zhou","Qihui Zhang","Zhigang He","Yilin Bai","Chujie Gao","Liuyi Chen","Yiqiang Li","Chenlong Wang","Yue Yu","Tianshuo Zhou","Zhen Li","Yi Gui","Yao Wan","Pan Zhou","Jianfeng Gao","Lichao Sun"]}},"tmdate":1736861080263,"pdate":1730081752279,"tcdate":1725755052652,"writers":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission17/Authors"],"signatures":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission17/Authors"],"forum":"SkZRB75Q3H","license":"CC BY 4.0","number":17,"cdate":1725755052652,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Post_Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission17/-/Full_Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Edit"],"mdate":1736861080263,"odate":1736861080245,"domain":"NeurIPS.cc/2024/Workshop/Video-Langauge_Models","id":"SkZRB75Q3H","version":2},{"content":{"summary":{"value":"This paper proposes a fast video editing method leveraging one-step text-to-image diffusion models.Starting from the motivation that three bottleneck, (1) time consuming multi-step inversion process, (2) spatial and (3) temporal inconsistency when using T2I model for video editing, the authors propose training inversion network with Structure-aware editing loss and unified frame editing. Qualitative and quantitaive experiments are conducted with various basline with Vbench metric."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1. Why use Gaussian noise for perturbing the prompt condition? I guess there are lots of design choices for perturbing the text condition, such as synonym substitution (e.g, dog-> cat), dropout, and randomly reordering the text."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"From the explicit problems in the video editing task, this paper is well motivated and proposes a good research direction that uses T2I model for video editing for efficiency. Training an inversion network with structure-aware editing loss, which uses prompt condition perturbation, is quite novel and impressive. The authors show a lot of quantitative experiments with various baseline comparisons which makes the proposed method confident."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. **Poor and Inefficient experiments**\nThe supplementary visual results reveal substantial flickering and noticeable artifacts throughout the provided videos. The generated videos exhibit clear temporal inconsistency with background distortions, and object-level artifacts are evident. For example, in the “dog on car” → “cat on car” editing example, the cat’s face intermittently appears and disappears at the lower-left corner of the frame.\nSuch a level of artifact raises serious concerns regarding the robustness of the proposed method and suggests that the model does not perform reliably. In addition, the **temporal flickering score** in the **VBench** evaluation is not reported, further limiting the quantitative assessment of temporal consistency. Moreover, the paper lacks sufficient video comparison results against baseline models, which are necessary for a fair and comprehensive evaluation of visual quality.\n\n2. **Lack of Inference Time Evaluation**\nOne of the paper’s main claimed contributions is fast video editing. However, there is no quantitative comparison or analysis of inference time against baseline models. Detailed inference time evaluations under varying numbers of frames are also required to substantiate the claimed efficiency.\n\n3. **Proposed Unified-Frame Editing: Not Novel and Questionable for Efficiency**\nThe proposed unified-frame editing design is not novel. The idea of injecting additional tokens into the attention mechanism is a well-known technique already proposed in several prior works [1,2]. Furthermore, when the video sequence becomes longer, the method may suffer from degraded global consistency due to the limited receptive range of the sliding window.\n\n4. **Insufficient Information on the Inversion Network Training**\nThe paper does not report sufficient details regarding the training of the inversion network, such as the dataset used or the number of training images. In addition, while the SAE loss introduces a noise weight term $\\lambda$, no quantitative analysis or ablation results on this parameter are provided. Such missing details make it difficult to reproduce and verify the claimed improvements.\n\nThe paper is well-motivated, and the idea of training an inversion network with SAE loss is novel and interesting. However, the overall qualitative performance is unsatisfactory: the generated videos exhibit **prominent artifacts** and **severe temporal flickering**. Given these significant issues in both robustness and visual quality, I am unable to recommend acceptance of the paper in its current state and must, regrettably, recommend rejection.\n\n---\n\n[1] Geyer, Michal, et al. \"Tokenflow: Consistent diffusion features for consistent video editing.\"\n\n[2] Wu, et al. \"Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation.\""}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762928446427,"tcdate":1761898925275,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission18735/Reviewer_Kcgp"],"signatures":["ICLR.cc/2026/Conference/Submission18735/Reviewer_Kcgp"],"forum":"DscflMFynS","number":3,"license":"CC BY 4.0","cdate":1761898925275,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission18735/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762928446427,"domain":"ICLR.cc/2026/Conference","replyto":"DscflMFynS","id":"AsV6sQTcX1","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Video Editing","Diffusion Models","Generative Models","Text editing"]},"supplementary_material":{"value":"/attachment/4105fea3522ed19959e01303b8da3ec2411c7f04.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Text-guided video editing with diffusion models is prohibitively slow, hindered by costly multi-step sampling and inversion. We present VIDES, the first framework to successfully adapt one-step text-to-image (T2I) models for high-quality video editing, addressing the core challenges of inversion, editability, and temporal consistency. To bypass slow iterative inversion, we train a learnable encoder that predicts the initial noise for each frame in a single forward pass. This encoder is trained with a novel Structure-Aware Editing (SAE) loss on a curated dataset of structurally-aligned image pairs, teaching it to preserve the source video's geometry during edits. For temporal coherence, we introduce Unified-Frame Editing (UFE), a technique that concatenates frame latents to facilitate cross-frame attention in a single generation step; for long videos, a sliding-window strategy with an anchor frame maintains global consistency. Our extensive experiments demonstrate that VIDES achieves editing quality comparable or superior to state-of-the-art multi-step methods, while operating approximately 155 times faster. This breakthrough paves the way for practical, real-time video editing applications."},"_bibtex":{"value":"@misc{\nlim2025vides,\ntitle={{VIDES}: {VIDEO} {EDITING} {IN} {SECONDS} {WITH} {ONE}-{STEP} {DIFFUSION} {MODELS}},\nauthor={Habin Lim and Gyeong-Moon Park},\nyear={2025},\nurl={https://openreview.net/forum?id=DscflMFynS}\n}"},"title":{"value":"VIDES: VIDEO EDITING IN SECONDS WITH ONE-STEP DIFFUSION MODELS"},"pdf":{"value":"/pdf/b776c3f6cdd85e19a19651574abda9a0b60de506.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"lim|vides_video_editing_in_seconds_with_onestep_diffusion_models"},"authorids":{"value":["~Habin_Lim2","~Gyeong-Moon_Park1"]},"authors":{"value":["Habin Lim","Gyeong-Moon Park"]}},"version":2},{"content":{"summary":{"value":"This paper explores various video token merging strategies in the context of long-form video classification and finally propose a learnable Video Token Merging algorithm that dynamically merges video tokens based on visual salient areas. The contributions are summarized as follow:\n1.  Explore various video token merging methods including the naïve VTM, the region concentrated VTM, and the motion-based VTM.\n2.  Propose the learnable video token merging algorithm, which estimates the saliency scores of each token and adaptively merge visual tokens based on those scores.\n3. The proposed algorithm achieves the best or competitive results on various datasets."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. Due to the two-paths design, has the training time doubled?\n2. The paper tries to learn saliency scores using matrix $U_s$. How about using $\\sum{QK^T}$ in Equation 8 as saliency scores for each visual token?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. This paper explores various video token merging methods including the naïve VTM, the region concentrated VTM, and the motion-based VTM.\n2. Compare with baseline and rule-based video token merging. The proposed learnable video token merging strategy has large improvement.\n3. The two-paths design to deal with non-differentiable problem in partitioning process is interesting."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. This paper proposes a leanable video token merging strategy. The similiar high-level idea can be found by CTS[1] in image domain. The novelty is insufficient。\n2. This paper focuses on video token merging. However, I do not observe any specific design tailored for the video domain in terms of the methodology. let alone long video.\n\n\n\n[1] Content-aware Token Sharing for Efficient Semantic Segmentation with Vision Transformers"},"limitations":{"value":"Yes"}},"nonreaders":[],"tmdate":1730878904554,"tcdate":1720612172795,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission4154/Reviewer_7bwU"],"signatures":["NeurIPS.cc/2024/Conference/Submission4154/Reviewer_7bwU"],"forum":"wduRaBDRBS","number":2,"license":"CC BY 4.0","cdate":1720612172795,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission4154/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730878904554,"domain":"NeurIPS.cc/2024/Conference","replyto":"wduRaBDRBS","id":"pyJ4gizUn1","forumContent":{"venue":{"value":"NeurIPS 2024 poster"},"keywords":{"value":["Long video understanding","token merging"]},"primary_area":{"value":"machine_vision"},"abstract":{"value":"As the scale of data and models for video understanding rapidly expand, handling long-form video input in transformer-based models presents a practical challenge. Rather than resorting to input sampling or token dropping, which may result in information loss, token merging shows promising results when used in collaboration with transformers. However, the application of token merging for long-form video processing is not trivial. We begin with the premise that token merging should not rely solely on the similarity of video tokens; the saliency of tokens should also be considered. To address this, we explore various video token merging strategies for long-form video classification, starting with a simple extension of image token merging, moving to region-concentrated merging, and finally proposing a learnable video token merging (VTM) algorithm that dynamically merges tokens based on their saliency. Extensive experimental results show that we achieve better or comparable performances on the LVU, COIN, and Breakfast datasets. Moreover, our approach significantly reduces memory costs by 84% and boosts throughput by approximately 6.89 times compared to baseline algorithms."},"_bibtex":{"value":"@inproceedings{\nlee2024video,\ntitle={Video Token Merging for Long Video Understanding},\nauthor={Seon-Ho Lee and Jue Wang and Zhikang Zhang and David Fan and Xinyu Li},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=wduRaBDRBS}\n}"},"title":{"value":"Video Token Merging for Long Video Understanding"},"pdf":{"value":"/pdf/02ca725ea4f0f64f7c82a9d5389359bf6b97e1bf.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"lee|video_token_merging_for_long_video_understanding"},"authorids":{"value":["~Seon-Ho_Lee1","~Jue_Wang8","~Zhikang_Zhang1","~David_Fan2","~Xinyu_Li4"]},"authors":{"value":["Seon-Ho Lee","Jue Wang","Zhikang Zhang","David Fan","Xinyu Li"]}},"version":2},{"content":{"summary":{"value":"This paper presents a interpolation method for extending Video-LLMs to process longer video sequences under training-free setting, called INTP-Video-LLMs. The approach leverages a video token rearrangement technique and a training-free LLM context window extension method to bypass the limitations of existing Video-LLMs, which typically only process short video clips. Furthermore, a training-free key-value (KV) cache compression mechanism is introduced to optimize memory usage during inference. The proposed INTP-Video-LLMs can comprehend longer video sequences (up to 32 frames) without additional training, and experimental results indicate that this approach provides a significant improvement in both video processing capabilities and inference efficiency."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"Q1. The authors perform 2-bit quantization of the KV cache to reduce storage overhead during the inference process of the video LLM. However, does quantizing the stored KV result in performance degradation?\n\nQ2. Could you provide a more detailed description of the video tokens rearrangement process (e.g., in the form of pseudocode) as well as an analysis of the effectiveness of this method?"},"rating":{"value":5},"details_of_ethics_concerns":{"value":"N/A"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"- The motivation for the proposed INTP-Video-LLM is clear and fundamental. The Introduction is well-crafted, making it easy to grasp the paper's concept.\n- The proposed methods allow existing Video-LLMs to process longer video (i.e. more video frames) **in a training-free manner**, effectively addressing computational constraints.\n- The paper comprehensively considers and analyzes whether the proposed **Video Tokens Rearrangement** and **Interpolating Video-LLM Backbone** will lead to additional computational overhead.\n- The **memory optimization** via KV-cache compression ensures that extended video sequences can be processed with minimal memory overhead, making the approach feasible for practical deployment."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The paper points out the existing limitations in the temporal consistency of encoders and projectors, but the description of the video token rearrangement method is unclear. It is not explained why rearrangement would help maintain consistency, and a more thorough explanation is needed.\n- Near Figure 2, at line 234, it describes “we obtain two subsequences, $X_{v,1}$ and $X_{v,1}$,” where $X_{v,1}$ appears twice. Is this a typographical error?\n- In Section 3.4.1, the authors analyze the inference cost and propose 2-bit quantization of the KV cache to reduce storage overhead during inference. However, the potential performance degradation due to quantization is not discussed in detail. It would be beneficial to include an analysis of the trade-offs between reduced storage and potential accuracy loss.\n- In the ablation study (Section 4.3), the authors compare the performance of using different numbers of frames in Video QA tasks. However, it is unclear what the individual contributions of each module are to the final performance. A detailed breakdown of the impact of each module would provide more insight into their effectiveness."}},"nonreaders":[],"tmdate":1732609526675,"tcdate":1730137517463,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission3931/Reviewer_qZSz"],"signatures":["ICLR.cc/2025/Conference/Submission3931/Reviewer_qZSz"],"forum":"QrTvFCa4nX","number":1,"license":"CC BY 4.0","cdate":1730137517463,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission3931/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732609526675,"domain":"ICLR.cc/2025/Conference","replyto":"QrTvFCa4nX","id":"kC8nfypjcL","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"TLDR":{"value":"Obtain a Video-LLM which can process more frames in a totally training-free manner."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video Understanding"]},"primary_area":{"value":"foundation or frontier models, including LLMs"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Advancements in Large Language Models (LLMs) inspire various strategies for integrating video modalities. \nA key approach is Video-LLMs, which incorporate an optimizable interface linking sophisticated video encoders to LLMs. \nHowever, due to computation and data limitations, these Video-LLMs are typically pre-trained to process only short videos, limiting their broader application for understanding longer video content. Additionally, fine-tuning Video-LLMs to handle longer videos is cost-prohibitive.\nConsequently, it becomes essential to explore the interpolation of Video-LLMs under a completely training-free setting. In this paper, we first identify the primary challenges in interpolating Video-LLMs: (1) the video encoder and modality alignment projector are fixed, preventing the integration of additional frames into Video-LLMs, and (2) the LLM backbone is limited in its content length capabilities, which complicates the processing of an increased number of video tokens.\nTo address these challenges, we propose a specific INTerPolation method for Video-LLMs (INTP-Video-LLMs). We introduce an alternative video token rearrangement technique that circumvents limitations imposed by the fixed video encoder and alignment projector. Furthermore, we introduce a training-free LLM context window extension method to enable Video-LLMs to understand a correspondingly increased number of visual tokens."},"_bibtex":{"value":"@misc{\nshang2025interpolating,\ntitle={Interpolating Video-{LLM}s:  Toward Longer-sequence {LMM}s in a Training-free Manner},\nauthor={Yuzhang Shang and Bingxin Xu and Weitai Kang and Mu Cai and Yuheng Li and Zehao Wen and Zhen Dong and Kurt Keutzer and Yong Jae Lee and Yan Yan},\nyear={2025},\nurl={https://openreview.net/forum?id=QrTvFCa4nX}\n}"},"title":{"value":"Interpolating Video-LLMs:  Toward Longer-sequence LMMs in a Training-free Manner"},"pdf":{"value":"/pdf/610d4d569fba56bc840ff0a6d6108535129e04d6.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"shang|interpolating_videollms_toward_longersequence_lmms_in_a_trainingfree_manner"},"authorids":{"value":["~Yuzhang_Shang1","~Bingxin_Xu1","~Weitai_Kang1","~Mu_Cai1","~Yuheng_Li1","~Zehao_Wen1","~Zhen_Dong3","~Kurt_Keutzer1","~Yong_Jae_Lee2","~Yan_Yan6"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Yuzhang Shang","Bingxin Xu","Weitai Kang","Mu Cai","Yuheng Li","Zehao Wen","Zhen Dong","Kurt Keutzer","Yong Jae Lee","Yan Yan"]}},"version":2},{"content":{"summary":{"value":"The paper studies a phenomenon the authors call surface learning: LLMs achieving high accuracy on benchmarks by exploiting superficial correlations rather than engaging in deeper reasoning. They propose ME-Test, a probing framework that pairs original questions with minimally-edited variants where surface cues are altered but core semantics remain. Performance drops under these variants are used to quantify surface learning. To mitigate the issue, the paper introduces two interventions: thinking-path correction (prompting models to revise or reflect on earlier reasoning) and behavior correction (explicitly penalizing shallow patterns during generation). Experiments on GSM8K, StrategyQA, and BBH tasks show large performance drops under ME-Test for several open and closed models; the proposed mitigations yield partial recovery."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1. How is “surface learning” fundamentally different from shortcut learning or pattern exploitation described extensively in prior literature?\n\n2. How do you ensure that ME-Test edits do not unintentionally change semantics or difficulty? Was human validation performed?\n\n3. Does ME-Test correlate with other robustness metrics or real user-facing failures?\n\n4. Can the mitigation methods cause degradation in standard accuracy or introduce verbosity without improving reasoning quality?\n\n5. Do models with explicit training for chain-of-thought (e.g., R1, DeepSeek-R1) still show the same pattern?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. ME-Test is simple, reproducible, and architecture-agnostic: minimal edits to input help separate form-matching from semantic reasoning.  \n\n2. The performance gaps are systematically measured and provide clear empirical evidence of shortcut reliance.\n\n3. Mitigation strategies are lightweight (prompt-level) and do not require retraining."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The concept of “surface learning” is not clearly distinguished from existing notions like pattern matching, or spurious correlations. The paper introduces a new term but does not establish conceptual novelty.\n\n2. ME-Test edits sometimes oversimplify meaning preservation, and no human validation is provided to confirm that all minimal edits preserve problem semantics fully.\n\n3. Mitigation methods resemble prompt engineering and chain-of-thought reflection; they are not fundamentally novel, and their effectiveness varies across tasks.\n\n4. Improvements after mitigation are moderate and do not eliminate the gap; no analysis is provided on failure cases or when mitigation backfires.\n\n5. No examination of whether ME-Test correlates with real-world robustness or downstream utility; unclear if this is just another adversarial dataset."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762927287493,"tcdate":1762103466248,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission17380/Reviewer_ga6M"],"signatures":["ICLR.cc/2026/Conference/Submission17380/Reviewer_ga6M"],"forum":"UVmeiLJbr4","number":3,"license":"CC BY 4.0","cdate":1762103466248,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission17380/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762927287493,"domain":"ICLR.cc/2026/Conference","replyto":"UVmeiLJbr4","id":"xIvRmgQ3U6","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"TLDR":{"value":"A novel finding on surface learning and a novel approach to mitigate surface learning."},"keywords":{"value":["Surface Learning","Shortcut Learning","Large Language Models"]},"primary_area":{"value":"alignment, fairness, safety, privacy, and societal considerations"},"abstract":{"value":"As Large Language Models (LLMs) continue to evolve, assessing their genuine comprehension of underlying knowledge is crucial to ensure the reliability in real-world applications. To evaluate what LLMs learn, we first introduce ME-Test suite, including Mathematical and English grammar examinations, where each question is equipped with relevant knowledge to guide the model. Building upon this, we construct a sequence of questions with increasing difficulty based on Cognitive Load theory, enabling the model to perform continuous problem-solving using the dialogue history. Through a comprehensive evaluation, we uncover a phenomenon of **Surface Learning** behavior on LLMs similar to student learning behavior in Education Psychology. The behavior indicates that although the models seem to know the formulas and strategies required to solve specific types of problems on the surface, they do not truly comprehend the essence of these concepts, resulting in surface-level short-term benefits rather than in-depth learning. Further to mitigate surface learning behavior of LLMs, we propose a long-term strategy for both training-free and post-training scenarios. In training-free scenario, inspired by Self-Concept theory, LLMs are prompted with goal-setting and planning beforehand as well as feedback afterward to improve the ability in reasoning process. To better activate the underlying knowledge during the post-training process, we propose behavior correction strategy to re-rank samples based on the designed self-cognition indicators of LLMs. This strategy prevents models from relying on easy-to-find paradigms to maximize rewards or minimize losses in the initial training stage, rather than undertaking actual reasoning. Extensive experiments of Supervised and Reinforcement Fine-Tuning (SFT, RFT) conducted on LLMs demonstrate the effectiveness of the strategy."},"_bibtex":{"value":"@misc{\nzhao2026beneath,\ntitle={Beneath the Surface: Exposing and Mitigating Surface Learning in Large Language Models},\nauthor={Lili Zhao and Yang Wang and Wei Chen and Yu Yuan and Qi Liu and Dongjie Guo and Kai Han and Shijin Wang},\nyear={2026},\nurl={https://openreview.net/forum?id=UVmeiLJbr4}\n}"},"title":{"value":"Beneath the Surface: Exposing and Mitigating Surface Learning in Large Language Models"},"pdf":{"value":"/pdf/6c4aa9315d2a5091de3f3953b5ab9b72d98fc992.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"zhao|beneath_the_surface_exposing_and_mitigating_surface_learning_in_large_language_models"},"authorids":{"value":["~Lili_Zhao3","~Yang_Wang69","~Wei_Chen54","~Yu_Yuan5","~Qi_Liu3","~Dongjie_Guo1","~Kai_Han8","~Shijin_Wang1"]},"authors":{"value":["Lili Zhao","Yang Wang","Wei Chen","Yu Yuan","Qi Liu","Dongjie Guo","Kai Han","Shijin Wang"]}},"version":2},{"content":{"summary":{"value":"This paper proposed a synthetic framework called VideoNIAH to construct the benchmark for evaluating video MLLMs. This framework aims to decouple video content from query-response pairs by corrupting the original videos with \"needles\" that are irrelevant to the videos.  Based on this frame, a video benchmark named VNBench is further built to evaluate different capabilities of proprietary and open-source video MLLMs, such as temporal perception, chronological ordering, and spatio-temporal coherence."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. The paper mentioned the used 'needles' are not related to the original videos. Did you conduct the similarity calculation between each video frame and the candidate \"needle\" before inserting them? \n2. Inserting the unrelated visual \"needle\" may not make the benchmark very challenging. Did you try the \"needles\" that are similar to the original video frame to a certain extent but are from a different category or situation? Maybe it will make the benchmark more challenging."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. The proposed VideoNIAH is a simple and flexible framework for constructing the video benchmark by avoiding expensive human annotation and leveraging existing video datasets. \n2. The VNBench seems to be the first synthetic video benchmark for evaluating video MLLMs through three skills including retrieval,\nordering and counting. \n3. Comprehensive evaluation experiments are conducted to evaluate close- and open-source MLLMs, further giving some insights for improving the model development and training."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Although the authors mentioned the proposed framework is inspired by language model evaluation. It is still not very clear about the motivation for proposing such a synthetic framework for video benchmarking and why the proposed one is necessary or important. \n2. The proposed framework mainly adopts subtitles and unrelated images as the \"needle\", it is better to discuss whether there are other types of \"needle\" that can be used. \n3. A small number of video samples and the types of selected video datasets are used to build the benchmark. In addition, only three tasks are adopted to evaluate the video MLLMs. It is better to make the benchmark more scalable and diverse with comprehensive evaluation capabilities."}},"nonreaders":[],"tmdate":1732555882142,"tcdate":1730667398720,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission8848/Reviewer_ozpA"],"signatures":["ICLR.cc/2025/Conference/Submission8848/Reviewer_ozpA"],"forum":"ZJo6Radbqq","number":2,"license":"CC BY 4.0","cdate":1730667398720,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission8848/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732555882142,"domain":"ICLR.cc/2025/Conference","replyto":"ZJo6Radbqq","id":"JYlxpLIkeS","forumContent":{"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["video MLLM"]},"primary_area":{"value":"datasets and benchmarks"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Video understanding is a crucial next step for multimodal large language models (MLLMs).\nVarious benchmarks are introduced for better evaluating the MLLMs.\nNevertheless, current video benchmarks are still inefficient for evaluating video models during iterative development due to the high cost of constructing datasets and the difficulty in isolating specific skills.\nIn this paper, we propose VideoNIAH (Video Needle in A Haystack), a benchmark construction framework through synthetic video generation. \nVideoNIAH decouples video content from their query-responses by inserting unrelated visual 'needles' into original videos. \nThe framework automates the generation of query-response pairs using predefined rules, minimizing manual labor.  The queries focus on specific aspects of video understanding, enabling more skill-specific evaluations. The separation between video content and the queries also allow for increased video variety and evaluations across different lengths.\nUtilizing VideoNIAH, we compile a video benchmark, VNBench, which includes tasks such as retrieval, ordering, and counting to evaluate three key aspects of video understanding: temporal perception, chronological ordering, and spatio-temporal coherence. We conduct a comprehensive evaluation of both proprietary and open-source models, uncovering significant differences in their video understanding capabilities across various tasks. Additionally, we perform an in-depth analysis of the test results and model configurations. Based on these findings, we provide some advice for improving video MLLM training, offering valuable insights to guide future research and model development."},"_bibtex":{"value":"@inproceedings{\nzhao2025needle,\ntitle={Needle In A Video Haystack: A Scalable  Synthetic Evaluator for Video {MLLM}s},\nauthor={Zijia Zhao and Haoyu Lu and Yuqi Huo and Yifan Du and Tongtian Yue and Longteng Guo and Bingning Wang and weipeng chen and Jing Liu},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=ZJo6Radbqq}\n}"},"title":{"value":"Needle In A Video Haystack: A Scalable  Synthetic Evaluator for Video MLLMs"},"pdf":{"value":"/pdf/875d019fb5070873ce0564e8a175d24d3df4dc3f.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"zhao|needle_in_a_video_haystack_a_scalable_synthetic_evaluator_for_video_mllms"},"authorids":{"value":["~Zijia_Zhao1","~Haoyu_Lu1","~Yuqi_Huo1","~Yifan_Du1","~Tongtian_Yue1","~Longteng_Guo1","~Bingning_Wang3","~weipeng_chen2","~Jing_Liu1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Zijia Zhao","Haoyu Lu","Yuqi Huo","Yifan Du","Tongtian Yue","Longteng Guo","Bingning Wang","weipeng chen","Jing Liu"]}},"version":2},{"content":{"summary":{"value":"This paper explores a critical challenge in deep learning for medical imaging: the use of 'shortcuts' in model predictions. By employing weight space correlation analysis, the authors investigate how models utilize non-target features. They conducted experiments on two specific datasets, the Fetal and Cervix,and demonstrate that this method can effectively identify whether metadata is being leveraged as a shortcut for predictions."},"justification_of_final_rating":{"value":"I am pleased with the planned edits to the PDF. This research topic is vital to the field of medical imaging, and further discussion is essential to help translate these findings into clinical practice. All the best!"},"confidence":{"value":5},"final_rating":{"value":4},"justification_of_the_preliminary_rating":{"value":"The paper addresses a highly relevant and timely problem in medical imaging: the lack of transparency regarding how deep learning models utilize metadata and shortcuts. The motivation is sound, the manuscript is well-written, and the use of clinically significant datasets (Fetal and Cervix) is a clear strength. However, the current version of the manuscript lacks the methodological depth and quantitative evidence necessary to fully support its conclusions. My preliminary rating is primarily driven by the weaknesses mentioned above."},"confidentiality_llm_acknowledgment":{"value":"Yes"},"strengths":{"value":"* The manuscript provides a clear motivation and a well-structured narrative.\n* Furthermore, the experimental results are supported by evaluations performed on two clinically relevant datasets.\n* The experimental framework includes baseline and multi-task configurations, utilizing induced bias to analyze model behavior and shortcut reliance."},"weaknesses":{"value":"* **Literature Review :** The introduction would benefit significantly from a more comprehensive discussion of existing literature regarding shortcut identification. Specifically, incorporating a discussion of PCA-based quantitative analysis and counterfactual studies would help the reader better position this study within the current state of the art.\n\n* **Methodological Rigor:** The methodological explanations require further detail. In particular, the rationale for the PCA projection of weights needs clarification. A more robust mathematical explanation is necessary to justify why projecting a weight vector onto the feature space of the backbone or data is an appropriate analytical approach.\n\n* **Quantitative Results:** The findings regarding shortcut utilization currently rely on qualitative assessments of color scales. Providing quantitative metrics would strengthen the results and allow for a more objective evaluation of the model's behavior.\n\n* **Support for Claims:** While the authors claim that the induced bias dataset demonstrates that metadata are utilized as shortcuts, this conclusion is not immediately evident from the provided figures. Additional evidence or clearer visualization is needed to support this assertion.\n\n* **Novelty and Comparative Analysis:** Given that this work builds upon the framework established by Glocker et al., the authors should explicitly clarify the advantages of their weight-based correlation analysis compared to the Kolmogorov-Smirnov (K-S) tests in the PCA space used in the original study."},"detailed_comments":{"value":"**Major Revisions**\n\nPlease refer to the \"Weaknesses\" section above. The primary concerns involve the need for a more rigorous mathematical justification of the weight projection method, a more comprehensive literature review to position the work, and the transition from qualitative visual assessments to quantitative metrics.\n\n**Minor Revisions**\n\nTypography: The quotation marks are inverted (flipped) in several instances throughout the manuscript. Please ensure that opening and closing marks follow standard LaTeX/typographic conventions (e.g., using `` and '').\n\nVisualization of Correlation Matrices: In Figure 1 and Figure 2, the visualization of the results could be significantly improved by removing the self-correlation diagonal (where $r = 1$)."},"questions_to_address_in_the_rebuttal":{"value":"1. Comparative Advantage: How does the proposed weight-based correlation analysis offer superior diagnostic power compared to the $K$-$S$ tests on the $PCA$ space utilized in the Glocker et al. framework?\n\n2. Could the authors provide quantitative metrics (such as correlation coefficients or $p$-values) to support the claims made regarding the figures? Relying on qualitative color-scale interpretations makes it difficult to assess the effect size of the shortcut utilization.\n\n3. Evidence of Induced Bias: could the authors provide a more granular breakdown or an alternative visualization that explicitly demonstrates the model's reliance on metadata? Currently, it is difficult to definitively conclude from the figures that metadata is the primary driver of the observed shortcuts."},"preliminary_rating":{"value":2}},"parentInvitations":"MIDL.io/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1771530043469,"tcdate":1767648842654,"writers":["MIDL.io/2026/Conference","MIDL.io/2026/Conference/Submission234/Reviewer_RHE6"],"signatures":["MIDL.io/2026/Conference/Submission234/Reviewer_RHE6"],"forum":"Os1ua32K56","number":1,"license":"CC BY 4.0","cdate":1767648842654,"readers":["everyone"],"invitations":["MIDL.io/2026/Conference/Submission234/-/Official_Review","MIDL.io/2026/Conference/-/Edit"],"mdate":1771530043469,"domain":"MIDL.io/2026/Conference","replyto":"Os1ua32K56","id":"lCurhRKQHS","forumContent":{"TLDR":{"value":"A method to quantify how much does a trained classifier model rely on irrelevant features in its prediction."},"venue":{"value":"MIDL 2026 Poster"},"midl_latex_submission_checklist":{"value":["The paper compiles correctly using the pdflatex compiler.","Created a single midl26_NNN.zip file with midl26_NNN.tex, midl26_NNN.bib, all necessary figures and files.","The LaTeX file includes the correct header commands before \\title.","Hyperref is preloaded; do not modify it.","The times package is not used.","All co-authors are listed with correct firstname, lastname, and any suffix/prefix.","Math in the title and abstract uses valid LaTeX.","References are provided through the .bib file only.","Tables and figures stay within the page margins.","The zip archive contains all necessary figures and no unused files.","Special formatting from rebuttal has been removed.","All special characters use LaTeX commands.","Appendices and supplementary material are included in the same PDF after references.","The main paper does not exceed 12 pages."]},"keywords":{"value":["shortcut learning","feature utilization","obstetric ultrasound"]},"reproducibility":{"value":"https://github.com/wong-ck/wsc-analysis"},"read_cfp_and_author_instructions":{"value":"Yes"},"originality_policy":{"value":"Yes"},"abstract":{"value":"Deep learning models in medical imaging are susceptible to shortcut learning, relying on confounding metadata (e.g. scanner model) that is often encoded in image embeddings. The crucial question is whether the model actively utilizes this encoded information for its final prediction. We introduce Weight Space Correlation analysis, an interpretable methodology that quantifies feature utilization by measuring the alignment between the classification heads of a primary clinical task and auxiliary metadata tasks. We first validate our method by successfully detecting artificially induced shortcut learning. We then apply it to probe the feature utilization of an SA-SonoNet model trained for Spontaneous Preterm Birth (sPTB) prediction. Our analysis confirmed that while the embeddings contain substantial metadata, the sPTB classifier's weight vectors were highly correlated with clinically relevant factors (e.g. cervical length) but decoupled from clinically irrelevant acquisition factors (e.g. scanner). Our methodology provides a tool for verifying model trustworthiness, by inspecting whether it utilizes features unrelated to the genuine clinical signal. Code available at https://github.com/wong-ck/wsc-analysis"},"_bibtex":{"value":"@inproceedings{\nwong2026weight,\ntitle={Weight Space Correlation Analysis: Quantifying Feature Utilization in Deep Learning Models},\nauthor={Chun Kit Wong and Paraskevas Pegios and Nina Weng and Emilie Pi Fogtmann Sejer and Martin Gr{\\o}nneb{\\ae}k Tolsgaard and Anders Nymark Christensen and Aasa Feragen},\nbooktitle={Medical Imaging with Deep Learning},\nyear={2026},\nurl={https://openreview.net/forum?id=Os1ua32K56}\n}"},"title":{"value":"Weight Space Correlation Analysis: Quantifying Feature Utilization in Deep Learning Models"},"latex_code":{"value":"/attachment/f0d19d5ca19899663f0003a6f73ab0e9f29d648b.zip"},"secondary_subject_area":{"value":"Integration of Imaging and Clinical Data"},"pdf":{"value":"/pdf/800af085a90aae8f9da5b0a8f0144eae85aa7c75.pdf"},"copyright_form":{"value":"/attachment/068ac4af1df9743d3f284cec1fa4b58b3db34226.pdf"},"visa":{"value":"No"},"single_blind_notice":{"value":"Yes"},"venueid":{"value":"MIDL.io/2026/Conference"},"paperhash":{"value":"wong|weight_space_correlation_analysis_quantifying_feature_utilization_in_deep_learning_models"},"primary_subject_area":{"value":"Fairness and Bias"},"authorids":{"value":["~Chun_Kit_Wong1","~Paraskevas_Pegios1","~Nina_Weng1","~Emilie_Pi_Fogtmann_Sejer1","~Martin_Grønnebæk_Tolsgaard1","~Anders_Nymark_Christensen1","~Aasa_Feragen2"]},"registration":{"value":"Yes"},"authors":{"value":["Chun Kit Wong","Paraskevas Pegios","Nina Weng","Emilie Pi Fogtmann Sejer","Martin Grønnebæk Tolsgaard","Anders Nymark Christensen","Aasa Feragen"]},"llm_policy_acknowledgment":{"value":"Yes"}},"version":2},{"content":{"summary":{"value":"The paper proposes Vid2World, a framework that turns a single video into a 4D (3D + time) dynamic scene representation. Unlike typical video diffusion models that just generate future frames, this method tries to reconstruct a consistent, camera-aware world that can be rendered from novel viewpoints over time. Technically, the authors build on top of DynamiCrafter and modify it for causal, autoregressive video generation with action conditioning. They introduce weight-transfer tricks and causal attention masking so that the model can generate frames one by one while maintaining spatial-temporal coherence."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"see weaknesses."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"The idea of turning raw videos into a dynamic 4D world is compelling. Most video models only hallucinate future frames from a fixed camera, so explicitly reconstructing a consistent world is an interesting step forward.\n\nI appreciate that they go beyond synthetic datasets and test on RT-1 real robot data and a large-scale human gameplay dataset (CS:GO). It shows they’re not only targeting clean, toy data.\n\nThe architectural changes to turn an offline video diffusion model into a causal and action-aware generator are reasonable. Especially, I think the “extrapolative weight transfer” idea for converting bidirectional attention layers into causal ones is practical.\n\nThe results (especially novel-view video rollout) are visually impressive and clearly better than standard video diffusion models that ignore camera motion."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"No true 3D ground-truth evaluation.\nThe method claims to build a 4D world, but there’s no quantitative evaluation comparing it to 3D scene reconstruction methods (like NeRF, BANMo, DynamicNeRF, etc.). Most results are still evaluated only on video quality metrics or visuals. It's hard to tell if the “world” is actually geometrically consistent or just looks plausible.\n\nMostly relies on pre-trained models.\nA large part of the pipeline inherits DynamiCrafter. The core novelty is in causalization and conditioning, but I sometimes felt like the paper is more of a clever adaptation rather than a fundamentally new representation of dynamic scenes.\n\nAction conditioning is not strongly validated.\nThey claim the model can take actions and simulate their effects in the generated world, but I didn’t see strong evidence that the generated world actually follows the action semantics accurately, especially in RT-1 tasks. It feels closer to action-guided video synthesis than true action-driven world modeling.\n\nNo comparison to video-to-3D methods.\nWorks like Nerfies, Vid2NeRF, DynIBaR also reconstruct dynamic scenes from monocular videos. It would help to position Vid2World against those rather than only comparing to video diffusion baselines.\n\nComputation is heavy and unclear.\nTraining takes 4×A100 for 100k steps with 3 FPS video clips. The paper doesn’t report inference speed or memory usage. If the goal is a usable world model, some efficiency discussion would help."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762917615849,"tcdate":1761964596216,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission4851/Reviewer_XNTH"],"signatures":["ICLR.cc/2026/Conference/Submission4851/Reviewer_XNTH"],"forum":"pFyzqbUiF9","number":2,"license":"CC BY 4.0","cdate":1761964596216,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission4851/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762917615849,"domain":"ICLR.cc/2026/Conference","replyto":"pFyzqbUiF9","id":"5XrtxttCgf","forumContent":{"TLDR":{"value":"We propose Vid2World,  a general approach for leveraging and transferring pre-trained video diffusion models into interactive world models."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["World Models","Video Diffusion Models"]},"primary_area":{"value":"foundation or frontier models, including LLMs"},"abstract":{"value":"World models, which predict future transitions from past observation and action sequences, have shown great promise for improving data efficiency in sequential decision-making. However, existing world models often require extensive domain-specific training and still produce low-fidelity, coarse predictions, limiting their usefulness in complex environments. In contrast, video diffusion models trained on large-scale internet data have demonstrated impressive capabilities in generating high-quality videos that capture diverse real-world dynamics. In this work, we present _Vid2World_, a general approach for leveraging and transferring pre-trained video diffusion models into interactive world models. To bridge the gap, Vid2World systematically explores _video diffusion causalization_, reshaping both the architecture and training objective of pre-trained models to enable autoregressive generation. Additionally, it incorporates a _causal action guidance_ mechanism to enhance action controllability in the resulting interactive world models. Extensive experiments across multiple domains, including robot manipulation, 3D game simulation, and open-world navigation, demonstrate that our method offers a scalable and effective pathway for repurposing highly capable video diffusion models into interactive world models."},"_bibtex":{"value":"@inproceedings{\nhuang2026vidworld,\ntitle={Vid2World: Crafting Video Diffusion Models to Interactive World Models},\nauthor={Siqiao Huang and Jialong Wu and Qixing Zhou and Shangchen Miao and Mingsheng Long},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=pFyzqbUiF9}\n}"},"title":{"value":"Vid2World: Crafting Video Diffusion Models to Interactive World Models"},"pdf":{"value":"/pdf/bb7427c741f3d9f284a8429b4fc25037b01b4e8a.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"huang|vid2world_crafting_video_diffusion_models_to_interactive_world_models"},"authorids":{"value":["~Siqiao_Huang1","~Jialong_Wu1","~Qixing_Zhou1","~Shangchen_Miao1","~Mingsheng_Long5"]},"authors":{"value":["Siqiao Huang","Jialong Wu","Qixing Zhou","Shangchen Miao","Mingsheng Long"]}},"version":2},{"content":{"summary":{"value":"The paper proposes the concept of minimal repair (MR) and almost minimal repair (AMR) for learning models over datasets with missing values. The core idea is to identify and impute a minimal subset of missing values that are sufficient to achieve an accurate model, thus reducing the time and computational resources typically required for full imputation. The authors provide theoretical foundations, proving that finding minimal repairs for SVM and linear regression is NP-hard, and propose efficient approximation algorithms with provable error bounds. Through experiments on real-world datasets, the authors demonstrate that their methods (MR and AMR) can reduce imputation time and manual effort without significantly compromising model accuracy."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"Q1. Can the authors show any significant performance improvements when applying MR/AMR to more complex models (e.g., neural networks) or in real-world domains such as healthcare or finance?\n\nQ2. How does the approach scale with high-dimensional datasets (e.g., millions of features or samples)? Can the authors provide additional scalability experiments?\n\nQ3. The paper mentions that the algorithms may be computationally expensive. Could the authors offer further justification for why MR/AMR would be preferred over simpler methods in scenarios where full imputation methods are already cheap to compute?\n\nQ4. What would happen to the performance if the imputation model is poor? Are there any safeguards or adjustments the authors propose to ensure model robustness?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"S1. Missing data is a critical issue in real-world datasets, and addressing this problem with minimal imputation is valuable.\n\nS2. The formal definitions of minimal and almost minimal repairs are clear and well-explained.\n\nS3. The paper introduces approximation algorithms with provable error bounds, which is an important step in dealing with the NP-hard problem of finding minimal repairs."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"W1. The core idea of minimal repair has been explored in prior work, especially in ActiveClean and Certain Model Learning, which already aim to reduce the cost of imputation without sacrificing model accuracy. The paper does not provide significant improvements or innovations to justify its claims as groundbreaking.\n\nW2. The experiments are based on a narrow set of models (SVM and linear regression) and imputation methods (KNN, MICE). While these are standard, the results don’t convincingly show why MR/AMR is preferable in practice. The paper would have been stronger with a broader evaluation that includes more recent models or deep learning applications, where missing data issues are also common.\n\nW3. Despite the claims of efficiency, the algorithms for finding MR and AMR can still be computationally expensive, especially in high-dimensional datasets. The claim that MR/AMR will always be faster than full imputation is questionable, particularly when simpler imputation methods like KNN are used. The paper does not sufficiently demonstrate the practical computational benefits for large-scale datasets.\n\nW4. While the NP-hardness results are a nice theoretical contribution, they detract from the practical usability of the method. The algorithms for minimal repair are highly complex, and the paper does not present sufficient evidence to suggest they would be widely applicable in real-world data cleaning scenarios.\n\nW5. The formalism and technical details make the paper hard to follow, especially for readers without a deep background in optimization and data imputation. The dense presentation of the experimental results in the tables, combined with jargon-heavy theoretical discussions, makes the paper less accessible and less impactful."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762921356752,"tcdate":1761967976394,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission9896/Reviewer_21Zj"],"signatures":["ICLR.cc/2026/Conference/Submission9896/Reviewer_21Zj"],"forum":"GCVnFrF9ok","number":4,"license":"CC BY 4.0","cdate":1761967976394,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission9896/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762921356752,"domain":"ICLR.cc/2026/Conference","replyto":"GCVnFrF9ok","id":"aDuqzBQWlP","forumContent":{"TLDR":{"value":"We demonstrate a new approach to learn accurate machine learning models over incomplete data with the minimal or almost imputation effort."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["ML over incomplete data","Data imputation for ML","Supervised ML"]},"primary_area":{"value":"unsupervised, self-supervised, semi-supervised, and supervised representation learning"},"abstract":{"value":"Missing data often exists in real-world datasets, requiring significant time and effort for data repair to learn accurate machine learning (ML) models. In this paper, we show that imputing all missing values is not always necessary to achieve an accurate ML model. We introduce concepts of minimal and almost minimal repair, which are subsets of missing data items in training data whose imputation delivers accurate and reasonably accurate models, respectively. Repairing these sets can significantly reduce the time, computational resources, and manual effort required for learning models. We show that finding these sets is NP-hard for SVM and linear regression and propose efficient approximation algorithms with provable error bounds. Our extensive experiments indicate that our proposed algorithms can substantially reduce the time and effort required to learn on incomplete datasets."},"_bibtex":{"value":"@misc{\nzhen2026minimal,\ntitle={Minimal Repairs for Learning Over Incomplete Data},\nauthor={Cheng Zhen and Prayoga and Nischal Aryal and Arash Termehchy and Garrett Biwer and Lubna Alzamil},\nyear={2026},\nurl={https://openreview.net/forum?id=GCVnFrF9ok}\n}"},"title":{"value":"Minimal Repairs for Learning Over Incomplete Data"},"pdf":{"value":"/pdf/0e0bd9b0bcd90113edee10af7511090574d26430.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"zhen|minimal_repairs_for_learning_over_incomplete_data"},"authorids":{"value":["~Cheng_Zhen2","~Prayoga1","~Nischal_Aryal1","~Arash_Termehchy3","~Garrett_Biwer1","~Lubna_Alzamil1"]},"authors":{"value":["Cheng Zhen","Prayoga","Nischal Aryal","Arash Termehchy","Garrett Biwer","Lubna Alzamil"]}},"version":2},{"content":{"TLDR":{"value":"We introduce domain-organized minimal-pair benchmarks for vision-language alignment and show that alignment varies substantially across semantic domains and levels of model training."},"venue":{"value":"ICLR 2027 Conference Submission"},"keywords":{"value":["Vision-language alignment","Cross-modal representation","Minimal-pair benchmarks","Multimodal models"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"To what extent does alignment between vision and language preserve specific semantic distinctions? We introduce a new benchmark that examines this question across shallow mappings between independently trained vision and language models, CLIP-style contrastive models, and generative vision-language models. Each benchmark item leverages a minimal pair (of pairs) design comprising two minimally contrasting images and two corresponding captions, with contrasts spanning object properties and relations, physical events, and social interactions. By testing whether similarity scores or matching decisions consistently favor the correct image–caption correspondences over the mismatched alternatives, the benchmark probes whether alignment preserves the specific semantic distinction targeted by each item. Across these settings, we find that alignment is highly conditional. \nShallow unimodal alignment recovers some shared object and scene structure, but performs poorly on minimal pairs that require preserving specific relational, causal, or social distinctions. Multimodal training improves performance, but gaps remain even for stronger models compared to human baselines. Further analyses show that implicit captions are consistently harder than explicit captions, suggesting that current models are better at matching named visual properties than inferring latent consequences from world knowledge. Together, our results suggest that broad semantic structure may support cross-modal matching, but fine-grained inferential distinctions needed to match closely related images and captions remain fragile even in multimodal systems."},"_bibtex":{"value":"@inproceedings{\nanonymous2026probing,\ntitle={Probing Fine-Grained Semantic Alignment Across Vision and Language with Minimal Pairs},\nauthor={Anonymous},\nbooktitle={Submitted to The Fifteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=S9DxRuoIcA},\nnote={under review}\n}"},"title":{"value":"Probing Fine-Grained Semantic Alignment Across Vision and Language with Minimal Pairs"},"pdf":{"value":"/pdf/ad8d6b3b6aa53d1f68f0a0cad43b4225a7f6fe25.pdf"},"venueid":{"value":"ICLR.cc/2027/Conference/Submission"}},"tmdate":1791235292777,"tcdate":1789807288251,"writers":["ICLR.cc/2027/Conference","ICLR.cc/2027/Conference/Submission57846/Authors"],"signatures":["ICLR.cc/2027/Conference/Submission57846/Authors"],"forum":"S9DxRuoIcA","license":"CC BY 4.0","number":57846,"cdate":1789807288251,"readers":["everyone"],"invitations":["ICLR.cc/2027/Conference/-/Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Bidding","ICLR.cc/2027/Conference/Submission57846/-/Full_Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Reviewing","ICLR.cc/2027/Conference/-/Edit"],"mdate":1791235292777,"odate":1791031328286,"domain":"ICLR.cc/2027/Conference","id":"S9DxRuoIcA","version":2},{"content":{"summary":{"value":"The paper proposes a zero-shot text-to-video generating pipeline called Free-Bloom, which first use LLM to generate prompt sequence decribing frames in a video, then generate frames according the prompts. To enhancing coherence, the authors proposed joint noise sampling, step-aware attention shift and dual-path interpolation. The authors compare their method with other video generator quantitatively (Clip metrics, user study) and qualitatively. "},"soundness":{"value":"2 fair"},"confidence":{"value":"4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work."},"questions":{"value":"1. Please provide more clarification of experiments, including prompt engineering to get frame description, hyperparamters (such as \\tau, \\tau*)\n\n2. The coherence is not good, but the idea of leveraging LLM is interesting. Is it possible to combine LLM director with other existing pretrained text-to-video methods (e.g.  make-a-video) to enhance conherence?"},"rating":{"value":"5: Borderline accept: Technically solid paper where reasons to accept outweigh reasons to reject, e.g., limited evaluation. Please use sparingly."},"code_of_conduct":{"value":"Yes"},"presentation":{"value":"2 fair"},"contribution":{"value":"2 fair"},"strengths":{"value":"1. explore zero-shot text-to-video generation LLM director, which leverage the story generation ability of LLM to generate semantic meaningful frame sequences. \n2. Good per-frame quality. \n3. tried several methods to enhancing coherence of zero-shot video generation. "},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. insufficient ablation study. \n\n   a. there's only two qualitative cases in ablation study, which is not convincing enough. More cases and quantitative results (e.g. CLIP metrics) should be provided. \n\n   b. given that joint noise sampling, step-aware attention shift and dual-path interpolation are universal technics for zero-shot video generation, experiment result of combining these technics and past zero-shot methods (e.g. text2video-zero) is needed. \n\n2. The coherence is not good enough. e.g. sudden changes in the color / indentity can be observed. There is a fatal flaw in technical design: authors set m(t) to 1 and discard attention shift when t is small, which may lead to strong incoherence. \n\n3. Insufficient clarification of experiments, e.g. prompts to LLM, hyperparameters. "},"limitations":{"value":"see weakness and questions."}},"nonreaders":[],"tmdate":1702410780511,"tcdate":1688634706607,"writers":["NeurIPS.cc/2023/Conference","NeurIPS.cc/2023/Conference/Submission1312/Reviewer_VS1k"],"signatures":["NeurIPS.cc/2023/Conference/Submission1312/Reviewer_VS1k"],"forum":"paa2OU5jN8","number":4,"license":"CC BY 4.0","cdate":1688634706607,"mdate":1702410780511,"readers":["everyone"],"invitations":["NeurIPS.cc/2023/Conference/Submission1312/-/Official_Review","NeurIPS.cc/2023/Conference/-/Edit"],"domain":"NeurIPS.cc/2023/Conference","replyto":"paa2OU5jN8","id":"RuGkkAOpQO","forumContent":{"venue":{"value":"NeurIPS 2023 poster"},"keywords":{"value":["Text-to-Video","Zero-Shot Generation","Large Language Model","Latent Diffusion Models"]},"supplementary_material":{"value":"/attachment/f2d0c5a77c34a820786ffe50cd192e361bcd6be7.zip"},"_bibtex":{"value":"@inproceedings{\nhuang2023freebloom,\ntitle={Free-Bloom: Zero-Shot Text-to-Video Generator with {LLM} Director and {LDM} Animator},\nauthor={Hanzhuo Huang and Yufan Feng and Cheng Shi and Lan Xu and Jingyi Yu and Sibei Yang},\nbooktitle={Thirty-seventh Conference on Neural Information Processing Systems},\nyear={2023},\nurl={https://openreview.net/forum?id=paa2OU5jN8}\n}"},"title":{"value":"Free-Bloom: Zero-Shot Text-to-Video Generator with LLM Director and LDM Animator"},"paperhash":{"value":"huang|freebloom_zeroshot_texttovideo_generator_with_llm_director_and_ldm_animator"},"TLDR":{"value":"This work introduces Free-Bloom, a zero-shot, training-free, semantic-coherent text-to-video generator that is capable of generating high-quality, temporally consistent, and semantically aligned videos."},"abstract":{"value":"Text-to-video is a rapidly growing research area that aims to generate a semantic, identical, and temporal coherence sequence of frames that accurately align with the input text prompt. This study focuses on zero-shot text-to-video generation considering the data- and cost-efficient. To generate a semantic-coherent video, exhibiting a rich portrayal of temporal semantics such as the whole process of flower blooming rather than a set of ``moving images'', we propose a novel Free-Bloom pipeline that harnesses large language models (LLMs) as the director to generate a semantic-coherence prompt sequence, while pre-trained latent diffusion models (LDMs) as the animator to generate the high fidelity frames. Furthermore, to ensure temporal and identical coherence while maintaining semantic coherence, we propose a series of annotative modifications to adapting LDMs in the reverse process, including joint noise sampling, step-aware attention shift, and dual-path interpolation. Without any video data and training requirements, Free-Bloom generates vivid and high-quality videos, awe-inspiring in generating complex scenes with semantic meaningful frame sequences.  In addition, Free-Bloom is naturally compatible with LDMs-based extensions."},"pdf":{"value":"/pdf/303a32c6185f4e35880879ec412f25eccd89a86c.pdf"},"venueid":{"value":"NeurIPS.cc/2023/Conference"},"authorids":{"value":["~Hanzhuo_Huang1","~Yufan_Feng1","~Cheng_Shi4","~Lan_Xu2","~Jingyi_Yu5","~Sibei_Yang1"]},"authors":{"value":["Hanzhuo Huang","Yufan Feng","Cheng Shi","Lan Xu","Jingyi Yu","Sibei Yang"]}},"version":2},{"content":{"summary":{"value":"This paper targets video editing by proposing a video-to-paragraph-to-video pipeline. The editing is achieved based on the modification of the paragraph.  \nThe contribution lies in (1) a novel video editing pipeline, and the editing is based on detailed text descriptions, which is user-friendly. (2) improved detailed video captioning performance by a novel multi-granular pooling. (3) A new dataset with high-detailed captions and object caption-mask pairs to facilitate this framework."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"The text encoder used for P2V is Clip? Does the max token length be smaller than the input token length? Since the paragraph description is very long, it may exceed the maximum token length of the text encoder. If so, will the information lost will affect the generation sematics?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"- The pipeline for video editing is novel and intuitive. The idea of leveraging the recent development of MLLMs to tackle video editing tasks is interesting.  \n- The experiments are comprehensive. The results of single object prediction are good and outperform some strong baselines. The results of video editing show remarkable improvement.  \n- The source code is provided in the supplementary materials.\n- The method can be integrated with an inversion-based video editing method, and the results are provided.   \n- Limitations are discussed."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The author can consider providing more video results to demonstrate the actual editing performance in different settings. For example, the video comparison results with Fatezero and Token flow. Also the results of the proposed method with VideoCrafter and DynamiCrafter."}},"nonreaders":[],"tmdate":1731427999713,"tcdate":1730718201331,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission5042/Reviewer_vHvK"],"signatures":["ICLR.cc/2025/Conference/Submission5042/Reviewer_vHvK"],"forum":"qnGir4dyu9","number":4,"license":"CC BY 4.0","cdate":1730718201331,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission5042/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427999713,"domain":"ICLR.cc/2025/Conference","replyto":"qnGir4dyu9","id":"kRQXKsgREg","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video Inpainting","Video Editing","Video Caption","Multimodal Large Language Model"]},"supplementary_material":{"value":"/attachment/9c47b8f3ae2d08423fd5f0d8fb34f9be2dd9a9e1.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Recent video generative models primarily rely on carefully written text prompts for specific tasks, like inpainting or style editing. They require labor-intensive textual descriptions for input videos, hindering their flexibility to adapt personal/raw videos to user specifications. This paper proposes RACCooN, a versatile and user-friendly video-to-paragraph-to-video generative framework that supports multiple video editing capabilities, such as removal, addition, and modification, through a unified pipeline. RACCooN consists of two principal stages: Video-to-Paragraph (V2P) and Paragraph-to-Video (P2V). In the V2P stage, we automatically describe video scenes in well-structured natural language, capturing both the holistic context and focused object details. Subsequently, in the P2V stage, users can optionally refine these descriptions to guide the video diffusion model, enabling various modifications to the input video, such as removing, changing subjects, and/or adding new objects. The proposed approach stands out from other methods through several significant contributions: (1) RACCooN suggests a multi-granular spatiotemporal pooling strategy to generate well-structured video descriptions, capturing both the broad context and object details without requiring complex human annotations, simplifying precise video content editing based on text for users. (2) Our video generative model incorporates auto-generated narratives or instructions to enhance the quality and accuracy of the generated content. (3) RACCooN also plans to imagine new objects in a given video, so users simply prompt the model to receive a detailed video editing plan for complex video editing. The proposed framework demonstrates impressive versatile capabilities in video-to-paragraph generation (up to 9.4% absolute improvement in human evaluations against the baseline), video content editing (relative 49.7% in FVD), and can be incorporated into other SoTA video generative models for further enhancement."},"_bibtex":{"value":"@misc{\nyoon2025raccoon,\ntitle={{RACC}ooN: A Versatile Instructional Video Editing Framework with Auto-Generated Narratives},\nauthor={Jaehong Yoon and Shoubin Yu and Mohit Bansal},\nyear={2025},\nurl={https://openreview.net/forum?id=qnGir4dyu9}\n}"},"title":{"value":"RACCooN: A Versatile Instructional Video Editing Framework with Auto-Generated Narratives"},"pdf":{"value":"/pdf/92793d9a9c07981b6886e1172d1568eebdd6d952.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"yoon|raccoon_a_versatile_instructional_video_editing_framework_with_autogenerated_narratives"},"authorids":{"value":["~Jaehong_Yoon1","~Shoubin_Yu1","~Mohit_Bansal2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Jaehong Yoon","Shoubin Yu","Mohit Bansal"]}},"version":2},{"content":{"summary":{"value":"This work introduces LaVie, a text-to-video generative model that builds upon a pre-trained text-to-image model.  LaVie consists of cascaded video latent diffusion models, including a base T2V model, a temporal interpolation model, and a video super-resolution model. The key insights include the use of temporal self-attentions and rotary positional encoding to capture temporal correlations in video data. Joint image-video fine-tuning is crucial for high-quality results. The authors also contribute a large and diverse video dataset called Vimeo25M. Experimental results demonstrate that LaVie achieves state-of-the-art performance in both quantitative and qualitative measuress."},"presentation":{"value":"2 fair"},"contribution":{"value":"2 fair"},"soundness":{"value":"2 fair"},"strengths":{"value":"The article provides some insights: \n1. The use of simple temporal self-attention mechanisms coupled with rotary positional encoding to capture temporal correlations in video data \n2. Joint image-video fine-tuning in producing high-quality and creative outcomes.\n3. The authors contribute the Vimeo25M dataset, which comprises 25 million text-video pairs, serving as a valuable resource for research and development in the field."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. I think the article lacks technical innovation. Its core contribution involves extending the pre-trained LDM to a video generation model, introducing temporal attention, and emphasizing joint image-video training. However, similar techniques have been mentioned in previous works, such as VDM and Align your latent. The authors should clarify the distinguishing aspects of these innovation. \n\n2. The article presents several assertions without strong empirical support or ablation study. For instance, the introduction of RoPE as a positional encoding lacks corresponding experimental evidence to demonstrate its effectiveness and advantages over other encoding methods. The statement, \"we validate that the process of joint image-video fine-tuning plays a pivotal role,\" lacks experimental substantiation in the article.\n\n3. The writing should be improved, and there exist some unclear explanations. For instance, the modules in Figure 2 are not adequately introduced in the text, causing confusion. For example,  I want to know whether 'E' denotes the encoder or denoiser? Additionally, the article mentions \"By applying rigorous filtering criteria\" to construct Vimeo25M, but the specific criteria are not outlined. Will this dataset be made publicly available?"},"confidence":{"value":"4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work."},"questions":{"value":"See Weaknesses."},"rating":{"value":"5: marginally below the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1701427649694,"tcdate":1698570451686,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission2758/Reviewer_9GLX"],"signatures":["ICLR.cc/2024/Conference/Submission2758/Reviewer_9GLX"],"forum":"p09XyFxZkc","number":1,"license":"CC BY 4.0","cdate":1698570451686,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission2758/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1701427649694,"domain":"ICLR.cc/2024/Conference","replyto":"p09XyFxZkc","id":"0vzyQ1peFX","forumContent":{"venue":{"value":"Submitted to ICLR 2024"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["text-to-video generation","diffusion models"]},"supplementary_material":{"value":"/attachment/9b7be39ae37699c99be3e7c7ae7b8d3b62214416.zip"},"primary_area":{"value":"generative models"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"This work aims to learn a high-quality text-to-video (T2V) generative model by leveraging a pre-trained text-to-image (T2I) model as a basis. It is a highly desirable yet challenging task to simultaneously a) accomplish the synthesis of visually realistic and temporally coherent videos while b) preserving the strong creative generation nature of the pre-trained T2I model. To this end, we propose LaVie, an integrated video generation framework that operates on cascaded video latent diffusion models, comprising a base T2V model, a temporal interpolation model, and a video super-resolution model. Our key insights are two-fold: 1) We reveal that the incorporation of simple temporal self-attentions, coupled with relative positional encoding, adequately captures the temporal correlations inherent in video data. 2) Additionally, we validate that the process of joint image-video fine-tuning plays a pivotal role in producing high-quality and creative outcomes. To enhance the performance of LaVie, we contribute a comprehensive and diverse video dataset named Vimeo25M, consisting of 25 million text-video pairs that prioritize quality, diversity, and aesthetic appeal. Extensive experiments demonstrate that LaVie achieves state-of-the-art performance both quantitatively and qualitatively. Furthermore, we showcase the versatility of pre-trained LaVie models in various long video generation and personalized video synthesis applications."},"_bibtex":{"value":"@misc{\nwang2024lavie,\ntitle={LaVie: High-Quality Video Generation with Cascaded Latent Diffusion Models},\nauthor={Yaohui Wang and Xinyuan Chen and Xin Ma and Shangchen Zhou and Ziqi Huang and Yi Wang and Ceyuan Yang and Yinan He and Jiashuo Yu and Peiqing Yang and Yuwei Guo and Tianxing Wu and Chenyang Si and Yuming Jiang and Cunjian Chen and Chen Change Loy and Bo Dai and Dahua Lin and Yu Qiao and Ziwei Liu},\nyear={2024},\nurl={https://openreview.net/forum?id=p09XyFxZkc}\n}"},"title":{"value":"LaVie: High-Quality Video Generation with Cascaded Latent Diffusion Models"},"pdf":{"value":"/pdf/f40f1ea778821a8640f919df133835970623a510.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Rejected_Submission"},"paperhash":{"value":"wang|lavie_highquality_video_generation_with_cascaded_latent_diffusion_models"},"authorids":{"value":["~Yaohui_Wang1","~Xinyuan_Chen1","~Xin_Ma3","~Shangchen_Zhou1","~Ziqi_Huang2","~Yi_Wang19","~Ceyuan_Yang2","~Yinan_He1","~Jiashuo_Yu1","~Peiqing_Yang1","~Yuwei_Guo1","~Tianxing_Wu2","~Chenyang_Si2","~Yuming_Jiang1","~Cunjian_Chen2","~Chen_Change_Loy2","~Bo_Dai2","~Dahua_Lin1","~Yu_Qiao1","~Ziwei_Liu1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Yaohui Wang","Xinyuan Chen","Xin Ma","Shangchen Zhou","Ziqi Huang","Yi Wang","Ceyuan Yang","Yinan He","Jiashuo Yu","Peiqing Yang","Yuwei Guo","Tianxing Wu","Chenyang Si","Yuming Jiang","Cunjian Chen","Chen Change Loy","Bo Dai","Dahua Lin","Yu Qiao","Ziwei Liu"]}},"version":2},{"content":{"TLDR":{"value":"We studied the effect of treating images from the same video as positive pairs for contrastive and non-contrastive learning with medical ultrasound, finding that it improved performance on some downstream tasks but not for others."},"venue":{"value":"NeurIPS 2024 Workshop SSL"},"pdf":{"value":"/pdf/34473962e72bc826f7124945df3f50be7e181d08.pdf"},"keywords":{"value":["self-supervised learning","contrastive learning","non-contrastive learning","ultrasound","video"]},"venueid":{"value":"NeurIPS.cc/2024/Workshop/SSL"},"paperhash":{"value":"vanberlo|intravideo_positive_pairs_in_selfsupervised_learning_for_ultrasound"},"authorids":{"value":["~Blake_VanBerlo1","~Alexander_Wong2","~Jesse_Hoey1","~Robert_Arntfield1"]},"abstract":{"value":"The videographic nature of ultrasound offers flexibility for defining the similarity relationship between pairs of images for self-supervised learning (SSL). In this study, we investigated the effect of utilizing proximal, distinct images from the same ultrasound video as pairs for joint embedding SSL. Additionally, we introduced a sample weighting scheme that increases the weight of closer image pairs and demonstrated how it can be integrated into SSL objectives. Named Intra-Video Positive Pairs (IVPP), the method surpassed previous ultrasound-specific contrastive learning methods’ average test accuracy on COVID-19 classification with the POCUS dataset by ≥ 1.3%. Investigations revealed that some combinations of IVPP hyperparameters can lead to improved or worsened performance, depending on the downstream task"},"_bibtex":{"value":"@inproceedings{\nvanberlo2024intravideo,\ntitle={Intra-video Positive Pairs in Self-Supervised Learning for Ultrasound},\nauthor={Blake VanBerlo and Alexander Wong and Jesse Hoey and Robert Arntfield},\nbooktitle={NeurIPS 2024 Workshop: Self-Supervised Learning - Theory and Practice},\nyear={2024},\nurl={https://openreview.net/forum?id=VpB4NrileF}\n}"},"title":{"value":"Intra-video Positive Pairs in Self-Supervised Learning for Ultrasound"},"authors":{"value":["Blake VanBerlo","Alexander Wong","Jesse Hoey","Robert Arntfield"]}},"tmdate":1733165447212,"pdate":1728801925653,"tcdate":1726084797204,"writers":["NeurIPS.cc/2024/Workshop/SSL","NeurIPS.cc/2024/Workshop/SSL/Submission38/Authors"],"signatures":["NeurIPS.cc/2024/Workshop/SSL/Submission38/Authors"],"forum":"VpB4NrileF","license":"CC BY 4.0","number":38,"cdate":1726084797204,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Workshop/SSL/-/Submission","NeurIPS.cc/2024/Workshop/SSL/-/Post_Submission","NeurIPS.cc/2024/Workshop/SSL/-/Edit"],"mdate":1733165447212,"odate":1733165447175,"domain":"NeurIPS.cc/2024/Workshop/SSL","id":"VpB4NrileF","version":2},{"content":{"venue":{"value":"Video-Langauge Models Poster"},"pdf":{"value":"/pdf/2c15c2b9b9c80c7a78a3c7401fa8e12b293d94a0.pdf"},"keywords":{"value":["Multi-Agent System","Action Anticipation","Vision Transformer","Memory Information Retrieval"]},"venueid":{"value":"NeurIPS.cc/2024/Workshop/Video-Langauge_Models"},"paperhash":{"value":"wang|himemformer_hierarchical_memoryaware_transformer_for_multiagent_action_anticipation"},"authorids":{"value":["~Zirui_Wang13","~Xinran_Zhao1","~Simon_Stepputtis1","~Woojun_Kim1","~Tongshuang_Wu1","~Katia_P._Sycara1","~Yaqi_Xie1"]},"abstract":{"value":"Understanding and predicting human actions has been a long-standing challenge and is a crucial measure of perception in robotics AI. While significant progress has been made in anticipating the future actions of individual agents, prior work has largely overlooked a key aspect of real-world human activity -- interactions. To address this gap in human-like forecasting within multi-agent environments, we present the Hierarchical Memory-Aware Transformer (HiMemFormer), a transformer-based model for online multi-agent action anticipation. HiMemFormer integrates and distributes global memory that captures joint historical information across all agents through a transformer framework, with a hierarchical local memory decoder that interprets agent-specific features based on these global representations using a coarse-to-fine strategy. In contrast to previous approaches, HiMemFormer uniquely hierarchically applies the global context with agent-specific preferences to avoid noisy or redundant information in multi-agent action anticipation. Extensive experiments on various multi-agent scenarios demonstrate the significant performance of HiMemFormer, compared with other state-of-the-art methods."},"_bibtex":{"value":"@inproceedings{\nwang2025himemformer,\ntitle={HiMemFormer:  Hierarchical Memory-Aware Transformer for Multi-Agent Action Anticipation},\nauthor={Zirui Wang and Xinran Zhao and Simon Stepputtis and Woojun Kim and Tongshuang Wu and Katia P. Sycara and Yaqi Xie},\nbooktitle={Workshop on Video-Language Models @ NeurIPS 2024},\nyear={2025},\nurl={https://openreview.net/forum?id=MIY7VeeMmq}\n}"},"title":{"value":"HiMemFormer:  Hierarchical Memory-Aware Transformer for Multi-Agent Action Anticipation"},"track":{"value":"Short Paper Track (up to 3 pages)"},"authors":{"value":["Zirui Wang","Xinran Zhao","Simon Stepputtis","Woojun Kim","Tongshuang Wu","Katia P. Sycara","Yaqi Xie"]}},"tmdate":1736861080869,"pdate":1730081753293,"tcdate":1726030315026,"writers":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission45/Authors"],"signatures":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission45/Authors"],"forum":"MIY7VeeMmq","license":"CC BY 4.0","number":45,"cdate":1726030315026,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Post_Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission45/-/Full_Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Edit","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission45/-/Camera-Ready_Revision"],"mdate":1736861080869,"odate":1736861080853,"domain":"NeurIPS.cc/2024/Workshop/Video-Langauge_Models","id":"MIY7VeeMmq","version":2},{"content":{"summary":{"value":"The authors of the paper propose a new video large language model, VideoGPT+, by (1) using two encoders—an image encoder and a video encoder—to extract visual features, and (2) introducing a new video instruction tuning dataset called VCG+112K, which contains 112K samples generated using a semi-automatic annotation pipeline based on PySceneDetect for keyframe extraction, LLaVA-v1.6 for frame description generation, GPT-4 for detailed video description generation, and finally GPT-3.5 for QA pair generation.\n\nAdditionally, they introduce a new benchmark called VCGBench-Diverse, consisting of 4,354 question-answer pairs for 877 videos sourced from HDVILA, MPII, YouCook2, UCF Crime, and STUD Traffic. A human annotation process assisted by GPT-3.5 is used to obtain the QA pairs for this benchmark."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"1. Since your main modeling contribution is the use of dual encoders, for fair comparisons, could you use the data recipe and model design from an existing work (e.g., VideoLLaVA) and apply your dual encoders to demonstrate the effectiveness of your dual-encoder design across evaluation benchmarks? Similarly, could you train another model architecture on VCG+112K to demonstrate the effectiveness of VCG+112K? That would make the paper's claims more reliable.\n\n2. What is the video source for VCG+112K? Is it ActivityNet? Are the videos in VCG+112K the same set of videos as those in VideoInstruction100K? How many videos are included in VCG+112K?\n\n3. Lines 267–269 mention that the detailed video descriptions generated by GPT-4 include a timeline of events, actions, object attributes, and scene settings. Could you qualitatively show some examples of video descriptions obtained after this step? Since a timeline of events is included, does this imply that a subset of the resulting VCG+112K data could be used to train models like Grounded-VideoLLM for tasks such as temporal grounding? Would it be advisable to choose the dense captions from VCG+112K over other existing video description/caption datasets for training a video description/caption model?\n\n4. Lines 473–475 mention that the video-only model performs better than the image-only model on MVBench action categories. How does the dual-encoder design perform in comparison?\n\n5. For the ablation 'without segment-wise sampling,' did you use uniform frame sampling instead? For the ablation 'without adaptive token pooling,' it is mentioned that this restricts the model to fewer frames. How many frames were used for this ablation? More implementation details are needed for the ablations.\n\n6. In the table showing the results of the VCG+112K ablation study, does the second row indicate the use of VCG+112K in training? The table design is problematic because the second row shows better performance but is marked with a cross mark, suggesting that VCG+112K was not used.\n\n7. Line 524 mentions a combined dataset. Does this mean that you combined all the training data listed in lines 353 to 364? For datasets with multiple splits mentioned in lines 353 to 364, did you use only the training split? Did you use the full dataset for each one mentioned in lines 353 to 364, or was any sampling performed when designing your training data recipe?\n\n8. How is VCGBench-Diverse unique compared to existing benchmarks?\n\n9. How do the new design choices affect the training and inference costs? How does the training data affect the training time?\n\nMinor comments:\n\nIn Figure 1 caption: \"VideoGPT+ permors better\" -> \"VideoGPT+ performs better\""},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. The model VideoGPT+ is a 3.8B-scale model. The VideoGPT+ model along with the VCG+112K dataset will be valuable resources once refined and released. \n\n2. The primary design choice in this work—using both an image encoder and a video encoder to develop a video large language model—is reasonable and straightforward."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. **The novelty and technical depth of this paper are limited. While the paper follows empirical best practices for constructing a multimodal large language model, it lacks the original and novel ideas expected for a conference like ICLR.** The main architectural change introduced in VideoGPT+ is the use of a dual-encoding scheme; however, this approach is neither particularly innovative nor original. In the multimodal large language model domain, the use of two or multiple encoders to extract visual features has been widely explored [1-5]. In terms of ideas, I don’t see a fundamental difference, notable innovation, or new insights.\n\n2. **(a) The introduced benchmark, VCGBench-Diverse, is limited: it is small (containing only 877 videos) and supports only open-ended question answering.** The authors claim that a multiple-choice question-answering format introduces bias and fails to capture the model's true understanding (Lines 136-138), which I disagree with, as open-ended QA, relying on LLM-as-a-judge, often leads to more problematic evaluations compared to multiple-choice QA. In multimodal large language model evaluation, the trend favors using multiple-choice QA over open-ended QA for robust and reliable benchmarking, or incorporating multiple task formats (e.g., TempCompass).\n**(b) There are many benchmarks for video conversation models nowadays. It is also unclear how VCGBench-Diverse is unique compared to existing benchmarks**, including those mentioned in this paper, such as VCGBench, MVBench, and VideoMME, as well as others not mentioned, such as TempCompass [6], CVRR-ES [7], Vinoground [8], VideoVista [9], and more [10-14].\n\n3. **State-of-the-art (SOTA) works**, such as LLaVA-NeXT-Video [15], LLaVA-OneVision [16], PLLaVA [17], SlowFast-LLaVA [18], Tarsier [19], MiniGPT4-Video [20], ShareGPT4Video [21], VideoLLaMA 2 [22], etc., **are missing**. The SOTA model comparisons are unconvincing due to the absence of these SOTA models and different training data recipes are used across models. Moreover, the **ablation studies are performed on a selectively chosen and inconsistent subset of evaluation benchmarks**.\n\n4. **Latency is not discussed.**\n\n5. **Clarity of the paper could be improved.**\n\n\nReferences:\n\n[1] BRAVE: Broadening the Visual Encoding of Vision-Language Models\n\n[2] Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders\n\n[3] SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models\n\n[4] Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs\n\n[5] Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models\n\n[6] TempCompass: Do Video LLMs Really Understand Videos?\n\n[7] How Good is my Video LMM? Complex Video Reasoning and Robustness Evaluation Suite for Video-LMMs\n\n[8] Vinoground: Scrutinizing LMMs over Dense Temporal Reasoning with Short Videos\n\n[9] VideoVista: A Versatile Benchmark for Video Understanding and Reasoning\n\n[10] AutoEval-Video: An Automatic Benchmark for Assessing Large Vision Language Models in Open-Ended Video Question Answering\n\n[11] VideoBench: https://github.com/PKU-YuanGroup/Video-Bench\n\n[12] MLVU: A Comprehensive Benchmark for Multi-Task Long Video Understanding\n\n[13] LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding\n\n[14] TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models\n\n[15] LLaVA-NeXT: Tackling Multi-Image, Video, and 3D in Large Multimodal Models\n\n[16] LLaVA-OneVision: Easy Visual Task Transfer\n\n[17] PLLaVA: Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning\n\n[18] SlowFast-LLaVA- A Strong Training-Free Baseline for Video Large Language Models\n\n[19] Tarsier: Recipes for Training and Evaluating Large Video Description Models\n\n[20] MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens\n\n[21] ShareGPT4Video- Improving Video Understanding and Generation with Better Captions\n\n[22] VideoLLaMA 2 Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs"}},"nonreaders":[],"tmdate":1731428662501,"tcdate":1730582850584,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission11214/Reviewer_yT2F"],"signatures":["ICLR.cc/2025/Conference/Submission11214/Reviewer_yT2F"],"forum":"YGWxpOI6Y0","number":3,"license":"CC BY 4.0","cdate":1730582850584,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission11214/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428662501,"domain":"ICLR.cc/2025/Conference","replyto":"YGWxpOI6Y0","id":"nwxvuqgDZx","forumContent":{"TLDR":{"value":"VideoGPT+: Spatiotemporal Aware Video Conversation Model"},"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["video-conversation-model","large multi-modal model","multi-modal","video-conversation","image-and-video","phi-3-min","vision-language","video-chatbot"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Building on the advances of language models, Large Multimodal Models (LMMs) have contributed significant improvements in video understanding. While the current video LMMs utilize advanced Large Language Models (LLMs), they rely on either image or video encoders to process visual inputs, each of which has its own limitations. Image encoders excel at capturing rich spatial details from frame sequences but lack explicit temporal context, which can be important in videos with intricate action sequences. On the other hand, video encoders provide temporal context but are often limited by computational constraints that lead to processing only sparse frames at lower resolutions, resulting in reduced contextual and spatial understanding. To this end, we introduce our model, which combines the complementary benefits of the image encoder (for detailed spatial understanding) and the video encoder (for global temporal context modeling). The model processes videos by dividing them into smaller segments and applies an adaptive pooling strategy on features extracted by both image and video encoders. Our architecture showcases improved performance across multiple video benchmarks, including VCGBench, MVBench and Zero-shot question-answering. Further, we develop 112K video-instruction set using a novel semi-automatic annotation pipeline which further improves the model performance. Additionally, to comprehensively evaluate video LMMs, we present our bench, covering 18 broad video categories such as lifestyle, sports, science, gaming, and surveillance videos. This benchmark with 4,354 question-answer pairs evaluates the generalization of existing LMMs on dense video captioning, spatial and temporal understanding, and complex reasoning, ensuring comprehensive assessment across diverse video types and dynamics. Our code, dataset, and pre-trained models will be publicly released."},"_bibtex":{"value":"@misc{\nmaaz2024videogpt,\ntitle={Video{GPT}+: Integrating Image and Video Encoders for Enhanced Video Understanding},\nauthor={Muhammad Maaz and Hanoona Abdul Rasheed and Salman Khan and Fahad Shahbaz Khan},\nyear={2024},\nurl={https://openreview.net/forum?id=YGWxpOI6Y0}\n}"},"title":{"value":"VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding"},"pdf":{"value":"/pdf/ec0f125e42fed1a2714c46de829865aaa0e50324.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"maaz|videogpt_integrating_image_and_video_encoders_for_enhanced_video_understanding"},"authorids":{"value":["~Muhammad_Maaz1","~Hanoona_Abdul_Rasheed1","~Salman_Khan4","~Fahad_Shahbaz_Khan1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Muhammad Maaz","Hanoona Abdul Rasheed","Salman Khan","Fahad Shahbaz Khan"]}},"version":2},{"content":{"summary":{"value":"The paper proposes Modality Composition Awareness, a training framework to mitigate the “modality shortcut” problem when using Multimodal Large Language Models as a unified encoder for composed multimodal retrieval. MCA has two complementary objectives: 1.Modality Composition Preference, a preference-style loss that enforces composed embeddings to be more discriminative than any unimodal counterpart. 2. Modality Composition Regularization, aligns the composed embedding with a prototype assembled from unimodal embeddings via a simple mixer and a contrastive objective.  Extensive experiments on IND and OOD benchmarks show MCA preserves IND accuracy while improving OOD robustness. The paper also studies sensitivity to image resolution, loss weighting, and mixer variants."},"soundness":{"value":2},"confidence":{"value":5},"questions":{"value":"1.\tPlease quantify additional time/memory cost of MCP/MCR during training compared to vanilla contrastive learning. \n2.\tCould you provide more qualitative examples and visualizations (score histograms, embedding distance changes) in the appendix to illustrate how MCA modifies representations?\n3.\tIt seems that an additional empirical experiment is needed to directly verify whether the proposed method effectively mitigates the modality shortcut problem. Moreover, the paper lacks sufficient discussion from the textual perspective (i.e., how the method affects or interacts with text representations).\n4.\tIs the modality shortcut problem empirically verified to exist in the current datasets? If it indeed exists, why does the proposed method improve performance only on out-of-domain benchmarks but not on in-domain ones? It would be more convincing if the authors could collect or annotate a dedicated benchmark that explicitly reflects this issue.\n5.\tDo existing open-source models also exhibit the modality shortcut problem? Why not include comparisons with these models to empirically demonstrate that the issue truly exists?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1.\tThe paper formalizes the “modality shortcut” problem for unified MLLM encoders under composed inputs\n2.\tMCP (preference) and MCR (consistency to mixed prototype) are conceptually simple and can be integrated as light-weight regularizers to contrastive training"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1.\tIt seems that an additional empirical experiment is needed to directly verify whether the proposed method effectively mitigates the modality shortcut problem. \n2.\tThe paper lacks sufficient discussion from the textual perspective (i.e., how the method affects or interacts with text representations).\n3.\tThe paper lacks a relevant benchmark to empirically validate the problem mentioned by the authors, eg modality shortcut.\n4.\tThe paper is not sufficiently comprehensive; additional text-related experiments or analyses of the model’s training and testing behaviors would make the work more complete. \n5.\tThe paper lacks comparisons with existing open-source models. Since the evaluation is conducted only against models trained by the authors themselves, it is difficult to verify the effectiveness of the proposed method and to confirm whether the targeted research problem truly exists."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762922966465,"tcdate":1761707616816,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission11968/Reviewer_jWKA"],"signatures":["ICLR.cc/2026/Conference/Submission11968/Reviewer_jWKA"],"forum":"EzBRI0Llxk","number":1,"license":"CC BY 4.0","cdate":1761707616816,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission11968/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762922966465,"domain":"ICLR.cc/2026/Conference","replyto":"EzBRI0Llxk","id":"sSFTyYNLHq","forumContent":{"TLDR":{"value":"Mitigating the modality shortcut problem in MLLM-based multimodal retrieval with modal composition awareness."},"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Multimodal Retrieval","Modality Shortcut","MLLMs","Modality Composition"]},"primary_area":{"value":"unsupervised, self-supervised, semi-supervised, and supervised representation learning"},"abstract":{"value":"Multimodal retrieval, which seeks to retrieve relevant content across modalities such as text or image, supports applications from AI search to contents production. Despite the success of separate-encoder approaches like CLIP align modality-specific embeddings with contrastive learning, recent multimodal large language models (MLLMs) enable a unified encoder that directly processes composed inputs. While flexible and advanced, we identify that unified encoders trained with conventional contrastive learning are prone to learn modality shortcut, leading to poor robustness under distribution shifts. We propose a modality composition awareness framework to mitigate this issue. Concretely, a preference loss enforces multimodal embeddings to outperform their unimodal counterparts, while a composition regularization objective aligns multimodal embeddings with prototypes composed from its unimodal parts. These objectives explicitly model structural relationships between the composed representation and its unimodal counterparts. Experiments on various benchmarks show gains in out-of-distribution retrieval, highlighting modality composition awareness as a effective principle for robust composed multimodal retrieval when utilizing MLLMs as the unified encoder."},"_bibtex":{"value":"@misc{\nwu2025mca,\ntitle={{MCA}: Modality Composition Awareness for Robust Composed Multimodal Retrieval},\nauthor={Qiyu Wu and Shuyang Cui and Satoshi Hayakawa and Wei-Yao Wang and Hiromi Wakaki and Yuki Mitsufuji},\nyear={2025},\nurl={https://openreview.net/forum?id=EzBRI0Llxk}\n}"},"title":{"value":"MCA: Modality Composition Awareness for Robust Composed Multimodal Retrieval"},"pdf":{"value":"/pdf/2c99bf21eae5f0db226f555f11aba423303addf3.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"wu|mca_modality_composition_awareness_for_robust_composed_multimodal_retrieval"},"authorids":{"value":["~Qiyu_Wu2","~Shuyang_Cui1","~Satoshi_Hayakawa1","~Wei-Yao_Wang1","~Hiromi_Wakaki1","~Yuki_Mitsufuji1"]},"authors":{"value":["Qiyu Wu","Shuyang Cui","Satoshi Hayakawa","Wei-Yao Wang","Hiromi Wakaki","Yuki Mitsufuji"]}},"version":2},{"content":{"summary":{"value":"VideoMathQA introduces a benchmark for math problem solving in instructional videos, bridging traditional text/image-based QA and full multimodal video reasoning. It includes a few thousand curated QA pairs from formula-rich YouTube videos, each with aligned frames, audio, transcripts, and formulas across 10 math domains (geometry, calculus, statistics, etc.). Each question features step-by-step solutions (chain-of-thought) for detailed evaluation. The benchmark targets three scenarios: Direct Problem Solving, Conceptual Transfer, and Deep Instructional Comprehension. Baselines with models like Math-LLaVA, MathBLIP, and Video-CoT reveal a large gap from human performance, underscoring the difficulty of integrating visual, textual, and temporal reasoning in mathematical contexts."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. Is there any analysis of the model performance without any video and only given the question?\n2. Is there any analysis on how reliable the evaluation scores are? is the model following the video at all, or having the video distracting the model from the correct solution?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. Introduces the first benchmark for math reasoning in instructional videos, capturing dynamic visual and spoken content absent in prior static text/image datasets.\n\n2. Built through 920+ hours of expert annotation, covering 420 problems (~4.5K QA pairs) with aligned formulas, video timestamps, and full step-by-step solutions.\n\n3. Spans 10 math domains and varied video styles (lectures, tutorials, handwritten, slides), testing both short-term perception and long-range reasoning.\n\n4. Measures both answer accuracy and chain-of-thought alignment, includes conditions with/without subtitles, and offers detailed error categorization for model failures.\n\n5. Evaluates multiple vision-language models (e.g., Qwen-VL, InternVL, Math-LLaVA), showing large performance gaps to human accuracy, confirming the benchmark’s difficulty and diagnostic value."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The five-option format simplifies scoring but allows guessing or elimination, limiting evaluation of free-form reasoning and creativity.\n2. Some Conceptual Transfer questions may be solvable without watching the video, letting models rely on prior knowledge rather than true video understanding."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762916662721,"tcdate":1761966833410,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission3317/Reviewer_fxx1"],"signatures":["ICLR.cc/2026/Conference/Submission3317/Reviewer_fxx1"],"forum":"VI4kGUfPio","number":4,"license":"CC BY 4.0","cdate":1761966833410,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission3317/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762916662721,"domain":"ICLR.cc/2026/Conference","replyto":"VI4kGUfPio","id":"Ye4hNF2BbS","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"TLDR":{"value":"Mathematical Reasoning using Video MLLM Benchmark"},"keywords":{"value":["Multimodal Reasoning","Video Question Answering","Mathematical Understanding","Temporal Reasoning","Visual Grounding"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Mathematical reasoning in real-world video presents a fundamentally different challenge than static images or text. It requires interpreting fine-grained visual information, accurately reading handwritten or digital text, and integrating spoken cues, often dispersed non-linearly over time. In such multimodal contexts, success hinges not just on perception, but on selectively identifying and integrating the right details from a rich and noisy stream of content. To this end, we introduce VideoMathQA, a benchmark designed to evaluate whether models can perform such temporally extended cross-modal reasoning on videos. The benchmark spans 10 diverse mathematical domains, covering videos from 10 seconds to over 1 hour. We employ graduate-level experts to ensure high quality, for over 920 man-hours of annotation. To reflect real-world scenarios, questions are designed around three core reasoning challenges: direct problem solving, conceptual transfer, which requires applying learned methods to new problems; and deep instructional comprehension, involving multi-step reasoning over extended explanations and partially worked-out solutions. Each question includes multi-step reasoning annotations, enabling fine-grained diagnosis of model capabilities. Through this benchmark, we establish an evaluation framework for models that must reason, rather than merely perceive, jointly ground concepts across visual, audio, and textual modalities, across temporally extended mathematical problem settings."},"_bibtex":{"value":"@inproceedings{\nrasheed2026videomathqa,\ntitle={VideoMath{QA}: Benchmarking Mathematical Reasoning via Multimodal Understanding in Video},\nauthor={Hanoona Abdul Rasheed and Abdelrahman M Shaker and Anqi Tang and Muhammad Maaz and Ming-Hsuan Yang and Salman Khan and Fahad Shahbaz Khan},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=VI4kGUfPio}\n}"},"title":{"value":"VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Video"},"pdf":{"value":"/pdf/ecfd49591f38470840444b642dfebfb19c5d0c39.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"rasheed|videomathqa_benchmarking_mathematical_reasoning_via_multimodal_understanding_in_video"},"authorids":{"value":["~Hanoona_Abdul_Rasheed1","~Abdelrahman_M_Shaker1","~Anqi_Tang2","~Muhammad_Maaz1","~Ming-Hsuan_Yang1","~Salman_Khan4","~Fahad_Shahbaz_Khan1"]},"authors":{"value":["Hanoona Abdul Rasheed","Abdelrahman M Shaker","Anqi Tang","Muhammad Maaz","Ming-Hsuan Yang","Salman Khan","Fahad Shahbaz Khan"]}},"version":2},{"content":{"summary":{"value":"The paper introduces VidEgoThink, a benchmark for evaluating egocentric video understanding in low-level control tasks within Embodied AI. VidEgoThink employs GPT-4o and prompt engineering to generate question-answer pairs from the Ego4D dataset, focusing on four tasks: video question-answering, hierarchical planning, visual grounding, and reward modeling. Human filtering is applied to ensure the quality of the benchmark. In subsequent evaluations, the benchmark reveals that all current multimodal large language models (MLLMs), including GPT-4o, perform poorly across these tasks."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"Refer to weaknesses."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"The paper has several strong aspects:\n\n1.This paper is the first to integrate four key tasks—video question-answering, hierarchical planning, visual grounding, and reward modeling—covering a broad spectrum from task planning and video understanding to spatial-temporal reasoning and fine-grained task completion awareness. This reflects the authors' deep understanding of the requirements for low-level control tasks in Embodied AI scenarios.\n\n2.For each task, the authors attempt to define clear evaluation dimensions and metrics, providing a usable approach to assessing performance.\n\n3.The experiments show that VidEgoThink is a robust benchmark that effectively highlights the limitations of current MLLMs in processing first-person perspective data."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"I see the following as weak points:\n1. The quality of VidEgoThink's benchmark data is not thoroughly analyzed.\n\nThe paper employs GPT-4o and prompt engineering to generate question-answer pairs from the Ego4D dataset. However, the quality of the benchmark data is not thoroughly analyzed. While it is mentioned that 3 human filters were applied, the filtering criteria are not clearly explained. The statistical analysis in the appendix only covers aspects such as question and answer length, and scene types, which are likely based on the Ego4D dataset's categories. These analyses provide little insight into the overall quality of the benchmark.\n\nI would have liked to see more detailed insights, such as which tasks and actions are included in the benchmark, information on long-term and mid-term tasks along with their difficulty,  as well as which scenes are more challenging, and which are easier. The varying video lengths, differences in activity density, and changing camera perspectives in Ego4D are factors that likely contribute to significant difficulties in video understanding, but these are not addressed.\n\n\n2. The excessive length of egocentric videos, combined with the inherent difficulty in understanding egocentric scenes, likely leads to the performance decline of all current MLLMs.\n\nTable 3 of the paper shows that the average length of videos in VidEgoThink is 270.74 seconds. However, in the experimental evaluation, only 32 frames, 8 frames, or even captions are used as input. Given that Ego4D has a frame rate of 30 fps and an average video length of around 270 seconds, selecting 32 frames results in sampling roughly one frame every 250 frames. This sparse sampling from such long videos significantly hampers the video perception and understanding capabilities of current MLLMs.\n\nFor egocentric videos with complex scenes, numerous background objects, rapid camera movements, and transitions between scenes, such limited input is unlikely to provide sufficient information about the video. This likely also contributes to the generally poor performance of MLLMs."}},"nonreaders":[],"tmdate":1732539563077,"tcdate":1730261376771,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission2168/Reviewer_qHFD"],"signatures":["ICLR.cc/2025/Conference/Submission2168/Reviewer_qHFD"],"forum":"Z5nqeTH24j","number":2,"license":"CC BY 4.0","cdate":1730261376771,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission2168/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732539563077,"domain":"ICLR.cc/2025/Conference","replyto":"Z5nqeTH24j","id":"LoypL5jaT1","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Multi-modal Large Language Models","Egocentric Video Understanding","Embodied AI","Benchmark"]},"primary_area":{"value":"datasets and benchmarks"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Recent advancements in Multi-modal Large Language Models (MLLMs) have opened new avenues for applications in Embodied AI.\nBuilding on previous work, EgoThink, we introduce VidEgoThink, a comprehensive benchmark for evaluating egocentric video understanding capabilities. To bridge the gap between MLLMs and low-level control in Embodied AI, we design four key interrelated tasks: video question-answering, hierarchy planning, visual grounding and reward modeling. To minimize manual annotation costs, we develop an automatic data generation pipeline based on the Ego4D dataset, leveraging the prior knowledge and multimodal capabilities of GPT-4o. Three human annotators then filter the generated data to ensure diversity and quality, resulting in the VidEgoThink benchmark. We conduct extensive experiments with three types of models: API-based MLLMs, open-source image-based MLLMs, and open-source video-based MLLMs. Experimental results indicate that all MLLMs, including GPT-4o, perform poorly across all tasks related to egocentric video understanding. These findings suggest that foundation models still require significant advancements to be effectively applied to first-person scenarios in Embodied AI. In conclusion, VidEgoThink reflects a research trend towards employing MLLMs for egocentric vision, akin to human capabilities, enabling active observation and interaction in the complex real-world environments."},"_bibtex":{"value":"@misc{\ncheng2024videgothink,\ntitle={VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied {AI}},\nauthor={Sijie Cheng and Kechen Fang and Yangyang Yu and Sicheng Zhou and Bohao Li and Ye Tian and Tingguang Li and Lei Han and Yang Liu},\nyear={2024},\nurl={https://openreview.net/forum?id=Z5nqeTH24j}\n}"},"title":{"value":"VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI"},"pdf":{"value":"/pdf/04b24f0a982536bd150df66a7ea0233cf41dc649.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"cheng|videgothink_assessing_egocentric_video_understanding_capabilities_for_embodied_ai"},"authorids":{"value":["~Sijie_Cheng1","~Kechen_Fang1","~Yangyang_Yu2","~Sicheng_Zhou2","~Bohao_Li1","~Ye_Tian1","~Tingguang_Li1","~Lei_Han1","~Yang_Liu19"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Sijie Cheng","Kechen Fang","Yangyang Yu","Sicheng Zhou","Bohao Li","Ye Tian","Tingguang Li","Lei Han","Yang Liu"]}},"version":2},{"content":{"summary":{"value":"This paper presents a joint audio-video generation model.\nStarting from a pre-trained video generation model (Wan2.2), the audio backbone is designed with an identical architecture, which simplifies the design of the cross-modal interaction modules.\nThe model is trained in two stages: the audio backbone is first trained from scratch on speech and sound-effect datasets, followed by joint training of the audio and video branches on audiovisual datasets.\nSubjective evaluations demonstrate that the proposed model outperforms existing open-source joint audio-video generation models."},"soundness":{"value":1},"confidence":{"value":4},"questions":{"value":"See weaknesses above, particularly regarding the novelty and experimental setup."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":1},"strengths":{"value":"- The proposed model achieves state-of-the-art quality compared to existing open-source joint audio-video generation models.\n- The architecture is simple yet effective, yielding improved generation performance.\n- A new audiovisual dataset construction pipeline is introduced, producing well-synchronized audio-video pairs with rich captions, which could serve as a valuable contribution to the community."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"**Lack of methodological novelty**.  \nThe proposed approach replicates prior work (especially JavisDiT), raising concerns about the method's originality. Specific overlaps include:\n\n- The overall modeling framework and two-stage training strategy closely follow JavisDiT. JavisDiT also employs an identical architecture for the audio backbone with video branch (while it is based on OpenSora rather than Wan2.2) and uses a similar two-step training procedure: audio pre-training followed by joint audiovisual training. \n- The paper explains the limitation of JavisDiT as requiring \"learned prior estimator ... in order to achieve synchronization\" (Sec 2.4), but it is unclear why this is problematic or how severe this limitation actually is.\n- The proposed unified prompt conditioning mechanism differs from JavisDiT, but there is no direct comparison demonstrating its benefit.\n- The proposed RoPE configuration appears to be identical to that of MMAudio (see Fig. A5 in the MMAudio paper), which does not solely establish the novelty.\n\nGiven these similarities and lack of justification of the proposed design choices, the contributions (2), (3), and (4) in the introduction are not convincingly supported.\n\n**Insufficient experimental evidence**.  \nThe experiments are limited and do not clearly demonstrate the advantages of the proposed model.\n\n- Details of the subjective evaluation are missing. How many samples were evaluated per participant? Which prompts (or Verse-Bench splits) were used for generation? What questions were employed for each evaluation criterion?\n- Objective evaluation is limited to T2A and TTS tasks and focuses only on audio quality. Since the proposed method jointly generates both audio and video, it would be better to include evaluation of video quality (e.g., FVD[1], CLIP score[2], or VBench[3]) and audiovisual alignment (e.g., ImageBind[4] similarity, AV-Align[5], or DeSync[6]).\n- Fig.4 indicates that Ovi underperforms Wan2.2 for video generation, while Table 2 shows that Ovi underperforms most existing T2A or TTS models for audio generation. Based solely on these results, it is difficult to identify the strengths of the proposed approach. It would be more convincing to compare Ovi with a sequential baseline (e.g., Wan2.2 for T2V followed by MMAudio-L for V2A, or FishSpeech for TTS) to demonstrate the benefit of joint modeling.\n\n[1] \"FVD: A new metric for video generation,\" ICLRW, 2019  \n[2] \"Learning transferable visual models from natural language supervision,\" ICML, 2021  \n[3] \"VBench: Comprehensive Benchmark Suite for Video Generative Models,\" CVPR 2024  \n[4] \"Imagebind: One embedding space to bind them all,\" CVPR, 2023  \n[5] \"Diverse and aligned audio-to-video generation via text-to-video model adaptation,\" AAAI, 2024  \n[6] \"Synchformer: Efficient synchronization from sparse cues,\" ICASSP, 2024"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762932036819,"tcdate":1761723590391,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission19868/Reviewer_XP5Y"],"signatures":["ICLR.cc/2026/Conference/Submission19868/Reviewer_XP5Y"],"forum":"nIrF8xF0uN","number":2,"license":"CC BY 4.0","cdate":1761723590391,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission19868/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762932036819,"domain":"ICLR.cc/2026/Conference","replyto":"nIrF8xF0uN","id":"FnZ3yKcDMP","forumContent":{"TLDR":{"value":"Ovi is a Twin-DiT generative model that jointly generates perfectly synchronized video and audio"},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Artificial Intelligence","Deep Learning","Computer Vision","Video Generation","Audio Generation"]},"supplementary_material":{"value":"/attachment/9678322f1ca480ad98473d17310c073b9efcb7c7.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Audio–video (AV) generation has often relied on complex multi-stage architectures or sequential synthesis of sound and visuals. We introduce \\textsc{Ovi}, a unified paradigm for audio–video generation that models the two modalities as a single generative process. By using blockwise cross-modal fusion of twin-DiT modules, \\textsc{Ovi} achieves natural synchronization and removes the need for separate pipelines or post hoc alignment. To facilitate fine-grained multimodal fusion modeling, we initialize an audio tower with an architecture identical to that of a strong pretrained video model. Trained from scratch on hundreds of thousands of hours of raw audio, the audio tower learns to generate realistic sound effects, as well as speech that conveys rich speaker identity and emotion. Fusion is obtained by jointly training the identical video and audio towers via blockwise exchange of timing (via scaled-RoPE embeddings) and semantics (through bidirectional cross-attention) on a vast video corpus. Our model enables cinematic storytelling with natural speech and accurate, context-matched sound effects, producing movie-grade video clips."},"_bibtex":{"value":"@misc{\nlow2026ovi,\ntitle={Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation},\nauthor={Chetwin Low and Weimin Wang and Calder Katyal},\nyear={2026},\nurl={https://openreview.net/forum?id=nIrF8xF0uN}\n}"},"title":{"value":"Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation"},"pdf":{"value":"/pdf/69c5eec8e1c78deae66eecc446831071c95d2958.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"low|ovi_twin_backbone_crossmodal_fusion_for_audiovideo_generation"},"authorids":{"value":["~Chetwin_Low1","~Weimin_Wang5","~Calder_Katyal1"]},"authors":{"value":["Chetwin Low","Weimin Wang","Calder Katyal"]}},"version":2},{"content":{"summary":{"value":"This paper proposes Vidarc, an autoregressive embodied video diffusion framework integrated with a masked inverse dynamics model, designed to address the challenges of low-latency closed-loop control and generalization in data-scarce robotic manipulation scenarios. By introducing an embodiment-aware diffusion loss guided by action-relevant masks and incorporating KV caching for real-time environmental feedback, Vidarc aims to enhance the alignment between video generation and robotic embodiment dynamics. Pre-trained on one million cross-embodiment episodes and fine-tuned on unseen platforms, the model demonstrates higher task success rates (15-17% higher than baselines like Vidar and Pi0.5) and lower latency in both simulated and real-world experiments, along with error correction capabilities."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"Overall, the concept of Vidarc is relatively straightforward, and many implementation details lack sufficient empirical or theoretical support. The insufficient experimental design raises concerns about whether the reported results are robust or merely coincidental, reducing the work’s ability to provide meaningful guidance for future research."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"The integration of autoregressive video diffusion with closed-loop control via KV caching and environmental feedback re-prefilling addresses the high-latency issue of traditional non-autoregressive video-based methods, enabling more responsive robotic manipulation.\nThe embodiment-aware loss, weighted by masks from the inverse dynamics model, effectively prioritizes action-relevant regions, mitigating the problem of irrelevant visual distractions and improving the actionable quality of generated videos.\nExtensive pre-training on large-scale cross-embodiment datasets ensures strong generalization to unseen robotic platforms and tasks, with experimental results validating superior performance over state-of-the-art baselines in both simulation and real-world settings."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Unclear Visualization and Ambiguous Definitions: The error correction example in Figure 1 lacks sufficient explanation, making it difficult to understand the model’s error correction mechanism. Additionally, the definition of the observation space O is ambiguous—if O is a single image, predicting actions from a single frame contradicts the fundamental definition of inverse dynamics models and poses an ill-posed problem.\nIncomplete Literature Review: The paper claims that existing video-based methods are inefficient but overlooks prominent efficient approaches such as Vidman and VPP, leading to an incomplete assessment of the current research landscape.\nPractical Efficiency Limitations: Despite adopting CausVid for optimization, Vidarc’s reliance on the Wan2.2 backbone results in excessive memory overhead during training and inference. With a per-step latency of 3 seconds, the model fails to meet the real-time requirements of embodied intelligence applications.\nThin Simulation Experiments: The simulation evaluations are only conducted on the RoboTwin benchmark, lacking validation on other widely used benchmarks like Libero or RLBench. This limits the generalizability of the reported performance.\nAmbiguous Contribution of Core Components: The performance advantage of Vidarc over baselines is primarily attributed to closed-loop control, but this component is not inherently tied to video generation—any autoregressive method could integrate such feedback. Moreover, removing closed-loop control leads to a significant performance drop (66.8% vs. 80.7% average success rate on RoboTwin), casting doubt on the necessity of the embodiment-aware loss. The weighted loss may also neglect background information critical for diffusion denoising, yet no theoretical justification or quantitative analysis of the loss’s feasibility is provided.\nInsufficient Analysis of Key Mechanisms: The masked inverse dynamics model uses different input sources during training and inference, raising concerns about the stability and controllability of its performance, especially with observations generated via embodiment-aware training. Additionally, there is a lack of case studies and quantitative analysis on how closed-loop control improves action accuracy over iterations, leaving the mechanism’s effectiveness underexplored."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762916473851,"tcdate":1761904485431,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission2976/Reviewer_k5Kh"],"signatures":["ICLR.cc/2026/Conference/Submission2976/Reviewer_k5Kh"],"forum":"gsvjCTIYPb","number":2,"license":"CC BY 4.0","cdate":1761904485431,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission2976/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762916473851,"domain":"ICLR.cc/2026/Conference","replyto":"gsvjCTIYPb","id":"VTialeA6ti","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Robotics","Video Diffusion Model","Computer Vision"]},"supplementary_material":{"value":"/attachment/8ce7790fc4e592a31bf85c4f6cc642b29e9534e9.zip"},"primary_area":{"value":"applications to robotics, autonomy, planning"},"abstract":{"value":"Robotic arm manipulation in data-scarce settings is a highly challenging task due to the complex embodiment dynamics and diverse contexts. Recent video-based approaches have shown great promise in capturing and transferring the temporal and physical interactions by pre-training on Internet-scale video data. However, such methods are often not optimized for the embodiment-specific closed-loop control, typically suffering from high latency and insufficient grounding. In this paper, we present Vidarc (Video Diffusion for Action Reasoning and Closed-loop Control), a novel autoregressive embodied video diffusion approach augmented by a masked inverse dynamics model. By grounding video predictions with action-relevant masks and incorporating real-time feedback through cached autoregressive generation, Vidarc achieves fast, accurate closed-loop control. Pre-trained on one million cross-embodiment episodes, Vidarc surpasses state-of-the-art baselines, achieving at least a 15% higher success rate in real-world deployment and a 91% reduction in latency. We also highlight its robust generalization and error correction capabilities across previously unseen robotic platforms."},"_bibtex":{"value":"@misc{\nfeng2026vidarc,\ntitle={Vidarc: Low Latency Embodied Video Diffusion Model with Closed-loop Control},\nauthor={Yao Feng and Chendong Xiang and Xinyi Mao and Hengkai Tan and Zuyue Zhang and Shuhe Huang and Kaiwen Zheng and Haitian Liu and Hang Su and Jun Zhu},\nyear={2026},\nurl={https://openreview.net/forum?id=gsvjCTIYPb}\n}"},"title":{"value":"Vidarc: Low Latency Embodied Video Diffusion Model with Closed-loop Control"},"pdf":{"value":"/pdf/b6ebd3124a61d8dfba2342153bc18da637196e92.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"feng|vidarc_low_latency_embodied_video_diffusion_model_with_closedloop_control"},"authorids":{"value":["~Yao_Feng2","~Chendong_Xiang1","~Xinyi_Mao1","~Hengkai_Tan1","~Zuyue_Zhang1","~Shuhe_Huang1","~Kaiwen_Zheng2","~Haitian_Liu2","~Hang_Su3","~Jun_Zhu2"]},"authors":{"value":["Yao Feng","Chendong Xiang","Xinyi Mao","Hengkai Tan","Zuyue Zhang","Shuhe Huang","Kaiwen Zheng","Haitian Liu","Hang Su","Jun Zhu"]}},"version":2},{"content":{"summary":{"value":"The paper proposes HOMO, a higher-order Shortcut diffusion model that augments first-order (velocity) supervision with acceleration (and optionally jerk) along $x_t = \\alpha_t x_0 + \\beta_t x_1$. Each order is predicted by its own network and trained with losses for first/second-order matching plus a self-consistency constraint from composing small steps. They provide training/sampling schemes and show experiments on synthetic and curve datasets."},"soundness":{"value":1},"confidence":{"value":3},"questions":{"value":"1) The approximation bounds include a non-vanishing term $\\mathbb{E} \\left[|\\dot x_{\\text{true}}-\\ddot{x}_{\\text{true}}|^2\\right]$. Could the authors clarify how this bound provides any guarantee that the learned model converges to the true generative process or that the proposed losses identify the correct flow?\n\n2) Why is the first-order M1 loss evaluated with $u_1(x_t,t,2d)$ rather than at the instantaneous argument $u_1(x_t,t,0)$? If this is intentional, what theoretical justification ensures that using a finite step size does not bias the learned dynamics?\n\n3) All reported experiments are on 2D toy datasets. Have the authors tested the approach on any image or high-dimensional generative tasks, and if not, what challenges prevent such evaluation?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":1},"contribution":{"value":2},"strengths":{"value":"The core idea of explicitly modeling higher-order terms along the transport path with separate networks is simple and potentially broadly applicable. The paper gives concrete training and sampling procedures that are easy to implement, and on standard 2D benchmarks it shows consistent improvements over a first-order shortcut baseline."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The writing quality materially hurts readability. For example, the text says \"we define Shortcut model compute next field\" which is ungrammatical and obscures meaning. On the theory side, the approximation bounds are not informative for learning: even with large models, the bound in 5.1 retains an additive term $\\mathbb{E} \\left[\\|\\dot x_{\\text{true}}-\\ddot{x}_{\\text{true}}\\|^2\\right]$ that does not vanish. The results do not show that the learned velocity and acceleration converge to the truth nor that the minimizer of the proposed loss recovers a correct generative model of the data. The loss design is also unconvincing. If the Shortcut first order objective is optimized to match the path velocity, no second order correction should be needed, and the paper does not explain why adding a second order term helps. In addition there is likely an error in the objective as written since the M1 loss appears to evaluate $u_1(x_t,t,2d)$ rather than the instantaneous argument. Finally, the empirical scope is narrow since there are no image or other high dimensional experiments, so it is unclear whether any gains extend beyond two dimensional toys."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762924356275,"tcdate":1762125393377,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission13834/Reviewer_r58d"],"signatures":["ICLR.cc/2026/Conference/Submission13834/Reviewer_r58d"],"forum":"Sv5Ubt3dFi","number":3,"license":"CC BY 4.0","cdate":1762125393377,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission13834/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762924356275,"domain":"ICLR.cc/2026/Conference","replyto":"Sv5Ubt3dFi","id":"uiDLv8zcws","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["High-Order Matching","Diffusion Model","One-Step Shortcut"]},"primary_area":{"value":"generative models"},"abstract":{"value":"One-step shortcut diffusion models [Frans, Hafner, Levine and Abbeel, ICLR 2025] have shown potential in vision generation, but their reliance on first-order trajectory supervision is fundamentally limited. The Shortcut model's simplistic velocity-only approach fails to capture intrinsic manifold geometry, leading to erratic trajectories, poor geometric alignment, and instability-especially in high-curvature regions. These shortcomings stem from its inability to model mid-horizon dependencies or complex distributional features, leaving it ill-equipped for robust generative modeling. In this work, we introduce HOMO (High-Order Matching for One-Step Shortcut Diffusion), a game-changing framework that leverages high-order supervision to revolutionize distribution transportation. By incorporating acceleration, jerk, and beyond, HOMO not only fixes the flaws of the Shortcut model but also achieves unprecedented smoothness, stability, and geometric precision. Theoretically, we prove that HOMO's high-order supervision ensures superior approximation accuracy, outperforming first-order methods. Empirically, HOMO dominates in complex settings, particularly in high-curvature regions where the Shortcut model struggles. Our experiments show that HOMO delivers smoother trajectories and better distributional alignment, setting a new standard for one-step generative models."},"_bibtex":{"value":"@misc{\nchen2025highorder,\ntitle={High-Order Matching for One-Step Shortcut Diffusion Models},\nauthor={Yubin Chen and Chengyue Gong and Xiaoyu Li and Yingyu Liang and Zhizhou Sha and Zhenmei Shi and Zhao Song},\nyear={2025},\nurl={https://openreview.net/forum?id=Sv5Ubt3dFi}\n}"},"title":{"value":"High-Order Matching for One-Step Shortcut Diffusion Models"},"pdf":{"value":"/pdf/c52172f317c0fcbf437a819feb095714446d4c8b.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"chen|highorder_matching_for_onestep_shortcut_diffusion_models"},"authorids":{"value":["~Yubin_Chen4","~Chengyue_Gong1","~Xiaoyu_Li12","~Yingyu_Liang1","~Zhizhou_Sha1","~Zhenmei_Shi1","~Zhao_Song3"]},"authors":{"value":["Yubin Chen","Chengyue Gong","Xiaoyu Li","Yingyu Liang","Zhizhou Sha","Zhenmei Shi","Zhao Song"]}},"version":2},{"content":{"summary":{"value":"This paper proposes LIFT, a framework for enhancing offline reinforcement learning by augmenting logged trajectories through learned “shortcuts.” LIFT has two core features: (1) a trajectory-level augmentation that replaces suboptimal action subsequences with shorter, higher-value transitions, and (2) a policy-time augmentation, which intermittently injects high-value actions during data collection.\nTheoretical analysis provides sufficient conditions under which these shortcut augmentations provably improve performance.\nExperiments on simulated positioning tasks show that LIFT and its shortcut variant (LIFT-SC) outperform standard offline RL baselines."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"1. If the environment is designed such that every state is reachable from every state with a single action, why not treat this task like a bandit problem? Why use RL?\n1. Can the authors walk me through the purpose of the theoretical arguments at an intuitive level?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":1},"contribution":{"value":2},"strengths":{"value":"1. The problem is interesting and properly motivated. \n1. The paper provides useful ablation. Ablations vary distortions (linear and non-linear, with and without LPE), observation models (state vs. image), logger expertness, and dimensionality, plus analyses of augmentation frequency/probability and shortcut sampling schemes."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"I currently vote to reject, though I hope to discuss potential misunderstandings with the authors during the rebuttal period.\n\n# Experiments\n\nThe experiments appear narrowly focused on a highly specialized class of problems, characterized by a particular reward structure and transition dynamics. While the reported results are consistent and demonstrate that LIFT outperforms relevant baselines within this setup, the scope of applicability remains somewhat limited. It is unclear whether the observed gains would transfer to more general offline RL settings where the assumptions underlying LIFT may not hold. Expanding the evaluation to include at least one less structured domain (e.g., a standard offline RL benchmark or a stochastic environment with partial observability) would strengthen the paper’s empirical claims and help demonstrate that LIFT’s benefits are not confined to this narrow task family.\n\n# Theory \n\nI found the theoretical section somewhat difficult to follow, particularly because it was not clear why the theoretical machinery was introduced or how it would later support the proposed method. Providing additional context on the purpose of these definitions and propositions (e.g., whether they justify shortcut validity, guarantee improvement, or simply formalize intuitive conditions) would help orient the reader.\n\nIt would also greatly improve readability to include intuitive explanations alongside the formal statements. For example, some of the definitions appear to correspond to well-known RL concepts, but this connection is not made explicit. In particular, I was initially confused by Definition 3.1 and the following propositions, but on closer inspection, they seem to express a familiar idea: that a “$\\pi$-shortcut” is an action a that yields higher return than an action sampled from the policy $\\pi$---in other words, an action with positive advantage. Formally, this interpretation follows from the inequality in Definition 3.1:\n\n$$\\gamma V^{\\pi}(s', W) - V^{\\pi}(s, W) \\ge \\| s' - s_W \\| \\ge 0$$\n\nWe can rewrite this as\n\n$$-\\| s' - s_W \\| \\gamma V^{\\pi}(s', W) \\ge V^{\\pi}(s, W)$$\n\nand then recognizing that $-\\| s' - s_W \\| $ is the reward for taking action $a$ in state $s$, we can write\n\n$$r(s,a) + \\gamma V^{\\pi}(s', W) \\ge V^{\\pi}(s, W)$$\n\nwhich is just the TD-error, an unbiased estimate of the advantage under $pi$. Proposition 3.2 then immediately follows, because choosing an action with positive advantage is by definition better than sampling from $\\pi$. \n\n## Related Works\n\n> it remains unclear how the data-generating logging policy limits what an offline learner can achieve.\n\nMany prior works have studied the importance of high-coverage and near-expert quality data for offline RL. For instance, Corrado et al [1] and Kumar et al [2] emphasize the importance of expert data, while Yarats et. al emphasize data diversity / coverage. Works such as these are worth discussing in the related work.\n\nThere's also a rich literature in data augmentation for non-visual, state-based RL tasks that is not discussed [e.g, 1, 4-7]. For instance, Van de Pohl [5] and Corrado & Hanna [5]  define different classes of augmentations that leverage symmetries in an environment. Corrado et. al [1] referenced in the preceding paragraph also leverages symmetries to generate augmented data. Pitis et. al [6] provides a general framework that captures a lot of these augmentations.\n\n> This illustrates that shortcut identification relies less on global guarantees and more on local structure along trajectory segments.\n\nPitis et. al [6,7] discuss this idea as well. They observe that causal independences in a task's features are not global but local, and use this information to generate augmented data.\n\n1. Corrado et. al. Guided Data Augmentation for Online Reinforcement Learning and Imitation Learning. RLC 2024. https://arxiv.org/abs/2310.18247\n2. Kumar et. al. When Should We Prefer Online Reinforcement Learning Over Behavioral Cloning? ICLR 2022. https://arxiv.org/abs/2204.05618\n3.  Yarats et. al. Don’t Change the Algorithm, Change the Data: Exploratory Data for Online Reinforcement Learning. https://arxiv.org/abs/2201.13425\n4. Van de Pol et. al. MDP Homomorphic Networks: Group Symmetries in Reinforcement Learning. https://arxiv.org/abs/2006.16908\n5. Corrado & Hanna. Understanding when Dynamics-Invariant Data Augmentations Benefit Model-Free Reinforcement Learning Updates. https://arxiv.org/abs/2310.17786\n7. Counterfactual Data Augmentation using Locally Factored Dynamics. Pitis et. al, NeurIPS 2020. https://arxiv.org/abs/2007.02863\n6. MoCoDA: Model-based Counterfactual Data Augmentation. Pitis et. al, NeurIPS 2022. https://arxiv.org/abs/2210.11287"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762927124745,"tcdate":1761986116300,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission17132/Reviewer_j18R"],"signatures":["ICLR.cc/2026/Conference/Submission17132/Reviewer_j18R"],"forum":"AsVH1FQGuR","number":4,"license":"CC BY 4.0","cdate":1761986116300,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission17132/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762927124745,"domain":"ICLR.cc/2026/Conference","replyto":"AsVH1FQGuR","id":"NWQK4TWKbw","forumContent":{"TLDR":{"value":"We introduce a trajectory-based data augmentation method that improves offline reinforcement learning in active positioning tasks by leveraging task structure and geometric properties of rewards, values, and logging policies."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Offline Reinforcement Learning","Reinforcement Learning","Active Position","Off-Policy Learning","Value Function Geometry"]},"supplementary_material":{"value":"/attachment/bc599ad161c10f4d3c9c0b2faba2a5e8b9f311e9.zip"},"primary_area":{"value":"reinforcement learning"},"abstract":{"value":"We propose a method for data augmentation in offline reinforcement learning applied to active positioning problems. \nThe approach enables the training of off-policy models from a limited number of trajectories generated by a suboptimal logging policy.  \nOur method is a trajectory-based augmentation technique that exploits task structure and quantify the effect of admissible perturbations on the data using the geometric\ninterplay of properties of the reward, the value function, and the logging policy.\nMoreover, we show that by training an off-policy model with our augmentation while collecting data, the suboptimal logging policy can be supported \nduring collection, leading to higher data quality and improved offline reinforcement learning performance.\nWe provide theoretical justification for these strategies and validate them empirically across positioning tasks of varying dimensionality and under partial observability."},"_bibtex":{"value":"@misc{\nschmahling2026augmentations,\ntitle={Augmentations in Offline Reinforcement Learning for Active Positioning},\nauthor={Tobias Schm{\\\"a}hling and Matthias Burkhardt and Tobias Windisch},\nyear={2026},\nurl={https://openreview.net/forum?id=AsVH1FQGuR}\n}"},"title":{"value":"Augmentations in Offline Reinforcement Learning for Active Positioning"},"pdf":{"value":"/pdf/457214e7d63837bd6841b8548c5854336912ba8c.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"schmähling|augmentations_in_offline_reinforcement_learning_for_active_positioning"},"authorids":{"value":["~Tobias_Schmähling1","~Matthias_Burkhardt1","~Tobias_Windisch1"]},"authors":{"value":["Tobias Schmähling","Matthias Burkhardt","Tobias Windisch"]}},"version":2},{"content":{"venue":{"value":"IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems"},"pdf":{"value":"https://ieeexplore.ieee.org/iel7/43/9928799/09656540.pdf"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"sabet|similarityaware_cnn_for_efficient_video_recognition_at_the_edge"},"html":{"value":"https://doi.org/10.1109/TCAD.2021.3136815"},"_bibtex":{"value":"<head>\n<META HTTP-EQUIV=\"Refresh\" CONTENT=\"0;URL=/servlet/useragent\">\n</head>\n"},"abstract":{"value":"Convolutional neural networks (CNNs) often extract similar features from successive video frames due to having identical appearances. In contrast, conventional CNNs for video recognition process individual frames with a fixed computational effort. Each video frame is independently processed, resulting in numerous redundant computations and an inefficient use of limited energy resources, particularly for edge computing applications. To alleviate the high energy requirements associated with video frame processing, this article presented similarity-aware CNNs that recognize similar feature pixels across frames and avoid computations on them. First, with a loss of less than 1% in recognition accuracy, a proposed similarity-aware quantization technique increases the average number of unchanged feature pixels across frame pairs by up to 85%. Then, a proposed similarity-aware dataflow improves energy consumption by minimizing redundant computations and memory accesses across frame pairs. According to simulation experiments, the proposed dataflow decreases the energy consumed by video frame processing by up to 30%."},"title":{"value":"Similarity-Aware CNN for Efficient Video Recognition at the Edge"},"authors":{"value":[{"fullname":"Amin Sabet","username":"https://orcid.org/orcid-search/search?searchQuery=Amin%20Sabet"},{"fullname":"Jonathon Hare","username":"~Jonathon_Hare1"},{"fullname":"Bashir M. Al-Hashimi","username":"https://orcid.org/orcid-search/search?searchQuery=Bashir%20M.%20Al-Hashimi"},{"fullname":"Geoff V. Merrett","username":"https://orcid.org/orcid-search/search?searchQuery=Geoff%20V.%20Merrett"}]}},"tmdate":1778118669371,"pdate":1667260800000,"externalIds":["doi:10.1109/tcad.2021.3136815"],"tcdate":1778118665581,"writers":["~","OpenReview.net/Public_Article/ORCID.org","OpenReview.net/Support"],"signatures":["~Jonathon_Hare1"],"forum":"z0t8kW7zoz","license":"CC BY-SA 4.0","number":68425,"cdate":1667520576847,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/ORCID.org/-/Record","OpenReview.net/Public_Article/-/Edit"],"mdate":1778118669371,"domain":"OpenReview.net/Public_Article","id":"z0t8kW7zoz","version":2},{"content":{"venue":{"value":"IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2022"},"pdf":{"value":"https://ieeexplore.ieee.org/iel7/43/9928799/09656540.pdf"},"venueid":{"value":"dblp.org/journals/TCAD/2022"},"paperhash":{"value":"sabet|similarityaware_cnn_for_efficient_video_recognition_at_the_edge"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Amin_Sabet:","~Jonathon_S._Hare1","https://dblp.org/search/pid/api?q=author:Bashir_M._Al-Hashimi:","https://dblp.org/search/pid/api?q=author:Geoff_V._Merrett:"]},"html":{"value":"https://doi.org/10.1109/TCAD.2021.3136815"},"_bibtex":{"value":"@article{DBLP:journals/tcad/SabetHAM22,\n  author={Amin Sabet and Jonathon S. Hare and Bashir M. Al-Hashimi and Geoff V. Merrett},\n  title={Similarity-Aware CNN for Efficient Video Recognition at the Edge},\n  year={2022},\n  cdate={1640995200000},\n  journal={IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.},\n  volume={41},\n  number={11},\n  pages={4901-4914},\n  url={https://doi.org/10.1109/TCAD.2021.3136815}\n}\n"},"abstract":{"value":"Convolutional neural networks (CNNs) often extract similar features from successive video frames due to having identical appearances. In contrast, conventional CNNs for video recognition process individual frames with a fixed computational effort. Each video frame is independently processed, resulting in numerous redundant computations and an inefficient use of limited energy resources, particularly for edge computing applications. To alleviate the high energy requirements associated with video frame processing, this article presented similarity-aware CNNs that recognize similar feature pixels across frames and avoid computations on them. First, with a loss of less than 1% in recognition accuracy, a proposed similarity-aware quantization technique increases the average number of unchanged feature pixels across frame pairs by up to 85%. Then, a proposed similarity-aware dataflow improves energy consumption by minimizing redundant computations and memory accesses across frame pairs. According to simulation experiments, the proposed dataflow decreases the energy consumed by video frame processing by up to 30%."},"title":{"value":"Similarity-Aware CNN for Efficient Video Recognition at the Edge"},"authors":{"value":["Amin Sabet","Jonathon S. Hare","Bashir M. Al-Hashimi","Geoff V. Merrett"]}},"tmdate":1727635302036,"pdate":1640995200000,"tcdate":1727635282241,"writers":["~"],"signatures":["~Jonathon_Hare1"],"forum":"Pg7gpw3dJC","license":"CC BY-SA 4.0","number":107065,"cdate":1640995200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1727635302036,"domain":"DBLP.org","id":"Pg7gpw3dJC","version":2},{"content":{"summary":{"value":"This paper proposes a novel approach called Progressive Autoregressive Video Diffusion Models (PA-VDM) to address the limitations of existing video diffusion models in generating long videos. Instead of using a single noise level for all frames, PA-VDM assigns increasingly higher noise levels to frames during the denoising process. This allows each frame to condition on all previous frames with lower noise levels and provide information for future frames with higher noise levels."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"See the weakness above."},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. PA-VDM extends the capabilities of existing video diffusion models to generate longer videos, up to 1 minute in length (1440 frames at 24 FPS), without compromising quality. This is achieved by assigning latent frames with progressively increasing noise levels, allowing for autoregressive generation without degradation.\n2. PA-VDM maintains temporal consistency throughout the generated video, ensuring smooth transitions and realistic motion dynamics. This is in contrast to other methods that struggle to preserve secondary motion attributes like velocity and acceleration, leading to unnatural movement.\n3. PA-VDM can be easily implemented on top of existing video diffusion models without changing their architectures. This allows for efficient training and generation of long videos, overcoming the limitations of previous approaches that require iterative generation of short clips."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"An important baseline is missed.  \n\nThe increasing-noise diffusion scheduler of PA-VDM is actually a multi-task training process, i.e., the model is trained on mixed types of conditions, and therefore the model can perform both video generation and video extending. Towards this goal, there is a more straightforward fine-tuning strategy than the increasing-noise diffusion scheduler, i.e., directly training the model on both text-to-video and video-to-video data. For example, if the maximum video length during training is L, and we can assign 10% training paris as generating L frames with only text without known frames, 10% pairs with text and one given frame, 10% pairs with text and two given frame, and so on. \n\nPA-VDM needs to be compared with this baseline to show the priority of the increasing-noise diffusion scheduler."}},"nonreaders":[],"tmdate":1731427174403,"tcdate":1730691372780,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission1/Reviewer_2p7H"],"signatures":["ICLR.cc/2025/Conference/Submission1/Reviewer_2p7H"],"forum":"WSze9IIN3d","number":4,"license":"CC BY 4.0","cdate":1730691372780,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission1/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427174403,"domain":"ICLR.cc/2025/Conference","replyto":"WSze9IIN3d","id":"lqu6y4wNJv","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"TLDR":{"value":"We empower video diffusion models to autoregressively generate long videos."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Long Video Generation","Diffusion Models","Transformer","Autoregression"]},"primary_area":{"value":"generative models"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Current frontier video diffusion models have demonstrated remarkable results at\ngenerating high-quality videos. However, they can only generate short video clips,\nnormally around 5 seconds or 120 frames, due to computation limitations during\ntraining. In this work, we show that existing models can be naturally adapted to\nautoregressive video diffusion models without changing the architectures. Our\nkey idea is to assign the latent frames with progressively increasing noise levels\nrather than a single noise level. Thus, each latent can condition on all the less\nnoisy latents before it and provide condition for all the more noisy latents after it.\nSuch progressive video denoising allows our models to autoregressively generate\nframes without quality degradation. We present state-of-the-art results on long\nvideo generation at 1 minute (1440 frames at 24 FPS). Our results are available\nat this anonymous url: https://progressive-autoregressive-vdm.github.io/."},"_bibtex":{"value":"@misc{\nxie2024progressive,\ntitle={Progressive Autoregressive Video Diffusion Models},\nauthor={Desai Xie and Yicong Hong and Hao Tan and Difan Liu and Zhan Xu and Feng Liu and Arie Kaufman and Yang Zhou},\nyear={2024},\nurl={https://openreview.net/forum?id=WSze9IIN3d}\n}"},"title":{"value":"Progressive Autoregressive Video Diffusion Models"},"pdf":{"value":"/pdf/36e86237d02dc4cf7c026ca9176a30309a69bc38.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"xie|progressive_autoregressive_video_diffusion_models"},"authorids":{"value":["~Desai_Xie1","~Yicong_Hong1","~Hao_Tan1","~Difan_Liu2","~Zhan_Xu1","~Feng_Liu6","~Arie_Kaufman1","~Yang_Zhou10"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Desai Xie","Yicong Hong","Hao Tan","Difan Liu","Zhan Xu","Feng Liu","Arie Kaufman","Yang Zhou"]}},"version":2},{"content":{"summary":{"value":"This paper proposes ViML, a large-scale multi-modality dataset of 3M video-caption pairs. A systemic captioning framework is proposed to generate music, video, and frame captions separately and a multimodal caption is then merged by uni-modal captions along with objects and background labels."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"My major concerns are listed in the weaknesses part above. I only have minor questions here.\n\n1. What is the data source? Are all videos selected from YouTube? Will the data source be overlapped with some existing datasets, such as HD-Vila and YT-Temporal-1b?\n\n2. The average length is 4.6s, which is a relatively short length compared with other datasets. Why not adopt the stitching procedure in Panda-70M to make the split video longer?  \n\n3. I believe the proposed dataset is a perfect data source for sounding video generation. Have the authors tested this capability using some SVG model, such as MM-Diffusion [1] or MM-LDM [2]?  \n\n4. For music generation capabilities, are there any comparable results? I believe the text-to-music models and video-to-music models should achieve better results after training on the proposed dataset.\n\nReference:   \n[1] Ruan, Ludan, et al. \"Mm-diffusion: Learning multi-modal diffusion models for joint audio and video generation.\" Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2023.  \n[2] Sun, Mingzhen, et al. \"MM-LDM: Multi-Modal Latent Diffusion Model for Sounding Video Generation.\" arXiv preprint arXiv:2410.01594 (2024)."},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":1},"strengths":{"value":"1. Large-scale music caption datasets are scarce. A large-scale video-language dataset with high-quality music narration benefits the community. The proposed dataset also features high image quality, aesthetic scores, and motion characteristics, which shows promising potential for training high-quality text-to-video/text-to-audio&video generation models.\n\n2. The data curation pipeline looks sound and intuitive. The data filtering strategy has proved to be effective through ablations.\n\n3. Experiments on several downstream tasks reveal the high-quality nature of the proposed dataset."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The entire caption pipeline is trivial and not inspiring. Several similar data curation pipelines have been proposed [1-2], and the authors selected the existing MusicCaps model as the music captioner. Therefore, the data pipeline is more likely to be an integration of several existing uni-modal captioners, and even the integration tool(an LLM) has been used by several previous methods, making the technical contribution insufficient. \n\n2. The manuscript lacks ablations. Why select Coca as the image captioner and LLaVA as the video captioner? Can LLaVa be applied to generate image captions since its performance is better than CoCa? Why not use video-encoder-based MLLM as the video captioner such as VideoChat [3] and Video-LLaVa [4]? More clarifications are needed.  \n\n3. Experiments are inadequate. For video generation, the authors selected WebVid as the baseline, which is a relatively earlier dataset. How about the comparison with some newer datasets such as LVD-2M [5] or Panda-70M [6]? Besides, VBench has 16 dimensions, yet the authors only report 9, what is the rationale? For caption, the authors should test more caption datasets such as Valor32k [7], FAVD [8], and DiDeMo [9] that include rich audio-visual content.\n\nReference:  \n[1] Chen, Sihan, et al. \"Vast: A vision-audio-subtitle-text omni-modality foundation model and dataset.\" Advances in Neural Information Processing Systems 36 (2023): 72842-72866.  \n[2] Wang, Yi, et al. \"Internvideo2: Scaling video foundation models for multimodal video understanding.\" arXiv preprint arXiv:2403.15377 (2024).  \n[3] Li, KunChang, et al. \"Videochat: Chat-centric video understanding.\" arXiv preprint arXiv:2305.06355 (2023).  \n[4] Lin, Bin, et al. \"Video-llava: Learning united visual representation by alignment before projection.\" arXiv preprint arXiv:2311.10122 (2023).  \n[5] Xiong, Tianwei, et al. \"LVD-2M: A Long-take Video Dataset with Temporally Dense Captions.\" arXiv preprint arXiv:2410.10816 (2024).  \n[6] Chen, Tsai-Shien, et al. \"Panda-70m: Captioning 70m videos with multiple cross-modality teachers.\" Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2024.  \n[7] Chen, Sihan, et al. \"Valor: Vision-audio-language omni-perception pretraining model and dataset.\" arXiv preprint arXiv:2304.08345 (2023).  \n[8] Shen, Xuyang, et al. \"Fine-grained audible video description.\" Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2023.  \n[9] Anne Hendricks, Lisa, et al. \"Localizing moments in video with natural language.\" Proceedings of the IEEE international conference on computer vision. 2017."}},"nonreaders":[],"tmdate":1731428759955,"tcdate":1730699269310,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission6844/Reviewer_Qes1"],"signatures":["ICLR.cc/2025/Conference/Submission6844/Reviewer_Qes1"],"forum":"Tgsc0KEkN6","number":4,"license":"CC BY 4.0","cdate":1730699269310,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission6844/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428759955,"domain":"ICLR.cc/2025/Conference","replyto":"Tgsc0KEkN6","id":"uivfV8BKfN","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video Understanding and Generation","Dataset and Benchmark","Multimodal","Music"]},"primary_area":{"value":"datasets and benchmarks"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Integrating multimodal understanding and generation into a unified framework can bridge the domain gap across different modalities. \nHowever, existing multimodal-language datasets predominantly offer text descriptions for a single modality, treating visual and audio as separate tasks. This approach neglects the inherent audio-visual correlations, resulting in annotations that are often monotonous and modality-specific rather than comprehensive and precise. Such oversight hampers the advancement of cross-modality research. To fulfill this gap, we present ViML, a large-scale multi-modality-to-language dataset incorporating 3M video clips with high-quality multimodal captions. \nIn ViML, we propose a systemic captioning framework, achieving various modality annotations with more than 12.2k hours of trailer videos. Here, to ensure the caption retains music perspective while preserving the authority of visual context, we leverage the advanced LLM to merge all annotations adaptively. In particular, the ViML has two main advantages: \n(1) the topics are diverse, and the content characters are of various types, \\eg, film,  news, and gaming.\n(2) the corresponding background music is custom-designed, making it more coherent with the visual context. \nIn this fashion, our ViML dataset potentially paves the path for fine-grained large multimodal-language model training. In experiments, we provide evaluation metrics and benchmark results on our dataset, demonstrating the high quality of our annotation and its effectiveness for model training. We include demo data in \\hyperlink{https://anonymous.4open.science/w/ViML-4C78}{https://anonymous.4open.science/w/ViML-4C78}"},"_bibtex":{"value":"@misc{\nchi2024viml,\ntitle={Vi{ML}: A Video, Music, Language Unified Dataset for Understanding and Generation},\nauthor={Xiaowei Chi and Aosong Cheng and Yatian Wang and Pengjun Fang and Zeyue Tian and Yingqing He and Xingqun Qi and Zhaoyang Liu and Rongyu Zhang and Qifeng Chen and Wenhan Luo and Qifeng Liu and Shanghang Zhang and Wei Xue and Yike Guo},\nyear={2024},\nurl={https://openreview.net/forum?id=Tgsc0KEkN6}\n}"},"title":{"value":"ViML: A Video, Music, Language Unified Dataset for Understanding and Generation"},"pdf":{"value":"/pdf/70f659dc0c3cc0709ee9eaa73fa8721394ebcb55.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"chi|viml_a_video_music_language_unified_dataset_for_understanding_and_generation"},"authorids":{"value":["~Xiaowei_Chi1","~Aosong_Cheng1","~Yatian_Wang1","~Pengjun_Fang1","~Zeyue_Tian2","~Yingqing_He1","~Xingqun_Qi1","~Zhaoyang_Liu1","~Rongyu_Zhang1","~Qifeng_Chen1","~Wenhan_Luo1","~Qifeng_Liu1","~Shanghang_Zhang4","~Wei_Xue5","~Yike_Guo1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Xiaowei Chi","Aosong Cheng","Yatian Wang","Pengjun Fang","Zeyue Tian","Yingqing He","Xingqun Qi","Zhaoyang Liu","Rongyu Zhang","Qifeng Chen","Wenhan Luo","Qifeng Liu","Shanghang Zhang","Wei Xue","Yike Guo"]}},"version":2},{"content":{"summary":{"value":"This paper studies safety in Video LLMs and introduces a new benchmark and defense pipeline. The authors build VideoSafetyBench (VSB-77k), a large, policy-grounded dataset that contains 77,646 video–query pairs, 19 subcategories across 6 risk categories and 10 language communities, and derive an 11.4k evaluation split VSB-Eval and a 46k post-training set with CoT (VSB-R1-46k). The author shows that adding the video modality substantially degrades safety, highlighting systemic vulnerabilities unique to video inputs. To address this, they propose VideoSafety-R1, combining Alarm Token-Guided Safety Fine-Tuning (AT-SFT) with Safety-Guided GRPO to inject modality-aware “alarm” tokens and then reinforce defensive reasoning with rule-based rewards tied to dual-modality verification. Extensive experiments validate VideoSafety-R1’s effectiveness on the adversarial VSE-HH benchmark, it improves VideoLLaMA3-2B’s DSR by 71.1% (from 18.4% to 89.5%). This approach also generalizes well to image safety benchmarks, boosting DSR by 59.1% on MMBench, 44.3% on VLGuard, and 15.0% on FigStep. It maintains strong utility and avoids over-defense while outperforming existing methods like SPA-VL and VLGuard."},"soundness":{"value":4},"confidence":{"value":4},"questions":{"value":"1. Could the current “dual-modality verification” be extended to a temporal–text–vision triad (e.g., segment-level harmful annotations/rewards, time-aware alarm tokens, or temporal grounding constraints)?\n\n2. Could you evaluate a control set where a single malicious image is duplicated to N frames (no temporal variation) and check if safety still collapses as N grows, that implicates token load, not temporal content."},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":4},"strengths":{"value":"Clear evidence that video hurts safety: The paper provides quantitative analyses showing that integrating video can sharply reduce DSR relative to text-only scenes (e.g., −79.4% for VideoLLaMA3-2B), and that harmful videos amplify adversarial effectiveness. This is an important and under-explored finding that justifies a video-specific safety agenda.\n\nWell-designed benchmark and pipeline: VSB-77k is large, multilingual, and aligned to platform policies. The construction pipeline that consists of filtered video collection, multi-agent LLM annotation, and template-driven query generation is also clear. For metrics, the paper also sets DSR, Helpfulness score and FRR to systematically evaluate.\n\nEffective two-stage defense with clear ablations: AT-SFT (learnable modality-specific alarm tokens + multitask alarm-classification objectives) and Safety-Guided GRPO are proved to be complementary. Ablations show each component helps and the combination yields the best overall safety, while keeping original video understanding essentially unchanged. The method substantially boosts DSR on VSB-Eval-HH and transfers across external safety benchmarks."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"No coverage of “dynamic adversarial attacks”: Although the task is video, the method/evaluation effectively treats it as a set of sampled frames, rather than modeling sequence-level risk. The benchmark’s “harmful video” cases are largely explicitly harmful (e.g., direct fights, weapon displays). It does not include more realistic implicit dynamic attacks, for example, first ~10 frames benign, later ~20 frames contain fragmented harmful content or the video shows a seemingly legal tool but frame order implies harmful use. These higher-frequency, real-world attack patterns are absent, so the reported safety may overestimate robustness in deployment.\n\nToken-load confound (not controlled): The paper attributes safety degradation to the video modality, but it does not rule out a simpler cause \"large visual token loads\". If one constructs a “pseudo-video” by repeating the same malicious image across many frames, the model receives far more visual tokens without adding temporal information. If defenses fail primarily because of attention/kv-cache saturation or context dilution under high visual token counts, then the reported vulnerability is not uniquely video-specific."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762925538995,"tcdate":1761872600579,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission15243/Reviewer_8QWg"],"signatures":["ICLR.cc/2026/Conference/Submission15243/Reviewer_8QWg"],"forum":"QuW5RUDwMo","number":3,"license":"CC BY 4.0","cdate":1761872600579,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission15243/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762925538995,"domain":"ICLR.cc/2026/Conference","replyto":"QuW5RUDwMo","id":"Eyz6PPe538","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Video Large Language Model","Safety of Multimodal Large Language Model","Safety Alignment","RLHF"]},"primary_area":{"value":"alignment, fairness, safety, privacy, and societal considerations"},"abstract":{"value":"While the safety risks of image-based large language models (Image LLMs) have been extensively studied, their video-based counterparts (Video LLMs) remain critically under-examined. To systematically study this problem, we introduce \\textbf{VideoSafetyEval} -- a large-scale, real-world benchmark for Video LLM safety, which comprises 11.4k video-query pairs and spans 19 principal risk categories. Based on this, \\textit{we reveal that integrating video modality degrades safety performance by an average of 34.2\\%, thereby exposing systemic risks in multimodal attack exploitation.}\nTo address this vulnerability, we propose \\textbf{VideoSafety-R1}, a dual-stage framework achieving unprecedented safety gains through three innovations: (1) VideoSafetyThinking dataset contains 46k video-query–thinking response triplets.  (2) Alarm Token-Guided Safety Fine-Tuning (AT-SFT) injects learnable alarm tokens into visual and textual sequences, enabling explicit harm perception across modalities via multitask objectives. (3) Safety-guided GRPO enhances defensive reasoning through dynamic policy optimization with rule-based rewards derived from dual-modality verification. These components synergize to shift safety alignment from harm perception to active reasoning. The framework achieves a 71.1\\% improvement on VSE-HH, and improves by 59.1\\%, 44.3\\%, and 15.0\\% on the image safety datasets MMBench, VLGuard, and FigStep, respectively. Our code and dataset are  available at \\url{https://github.com/Emiya-syw/VideoSafety-R1.git}.\n\\textcolor{red}{Note: This paper contains harmful language and image examples, and reader discretion is recommended.}"},"_bibtex":{"value":"@inproceedings{\nsun2026from,\ntitle={From Evaluation to Defense: Advancing Safety in Video Large Language Models},\nauthor={Yiwei Sun and Peiqi Jiang and Chuanbin Liu and Luohao Lin and Zhiying Lu and Hongtao Xie},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=QuW5RUDwMo}\n}"},"title":{"value":"From Evaluation to Defense: Advancing Safety in Video Large Language Models"},"pdf":{"value":"/pdf/02077ad527117ebc8263104650597329d84c5fe2.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"sun|from_evaluation_to_defense_advancing_safety_in_video_large_language_models"},"authorids":{"value":["~Yiwei_Sun3","~Peiqi_Jiang3","~Chuanbin_Liu2","~Luohao_Lin1","~Zhiying_Lu1","~Hongtao_Xie2"]},"authors":{"value":["Yiwei Sun","Peiqi Jiang","Chuanbin Liu","Luohao Lin","Zhiying Lu","Hongtao Xie"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a text-to-video method by using a pre-trained text-to-image model as the initialization. The authors present two key findings: 1) temporal self-attention with relative position encoding could maintain temporal consistency well. 2) joint image-video fine-tuning is crucial for good performance. This paper also contributes a video dataset named Vimeo25M, with 25 million text-video pairs."},"presentation":{"value":"3 good"},"contribution":{"value":"2 fair"},"soundness":{"value":"3 good"},"strengths":{"value":"1. The paper is well-written and easy to follow, and the visualization results seem appealing. \n2. The paper proposes a new video-text datasets named Vimeo25M."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The technical novelty is limited. The authors present two findings which have both been proposed in previous works. 1) temporal attention over a pretrained text-to-image model could address the temporal consistency model well. This idea is first stated in “AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning” and then pointed out in \"MagicEdit: High-Fidelity and Temporally Coherent Video Editing\". 2) Joint image-video training is a well-known technique to improve the effectiveness of training text-to-video models. It was first proposed by Jonathan Ho et al. in \"Video Diffusion Models\". \n2.  The idea of the cascaded diffusion model, which first trains a base model then temporally interpolates it, and finally performs super-resolution spatially, is a standard procedure to learn a text-to-video model. Similar ideas are also proposed in \"Make-A-Video: Text-to-video Generation Without Text-Video Data\", \"Imagen Video: High Definition Video Generation with Diffusion Models\" and \"Align your Latents: High-Resolution Video Synthesis with Latent Diffusion Models\"."},"confidence":{"value":"5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully."},"questions":{"value":"Will the authors release the newly proposed video dataset Vimeo25M? Since it seems that there is no promise from the authors that they will release the dataset. I would say that the major contribution of the paper is the dataset."},"rating":{"value":"5: marginally below the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636218413,"tcdate":1698724844246,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission2758/Reviewer_aq4u"],"signatures":["ICLR.cc/2024/Conference/Submission2758/Reviewer_aq4u"],"forum":"p09XyFxZkc","number":2,"license":"CC BY 4.0","cdate":1698724844246,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission2758/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636218413,"domain":"ICLR.cc/2024/Conference","replyto":"p09XyFxZkc","id":"d2jBwSaWTJ","forumContent":{"venue":{"value":"Submitted to ICLR 2024"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["text-to-video generation","diffusion models"]},"supplementary_material":{"value":"/attachment/9b7be39ae37699c99be3e7c7ae7b8d3b62214416.zip"},"primary_area":{"value":"generative models"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"This work aims to learn a high-quality text-to-video (T2V) generative model by leveraging a pre-trained text-to-image (T2I) model as a basis. It is a highly desirable yet challenging task to simultaneously a) accomplish the synthesis of visually realistic and temporally coherent videos while b) preserving the strong creative generation nature of the pre-trained T2I model. To this end, we propose LaVie, an integrated video generation framework that operates on cascaded video latent diffusion models, comprising a base T2V model, a temporal interpolation model, and a video super-resolution model. Our key insights are two-fold: 1) We reveal that the incorporation of simple temporal self-attentions, coupled with relative positional encoding, adequately captures the temporal correlations inherent in video data. 2) Additionally, we validate that the process of joint image-video fine-tuning plays a pivotal role in producing high-quality and creative outcomes. To enhance the performance of LaVie, we contribute a comprehensive and diverse video dataset named Vimeo25M, consisting of 25 million text-video pairs that prioritize quality, diversity, and aesthetic appeal. Extensive experiments demonstrate that LaVie achieves state-of-the-art performance both quantitatively and qualitatively. Furthermore, we showcase the versatility of pre-trained LaVie models in various long video generation and personalized video synthesis applications."},"_bibtex":{"value":"@misc{\nwang2024lavie,\ntitle={LaVie: High-Quality Video Generation with Cascaded Latent Diffusion Models},\nauthor={Yaohui Wang and Xinyuan Chen and Xin Ma and Shangchen Zhou and Ziqi Huang and Yi Wang and Ceyuan Yang and Yinan He and Jiashuo Yu and Peiqing Yang and Yuwei Guo and Tianxing Wu and Chenyang Si and Yuming Jiang and Cunjian Chen and Chen Change Loy and Bo Dai and Dahua Lin and Yu Qiao and Ziwei Liu},\nyear={2024},\nurl={https://openreview.net/forum?id=p09XyFxZkc}\n}"},"title":{"value":"LaVie: High-Quality Video Generation with Cascaded Latent Diffusion Models"},"pdf":{"value":"/pdf/f40f1ea778821a8640f919df133835970623a510.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Rejected_Submission"},"paperhash":{"value":"wang|lavie_highquality_video_generation_with_cascaded_latent_diffusion_models"},"authorids":{"value":["~Yaohui_Wang1","~Xinyuan_Chen1","~Xin_Ma3","~Shangchen_Zhou1","~Ziqi_Huang2","~Yi_Wang19","~Ceyuan_Yang2","~Yinan_He1","~Jiashuo_Yu1","~Peiqing_Yang1","~Yuwei_Guo1","~Tianxing_Wu2","~Chenyang_Si2","~Yuming_Jiang1","~Cunjian_Chen2","~Chen_Change_Loy2","~Bo_Dai2","~Dahua_Lin1","~Yu_Qiao1","~Ziwei_Liu1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Yaohui Wang","Xinyuan Chen","Xin Ma","Shangchen Zhou","Ziqi Huang","Yi Wang","Ceyuan Yang","Yinan He","Jiashuo Yu","Peiqing Yang","Yuwei Guo","Tianxing Wu","Chenyang Si","Yuming Jiang","Cunjian Chen","Chen Change Loy","Bo Dai","Dahua Lin","Yu Qiao","Ziwei Liu"]}},"version":2},{"content":{"summary":{"value":"This paper introduces Vinoground, a benchmark designed to evaluate the temporal reasoning abilities of large multimodal models (LMMs) within the context of short videos. Vinoground comprises counterfactual video-caption pairs, where each pair of captions, generated by GPT-4, contains identical words arranged in different orders. The corresponding videos are collected from two sources: the VATEX dataset and the YouTube platform. The data in Vinoground is systematically categorized into three major categories: *object, action*, and *viewpoint*, along with four fine-grained subcategories: *interaction, cyclical, spatial*, and *contextual*. LMMs are tasked with distinguishing between a pair of counterfactual captions given a video (referred to as text scoring) or between a pair of counterfactual videos given a caption (referred to as video scoring). Empirical results show that current LMMs, including state-of-the-art models like GPT-4, still lag far behind human performance in understanding the temporal dynamics of short videos. They particularly struggle with recognizing fine-grained temporal differences, such as *interaction*, *spatial* (direction), and *cyclical* actions."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"N/A"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- Vinoground is specifically designed to expose the true temporal reasoning capabilities of LMMs. By using counterfactual video-caption pairs, it prevents LMMs from relying on single-frame information or language biases, as evidenced by the near random-guessing performance of GPT-4o, when provided with zero or only a single video frame.\n- The experiments are comprehensive, encompassing 12 advanced generative LMMs and three CLIP-based models. A detailed analysis is provided on the impact of input frames, the effectiveness of the Chain-of-Thought (CoT) strategy, and the fine-grained performance across different categories.\n- The paper is generally well-written and easy to follow."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Despite its strengths, the contribution of this work may not be substantial enough for a top-tier conference like ICLR. Compared to existing temporal reasoning benchmarks [1,2,3], the primary innovation of Vinoground lies in the introduction of natural negative videos and a more effective mitigation of single-frame and language biases. While these features do increase the difficulty of the benchmark, current LMMs still show significant room for improvement on existing benchmarks [1,2,3].\n- Human performance on the group score is only 90%, indicating potential issues with the quality of the data examples.\n- It seems that some critical details are missing when it comes to the generation of counterfactual captions, such as the prompt being used and how many captions are generated.\n- A related benchmark [3] is not mentioned or discussed in the paper.\n\n[1] TempCompass: Do video LLMs really understand videos?\n\n[2] Vitatecs: A diagnostic dataset for temporal concept understanding of video-language models.\n\n[3] MVBench: A Comprehensive Multi-modal Video Understanding Benchmark"}},"nonreaders":[],"tmdate":1732714429143,"tcdate":1729934671730,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission1636/Reviewer_F3rC"],"signatures":["ICLR.cc/2025/Conference/Submission1636/Reviewer_F3rC"],"forum":"a1P5kh2oo8","number":2,"license":"CC BY 4.0","cdate":1729934671730,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission1636/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732714429143,"domain":"ICLR.cc/2025/Conference","replyto":"a1P5kh2oo8","id":"nB8ad67CA4","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"TLDR":{"value":"Modern SoTA LMMs still demonstrates subpar performance at temporal reasoning with our temporal counterfactual benchmark composed of natural videos."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["temporal reasoning; counterfactual reasoning; short video comprehension"]},"supplementary_material":{"value":"/attachment/fd093e5ac3c36911b68e297a2932dcf6ce1ed019.zip"},"primary_area":{"value":"datasets and benchmarks"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"There has been growing sentiment recently that modern large multimodal models (LMMs) have addressed most of the key challenges related to short video comprehension. As a result, both academia and industry are gradually shifting their attention towards the more complex challenges posed by understanding long-form videos. \nHowever, is this really the case?  Our studies indicate that LMMs still lack many fundamental reasoning capabilities even when dealing with short videos.  We introduce Vinoground, a temporal counterfactual LMM evaluation benchmark encompassing 1000 short and natural video-caption pairs. We demonstrate that existing LMMs severely struggle to distinguish temporal differences between different actions and object transformations.  For example, the best model GPT-4o only obtains $\\sim$50\\% on our text and video scores, showing a large gap compared to the human baseline of $\\sim$90\\%. All open-source multimodal models and CLIP-based models perform much worse, producing mostly random chance performance. Through this work, we shed light onto the fact that temporal reasoning in short videos is a problem yet to be fully solved. We will make our benchmark publicly available."},"_bibtex":{"value":"@misc{\nzhang2025vinoground,\ntitle={Vinoground: Scrutinizing {LMM}s over Dense Temporal Reasoning with Short Videos},\nauthor={Jianrui Zhang and Mu Cai and Yong Jae Lee},\nyear={2025},\nurl={https://openreview.net/forum?id=a1P5kh2oo8}\n}"},"title":{"value":"Vinoground: Scrutinizing LMMs over Dense Temporal Reasoning with Short Videos"},"pdf":{"value":"/pdf/2e0dd590d0c11d11d7071fa71d7eaea83ded91ac.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"zhang|vinoground_scrutinizing_lmms_over_dense_temporal_reasoning_with_short_videos"},"authorids":{"value":["~Jianrui_Zhang1","~Mu_Cai1","~Yong_Jae_Lee2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Jianrui Zhang","Mu Cai","Yong Jae Lee"]}},"version":2},{"content":{"summary":{"value":"This paper presents LoopFormer, which allows reasoning under variable compute budgets. It extends prior looped Transformers by introducing time- and step-size modulation, where each iteration receives sinusoidal embeddings of layer index and step size to dynamically modulate RMSNorm and residual scaling.\nA shortcut-consistency loss aligns representations across different loop lengths, enabling stable performance even with fewer inference steps."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"NA"},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"* The paper is well written and easy to follow, with clear motivation and setups\n* The motivation is clearly presented, and the transition from fixed-depth looped Transformers to elastic-depth design feels natural.\n* Experiments are reasonably comprehensive, evaluating both variable loop lengths and the effect of the proposed shortcut-consistency loss.\n* The main claims are well supported"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"* The degree of novelty is not bad but moderate. While the proposed elastic-depth formulation and shortcut-consistency loss are well designed, they extend existing time-modulated looped Transformer frameworks rather than introducing a fundamentally new paradigm.\n* the paper does not provide theoretical intuition or analysis explaining why combining t and $\\Delta t$ through sinusoidal modulation is a good choice here"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762941848391,"tcdate":1761459340372,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission21596/Reviewer_xXws"],"signatures":["ICLR.cc/2026/Conference/Submission21596/Reviewer_xXws"],"forum":"RzYXb5YWBs","number":2,"license":"CC BY 4.0","cdate":1761459340372,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission21596/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762941848391,"domain":"ICLR.cc/2026/Conference","replyto":"RzYXb5YWBs","id":"j2942wDN1I","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["large language models","LLMs","reasoning","looped transformers","efficient inference","parameter sharing"]},"primary_area":{"value":"foundation or frontier models, including LLMs"},"abstract":{"value":"Looped Transformers have emerged as an efficient and powerful class of models for reasoning in the language domain. Recent studies show that these models achieve strong performance on algorithmic and reasoning tasks, suggesting that looped architectures possess an inductive bias toward latent reasoning. However, prior approaches fix the number of loop iterations during training and inference, leaving open the question of whether these models can flexibly adapt their computational depth under variable compute budgets. We introduce LoopFormer, a looped Transformer trained on variable-length trajectories to enable budget-conditioned reasoning. Our core contribution is a shortcut-consistency training scheme that aligns trajectories of different lengths, ensuring that shorter loops yield informative representations while longer loops continue to refine them. LoopFormer conditions each loop on the current time and step size, enabling representations to evolve consistently across trajectories of varying length rather than drifting or stagnating. Empirically, LoopFormer demonstrates robust performance on language modeling and reasoning benchmarks even under aggressive compute constraints, while scaling gracefully with additional budget. These results show that looped Transformers are inherently suited for adaptive language modeling, opening a path toward controllable and budget-aware large language models."},"_bibtex":{"value":"@inproceedings{\njeddi2026loopformer,\ntitle={LoopFormer: Elastic-Depth Looped Transformers for Latent Reasoning via Shortcut Modulation},\nauthor={Ahmadreza Jeddi and Marco Ciccone and Babak Taati},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=RzYXb5YWBs}\n}"},"title":{"value":"LoopFormer: Elastic-Depth Looped Transformers for Latent Reasoning via Shortcut Modulation"},"pdf":{"value":"/pdf/8e5f55b4cd00792565dbc6e38638fe4c3eafe7c6.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"jeddi|loopformer_elasticdepth_looped_transformers_for_latent_reasoning_via_shortcut_modulation"},"authorids":{"value":["~Ahmadreza_Jeddi2","~Marco_Ciccone1","~Babak_Taati1"]},"authors":{"value":["Ahmadreza Jeddi","Marco Ciccone","Babak Taati"]}},"version":2},{"content":{"summary":{"value":"This paper introduces Ref-Adv, a new benchmark for Referring Expression Comprehension (REC) designed to address the limitations of existing datasets like RefCOCO, RefCOCO+, and RefCOCOg. The authors argue that these benchmarks are overly simplistic and allow models to exploit shortcuts rather than perform genuine visual reasoning. Ref-Adv is constructed using a semi-automated pipeline combining LLM-generated expressions and human verification, with a focus on hard distractors, minimal sufficiency, and negation. The paper includes comprehensive ablation studies and evaluations of modern MLLMs, showing a significant performance drop on Ref-Adv compared to traditional benchmarks."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Please refer to Weaknesses."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The paper addresses a well-known and important issue in the REC community — the overestimation of model capabilities due to dataset shortcuts. The motivation is clear and well-argued.\n2. The two-stage pipeline (LLM-authored + human verification) is well-designed. The inclusion of hard distractors, minimal sufficiency, and negation adds meaningful complexity.\n3. The evaluation covers a wide range of MLLMs (both open and closed-source), includes ablation studies, and uses multiple IoU thresholds, which strengthens the empirical analysis."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. While the dataset is more challenging, the core task (REC) remains unchanged. The paper does not propose a new task formulation or evaluation protocol beyond traditional bounding box accuracy. The idea of “hard distractors” and “minimal sufficiency” is not entirely new — similar ideas have been explored in prior work (e.g., Cops-Ref, FineCops-Ref).\n2. The dataset is curated from COCO and OpenImages. It would be beneficial to explore how well Ref-Adv generalizes to more diverse or out-of-domain images\n3. The reliance on IoU-based accuracy is standard but limited. It does not capture partial correctness or reasoning steps. For example, the evaluation should reflect that the performance drop is only from the difficulty of object localization or distractors. The model locates the distractor or cannot localize the object correctly.\n4. As shown in Table 3 and 4, the word order removal and descriptor deletion are used to demonstrate that the proposed benchmark avoids the model from relying on the Grounding Shortcut. However, the performance gap is marginal compared with RefCOCO.\n5. The integration of COT also brings marginal performance gain, e.g., the performance of Qwen vl-3B decreases with COT. Therefore, we cannot conclude that the proposed benchmark relies on the reasoning.\n6. This paper only gives a conclusion and a benchmark. However, no solution is presented, which limits the contribution of this paper. In addition, only the evaluation data is provided. Even though we are aware of this conclusion, what can the community do to address such a problem?"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1764358755700,"tcdate":1761463377139,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission16232/Reviewer_EUYJ"],"signatures":["ICLR.cc/2026/Conference/Submission16232/Reviewer_EUYJ"],"forum":"iEBgrepR9i","number":1,"license":"CC BY 4.0","cdate":1761463377139,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission16232/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1764358755700,"domain":"ICLR.cc/2026/Conference","replyto":"iEBgrepR9i","id":"q5rZJ1DHaF","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["MLLM","Referring Expression Comprehensions"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Referring Expression Comprehension (REC) links language to region level visual\nperception. Standard benchmarks (RefCOCO, RefCOCO+, RefCOCOg) have\nprogressed rapidly with multimodal LLMs but remain weak tests of visual reasoning and grounding: (i) many expressions are very short, leaving little reason-\ning demand; (ii) images often contain few distractors, making the target easy to\nfind; and (iii) redundant descriptors enable shortcut solutions that bypass genuine\ntext understanding and visual reasoning. We introduce Ref-Adv, a modern REC\nbenchmark that suppresses shortcuts by pairing linguistically nontrivial expressions with only the information necessary to uniquely identify the target. The\ndataset contains various expressions on real images, curated with hard distractors and annotated with reasoning facets including negation. We conduct comprehensive ablations (word order perturbations and\ndescriptor deletion sufficiency) to show that solving Ref-Adv requires reasoning\nbeyond simple cues, and we evaluate a broad suite of contemporary multimodal\nLLMs on Ref-Adv. Despite strong results on RefCOCO, RefCOCO+, and RefCOCOg, models drop markedly on Ref-Adv, revealing reliance on shortcuts and\ngaps in visual reasoning and grounding. We provide an in depth failure analysis\nand aim for Ref-Adv to guide future work on visual reasoning and grounding in\nMLLMs.\nThe dataset is available at \\url{https://ref-adv.github.io/}."},"_bibtex":{"value":"@inproceedings{\ndong2026refadv,\ntitle={Ref-Adv: Exploring {MLLM} Visual Reasoning in Referring Expression Tasks},\nauthor={Qihua Dong and Kuo Yang and Lin Ju and Handong Zhao and Yitian Zhang and Yizhou Wang and Huimin Zeng and Jianglin Lu and Yun Fu},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=iEBgrepR9i}\n}"},"title":{"value":"Ref-Adv: Exploring MLLM Visual Reasoning in Referring Expression Tasks"},"pdf":{"value":"/pdf/9be59c3a7f7bd55caf7a10db58375ffb778a8649.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"dong|refadv_exploring_mllm_visual_reasoning_in_referring_expression_tasks"},"authorids":{"value":["~Qihua_Dong2","~Kuo_Yang6","~Lin_Ju2","~Handong_Zhao3","~Yitian_Zhang1","~Yizhou_Wang3","~Huimin_Zeng2","~Jianglin_Lu2","~Yun_Fu1"]},"authors":{"value":["Qihua Dong","Kuo Yang","Lin Ju","Handong Zhao","Yitian Zhang","Yizhou Wang","Huimin Zeng","Jianglin Lu","Yun Fu"]}},"version":2},{"content":{"review":{"value":"CorrMask mitigates shortcut learning in scFMs by jointly masking covariance-derived gene cliques instead of independent tokens. Improves sample efficiency (~3×), gene-level generalization, and rare cell-type performance with minimal overhead.\n\n### Pros\n- Masking strategy explicitly aligned with gene co-regulation structure.\n- Architecture-agnostic, simple masking modification.\n- Strong gains on underrepresented populations.\n\n### Cons\n- No ablation on dependency graph noise sensitivity or graph construction hyperparameters.\n- Empirical validation restricted primarily to Geneformer backbone."},"confidence":{"value":3},"rating":{"value":8},"title":{"value":"Mitigating Shortcut Learning in scFMs with Correlation-Guided Masking"}},"parentInvitations":"ICLR.cc/2026/Workshop/LMRL/-/Official_Review","nonreaders":[],"tmdate":1772451821967,"tcdate":1771919097351,"writers":["ICLR.cc/2026/Workshop/LMRL","ICLR.cc/2026/Workshop/LMRL/Submission60/Reviewer_jZga"],"signatures":["ICLR.cc/2026/Workshop/LMRL/Submission60/Reviewer_jZga"],"forum":"DmEPWRT0oE","number":2,"license":"CC BY 4.0","cdate":1771919097351,"readers":["everyone"],"invitations":["ICLR.cc/2026/Workshop/LMRL/Submission60/-/Official_Review","ICLR.cc/2026/Workshop/LMRL/-/Edit"],"mdate":1772451821967,"domain":"ICLR.cc/2026/Workshop/LMRL","replyto":"DmEPWRT0oE","id":"A6dixf8PCF","forumContent":{"TLDR":{"value":"Addressing a limitation of random masking to capture meaningful representations in single-cell foundation models, we present CorrMask, a correlation-based masking strategy that demonstrates better generalization on both gene and cell-level tasks."},"venue":{"value":"ICLR 2026 Workshop LMRL Poster"},"keywords":{"value":["Single-Cell Foundation Models","Single-Cell RNA Sequencing","Masked Language Modeling","Inductive Bias in Biology","Representation Learning","Self-Supervised Learning"]},"confirmation":{"value":"I have read and agree with the workshop's policy on behalf of myself and my co-authors."},"abstract":{"value":"Many single-cell foundation models (scFMs) learn representations of cellular identity through masked modeling of gene expression, yet standard random masking treats genes as independent tokens, a poor match for the modular, co-regulated structure of gene regulatory networks. In this work, we show that this mismatch enables shortcut learning: a model may reconstruct masked genes from locally correlated partners rather than capturing global cellular state, yielding representations that underserve underrepresented cell populations. \nWe introduce CorrMask, a data-driven masking strategy that constructs a gene dependency graph from expression covariance and masks correlated gene groups jointly, forcing the model to rely on higher-order biological context. Evaluating on tissue-specific corpora, CorrMask produces representations that improve cell type annotation, particularly for underrepresented populations, and gene-level generalization, while matching standard baselines with up to $3{\\times}$ less pre-training data. Our results suggest that meaningful single-cell representations require pre-training objectives that respect the dependency structure of the transcriptome."},"_bibtex":{"value":"@inproceedings{\nhacohen2026hiding,\ntitle={Hiding in Plain Sight: Visible Gene Correlations Undermine Single-Cell Representations},\nauthor={Alon Hacohen and Joseph Charles Bingham and Binyamin Perets and Dvir Aran},\nbooktitle={Learning Meaningful Representations of Life (LMRL) Workshop at ICLR 2026},\nyear={2026},\nurl={https://openreview.net/forum?id=DmEPWRT0oE}\n}"},"title":{"value":"Hiding in Plain Sight: Visible Gene Correlations Undermine Single-Cell Representations"},"Anonymization":{"value":"This submission has been anonymized for double-blind review via the removal of identifying information such as names, affiliations, and identifying URLs."},"pdf":{"value":"/pdf/342d8a008bdf7fe9ef95f14d023b5e84195ae5b5.pdf"},"venueid":{"value":"ICLR.cc/2026/Workshop/LMRL"},"paperhash":{"value":"hacohen|hiding_in_plain_sight_visible_gene_correlations_undermine_singlecell_representations"},"authorids":{"value":["~Alon_Hacohen1","~Joseph_Charles_Bingham1","~Binyamin_Perets1","~Dvir_Aran1"]},"Track":{"value":"long paper (4–8 pages excluding references)"},"authors":{"value":["Alon Hacohen","Joseph Charles Bingham","Binyamin Perets","Dvir Aran"]}},"version":2},{"content":{"summary":{"value":"The study addresses graph rationalization using shortcut guidance. It assumes that the graph neural network (GNN) might become biased towards shortcuts during early training. To counter this, an extra GNN is trained to recognize these shortcuts. This auxiliary GNN helps the main graph rationalization method avoid confusing shortcuts with rationale subgraphs. The method is tested using both simple and real-world molecular graphs."},"presentation":{"value":"2 fair"},"contribution":{"value":"2 fair"},"soundness":{"value":"2 fair"},"strengths":{"value":"1. The paper focuses on a crucial topic for graph neural networks: interpretability and generalization.\n\n2. The concept of enhancing graph rationalization with shortcut learning is attractive.\n\n3. The experiments seem to validate the promise of the proposed method."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"First, an important assumption mentioned several times in the paper (quoted below) is not well-justified.\n> ... previous research (Li et al., 2021; Nam et al., 2020; Fan et al., 2022) suggests that shortcut features are easier to learn than rationale features, indicating that the features learned in the initial training stages are more inclined to shortcuts ...\n>\nExperiments on the toy datasets (on page 6) show that a higher degree of bias might make the shortcut easier for GCN to learn in a specific example. To what extent can this conclusion be generalized? For instance, does this observation still hold for real graph classification and regression tasks (do they truly suffer from significant bias in model training)? Is there any theoretical proof or more compelling empirical evidence? Do existing graph rationalization methods face the same problems? The authors also cited a few papers here [1,3,4]. [1] focused on the backdoor attack, while the standard training data for graph learning is curated and is intentionally designed to be clean [2]. [3] claimed that \"neural networks learn to rely on the spurious correlation only **when it is “easier” to learn** than the desired knowledge\" but it remains unclear whether the confounding factors are truly easier for GNN to learn in real-world examples. [4] mostly focused on the graph from images\n\nSecond, the motivation behind the model designs is unclear. For Eq. (6), why doesn't the environment affect the task prediction? And why can an environment shift be achieved by simply adding two representation vectors?\n\n\nRef. \n\n[1] Anti-Backdoor Learning: Training Clean Models on Poisoned Data. NeurIPS 2021.\n\n[2] MoleculeNet: a benchmark for molecular machine learning. Chemical Science.\n\n[3] Learning from failure: De-biasing classifier from biased classifier. NeurIPS 2020.\n\n[4] Debiasing Graph Neural Networks via Learning Disentangled Causal Substructure. NeurIPS 2022."},"confidence":{"value":"5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully."},"questions":{"value":"1. In the ablation studies, when we remove the shortcut loss, why does the model performance seem to underperform compared to other graph rationalization methods? Does this imply that other methods can also avoid shortcuts?\n\n2. Is the proposed method applicable to graph regression tasks? Or it is designed for classification?\n\n3. All the rationales in Figure 6 seem similar. What is the advantage of the proposed method?\n\n4. The reported GCN performance of 0.7128 differs from the implementation in the OGBG leaderboard by the OGB official team."},"rating":{"value":"5: marginally below the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636812341,"tcdate":1698466708944,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission6956/Reviewer_enA8"],"signatures":["ICLR.cc/2024/Conference/Submission6956/Reviewer_enA8"],"forum":"XcwHDoKvVg","number":1,"license":"CC BY 4.0","cdate":1698466708944,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission6956/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636812341,"domain":"ICLR.cc/2024/Conference","replyto":"XcwHDoKvVg","id":"NDV9CDDKd9","forumContent":{"venue":{"value":"ICLR 2024 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["graph rationalization","shortcut learning"]},"primary_area":{"value":"representation learning for computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"The remarkable success in graph neural networks (GNNs) promotes the Graph Rationalization methods that aim to provide explanations to support the prediction results by identifying a small subset of the original graph (i.e., rationale). Although existing methods have achieved promising results, recent studies have proved that these methods still suffer from  exploiting shortcuts in the data to yield task results and compose rationales. Different from previous methods plagued by shortcuts, in this paper, we propose a Shortcut-guided Graph Rationalization (SGR) method, which identifies rationales by learning from shortcuts. Specifically, SGR consists of two training stages. In the first stage, we train a shortcut guider with an early stop strategy to obtain shortcut information. During the second stage, SGR separates the graph into the rationale and non-rationale subgraphs and lets them learn from the shortcut information generated by the frozen shortcut guider to identify which information belongs to shortcuts and which does not. Finally, we employ the non-rationale subgraphs as environments and identify the invariant rationales which filter out the shortcuts under environment shifts. Extensive experimental results on both synthetic and real-world datasets clearly validate the effectiveness of our proposed method."},"_bibtex":{"value":"@misc{\nyue2024learning,\ntitle={Learning from Shortcut: A Shortcut-guided Approach for Graph Rationalization},\nauthor={Linan Yue and Qi Liu and Ye Liu and Weibo Gao and Chao Song},\nyear={2024},\nurl={https://openreview.net/forum?id=XcwHDoKvVg}\n}"},"title":{"value":"Learning from Shortcut: A Shortcut-guided Approach for Graph Rationalization"},"pdf":{"value":"/pdf/40f49261208b6613b6cba47b97c7d708165772cf.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Withdrawn_Submission"},"paperhash":{"value":"yue|learning_from_shortcut_a_shortcutguided_approach_for_graph_rationalization"},"authorids":{"value":["~Linan_Yue1","~Qi_Liu3","~Ye_Liu10","~Weibo_Gao1","~Chao_Song2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Linan Yue","Qi Liu","Ye Liu","Weibo Gao","Chao Song"]}},"version":2},{"content":{"summary":{"value":"This paper presents MBDA, a method for text–video retrieval that aims to reduce the semantic gap caused by the information imbalance between visual and textual modalities.\nMBDA includes two main components:1) MPA (Modality Projection Alignment), which aligns video and text embeddings in a shared space.2) VROD (Video Representation Orthogonal Decomposition), which divides the video embedding into two orthogonal parts to balance its information capacity with the corresponding text.\nThe decoupling is achieved through a low-rank decomposition process and analyzed in the frequency domain using the Discrete Fourier Transform (DFT).\nExtensive experiments on four public benchmarks demonstrate that MBDA delivers competitive performance and validates the effectiveness of its design."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"refer to weakness"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":2},"strengths":{"value":"S1. The paper identifies two key challenges in text–video retrieval: the imbalance in information capacity between modalities and the strong coupling of video information. These insights motivate the proposed method effectively.\n\nS2.Comprehensive experiments are conducted on four benchmarks, supported by detailed ablation studies that confirm the effectiveness of each module."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"W1. Although the DFT proofs claim orthogonality between decomposed components, no quantitative metrics are provided to show that feature redundancy is actually reduced.\n\nW2. The paper argues that orthogonality leads to balanced text–video representations, but text features remain monolithic. There is no evidence that decomposed video features align meaningfully with text semantics.\n\nW3. The paper states that video representations become comparable in scale to text after decomposition, but this is not supported by evidence. The t-SNE visualization still shows that video features are much more spread out.\n\nW4. The method relies on text–video pairs for fine-tuning. It does not support zero-shot or few-shot retrieval, limiting its scalability and adaptability."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762942370851,"tcdate":1761922410785,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission22750/Reviewer_sfyV"],"signatures":["ICLR.cc/2026/Conference/Submission22750/Reviewer_sfyV"],"forum":"POTch0RnWL","number":2,"license":"CC BY 4.0","cdate":1761922410785,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission22750/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762942370851,"domain":"ICLR.cc/2026/Conference","replyto":"POTch0RnWL","id":"AGICTsndr7","forumContent":{"TLDR":{"value":"In text-video retrieval, we propose Modality-Balanced Decoupling Alignment to address the challenge posed by the imbalance in multimodal representation space."},"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Video Understanding","Text-Video Retrieval","Modality-Balanced Decoupling Alignment"]},"supplementary_material":{"value":"/attachment/0bd5b450ada295622ebd6e8ba2b4f2a61294fd2d.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Text-video retrieval, the task of retrieving videos given a text query or vice versa, plays a significant role in video understanding. A significant challenge in this task is the semantic gap between video and text, primarily caused by the disparity in information capacity and the highly coupled nature of video information. Existing alignment methods mainly focus on multi-grained alignment between videos and text, which fails to address the capacity imbalance between video and text feature space. To address these issues, we propose Modality-Balanced Decoupling Alignment (MBDA) , a novel method that align the two modalities with closer distribution and more balanced information capacity in the feature space. Specifically, our model consists of two modules. The Modality Proximity Alignment module brings the video embedding closer to the text embedding, while the Video Representation Orthogonal Decoupling module separates the aligned video embedding into two orthogonal components, achieving better balance with their textual counterparts. Furthermore, we demonstrate that our decoupling approach achieves orthogonality while eliminating information redundancy among components through low-rank decomposition and frequency-domain analysis via Discrete Fourier Transform. The proposed method improves the baseline by a large margin. Extensive experiments demonstrate that MBDA achieves state-of-the-art performance on four most widely used public benchmarks, MSR-VTT(52.4%), DiDeMo(53.1%), MSVD(54.0%), and ActivityNet(49.6%)."},"_bibtex":{"value":"@misc{\nwang2025modalitybalanced,\ntitle={Modality-Balanced Decoupling Alignment for Text-Video Retrieval},\nauthor={Feng Wang and Ruyang Liu and Xinpeng Liu and Shiqiang Long and Ge Li},\nyear={2025},\nurl={https://openreview.net/forum?id=POTch0RnWL}\n}"},"title":{"value":"Modality-Balanced Decoupling Alignment for Text-Video Retrieval"},"pdf":{"value":"/pdf/0df1834fdd94180f6d351fbfe36be7ebdf8de877.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"wang|modalitybalanced_decoupling_alignment_for_textvideo_retrieval"},"authorids":{"value":["~Feng_Wang37","~Ruyang_Liu1","~Xinpeng_Liu4","~Shiqiang_Long1","~Ge_Li2"]},"authors":{"value":["Feng Wang","Ruyang Liu","Xinpeng Liu","Shiqiang Long","Ge Li"]}},"version":2},{"content":{"summary":{"value":"The paper studies video-driven portrait animation, aiming to generate temporally coherent videos that reenact a source identity from a driving video. It proposes Full-Video Attention for spatiotemporal interaction and History-frame conditioning for continuity. These designs improve temporal smoothness and identity preservation while efficiently adapting an image diffusion backbone (SD3.5) to video. The model produces visually coherent videos with realistic facial motion, showing strong qualitative results compared to prior methods."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1. Please provide explicit temporal-consistency metrics (e.g., tOF, tLPIPS) to directly measure frame-to-frame smoothness. \n\n2. A side-by-side comparison with a recent video diffusion baseline would clarify whether the proposed architecture achieves competitive coherence and motion fidelity. \n\n3. Evaluate stability under synthetic perturbations (e.g., random mask jitter, partial occlusion, dropout)."},"rating":{"value":4},"details_of_ethics_concerns":{"value":"No ethics review needed."},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. The paper proposes an elegant adaptation of an image diffusion backbone (SD3.5) to a video reenactment task using minimal additional parameters. This design demonstrates parameter efficiency and training stability. \n2. The generated videos exhibit clear identity retention and stylistic fidelity, benefiting from SD3.5’s strong image priors. Visual comparisons indicate improved fine-grained motion transfer and facial realism. \n3.  The combination of full-video attention and history conditioning is intuitively sound and leads to visibly smoother transitions between segments."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. While the paper introduces full-video attention and history-frame conditioning, temporal smoothness is evaluated only indirectly through the **FVD** metric. No explicit temporal-consistency measures (e.g., **tOF**, **tLPIPS**) are reported, and there is no comparison with true video diffusion backbones (e.g., Wan2.1). It remains unclear to what extent the proposed method approaches the temporal coherence achievable by video foundation models. \n\n2. The model uses masked eye–nose–mouth regions from the driving video as motion cues, yet these masks are prone to temporal jitter due to landmark noise, pose variation, or occlusion. The paper does not include robustness or ablation study under imperfect or noisy masks, which is critical for real-world deployment. \n\n3. The method is built upon an image diffusion model, without exploring fine-tuning from available video foundation models that already encode long-term motion priors. Since such models (e.g., Wan2.1) are publicly available, adopting them could provide stronger temporal consistency with little additional cost. The decision to stay with an image-only base seems conservative and possibly leaves performance unrealized."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762923945039,"tcdate":1761964885963,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission13269/Reviewer_Gnvm"],"signatures":["ICLR.cc/2026/Conference/Submission13269/Reviewer_Gnvm"],"forum":"dpVQPM6P3q","number":3,"license":"CC BY 4.0","cdate":1761964885963,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission13269/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762923945039,"domain":"ICLR.cc/2026/Conference","replyto":"dpVQPM6P3q","id":"cuC8qExSjV","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Diffusion model","face","reenactment"]},"supplementary_material":{"value":"/attachment/67dc3f56d8f55ac9b10fcbd2e3e8c5dbc0ed51e1.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Portrait animation aims to generate photo-realistic videos from a single source image by reenacting the expression and pose from a driving video. While early methods relied on 3D morphable models or feature warping techniques, they often suffered from limited expressivity, temporal inconsistency, and poor generalization to unseen identities or large pose variations. Recent advances using diffusion models have demonstrated improved quality but remain constrained by weak control signals and architectural limitations. In this work, we propose a novel diffusion-based framework that leverages masked facial regions—specifically the eyes, nose, and mouth—from the driving video as strong motion control cues. To enable robust training without appearance leakage, we adopt cross-identity supervision. To leverage the strong prior from the pre-trained diffusion model, our novel architecture introduces minimal new parameters that converge faster and help in better generalization. We introduce spatial-temporal attention mechanisms that allow inter-frame and intra-frame interactions, effectively capturing subtle motions and reducing temporal artifacts. Our model uses history frames to ensure continuity across segments. At inference, we propose a novel signal fusion strategy that balances motion fidelity with identity preservation. Our approach achieves superior temporal consistency and accurate expression control, enabling high-quality, controllable portrait animation suitable for real-world applications."},"_bibtex":{"value":"@misc{\nreddy2026stable,\ntitle={Stable Video-Driven Portraits},\nauthor={Mallikarjun Byrasandra Ramalinga Reddy and Fei Yin and Vikram Voleti and Nikita Drobyshev and Maksim Lapin and Aaryaman Vasishta and Varun Jampani},\nyear={2026},\nurl={https://openreview.net/forum?id=dpVQPM6P3q}\n}"},"title":{"value":"Stable Video-Driven Portraits"},"pdf":{"value":"/pdf/a16ccb4e94a4164853b85295ae2cbb7acd5e9c39.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"reddy|stable_videodriven_portraits"},"authorids":{"value":["~Mallikarjun_Byrasandra_Ramalinga_Reddy2","~Fei_Yin3","~Vikram_Voleti1","~Nikita_Drobyshev1","~Maksim_Lapin1","~Aaryaman_Vasishta1","~Varun_Jampani2"]},"authors":{"value":["Mallikarjun Byrasandra Ramalinga Reddy","Fei Yin","Vikram Voleti","Nikita Drobyshev","Maksim Lapin","Aaryaman Vasishta","Varun Jampani"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a new complex video reasoning and robustness benchmark, CVRR-ES, to assess Video-LLMs. CVRR-ES includes 2,400 high-quality open-ended question-answer pairs, spanning 214 high-quality videos, covering 11 video evaluation dimensions. This work evaluates 11 models, including closed-source and open-source Video-LLMs and establishes a human baseline. It also introduces a training-free dual-step contextual prompting method, DSCP, to enhance the performance of Video-LLMs."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"Please see the \"Weaknesses\"."},"rating":{"value":3},"details_of_ethics_concerns":{"value":"None"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":1},"strengths":{"value":"- The motivation and issues discussed in the paper are valuable. By constructing the CVRR-ES to assess the reasoning ability and robustness of existing Video-LLMs on real-world videos, the benchmark is very detailed in the design of video evaluation dimensions.\n\n- The construction of a human baseline provides a good reference for the evaluation of Video-LLMs."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- I express skepticism about whether the number of videos in the benchmark can achieve a robust assessment. The CVRR-ES benchmark includes only 214 videos, with the shortest video being just 2 seconds. Upon reviewing several videos from the anonymous link, I noticed a significant proportion of short videos. I question whether such short videos can adequately cover 11 categories. Moreover, current work that focuses solely on designing Video-LLMs, without specifically constructing evaluation benchmarks, provides a much larger number of assessment videos than the 214 included in CVRR-ES, for example, Tarsier [1].\n\n- As mentioned in the previous question, the distribution of videos of different lengths within the benchmark is crucial for the assessment of reasoning ability and robustness, and the paper does not provide relevant explanations. The authors should include a table showing the distribution of video lengths across the dataset, and explain how they ensured a balanced representation of different video lengths across the 11 categories.\n\n- In the motivation, it is mentioned that the goal is to build human-centric AI systems. Does the paper's reflection on this point merely consist of providing a human baseline? I think that offering more fine-grained visual examples would be more helpful for human-AI comparisons.\n\n- I think that the contribution of the DSCP is somewhat overstated and lacks novelty. Such prompt engineering-based methods have already been applied in many works for data generation, model evaluation, and other stages. The introduction and ablation experiments of this technology in the paper seem redundant.\n\n- The discussion on DSCP occupies a significant portion of the experimental analysis. I think that the current analysis provided in the paper lacks insight and does not fully reflect the value of CVRR-ES, especially in terms of human-machine comparison.\n\n- The phrase should be \"there exist a few limitations\" instead of \"there exist few limitations\" in line 520.\n\n- The paper does not provide prompt templates for all the closed-source and open-source Video-LLMs used, which will influence the reproducibility.\n\nThe problems discussed in this paper are valuable, but the most crucial aspects of benchmark construction and evaluation are not entirely convincing. Instead, a significant amount of space is dedicated to introducing the DSCP method. I don't think it meets the acceptance standards of ICLR yet. I will consider modifying the score based on the feedback from other reviewers and the authors' responses.\n\n***\n[1] Wang J, Yuan L, Zhang Y. Tarsier: Recipes for Training and Evaluating Large Video Description Models[J]. arXiv preprint arXiv:2407.00634, 2024."}},"nonreaders":[],"tmdate":1731428724182,"tcdate":1729499598236,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission6697/Reviewer_yDR5"],"signatures":["ICLR.cc/2025/Conference/Submission6697/Reviewer_yDR5"],"forum":"BTr3PSlT0T","number":1,"license":"CC BY 4.0","cdate":1729499598236,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission6697/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428724182,"domain":"ICLR.cc/2025/Conference","replyto":"BTr3PSlT0T","id":"Rf9wZVzVHl","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"TLDR":{"value":"Complex Video Reasoning and Robustness benchmark for Video-LMMs, and a test-time dual stage prompting technique for eliciting reasoning and robustness in Video-LMMs"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video Large Multi-modal Models","Complex Reasoning","Prompting for Multi-modal models"]},"supplementary_material":{"value":"/attachment/c684c658fbd5ec759b72f301664ca3414a7f27ef.zip"},"primary_area":{"value":"datasets and benchmarks"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Recent advancements in Large Language Models (LLMs) have led to the development of Video Large Multi-modal Models (Video-LMMs) that can handle a wide range of video understanding tasks. These models have the potential to be deployed in real-world applications such as robotics, AI assistants, medical surgery, and autonomous vehicles. The widespread adoption of Video-LMMs in our daily lives underscores the importance of ensuring and evaluating their robust performance in mirroring human-like reasoning and interaction capabilities in complex, real-world contexts. However, existing benchmarks for Video-LMMs primarily focus on general video comprehension abilities and neglect assessing their reasoning capabilities over complex videos in the real-world context, and the robustness of these models through the lens of user prompts as text queries. In this paper, we present the Complex Video Reasoning and Robustness Evaluation Suite (CVRR-ES), a novel benchmark that comprehensively assesses the performance of Video-LMMs across 11 diverse real-world video dimensions. We evaluate 11 recent models, including both open-source and closed-source variants, and find that most of the Video-LMMs, especially open-source ones, struggle with robustness and reasoning when dealing with complex videos. Based on our analysis, we develop a training-free Dual-Step Contextual Prompting (DSCP) technique to effectively enhance the performance of existing Video-LMMs on CVRR-ES benchmark. Our findings provide valuable insights for building the next generation of human-centric AI systems with advanced robustness and reasoning capabilities. Our dataset and code will be made publicly available."},"_bibtex":{"value":"@misc{\nkhattak2024how,\ntitle={How Good is my Video {LMM}? Complex Video Reasoning and Robustness Evaluation Suite for Video-{LMM}s},\nauthor={Muhammad Uzair Khattak and Muhammad Ferjad Naeem and Jameel Hassan Abdul Samadh and Muzammal Naseer and Federico Tombari and Fahad Shahbaz Khan and Salman Khan},\nyear={2024},\nurl={https://openreview.net/forum?id=BTr3PSlT0T}\n}"},"title":{"value":"How Good is my Video LMM? Complex Video Reasoning and Robustness Evaluation Suite for Video-LMMs"},"pdf":{"value":"/pdf/28315ce309cf834666ab34bd08d742c3867bd42f.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"khattak|how_good_is_my_video_lmm_complex_video_reasoning_and_robustness_evaluation_suite_for_videolmms"},"authorids":{"value":["~Muhammad_Uzair_Khattak1","~Muhammad_Ferjad_Naeem1","~Jameel_Hassan_Abdul_Samadh1","~Muzammal_Naseer1","~Federico_Tombari1","~Fahad_Shahbaz_Khan1","~Salman_Khan4"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Muhammad Uzair Khattak","Muhammad Ferjad Naeem","Jameel Hassan Abdul Samadh","Muzammal Naseer","Federico Tombari","Fahad Shahbaz Khan","Salman Khan"]}},"version":2},{"content":{"summary":{"value":"This paper focuses on video reasoning with multimodal large language models. Existing methods rely on coarse sequence level rewards or single-factor token selection for reinforcement learning (RL). This work Video-KTR is a policy shaping framework for token-level RL that only look at three types of tokens, which are visual-aware tokens (i.e., tokens closely associated with visual input like appear, show, etc.), temporal-aware tokens (i.e., tokens sensitive to the temporal structure of videos like finally, first, etc.), and entropy tokens (i.e.,  reasoning-critical tokens like however, wait, seem, now, etc.). Therefore, Video-KTR learns semantically informative, modality-sensitive, and filters low-value tokens, which contributing to the strong performance across five benchmarks."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"Is the Video-KTR model initialized from Qwen2.5-VL? It's better to make clear about the base model for implementation in the paper."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The general idea is simple and insightful by letting the model reinforce on 3 types of tokens that are crucial to video reasoning, but effective that has performance gain over 5 video reasoning benchmarks.\n2. Clear ablation study of all combination of 3 types of tokens to show clearly what type of token matters the most and how each type of tokens contribute to the performance gain.\n3. 7 research questions are insightful. For examples, by using the same dataset and training recipe, they made sure that the comparisons are fair. There are linguistic insights like visual-aware tokens are mainly nouns, temporal aware tokens emphasize verbs and pronouns, and entropy-aware tokens has a higher share of adjectives. A general insight for large language models is that the log-probability differences is a reliable and efficient signal for tracking prediction-confidence shifts.\n4. Writing and images are clear and easy to understand."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The performance gain over the Vanilla GRPO for three types of tokens, despite being higher, but it's a small increase. When enabled all three types, the average delta performance gain is just 2.4% for the average of three benchmark, while other ablations show even smaller differences. Therefore, the effectiveness of the Video-KTR is limited from the results."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762925858173,"tcdate":1761673086596,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission15586/Reviewer_DmeU"],"signatures":["ICLR.cc/2026/Conference/Submission15586/Reviewer_DmeU"],"forum":"p0sDIEsYG3","number":3,"license":"CC BY 4.0","cdate":1761673086596,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission15586/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762925858173,"domain":"ICLR.cc/2026/Conference","replyto":"p0sDIEsYG3","id":"glciyiQt61","forumContent":{"TLDR":{"value":"Video-KTR applies modality-aware token-level RL for video reasoning, reinforcing only visual, temporal, and uncertain tokens. It boosts accuracy and interpretability, reaching 42.7% on Video-Holmes (above GPT-4o) with broad benchmark gains."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Video Reasoning","Modality-aware Attribution","Reinforcement Learning","Multimodal Large Language Models"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Reinforcement learning (RL) has shown strong potential for enhancing reasoning in multimodal large language models (MLLMs), yet existing video reasoning methods often rely on coarse sequence-level rewards or single-factor token selection. Such approaches neglect fine-grained links among visual inputs, temporal dynamics, and linguistic outputs, limiting both accuracy and interpretability. We propose Video-KTR, a modality-aware policy shaping framework that performs selective, token-level RL by combining three attribution signals: (1) visual-aware tokens identified via counterfactual masking to reveal perceptual dependence; (2) temporal-aware tokens detected through frame shuffling to expose causal and temporal sensitivity; and (3) high-entropy tokens signaling predictive uncertainty. By reinforcing only the union of key tokens, Video-KTR focuses learning on semantically informative, modality-sensitive content while filtering out low-value tokens. Across five challenging benchmarks, Video-KTR achieves state-of-the-art or highly competitive results—42.7% on Video-Holmes, surpassing GPT-4o—with consistent gains on both reasoning-centric and general video understanding tasks. Ablation studies verify the complementary roles of the attribution signals and the robustness of targeted token-level updates. Overall, Video-KTR improves accuracy and interpretability, offering a simple, drop-in extension to RL for complex video reasoning."},"_bibtex":{"value":"@inproceedings{\nwang2026videoktr,\ntitle={Video-{KTR}: Reinforcing Video Reasoning via Key Token Attribution},\nauthor={Ziyue Wang and Sheng Jin and ZHONGRONG ZUO and Jiawei Wu and Han Qiu and Qi She and Hao Zhang and Xudong Jiang},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=p0sDIEsYG3}\n}"},"title":{"value":"Video-KTR: Reinforcing Video Reasoning via Key Token Attribution"},"pdf":{"value":"/pdf/edcb920404dd7767bbb236d8183baf33c7f4a055.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"wang|videoktr_reinforcing_video_reasoning_via_key_token_attribution"},"authorids":{"value":["~Ziyue_Wang18","~Sheng_Jin3","~ZHONGRONG_ZUO1","~Jiawei_Wu8","~Han_Qiu2","~Qi_She1","~Hao_Zhang3","~Xudong_Jiang1"]},"authors":{"value":["Ziyue Wang","Sheng Jin","ZHONGRONG ZUO","Jiawei Wu","Han Qiu","Qi She","Hao Zhang","Xudong Jiang"]}},"version":2},{"content":{"summary":{"value":"This paper introduces Youku Dense Caption, a large-scale Chinese dense video captioning dataset. Dataset addresses the scarcity of high-quality Chinese video captioning resources, containing 31,466 short videos with 311,921 Chinese captions. A strategies is proposed to improve benchmark quality by filtering out redundant or low-quality annotations. The authors establish several benchmarks for Chinese video-language tasks and conduct extensive experiments demonstrating the dataset's utility and potential for research. They also discuss challenges related to the linguistic and cultural differences between Chinese and English video data."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"see weakness."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. A Chinese video captioning dataset is proposed to fill the research gap in the Chinese community for video captioning data.\n2. A embedding-based similarity and a Non-Maximum Suppression method is used to set up a Chinese PRVR benchmark that effectively reduces annotation redundancy.\n3. The work reduces redundancy in video captioning and grounding by filtering out videos with high self-BLEU scores and minimal scene changes,  which is measured through color histogram correlation, ensuring a diverse and representative dataset."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The statement in the section “Chinese Characteristics” seems unclear. The English translation of the so-called “fine-grained Chinese captions” could also serve as fine-grained English captions in an English-language context. For me, only the localized data part is valuable, as it highlights a major difference between Chinese and English video captions. Adding more data and statistics to support this distinction would strengthen the paper.\n2. In the experiment, when translated back to Chinese. It's kind of blur that which attributes of Chinese dataset lead to the poor performance, the analysis failed to state clearly about the language differences between Chinese and English.\n3. In ablation study, mixing of different datasets only strike a balance between different tasks but failed to achieve idealized performance across different tasks. And the best performance comes from larger data scale rather than data distribution and video-caption pair. \n4. Overall, the dataset serve as a valuable data source for Chinese community in video caption domain, but the value and key attributes of the dataset remain unclear and is not fully proved by the experiment."}},"nonreaders":[],"tmdate":1732716078929,"tcdate":1730595078616,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission2869/Reviewer_C5mS"],"signatures":["ICLR.cc/2025/Conference/Submission2869/Reviewer_C5mS"],"forum":"vvi5OjPhbu","number":4,"license":"CC BY 4.0","cdate":1730595078616,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission2869/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732716078929,"domain":"ICLR.cc/2025/Conference","replyto":"vvi5OjPhbu","id":"BBlvTe5lvU","forumContent":{"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Chinese Video Datasets","Retrieval","Grounding","Generation"]},"supplementary_material":{"value":"/attachment/49cc7ec7758feac692eca721ff6deaaace182f53.zip"},"primary_area":{"value":"datasets and benchmarks"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"With the explosive growth of video content, video captions have emerged as a crucial tool for video comprehension, significantly enhancing the ability to understand and retrieve information from videos. However, most publicly available dense video captioning datasets are in English, resulting in a scarcity of large-scale and high-quality Chinese dense video captioning datasets. To address this gap within the Chinese community and to promote the advancement of Chinese multi-modal models, we develop the first, large-scale, and high-quality Chinese dense video captioning dataset, named Youku Dense Caption. This dataset is sourced from Youku, a prominent Chinese video-sharing website. Youku Dense Caption includes 31,466 complete short videos annotated by 311,921 Chinese captions. To the best of our knowledge, it is currently the largest publicly available dataset for fine-grained Chinese video descriptions. Additionally, we establish several benchmarks for Chinese video-language tasks based on the Youku Dense Caption, including retrieval, grounding, and generation tasks. Extensive experiments and evaluations are conducted on existing state-of-the-art multi-modal models, demonstrating the dataset's utility and the potential for further research."},"_bibtex":{"value":"@inproceedings{\nxiong2025youku,\ntitle={Youku Dense Caption: A Large-scale Chinese Video Dense Caption Dataset and Benchmarks},\nauthor={Zixuan Xiong and Guangwei Xu and Wenkai Zhang and Yuan Miao and Xuan Wu and Lin Hai and Ruijie Guo and Hai-Tao Zheng},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=vvi5OjPhbu}\n}"},"title":{"value":"Youku Dense Caption: A Large-scale Chinese Video Dense Caption Dataset and Benchmarks"},"pdf":{"value":"/pdf/5001de295fcc1b2cb6ad3d41bcca36153672604f.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"xiong|youku_dense_caption_a_largescale_chinese_video_dense_caption_dataset_and_benchmarks"},"authorids":{"value":["~Zixuan_Xiong1","~Guangwei_Xu2","~Wenkai_Zhang3","~Yuan_Miao1","~Xuan_Wu5","~Lin_Hai1","~Ruijie_Guo1","~Hai-Tao_Zheng2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Zixuan Xiong","Guangwei Xu","Wenkai Zhang","Yuan Miao","Xuan Wu","Lin Hai","Ruijie Guo","Hai-Tao Zheng"]}},"version":2},{"content":{"summary":{"value":"The paper introduces a novel surgical multimodal dataset, which consists of over 102,000 video-instruction pairs generated through a two-stage pipeline, aimed at enhancing the understanding and conversational capabilities of surgical videos."},"soundness":{"value":2},"confidence":{"value":5},"questions":{"value":"1. How can the quality of the data be ensured? The data collected may already contain a lot of noise and has been reprocessed by an LLM. Is there any person or clinician reviewing these raw data?\n2. Can the data be released? Are there privacy and permission risks associated with the collected data?\n3. The authors need to conduct more zero-shot evaluations on downstream tasks relevant to the surgical field, such as phase recognition, action/instrument classification, and other surgical domain VQA data to demonstrate the clinical usability of their method.\n4. The authors need to compare with more state-of-the-art methods. The comparison methods in Table 3 were all first released in 2023.\n5. The authors may verify their dataset on more benchmarks of SOTA Video MLLM architectures.\n6. Also, the authors need more zero-shot comparisons with the same VLM trained on other surgical datasets, to showcase the generalizability of their proposed dataset.\n7. The authors may evaluate the visual quality of the surgical videos themselves, as they are obtained from the website."},"rating":{"value":3},"details_of_ethics_concerns":{"value":"Potential copyright problem for online data."},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"1. With over 102,000 video-instruction pairs, this dataset is the largest in the surgical field.\n2. Structured data annotation pipeline using LLMs minimizes the risk of generating inaccurate or nonsensical content, improving dataset reliability.\n3. Releasing the dataset, model, and code publicly fosters further research and development in the surgical AI domain.\n4. The dataset can be a valuable resource for training and education, helping surgical trainees learn through interactive Q&A about real procedures."},"flag_for_ethics_review":{"value":["Yes, Legal compliance (e.g., GDPR, copyright, terms of use)"]},"weaknesses":{"value":"1. The paper does not address how the data's quality is maintained as the videos are obtained from the web. The clinicians have reviewed the output of their MLLM model, but the paper does not confirm whether clinicians or domain experts have reviewed the raw data to ensure accuracy and reliability.\n2. Concerns regarding the release, privacy, and permission risks associated with using sensitive surgical videos are not adequately discussed.\n3. The paper lacks comprehensive validation across essential surgical downstream tasks and other surgical QA datasets, which are crucial for demonstrating clinical usability. There is also a need for more rigorous benchmarking against a broader range of state-of-the-art video MLLM architectures to establish the dataset's utility and the model's performance more robustly.\n4. The comparison of the proposed methods with SOTA methods is limited and does not include the latest works. The manuscript also lacks evaluations with models trained on other surgical datasets, limiting the assessment of the proposed model's generalizability across different surgical scenarios.\n5. The paper may need to evaluate the visual quality of the surgical videos."}},"nonreaders":[],"tmdate":1731428251868,"tcdate":1730499784644,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission4896/Reviewer_QjG9"],"signatures":["ICLR.cc/2025/Conference/Submission4896/Reviewer_QjG9"],"forum":"063FuFYQQd","number":2,"license":"CC BY 4.0","cdate":1730499784644,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission4896/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428251868,"domain":"ICLR.cc/2025/Conference","replyto":"063FuFYQQd","id":"dkbJzxbH3G","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Multimodal assistant","surgical","multimodal instruction-following data","dataset"]},"primary_area":{"value":"foundation or frontier models, including LLMs"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Multimodal large language models (LLMs) have achieved notable success across various domains, while research in the medical field has largely focused on unimodal images. Meanwhile, current general-domain multimodal models for videos still lack the capabilities to understand and engage in conversations about surgical videos. One major contributing factor is the absence of datasets in the surgical field. In this paper, we create a new dataset, Surg-QA, consisting of 102,000 surgical video-instruction pairs, the largest of its kind so far. To build such a dataset, we propose a novel two-stage question-answer generation pipeline with LLM to learn surgical knowledge in a structured manner from the publicly available surgical lecture videos. The pipeline breaks down the generation process into two stages to significantly reduce the task complexity, allowing us to use a more affordable, locally deployed open-source LLM than the premium paid LLM services. It also mitigates the risk of LLM hallucinations during question-answer generation, thereby enhancing the overall quality of the generated data. We further train LLaVA-Surg, a novel vision-language conversational assistant capable of answering open-ended questions about surgical videos, on this Surg-QA dataset, and conduct comprehensive evaluations on zero-shot surgical video question-answering tasks. We show that LLaVA-Surg significantly outperforms all previous general-domain models, demonstrating exceptional multimodal conversational skills in answering open-ended questions about surgical videos. We will release our code, model, and the instruction-tuning dataset."},"_bibtex":{"value":"@misc{\nli2025llavasurg,\ntitle={{LL}a{VA}-Surg: Towards Multimodal Surgical Assistant via Structured Lecture Learning},\nauthor={Jiajie Li and Garrett Skinner and Brian R Quaranto and Gene Yang and Steven D Schwaitzberg and Peter C W Kim and Jinjun Xiong},\nyear={2025},\nurl={https://openreview.net/forum?id=063FuFYQQd}\n}"},"title":{"value":"LLaVA-Surg: Towards Multimodal Surgical Assistant via Structured Lecture Learning"},"pdf":{"value":"/pdf/04d73daf100581d96e3a971dd358d0aad68ebdd1.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"li|llavasurg_towards_multimodal_surgical_assistant_via_structured_lecture_learning"},"authorids":{"value":["~Jiajie_Li2","~Garrett_Skinner1","~Brian_R_Quaranto1","~Gene_Yang1","~Steven_D_Schwaitzberg1","~Peter_C_W_Kim1","~Jinjun_Xiong1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Jiajie Li","Garrett Skinner","Brian R Quaranto","Gene Yang","Steven D Schwaitzberg","Peter C W Kim","Jinjun Xiong"]}},"version":2},{"content":{"summary":{"value":"This paper introduces OVID, a massive open dataset comprising 10 million hours of video, positioning it as a scalable source for image-text data. The authors demonstrate that frames extracted from these videos, paired with machine-generated captions, constitute an image-text dataset with a distribution distinct from existing web-crawled image collections. They validate this by training CLIP models in image-text retrieval tasks."},"soundness":{"value":2},"confidence":{"value":2},"questions":{"value":"- Could you provide a more quantitative or qualitative analysis of how the visual distribution of OVID frames differs from that of web-crawled image datasets?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- Unprecedented Scale: 10 million hours of video and 300 million frame-caption pairs.\n- Notable Diversity: The dataset exhibits strong diversity across topics, languages, and video lengths, supporting a wide range of potential training scenarios."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Potential Circularity in Evaluation: The reliance on CLIP-based metrics (CLIPScore) for captioner selection and CLIP-based training for downstream validation may introduce a bias towards captions that align with the CLIP embedding space, rather than capturing broader semantic or human-aligned quality.\n2. Insufficient Evidence for Distributional Difference: The central claim that video frames constitute a distributionally distinct and high-quality source of image-text data is only partially supported. The evidence rests heavily on the downstream performance of CLIP models, particularly on retrieval tasks. A more direct analysis—for instance, quantifying the domain shift between OVID frames and existing image corpora (e.g., using FID or other divergence measures)—would substantially strengthen this claim.\n3. Limited Validation for Video Tasks: As acknowledged in the limitations, the dataset does not consider temporal or audio-visual aspects. While the frame-level data is validated for image-text tasks, the potential of OVID for video-language modeling (e.g., training ViCLIP or other video-text models) remains unexplored. This significantly reduces the demonstrated applicability of what is, fundamentally, a video dataset."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762933737052,"tcdate":1761899965757,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission20250/Reviewer_c4tE"],"signatures":["ICLR.cc/2026/Conference/Submission20250/Reviewer_c4tE"],"forum":"etFOgs8vIb","number":3,"license":"CC BY 4.0","cdate":1761899965757,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission20250/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762933737052,"domain":"ICLR.cc/2026/Conference","replyto":"etFOgs8vIb","id":"aDZkkQm5Gw","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["video dataset","recaptioning","clip","open foundation models","open datasets"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"We present OVid, a large open video dataset comprising _10 million hours_ of diverse content collected from CommonCrawl. To complement the raw data, we generate image captions for scene-changing frames and video-level captions for a 300M frame–caption subset. Using this subset, we train CLIP models at multiple scales and benchmark them against reference CLIP models trained on DataComp, Re-LAION and DataComp recaptioned with the same captioning pipeline. Observed scaling trends for classification and retrieval show evidence that OVid can be another valuable and scalable source of image-text data, in addition to image-text pairs from public webpages. OVid marks a significant step towards democratizing access to large-scale video data and fostering the development of open multimodal foundation models. To this end, all the data will be freely available to research institutions."},"_bibtex":{"value":"@misc{\nhochlehnert2026ovid,\ntitle={{OV}id: Open Large-Scale Video Dataset as a Novel Source for Image-Text Data},\nauthor={Andreas Hochlehnert and Marianna Nezhurina and Thadd{\\\"a}us Wiedemer and Christoph Schuhmann and Mehdi Cherti and Romain Beaumont and Andrii Matiuk and Andrej Radonjic and Bernhard Sch{\\\"o}lkopf and Wieland Brendel and A. Sophia Koepke and Jenia Jitsev and Matthias Bethge},\nyear={2026},\nurl={https://openreview.net/forum?id=etFOgs8vIb}\n}"},"title":{"value":"OVid: Open Large-Scale Video Dataset as a Novel Source for Image-Text Data"},"pdf":{"value":"/pdf/7a5482f7f60923f65052b07107b7cc4fc627e148.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"hochlehnert|ovid_open_largescale_video_dataset_as_a_novel_source_for_imagetext_data"},"authorids":{"value":["~Andreas_Hochlehnert1","~Marianna_Nezhurina1","~Thaddäus_Wiedemer1","~Christoph_Schuhmann1","~Mehdi_Cherti2","~Romain_Beaumont1","~Andrii_Matiuk2","~Andrej_Radonjic1","~Bernhard_Schölkopf1","~Wieland_Brendel1","~A._Sophia_Koepke1","~Jenia_Jitsev1","~Matthias_Bethge1"]},"authors":{"value":["Andreas Hochlehnert","Marianna Nezhurina","Thaddäus Wiedemer","Christoph Schuhmann","Mehdi Cherti","Romain Beaumont","Andrii Matiuk","Andrej Radonjic","Bernhard Schölkopf","Wieland Brendel","A. Sophia Koepke","Jenia Jitsev","Matthias Bethge"]}},"version":2},{"content":{"venue":{"value":"CoRR 2024"},"pdf":{"value":"http://arxiv.org/pdf/2407.00280v1"},"venueid":{"value":"dblp.org/journals/CORR/2024"},"paperhash":{"value":"liao|ivca_interrelationaware_video_complexity_analyzer"},"authorids":{"value":["~Junqi_Liao1","https://dblp.org/search/pid/api?q=author:Yao_Li:","~Zhuoyuan_Li2","~Li_Li1","https://dblp.org/search/pid/api?q=author:Dong_Liu:"]},"html":{"value":"https://doi.org/10.48550/arXiv.2407.00280"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2407-00280,\n  publtype={informal},\n  author={Junqi Liao and Yao Li and Zhuoyuan Li and Li Li and Dong Liu},\n  title={IVCA: Inter-Relation-Aware Video Complexity Analyzer},\n  year={2024},\n  cdate={1704067200000},\n  journal={CoRR},\n  volume={abs/2407.00280},\n  url={https://doi.org/10.48550/arXiv.2407.00280}\n}\n"},"abstract":{"value":"To meet the real-time analysis requirements of video streaming applications, we propose an inter-relation-aware video complexity analyzer (IVCA) as an extension to VCA. The IVCA addresses the limitation of VCA by considering inter-frame relations, namely motion and reference structure. First, we enhance the accuracy of temporal features by introducing feature-domain motion estimation into the IVCA. Next, drawing inspiration from the hierarchical reference structure in codecs, we design layer-aware weights to adjust the majorities of frame complexity in different layers. Additionally, we expand the scope of temporal features by considering frames that be referred to, rather than relying solely on the previous frame. Experimental results show the significant improvement in complexity estimation accuracy achieved by IVCA, with minimal time complexity increase."},"title":{"value":"IVCA: Inter-Relation-Aware Video Complexity Analyzer"},"authors":{"value":["Junqi Liao","Yao Li","Zhuoyuan Li","Li Li","Dong Liu"]}},"tmdate":1747320730950,"pdate":1704067200000,"tcdate":1724921258287,"writers":["~"],"signatures":["~Zhuoyuan_Li2"],"forum":"Zxk15ib12H","license":"CC BY-SA 4.0","number":76819,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1747320730950,"domain":"DBLP.org","id":"Zxk15ib12H","version":2},{"content":{"summary":{"value":"This paper proposes TGT (Text-Grounded Trajectories), a framework for controllable text-to-video generation that combines point-based trajectories with localized text descriptions. It introduces a lightweight Location-Aware Cross-Attention (LACA) module and a dual classifier-free guidance (dual-CFG) scheme to separately control motion and appearance via global and local prompts. To support training, the authors also construct a large-scale dataset by automatically annotating trajectories with localized captions. Experimental results on standard benchmarks demonstrate that TGT achieves stronger motion controllability and text alignment than existing baselines, with improved visual quality and flexibility in applications such as video editing."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"See weakness"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- The paper is clearly written and well-structured, with a good motivation and comprehensive methodology.\n\n- LACA and Dual CFG designs are intuitive and proven to be effective.\n\n- The experimental section is thorough.\n\n- The proposed approach demonstrates good results, outperforming prior methods in both motion control and alignment with local text.\n\n- The framework is practically useful, enabling applications such as video-to-video mirroring and localized text-driven editing."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"**Overstatement of the Problem Framing and Contribution Gap**\n\nThe paper frames its main contribution as introducing explicit text-grounded trajectories in text-to-video generation, suggesting this has not been addressed in prior literature. However, this characterization overlooks a range of recent works that already leverage trajectory-like spatial control in close association with local text descriptions. For example, PEEKABOO [1], BlobGEN-Vid [2], CineMaster [3], and 3DTrajMaster [4] all incorporate entity-specific prompts that are tied to either 2D or 3D trajectory representations such as blobs, layout paths, or pose streams. Therefore, the novelty of the “text-to-trajectory grounding” problem, as presented, is a bit overstated.\n\n**Limited Technical Novelty in Model Design**\n\nThe core technical component, Location-Aware Cross-Attention, amounts to spatially masking attention based on trajectory proximity and injecting local text features into those neighborhoods. While effective, this is a relatively standard design pattern in compositional generation models. Similar masked or region-focused local cross attention designs have already been used in Videotetris [5], DreamRunner [6], PEEKABOO [1], BlobGEN-Vid [2], etc, for grounding local texts to specific spatial-temporal regions.\n\nThe dual CFG strategy also closely mirrors prior work such as InstructPix2Pix [7], where independent CFG weights are applied to different input modalities (e.g., source image vs. edit prompt). TGT adapts this paradigm to global and local text, but the structural idea and formulation remain unchanged. As such, both the attention design and the guidance scheme may represent incremental improvements rather than conceptual innovations.\n\n**Missing Baseline for Data Pipeline**\n\nThe paper presents the data pipeline as a contribution. While interesting,  it is not compared with a straightforward baseline that uses existing trajectory annotation methods (that have already been introduced in many video trajectory control papers) with an additional captioning step. With modern tools like SAM2 and image captioning models, assigning a localized caption to each trajectory is relatively easy. A simple pipeline using “SAM2 + caption model” should be included as a baseline. Without such a comparison, it remains unclear whether the proposed pipeline offers any meaningful advantage.\n\n\n---\n[1] PEEKABOO: Interactive Video Generation via Masked-Diffusion, CVPR 2024\n\n[2] BlobGEN-Vid: Compositional Text-to-Video Generation with Blob Video Representations, CVPR 2025\n\n[3] CineMaster: A 3D-Aware and Controllable Framework for Cinematic Text-to-Video Generation, SIGGRAPH 2025\n\n[4] 3DTrajMaster: Mastering 3D Trajectory for Multi-Entity Motion in Video Generation, ICLR 2025\n\n[5] Videotetris: Towards Compositional Text-to-Video Generation, NeurIPS 2024\n\n[6] DreamRunner: Fine-Grained Compositional Story-to-Video Generation with Retrieval-Augmented Motion Adaptation, 2024\n\n[7] InstructPix2Pix: Learning to Follow Image Editing Instructions, CVPR 2023"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762917240722,"tcdate":1761998977485,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission4232/Reviewer_TNvT"],"signatures":["ICLR.cc/2026/Conference/Submission4232/Reviewer_TNvT"],"forum":"qUwOlwao20","number":4,"license":"CC BY 4.0","cdate":1761998977485,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission4232/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762917240722,"domain":"ICLR.cc/2026/Conference","replyto":"qUwOlwao20","id":"FuqGmN2Hvh","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Text-to-Video Generation; Motion Control"]},"supplementary_material":{"value":"/attachment/6e1e7fccbea1bb7e72815737a616744b7ae96035.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Text-to-video generation has advanced rapidly in visual fidelity, whereas standard methods still have limited ability to control the subject composition of generated scenes. Prior work shows that adding localized text control signals, such as bounding boxes or segmentation masks, can help. However, these methods struggle in complex scenarios and degrade in multi-object settings, offering limited precision and lacking a clear correspondence between individual trajectories and visual entities as the number of controllable objects increases. We introduce Text-Grounded Trajectories (TGT), a framework that conditions video generation on trajectories paired with localized text descriptions. We propose $\\textit{Location-Aware Cross-Attention}$ (LACA) to integrate these signals and adopt a dual-CFG scheme to separately modulate local and global text guidance. In addition, we develop a data processing pipeline that produces trajectories with localized descriptions of tracked entities, and we annotate two million high quality video clips to train TGT. Together, these components enable TGT to use point trajectories as intuitive motion handles, pairing each trajectory with text to control both appearance and motion. Extensive experiments show that TGT achieves higher visual quality, more accurate text alignment, and improved motion controllability compared with prior approaches. Website: https://textgroundedtraj.github.io."},"_bibtex":{"value":"@misc{\nzhang2025tgt,\ntitle={{TGT}: Text-Grounded Trajectories for Locally Controlled Video Generation},\nauthor={Guofeng Zhang and Angtian Wang and Jacob Zhiyuan Fang and Liming Jiang and Haotian Yang and Bo Liu and Yiding Yang and Guang Chen and Longyin Wen and Alan Yuille and Chongyang Ma},\nyear={2025},\nurl={https://openreview.net/forum?id=qUwOlwao20}\n}"},"title":{"value":"TGT: Text-Grounded Trajectories for Locally Controlled Video Generation"},"pdf":{"value":"/pdf/bb1243cca895d8c8ed42d7a0e645cf41492a66a2.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"zhang|tgt_textgrounded_trajectories_for_locally_controlled_video_generation"},"authorids":{"value":["~Guofeng_Zhang4","~Angtian_Wang2","~Jacob_Zhiyuan_Fang1","~Liming_Jiang1","~Haotian_Yang1","~Bo_Liu16","~Yiding_Yang1","~Guang_Chen5","~Longyin_Wen1","~Alan_Yuille1","~Chongyang_Ma1"]},"authors":{"value":["Guofeng Zhang","Angtian Wang","Jacob Zhiyuan Fang","Liming Jiang","Haotian Yang","Bo Liu","Yiding Yang","Guang Chen","Longyin Wen","Alan Yuille","Chongyang Ma"]}},"version":2},{"content":{"comment":{"value":"**Q5: Thank you for the additional ablations and for improving your the presentation and content of your manuscript. Your comments do not address one of the points I raised in my review though: how much does changing the noise schedule help here, compared to simply aligning the input and output sequences?**\n\n**I think this is an important discussion to have. What happens if we use the alignment set you define to create (linearized graph, output sequence pairs) with minimal edit distance, augment the data with multiple such pairs, and keep a regular noise schedule? what are the pros and cons of this method compared to modifying the schedule?**\n\n***\n\n**A5:** We thank the reviewer for highlighting this important point. We address (i) what “alignment-only + minimal edit distance” would mean in our setting, (ii) why this is not equivalent to changing the noise schedule, and (iii) the concrete pros and cons of the reviewer’s proposed alternative.\n\n**(1) Feasibility of “Alignment-Only + Minimal Edit Distance” in G2S.**\n\nWhile minimal-edit canonicalization is highly effective in domains like chemical reaction prediction (e.g., *Root-aligned SMILES* [Zhong et al., 2022]), it is fundamentally ill-suited for Graph-to-Sequence (G2S) generation due to the lack of a one-to-one mapping.\n\n* **Domain Contrast:** In reaction prediction, the molecular graph topology is largely unaltered between reactants and products, allowing for a strict one-to-one mapping and minimized edit distance. In contrast, natural language targets contain extensive syntactic elements (e.g., function words) that have no counterparts in the input graph nodes.\n* **Consequence for G2S:** Rewriting the target sentence to minimize edit distance to the linearized graph would force the output to mimic the graph's topology. This would likely hurt fluency.\n\nTherefore, we use alignment $\\mathcal{A}$ strictly to *modulate the noise schedule* for factual tokens, rather than to *canonicalize the data* which would sacrifice natural language quality.\n\n**(2) Why alignment/canonicalization is not equivalent to changing the schedule**\n\nSuppose we construct augmentation $(\\tilde{\\mathcal{G}}, S')$ pairs with minimal edit distance, an alignment-only approach operates only at the endpoint $\\mathbf{z}_0$: it changes the clean sequence and possibly augments the data, but it does not change how noise is injected over time in the diffusion process. Under a standard DDPM with a global schedule $\\bar{\\alpha}_t$, each token embedding $\\mathbf{x}_0^i$ is corrupted as:\n\n$$\n\\mathbf{z}_t^i = \\sqrt{\\bar{\\alpha}_t}\\,\\mathbf{x}_0^i + \\sqrt{1-\\bar{\\alpha}_t}\\,\\boldsymbol\\epsilon,\\quad \\forall i.\n$$\n\nThus, the signal-to-noise ratio (SNR) at timestep $t$ is *identical* for all tokens, regardless of whether $i$ is a high-frequency function word or a rare factual entity. This has two consequences:\n\n* **Common syntactic tokens** occupy dense regions of the embedding space and can often be recovered reliably even at relatively low SNR.\n* **Rare entity tokens** are much sparser and typically require higher SNR to be distinguishable from their neighbors.\n\nWith a uniform schedule, the embeddings of rare entities are heavily obscured; the model receives weak, noisy gradients on the tokens that are responsible for factual grounding and edit sensitivity. Minimizing edit distance in the discrete space (even with multiple augmented variants) changes $\\mathbf{x}_0^i$ but leaves $\\bar{\\alpha}_t$—and hence the per-token SNR profile over time—remains unchanged.\n\nOur graph-aware schedule is designed precisely to address this limitation. Using $\\mathcal{A}$, we construct token-specific schedules $\\{\\bar{\\alpha}^i_t\\}_{t=1}^T$ such that aligned tokens (entities/relations) retain higher SNR for longer along the trajectory, while unaligned syntactic tokens follow a more aggressive, standard schedule.\n\n**(3) Relation to our ablations: regular schedule vs. graph-aware schedule**\n\nWe do not claim to have implemented a full “minimal-edit augmentation” baseline. However, our ablations already isolate the effect of modifying the schedule while keeping the rest of the model and data pipeline fixed.\nConcretely, the `sqrt (baseline)` row in Table 6 uses the same encoder–decoder architecture, graph linearization, and training objective as `DLM4G`, but applies a standard *sqrt* schedule uniformly to all tokens.\nEmpirically, moving from `sqrt (baseline)` to `Graph-aware` ($\\mathcal{A}$) improves BLEU from $0.60$ to $0.65$ ($+0.05$ absolute). In Appendix A.7 (Table 13), we report the corresponding changes in our grounding and edit-sensitivity metrics.\n\nIn summary, we view alignment-only + minimal-edit canonicalization as a *complementary* improvement to baseline models, but not a substitute for schedule modification."},"title":{"value":"Clarifications on Minimal Edit Distance based Alignment"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Comment","tmdate":1770917943936,"tcdate":1764334614345,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission17569/Authors"],"signatures":["ICLR.cc/2026/Conference/Submission17569/Authors"],"forum":"I6lEri0e2K","number":37,"license":"CC BY 4.0","cdate":1764334614345,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission17569/-/Official_Comment","ICLR.cc/2026/Conference/-/Edit"],"mdate":1770917943936,"domain":"ICLR.cc/2026/Conference","replyto":"sZySL01Ry1","id":"GHCWaRLoa0","forumContent":{"TLDR":{"value":"Graph-to-Sequence generation with Graph-Aware Diffusion Framework"},"venue":{"value":"ICLR 2026 Conference Desk Rejected Submission"},"keywords":{"value":["Graph-to-Sequence","Autoregressive models","Non-autoregressive models","Diffusion Models"]},"primary_area":{"value":"learning on graphs and other geometries & topologies"},"abstract":{"value":"Pre-trained language models (PLMs) remain unreliable for graph-to-sequence (G2S) generation, where two challenges are particularly acute: (i) factual grounding, ensuring all entities are faithfully realized, and (ii) edit sensitivity, ensuring small, local graph edits to propagate consistently in the output. We propose Diffusion Language Models for Graphs (DLM4G), a non-autoregressive framework for iterative refinement conditioned on the graph input. Central to DLM4G is a graph-aware adaptive noising strategy, where noise is applied to the output sequence aligned with the graph components (entities and relations) using a learnable component-wise schedule. We learn a component-wise schedule by linearly mapping between per-component denoising loss and noise schedule. This ensures entities are generated faithfully and keeps graph edits localized in the text. Through extensive experiments on three benchmark datasets, DLM4G outperforms state-of-the-art autoregressive baselines that are 12–127× larger, achieving 10–15\\% relative gains on standard surface-level metrics (BLEU, ChrF++, METEOR) and embedding-based metrics (BERTScore-F1, MAUVE). More importantly, DLM4G improves factual grounding (FGT, $\\uparrow$) by +$\\Delta_{\\text{FGT}}$ 4.7 \\% and edit sensitivity (ESR, $\\uparrow$) by +$\\Delta_{\\text{ESR}}$ 7.9 \\% on average compared to comparably sized autoregressive baselines. Finally, we evaluate on molecule captioning, where molecular graphs are verbalized into textual descriptions, demonstrating the applicability of DLM4G to biomedical G2S tasks."},"_bibtex":{"value":"@misc{\nanonymous2026graphtosequence,\ntitle={Graph-to-Sequence Generation Beyond Autoregressive Models: A Graph-Aware Diffusion Framework},\nauthor={Anonymous},\nyear={2026},\nurl={https://openreview.net/forum?id=I6lEri0e2K}\n}"},"title":{"value":"Graph-to-Sequence Generation Beyond Autoregressive Models: A Graph-Aware Diffusion Framework"},"pdf":{"value":"/pdf/1e6820a81275ffe1d02cde5c74a0de0b67833b7b.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Desk_Rejected_Submission"},"paperhash":{"value":"shahane|graphtosequence_generation_beyond_autoregressive_models_a_graphaware_diffusion_framework"},"authorids":{"value":["~Aditya_Hemant_Shahane1","~Anuj_Kumar_Sirohi1","~Tanmoy_Chakraborty2","~Prathosh_AP1","~Sandeep_Kumar8"]},"authors":{"value":["Aditya Hemant Shahane","Anuj Kumar Sirohi","Tanmoy Chakraborty","Prathosh AP","Sandeep Kumar"]}},"version":2},{"content":{"summary":{"value":"The paper introduces Text‑Grounded Trajectories (TGT), a controllable text‑to‑video (T2V) framework that ties sparse 2D point trajectories to localized text prompts, enabling joint control of what moves (appearance/identity) and how/where it moves (trajectory). Architecturally, the method adds a lightweight Location‑Aware Cross‑Attention (LACA) branch to each DiT block and uses a dual classifier‑free guidance (dual‑CFG) scheme with separate scales for the global caption and the local text. Because no dataset exists with point‑trajectory ↔ local‑text pairs, the authors build a two‑stage data pipeline: (i) use Grounded‑SAM to segment a key frame, choose representative points, and distill a coordinate‑aware local captioner by prompting a teacher VLM (GPT‑4o) and fine‑tuning Qwen2.5‑VL‑3B; (ii) propagate points with TAP to get long trajectories and visibility flags. Empirically, TGT achieves the lowest EPE and the highest local CLIP‑T, outperforming WanT2V, MotionCtrl, TrailBlazer, and Tora."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"N/A"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. Strong empirical results compared with previous SoTA T2V generators across multiple evaluation metrics (e.g., EPE, CLIP-T). \n\n2. Detailed ablations demonstrate the effectiveness of each component in the model (e.g., CFG design, LACA component). \n\n3. Potential applications to other tasks like video-to-video generation / video editing tasks."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Utilizing attention for better grounding with the prompt has been explored in many previous approaches (e.g., Video-Editing-Attention[1], DreamRunner[2], Presto[3]). Though these approaches are not directly for trajectory control, but they can be easily extended for trajectory control. \n\n2. Evaluation is only on the curated dataset. It's worth checking the performance on how this approach can be extended to general T2V generation (e.g., PhyGenBench[4], T2VCompBench[5]), where these benchmarks also contain some examples that actually can be done with trajectory control (e.g., motion class in T2VCompBench).\n\n[1] Investigating the Effectiveness of Cross-Attention to Unlock Zero-Shot Editing of Text-to-Video Diffusion Models\n\n[2] DreamRunner: Fine-Grained Compositional Story-to-Video Generation with Retrieval-Augmented Motion Adaptation\n\n[3] Long Video Diffusion Generation with Segmented Cross-Attention and Content-Rich Video Data Curation\n\n[4] Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation\n\n[5] T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762917241108,"tcdate":1761954859955,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission4232/Reviewer_7VsE"],"signatures":["ICLR.cc/2026/Conference/Submission4232/Reviewer_7VsE"],"forum":"qUwOlwao20","number":3,"license":"CC BY 4.0","cdate":1761954859955,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission4232/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762917241108,"domain":"ICLR.cc/2026/Conference","replyto":"qUwOlwao20","id":"Jl2KYOkY5e","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Text-to-Video Generation; Motion Control"]},"supplementary_material":{"value":"/attachment/6e1e7fccbea1bb7e72815737a616744b7ae96035.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Text-to-video generation has advanced rapidly in visual fidelity, whereas standard methods still have limited ability to control the subject composition of generated scenes. Prior work shows that adding localized text control signals, such as bounding boxes or segmentation masks, can help. However, these methods struggle in complex scenarios and degrade in multi-object settings, offering limited precision and lacking a clear correspondence between individual trajectories and visual entities as the number of controllable objects increases. We introduce Text-Grounded Trajectories (TGT), a framework that conditions video generation on trajectories paired with localized text descriptions. We propose $\\textit{Location-Aware Cross-Attention}$ (LACA) to integrate these signals and adopt a dual-CFG scheme to separately modulate local and global text guidance. In addition, we develop a data processing pipeline that produces trajectories with localized descriptions of tracked entities, and we annotate two million high quality video clips to train TGT. Together, these components enable TGT to use point trajectories as intuitive motion handles, pairing each trajectory with text to control both appearance and motion. Extensive experiments show that TGT achieves higher visual quality, more accurate text alignment, and improved motion controllability compared with prior approaches. Website: https://textgroundedtraj.github.io."},"_bibtex":{"value":"@misc{\nzhang2025tgt,\ntitle={{TGT}: Text-Grounded Trajectories for Locally Controlled Video Generation},\nauthor={Guofeng Zhang and Angtian Wang and Jacob Zhiyuan Fang and Liming Jiang and Haotian Yang and Bo Liu and Yiding Yang and Guang Chen and Longyin Wen and Alan Yuille and Chongyang Ma},\nyear={2025},\nurl={https://openreview.net/forum?id=qUwOlwao20}\n}"},"title":{"value":"TGT: Text-Grounded Trajectories for Locally Controlled Video Generation"},"pdf":{"value":"/pdf/bb1243cca895d8c8ed42d7a0e645cf41492a66a2.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"zhang|tgt_textgrounded_trajectories_for_locally_controlled_video_generation"},"authorids":{"value":["~Guofeng_Zhang4","~Angtian_Wang2","~Jacob_Zhiyuan_Fang1","~Liming_Jiang1","~Haotian_Yang1","~Bo_Liu16","~Yiding_Yang1","~Guang_Chen5","~Longyin_Wen1","~Alan_Yuille1","~Chongyang_Ma1"]},"authors":{"value":["Guofeng Zhang","Angtian Wang","Jacob Zhiyuan Fang","Liming Jiang","Haotian Yang","Bo Liu","Yiding Yang","Guang Chen","Longyin Wen","Alan Yuille","Chongyang Ma"]}},"version":2},{"content":{"summary":{"value":"The paper trains SNNs using surrogate gradient learning. In order to mitigate the gradient vanishing problem, the paper proposed the Shortcut Back-propagation method and utilizes an evolutionary algorithm framework to balance the training of shallow and deep layers. The effectiveness of the proposed method is demonstrated through many experiments."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"1)\tWhy are the bolded values not always the best values?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1)\tThe shortcut backpropagation method and the evolutionary training method are novel. \n2)\tThis paper can well handle the gradient vanishing problem.\n3)\tThe paper is well-written.\n4)\tThe paper shows the effectiveness of the proposed methods through many experiments."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1)\tThe author should add more mathematical proof to demonstrate that the mentioned residual structure in SNN is not very effective? The introduction of shortcut branches might add complexity to the network architecture, which could affect the interpretability of the model.\n2)\tSome recent SOTA works should be compared with too.  The authors can also compare with paper [1][2] which obtains really good results by MS-ResNet-18 backbone with 1 or 6 timesteps on large imageNet datasets.\n\n\n[1]Yao M, Zhao G, Zhang H, et al. Attention spiking neural networks[J]. IEEE transactions on pattern analysis and machine intelligence, 2023.\n\n[2] Qiu X, Zhu R J, Chou Y, et al. Gated attention coding for training high-performance and efficient spiking neural networks[C]. Proceedings of the AAAI Conference on Artificial Intelligence. 2024, 38(1): 601-610."},"limitations":{"value":"I find no limitation about the paper."}},"nonreaders":[],"tmdate":1730879027772,"tcdate":1718702437963,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission5748/Reviewer_Jb1s"],"signatures":["NeurIPS.cc/2024/Conference/Submission5748/Reviewer_Jb1s"],"forum":"xjyU6zmZD7","number":1,"license":"CC BY 4.0","cdate":1718702437963,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission5748/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730879027772,"domain":"NeurIPS.cc/2024/Conference","replyto":"xjyU6zmZD7","id":"ofpAsTv15H","forumContent":{"venue":{"value":"NeurIPS 2024 poster"},"keywords":{"value":["Spiking Neural Network","Training SNNs","Surrogate gradient"]},"supplementary_material":{"value":"/attachment/219213773c28852413f852e0755c33b1f50db66d.zip"},"primary_area":{"value":"neuroscience_and_cognitive_science"},"abstract":{"value":"The Spiking Neural Network (SNN) is a biologically inspired neural network infrastructure that has recently garnered significant attention. It utilizes binary spike activations to transmit information, thereby replacing multiplications with additions and resulting in high energy efficiency. However, training an SNN directly poses a challenge due to the undefined gradient of the firing spike process. Although prior works have employed various surrogate gradient training methods that use an alternative function to replace the firing process during back-propagation, these approaches ignore an intrinsic problem: gradient vanishing. To address this issue, we propose a shortcut back-propagation method in the paper, which advocates for transmitting the gradient directly from the loss to the shallow layers. This enables us to present the gradient to the shallow layers directly, thereby significantly mitigating the gradient vanishing problem. Additionally, this method does not introduce any burden during the inference phase.\nTo strike a balance between final accuracy and ease of training, we also propose an evolutionary training framework and implement it by inducing a balance coefficient that dynamically changes with the training epoch, which further improves the network's performance. Extensive experiments conducted over static and dynamic datasets using several popular network structures reveal that our method consistently outperforms state-of-the-art methods."},"_bibtex":{"value":"@inproceedings{\nguo2024take,\ntitle={Take A Shortcut Back: Mitigating the Gradient Vanishing for Training Spiking Neural Networks},\nauthor={Yufei Guo and Yuanpei Chen and Zecheng Hao and Weihang Peng and Zhou Jie and Yuhan Zhang and Xiaode Liu and Zhe Ma},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=xjyU6zmZD7}\n}"},"title":{"value":"Take A Shortcut Back: Mitigating the Gradient Vanishing for Training Spiking Neural Networks"},"pdf":{"value":"/pdf/6dfff0aec6f93d33ffa638873f008d9ca6857190.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"guo|take_a_shortcut_back_mitigating_the_gradient_vanishing_for_training_spiking_neural_networks"},"authorids":{"value":["~Yufei_Guo1","~Yuanpei_Chen1","~Zecheng_Hao1","~Weihang_Peng2","~Zhou_Jie3","~Yuhan_Zhang1","~Xiaode_Liu1","~Zhe_Ma2"]},"authors":{"value":["Yufei Guo","Yuanpei Chen","Zecheng Hao","Weihang Peng","Zhou Jie","Yuhan Zhang","Xiaode Liu","Zhe Ma"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a DiT-based human animation framework which focuses on generating high-fidelity and long-duration human videos. Firstly, a set of hybrid implicit guidance signals is incorporated to enhance facial and hand feature details. Next, the time-aware position shift fusion module is incorporated to enable long video generation. Finally, a novel data augmentation strategy and a skeleton alignment model are proposed to reduce the impact of human shape variations across different identities. Experimental results demonstrate that the proposed method outperforms existing approaches and achieves superior performance in both high-fidelity and long-duration human image animation."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1.\tThe authors present comparative results on the TikTok test set and a self-constructed test set. However, they do not elaborate on the distinctions between the two in terms of data types and content. To ensure the method's reproducibility, the construction process for their custom dataset should be detailed.\n2.\tFor the comparison with other methods in Figure 5, it is recommended that the authors provide more examples to better demonstrate the superiority of the proposed method. In Figure 6, both the driving images and the reference images should be provided. Currently, Figures 5 and 6 seem to exhibit poor performance in terms of face expression following.\n3.\tRegarding the statement in Line 374 about the potential inclusion of TikTok test data in UniAnimate-DiT's training set, it would be helpful if the authors could clarify the basis for this observation. Alternatively, if this is intended as a hypothesis, it might be more constructive to focus the discussion on a deeper analysis of why the proposed method shows a performance gap on metrics like FID."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1．\tThe authors propose a comprehensive framework for human animation that enhances the clarity and fidelity of faces and hands in the generated video frames.\n2．\tThe authors incorporate the time-aware position shift fusion module and adapt it to video scenarios, enabling the generation of long videos.\n3．\tThe authors present a data augmentation strategy and a pose alignment module to eliminate body shape discrepancies across different human identities."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1.\tThe novelty of this paper appears limited, as the various modules are largely based on related existing works. For example, the appearance feature extractor is derived from LivePortrait [1] (Lines 208-209), the hand representation follows that of RealisDance [2] (Lines 219-220), and the highlighted contribution for long video generation—the time-aware position shift fusion module—originates from Sonic [3] (Line 240). The distinctions and improvements currently described in the paper are insufficient to support its novelty.\n2.\tThe paper fails to sufficiently explain its adopted modules, making them hard to understand. For instance, in Section 3.3, it is unclear how the time-aware position shift fusion module functions. The meaning of the shifted windows, denoted as [s, e] (Line 246), is also not clearly defined.\n3.\tThe paper's ablation study is not specific or sufficient enough: 1) As a primary contribution, the Laplacian Sharpness Factor for the hand guidance signal warrants a separate ablation study to justify its inclusion. 2) For the \"Ablation on Pose Alignment,\" the proposed Data Augmentation Strategy and the Alignment and Smoothness Model should be ablated separately.  \n\n[1] Guo J, Zhang D, Liu X, et al. Liveportrait: Efficient portrait animation with stitching and retargeting control[J]. arXiv preprint arXiv:2407.03168, 2024.\n[2] Zhou J, Wang B, Chen W, et al. Realisdance: Equip controllable character animation with realistic hands[J]. arXiv preprint arXiv:2409.06202, 2024.\n[3] Ji X, Hu X, Xu Z, et al. Sonic: Shifting focus to global audio perception in portrait animation[C]//Proceedings of the Computer Vision and Pattern Recognition Conference. 2025: 193-203."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762919044191,"tcdate":1761818885190,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission6767/Reviewer_d7aK"],"signatures":["ICLR.cc/2026/Conference/Submission6767/Reviewer_d7aK"],"forum":"KgyyIk59Nx","number":3,"license":"CC BY 4.0","cdate":1761818885190,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission6767/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762919044191,"domain":"ICLR.cc/2026/Conference","replyto":"KgyyIk59Nx","id":"VTcGyFsquo","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Human Image Animation","Diffusion Model","Video Generation"]},"supplementary_material":{"value":"/attachment/6aa7d32a85883e9f78a1fa9c40ec3774ee65c4fa.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Recent progress in diffusion models has significantly advanced the field of human image animation. While existing methods can generate temporally consistent results for short or regular motions, significant challenges remain, particularly in generating long-duration videos. Furthermore, the synthesis of fine-grained facial and hand details remains under-explored, limiting the applicability of current approaches in real-world, high-quality applications. To address these limitations, we propose a diffusion transformer (DiT)-based framework which focuses on generating high-fidelity and long-duration human animation videos. First, we design a set of hybrid implicit guidance signals, enabling our framework to additionally incorporate detailed face and hand features as guidance. Next, we incorporate the time-aware position shift fusion module, modify the input format within the DiT backbone, and refer to this mechanism as the Position Shift Adaptive Module, which enables video generation of arbitrary length. Finally, we introduce a novel data augmentation strategy and a skeleton alignment model to reduce the impact of human shape variations across different identities. Experimental results demonstrate that our method outperforms existing state-of-the-art approaches, achieving superior performance in both high-fidelity and long-duration human image animation."},"_bibtex":{"value":"@misc{\nzheng2026highfidelity,\ntitle={High-Fidelity and Long-Duration Human Image Animation with Diffusion Transformer},\nauthor={Shen zheng and Jiaran Cai and Yuansheng Guan and Shenneng Huang and Xingpei Ma and Junjie Cao and hanfeng Zhao and Qiang Zhang and Shunsi Zhang and Xiao-Ping Zhang},\nyear={2026},\nurl={https://openreview.net/forum?id=KgyyIk59Nx}\n}"},"title":{"value":"High-Fidelity and Long-Duration Human Image Animation with Diffusion Transformer"},"pdf":{"value":"/pdf/c24ce48ce6a16d555a1834dd44a4023cacb52cb8.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"zheng|highfidelity_and_longduration_human_image_animation_with_diffusion_transformer"},"authorids":{"value":["~Shen_zheng4","~Jiaran_Cai2","~Yuansheng_Guan1","~Shenneng_Huang1","~Xingpei_Ma1","~Junjie_Cao5","~hanfeng_Zhao2","~Qiang_Zhang20","~Shunsi_Zhang1","~Xiao-Ping_Zhang1"]},"authors":{"value":["Shen zheng","Jiaran Cai","Yuansheng Guan","Shenneng Huang","Xingpei Ma","Junjie Cao","hanfeng Zhao","Qiang Zhang","Shunsi Zhang","Xiao-Ping Zhang"]}},"version":2},{"content":{"summary":{"value":"The paper presents Q-Bench-Video, a comprehensive benchmark designed to assess the video quality understanding capabilities of Large Multi-modal Models (LMMs). It focuses specifically on video quality, incorporating various video types, including natural, AI-generated content (AIGC), and computer graphics (CG). The benchmark evaluates models across multiple question types (Yes-or-No, What-How, and Open-ended) and quality concerns (Technical, Aesthetic, Temporal, and AIGC distortions), providing a new assessment framework. Results show that there is still a notable gap between LMM and human performance."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. The phrasing in line 465 is unclear, potentially causing confusion for readers. Could you please explain it more.\n2. In the video pair comparisons, it is unclear if the benchmark includes comparisons of the same video at different quality levels or comparisons across entirely different videos with varying quality levels. \n3. Could you please provide the selected video information about the formats and resolution and so on."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":3},"strengths":{"value":"1. The benchmark includes a broad range of video types, including natural scenes, AIGC, and CG, enhancing the evaluation’s comprehensiveness.\n2. By integrating various question types and quality concerns, It offers a novel framework for assessing different aspects of video quality.\n3. The benchmark includes 2,378 question-answer pairs curated by experts, providing a reliable evaluation dataset."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Although the benchmark includes annotations reviewed by three additional participants, the performance on human evaluation is only 81.56%. This raises concerns about whether the annotations meet a robust confidence interval and whether there may be inherent biases in the data. The reliance on subjective assessments for video quality (such as aesthetics) could introduce inconsistencies.\n2. The technical and aesthetic evaluations in the dataset are annotated by eight experts, but the paper does not validate these annotations against established subjective scores, such as those from the MaxWell dataset. \n3. The paper provides no details on the video format and resolution used in the benchmark. It is unclear if any downsampling was applied, which may affect the quality assessment and limit the applicability of the results to different resolutions.\n4. The benchmark’s evaluation of open-ended questions relies on GPT-based scoring, repeated five times for consistency. However, this method may introduce biases, as GPT may favor responses with longer outputs or struggle with certain scenarios where perfect answers are difficult to generate. This raises concerns about the objectivity and reliability of the scoring for open-ended questions.\n5. The paper does not provide a comparison with other established video quality datasets to validate the reliability of its annotations."}},"nonreaders":[],"tmdate":1731427490192,"tcdate":1730701661712,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission1786/Reviewer_Wt97"],"signatures":["ICLR.cc/2025/Conference/Submission1786/Reviewer_Wt97"],"forum":"VaUy5GZO3f","number":3,"license":"CC BY 4.0","cdate":1730701661712,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission1786/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427490192,"domain":"ICLR.cc/2025/Conference","replyto":"VaUy5GZO3f","id":"WR39xroph0","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Large multi-modal model","benchmark","video quality assessment"]},"primary_area":{"value":"datasets and benchmarks"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"With the rising interest in research on Large Multi-modal Models (LMMs) for video understanding, many studies have emphasized general video comprehension capabilities, neglecting the **systematic exploration into video quality understanding**. To address this oversight, we introduce **Q-Bench-Video** in this paper, a new benchmark specifically designed to evaluate LMMs' proficiency in discerning video quality. **a)** To ensure the diversity of video sources, Q-Bench-Video encompasses videos from natural scenes, computer graphics (CG), and AI-generated content (AIGC). **b)** Building on the traditional multiple-choice questions format with the *Yes-or-No* and *What-How* categories, we include *Open-ended* questions to better evaluate complex scenarios. Additionally, we incorporate the **video pair quality comparison** question to enhance comprehensiveness. **c)** Beyond the traditional *Technical*, *Aesthetic*, and *Temporal* distortions, we have expanded our evaluation aspects to include the dimension of *AIGC* distortions, which addresses the increasing demand for video generation. Finally, we collect a total of 2,378 question-answer pairs and test them on 12 open-source & 5 proprietary LMMs. Our findings indicate that while LMMs have a foundational understanding of video quality, their performance remains incomplete and imprecise, with a notable discrepancy compared to human-level performance. Through **Q-Bench-Video**, we seek to catalyze community interest, stimulate further research, and unlock the untapped potential of LMMs to close the gap in video quality understanding."},"_bibtex":{"value":"@misc{\nzhang2024qbenchvideo,\ntitle={Q-Bench-Video: Benchmarking the Video Quality Understanding of {LMM}s},\nauthor={Zicheng Zhang and Ziheng Jia and Haoning Wu and Chunyi Li and Zijian Chen and Yingjie Zhou and Wei Sun and Xiaohong Liu and Xiongkuo Min and Weisi Lin and Guangtao Zhai},\nyear={2024},\nurl={https://openreview.net/forum?id=VaUy5GZO3f}\n}"},"title":{"value":"Q-Bench-Video: Benchmarking the Video Quality Understanding of LMMs"},"pdf":{"value":"/pdf/b6665a67a0e193ca939b3b34f488e5e0b380c738.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"zhang|qbenchvideo_benchmarking_the_video_quality_understanding_of_lmms"},"authorids":{"value":["~Zicheng_Zhang7","~Ziheng_Jia1","~Haoning_Wu1","~Chunyi_Li1","~Zijian_Chen1","~Yingjie_Zhou1","~Wei_Sun12","~Xiaohong_Liu2","~Xiongkuo_Min1","~Weisi_Lin1","~Guangtao_Zhai1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Zicheng Zhang","Ziheng Jia","Haoning Wu","Chunyi Li","Zijian Chen","Yingjie Zhou","Wei Sun","Xiaohong Liu","Xiongkuo Min","Weisi Lin","Guangtao Zhai"]}},"version":2},{"content":{"venue":{"value":"CoRR 2024"},"pdf":{"value":"http://arxiv.org/pdf/2403.07715v1"},"venueid":{"value":"dblp.org/journals/CORR/2024"},"paperhash":{"value":"vanberlo|intravideo_positive_pairs_in_selfsupervised_learning_for_ultrasound"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Blake_VanBerlo:","https://dblp.org/search/pid/api?q=author:Alexander_Wong:","~Jesse_Hoey1","https://dblp.org/search/pid/api?q=author:Robert_Arntfield:"]},"html":{"value":"https://doi.org/10.48550/arXiv.2403.07715"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2403-07715,\n  publtype={informal},\n  author={Blake VanBerlo and Alexander Wong and Jesse Hoey and Robert Arntfield},\n  title={Intra-video Positive Pairs in Self-Supervised Learning for Ultrasound},\n  year={2024},\n  cdate={1704067200000},\n  journal={CoRR},\n  volume={abs/2403.07715},\n  url={https://doi.org/10.48550/arXiv.2403.07715}\n}\n"},"abstract":{"value":"Self-supervised learning (SSL) is one strategy for addressing the paucity of labelled data in medical imaging by learning representations from unlabelled images. Contrastive and non-contrastive SSL methods produce learned representations that are similar for pairs of related images. Such pairs are commonly constructed by randomly distorting the same image twice. The videographic nature of ultrasound offers flexibility for defining the similarity relationship between pairs of images. In this study, we investigated the effect of utilizing proximal, distinct images from the same B-mode ultrasound video as pairs for SSL. Additionally, we introduced a sample weighting scheme that increases the weight of closer image pairs and demonstrated how it can be integrated into SSL objectives. Named Intra-Video Positive Pairs (IVPP), the method surpassed previous ultrasound-specific contrastive learning methods' average test accuracy on COVID-19 classification with the POCUS dataset by $\\ge 1.3\\%$. Detailed investigations of IVPP's hyperparameters revealed that some combinations of IVPP hyperparameters can lead to improved or worsened performance, depending on the downstream task. Guidelines for practitioners were synthesized based on the results, such as the merit of IVPP with task-specific hyperparameters, and the improved performance of contrastive methods for ultrasound compared to non-contrastive counterparts."},"title":{"value":"Intra-video Positive Pairs in Self-Supervised Learning for Ultrasound"},"authors":{"value":["Blake VanBerlo","Alexander Wong","Jesse Hoey","Robert Arntfield"]}},"tmdate":1724937033763,"pdate":1704067200000,"tcdate":1724936988163,"writers":["~"],"signatures":["~Jesse_Hoey1"],"forum":"zLFFWWtYF6","license":"CC BY-SA 4.0","number":77040,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1724937033763,"domain":"DBLP.org","id":"zLFFWWtYF6","version":2},{"content":{"TLDR":{"value":"Minimal-pair interventions reveal that CNNs become sensitive to distant visual structure before they learn connectivity itself, while path-conditioned AUC can mistake response strength for the extent of learned competence."},"venue":{"value":"ICLR 2027 Conference Submission"},"keywords":{"value":["visual reasoning","connectivity","convolutional neural networks","minimal pairs","counterfactual evaluation","algorithmic reasoning","long-range dependencies","mechanistic analysis","sample complexity"]},"primary_area":{"value":"learning theory"},"abstract":{"value":"A network can see every pixel a visual rule depends on and still not use them. We study this gap with seeded connectivity on near-critical $64\\times64$ grids, where a full-access convolutional network faces zero Bayes risk. Path-conditioned AUC suggests that competence spreads from short to long paths during learning, but this picture is misleading. We introduce minimal pairs for spatial rules: two images that differ in a single pixel outside the query's local window, where closing that pixel disconnects the seed from the query. Size-matched placebo edits, which remove at least as much of the query's component without disconnecting it, separate connectivity from cluster size. Minimal pairs reveal three stages. With $64$k training examples, the model's discrimination does not exceed what the seed position alone provides, and it ignores cuts. With $256$k, it responds to distant cuts at every path length, but mostly in proportion to how many pixels an edit removes. With $1$M, it computes connectivity itself, including on paths that need more certified propagation steps than it has layers. Within individual training runs, the response to cuts appears at all path lengths at once. At first, it depends about equally on how many pixels an edit removes and on whether it disconnects. The weight on disconnection then grows until removal size no longer matters. Path length orders the strength of the response rather than its presence, and response strength is what path-conditioned AUC detects."},"_bibtex":{"value":"@inproceedings{\nanonymous2026sensitivity,\ntitle={Sensitivity Before Connectivity: Minimal Pairs Reveal How {CNN}s Learn a Global Visual Rule},\nauthor={Anonymous},\nbooktitle={Submitted to The Fifteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=xlL4bjWnnH},\nnote={under review}\n}"},"title":{"value":"Sensitivity Before Connectivity: Minimal Pairs Reveal How CNNs Learn a Global Visual Rule"},"pdf":{"value":"/pdf/ca724eb8a0bd34fe42b214cc3e88a02be8b5c68a.pdf"},"venueid":{"value":"ICLR.cc/2027/Conference/Submission"}},"tmdate":1791231485916,"tcdate":1789654500558,"writers":["ICLR.cc/2027/Conference","ICLR.cc/2027/Conference/Submission32088/Authors"],"signatures":["ICLR.cc/2027/Conference/Submission32088/Authors"],"forum":"xlL4bjWnnH","license":"CC BY 4.0","number":32088,"cdate":1789654500558,"readers":["everyone"],"invitations":["ICLR.cc/2027/Conference/-/Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Bidding","ICLR.cc/2027/Conference/Submission32088/-/Full_Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Reviewing","ICLR.cc/2027/Conference/-/Edit"],"mdate":1791231485916,"odate":1791031328286,"domain":"ICLR.cc/2027/Conference","id":"xlL4bjWnnH","version":2},{"content":{"summary":{"value":"The paper proposes ZSPAPrune, a zero-shot, prompt-aware token pruning framework for Vision-Language Models (VLMs). Existing pruning methods are often prompt-agnostic, ignoring text guidance and thus failing to prioritize task-relevant visual information. ZSPAPrune addresses this by reframing pruning as a balance between task relevance and information diversity, achieved through a hierarchical process: Prompt Simplification, Prompt-Aware Selection, and Diversity Balance. The method selects core visual tokens most relevant to the prompt and augments them with diverse tokens to retain global context. Experiments on multiple benchmarks and models show that ZSPAPrune achieves state-of-the-art or comparable performance with minimal accuracy loss even when pruning up to 90% of tokens, while significantly reducing GPU memory usage and inference latency."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"null"},"rating":{"value":4},"details_of_ethics_concerns":{"value":"no"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. From a perspective of prompt-aware token selection to balance task relevance and information diversity in visual representations.\n2. Introducing a hierarchical pruning mechanism composed of Prompt Simplification, Prompt-Aware Selection, and Diversity Balance to achieve controllable token reduction.\n3. Achieving significant inference efficiency improvements with minimal accuracy loss under zero-shot settings across multiple Vision-Language Models and benchmarks."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The paper lacks comparison with other methods that explicitly address the trade-off between task relevance and information diversity. Without such comparison, it remains unclear whether the proposed balance strategy is superior or merely heuristic.\n2. As a plug-and-play method, ZSPAPrune should be validated on more models with different parameter scales to confirm its general applicability. The current experiments are limited to a narrow range of architectures, reducing the evidence of scalability.\n3. The comparison with task-relevance-based approaches appears potentially unfair. Some baselines are reimplemented without clear alignment in training setup or hyperparameter tuning, which may bias the reported results.\n4. The proposed method is overly simple and lacks crucial theoretical analysis. No formal justification or complexity discussion is provided to explain why the hierarchical prompt-aware pruning mechanism should work effectively.\n5. The framework figure (i.e., Figure 2) is overly general and resembles a process diagram rather than an architectural framework. It fails to visually highlight the innovation and importance of the proposed components, and a more informative figure is recommended."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762927827516,"tcdate":1762239034685,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission18041/Reviewer_85xg"],"signatures":["ICLR.cc/2026/Conference/Submission18041/Reviewer_85xg"],"forum":"5Y8PMEeAkv","number":2,"license":"CC BY 4.0","cdate":1762239034685,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission18041/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762927827516,"domain":"ICLR.cc/2026/Conference","replyto":"5Y8PMEeAkv","id":"mA02kDIBIO","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["VLM","Prompt-Aware","Zero-Shot","Visual token pruning"]},"supplementary_material":{"value":"/attachment/d5f400bfcf8521072140b3aaa9f32a1675042e58.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"As the capabilities of Vision-Language Models (VLMs) advance, they can process increasingly large inputs, which, unlike in LLMs, generates significant visual token redundancy and leads to prohibitive inference costs. While many methods aim to reduce these costs by pruning visual tokens, existing approaches, whether based on attention or diversity, typically neglect the guidance of the text prompt and thus fail to prioritize task relevance. In this work, we propose a novel, zero-shot method that reframes the problem by introducing a prompt-aware perspective, explicitly modeling visual token pruning as a balance between task relevance and information diversity. Our hierarchical approach first selects a core set of task-relevant visual tokens and then supplements them with diversity tokens to preserve broader context. Experiments across multiple models and benchmarks show that our method achieves performance that matches or surpasses the state-of-the-art with only minimal accuracy loss, even when pruning up to 90\\% of the tokens. Furthermore, these gains are accompanied by significant reductions in GPU memory footprint and inference latency."},"_bibtex":{"value":"@misc{\nzhang2026zspaprune,\ntitle={{ZSPAP}rune: Zero-Shot Prompt-Aware Token Pruning for Vision-Language Models},\nauthor={Pu Zhang and Yuwei Li and Xingyuan XIAN and Guoming Tang},\nyear={2026},\nurl={https://openreview.net/forum?id=5Y8PMEeAkv}\n}"},"title":{"value":"ZSPAPrune: Zero-Shot Prompt-Aware Token Pruning for Vision-Language Models"},"pdf":{"value":"/pdf/8efa76bee829dd3de9e6813aa96412c3cd10132c.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"zhang|zspaprune_zeroshot_promptaware_token_pruning_for_visionlanguage_models"},"authorids":{"value":["~Pu_Zhang7","~Yuwei_Li4","~Xingyuan_XIAN1","~Guoming_Tang1"]},"authors":{"value":["Pu Zhang","Yuwei Li","Xingyuan XIAN","Guoming Tang"]}},"version":2},{"content":{"summary":{"value":"The paper investigates the learning of shortcut / spurious features. Mainly, the paper looks at the interplay between predictability and availability (that defines how easy it is for a model to extract a feature). First, experiments on simple, synthetic data, generated from two variables for which we can control the predictability and availability are shown. A notion of shortcut bias is defined as the additional reliance on spurious features for a learned model, compared to an optimal model. It is shown that availability determines the learning of spurious features, even when the predictivity of spurious features is lower than that of the core features. Non-linearities are shown to induce more bias for shortcuts. Theoretical analysis shows that linear networks are not biased to feature availability while ReLU networks are. Experiments on image datasets show that they are biased for background and object size."},"presentation":{"value":"2 fair"},"contribution":{"value":"3 good"},"soundness":{"value":"3 good"},"strengths":{"value":"- S1. The paper deals with an important aspect of understanding neural networks, the bias for learning shortcuts.\n\n- S2. Using the notion of shortcut bias, based on the reliance of an optimal predictor is a good idea.\n\n- S3. Analysing based on a notion of availability, that can be computed produces some good observations. \n\n- S4. Interesting to see that theoretically, ReLU networks are more biased than linear networks."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- W1. The notion of availability is very closely connected with the concept of simplicity of neural networks. The simplicity bias has been pointed out as a cause of non-robust learning. The paper gives multiple references (e.g. Shah et al., 2020, etc.) for works dealing with simplicity bias, but they are not discussed in detail. The authors must clarify what is the difference between the notion of availability presented in this work, and the simplicity bias previously introduced. What new observations does the proposed framework bring?\n\n\n- W2.The experiments are not that strong. The image datasets do not seem to bring any interesting observations. It is already known that the background influences the prediction to a high degree. It is expected that alterating the background will improve the reliance on the core features. All methods that use these datasets, try to learn how to not rely on the background. There doesn’t seem to be any novel observation in this regard. \n\n- W3. For image datasets: “a Bayes optimal classifier is not comparably sensitive to the predictivity of the non-core features” How do we know this? What experiments give this conclusion? \n\n- W3.2 Also, how is the Bayes optimal classifier created in the case of real image datasets (WaterBirds, CelebA)?\n\n- W4. The paper defined a shortcut as a feature that is more available, but less predictive. Differently, a shortcut is usually defined by good predictivity in distribution, but poor predictivity in some other distributions (OOD)."},"confidence":{"value":"3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked."},"questions":{"value":"Could the curves given by making an intervention on the background, be use to benchmark the robustness of different method? E.g. more robust models would have curves that are less sensitive to interventions on the background. Whould this offer any additional insight as opposed to comparing the accuracy of the model on balanced data, without foreground-background spurious correlations?"},"rating":{"value":"6: marginally above the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636702308,"tcdate":1698878107457,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission6363/Reviewer_mfV7"],"signatures":["ICLR.cc/2024/Conference/Submission6363/Reviewer_mfV7"],"forum":"Tj3xLVuE9f","number":4,"license":"CC BY 4.0","cdate":1698878107457,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission6363/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636702308,"domain":"ICLR.cc/2024/Conference","replyto":"Tj3xLVuE9f","id":"eKgVo54Ef8","forumContent":{"venue":{"value":"ICLR 2024 spotlight"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["shortcut learning","spurious correlations","architectural inductive bias"]},"primary_area":{"value":"visualization or interpretation of learned representations"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Deep-learning models can extract a rich assortment of features from data. Which features a model uses depends not only on *predictivity*---how reliably a feature indicates training-set labels---but also on *availability*---how easily the feature can be extracted from inputs. The literature on shortcut learning has noted examples in which models privilege one feature over another, for example texture over shape and image backgrounds over foreground objects. Here, we test hypotheses about which input properties are more available to a model, and systematically study how predictivity and availability interact to shape models' feature use. We construct a minimal, explicit generative framework for synthesizing classification datasets with two latent features that vary in predictivity and in factors we hypothesize to relate to availability, and we quantify a model's shortcut bias---its over-reliance on the shortcut (more available, less predictive) feature at the expense of the core (less available, more predictive) feature. We find that linear models are relatively unbiased, but introducing a single hidden layer with ReLU or Tanh units yields a bias. Our empirical findings are consistent with a theoretical account based on Neural Tangent Kernels. Finally, we study how models used in practice trade off predictivity and availability in naturalistic datasets, discovering availability manipulations which increase models' degree of shortcut bias. Taken together, these findings suggest that the propensity to learn shortcut features is a fundamental characteristic of deep nonlinear architectures warranting systematic study given its role in shaping how models solve tasks."},"_bibtex":{"value":"@inproceedings{\nhermann2024on,\ntitle={On the Foundations of Shortcut Learning},\nauthor={Katherine Hermann and Hossein Mobahi and Thomas FEL and Michael Curtis Mozer},\nbooktitle={The Twelfth International Conference on Learning Representations},\nyear={2024},\nurl={https://openreview.net/forum?id=Tj3xLVuE9f}\n}"},"title":{"value":"On the Foundations of Shortcut Learning"},"pdf":{"value":"/pdf/3f47b29f0e35691e7047d9fbfa0e4c47ea966e49.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference"},"paperhash":{"value":"hermann|on_the_foundations_of_shortcut_learning"},"authorids":{"value":["~Katherine_Hermann1","~Hossein_Mobahi2","~Thomas_FEL1","~Michael_Curtis_Mozer1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Katherine Hermann","Hossein Mobahi","Thomas FEL","Michael Curtis Mozer"]}},"version":2},{"content":{"venue":{"value":"NLPCC 2017"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-319-73618-1_17.pdf"},"venueid":{"value":"dblp.org/conf/NLPCC/2017"},"paperhash":{"value":"wu|shortcut_sequence_tagging"},"authorids":{"value":["~Huijia_Wu1","",""]},"html":{"value":"https://doi.org/10.1007/978-3-319-73618-1_17"},"_bibtex":{"value":"@inproceedings{DBLP:conf/nlpcc/WuZZ17,\n  author={Huijia Wu and Jiajun Zhang and Chengqing Zong},\n  title={Shortcut Sequence Tagging},\n  year={2017},\n  cdate={1483228800000},\n  pages={196-207},\n  url={https://doi.org/10.1007/978-3-319-73618-1_17},\n  booktitle={NLPCC},\n  crossref={conf/nlpcc/2017}\n}\n"},"abstract":{"value":"Deep stacked RNNs are usually hard to train. Recent studies have shown that shortcut connections across different RNN layers bring substantially faster convergence. However, shortcuts increase the computational complexity of the recurrent computations. To reduce the complexity, we propose the shortcut block, which is a refinement of the shortcut LSTM blocks. Our approach is to replace the self-connected parts (\\(c_t^l\\)) with shortcuts (\\(h_t^{l-2}\\)) in the internal states. We present extensive empirical experiments showing that this design performs better than the original shortcuts. We evaluate our method on CCG supertagging task, obtaining a 8% relatively improvement over current state-of-the-art results."},"title":{"value":"Shortcut Sequence Tagging"},"authors":{"value":["Huijia Wu","Jiajun Zhang","Chengqing Zong"]}},"tmdate":1773322978305,"pdate":1514678400000,"externalIds":["dblp:conf/nlpcc/WuZZ17"],"tcdate":1773322963910,"writers":["~"],"signatures":["~Huijia_Wu1"],"forum":"GYD4jbE2dR","license":"CC BY-SA 4.0","number":851098,"cdate":1483228800000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1773322978305,"domain":"DBLP.org","id":"GYD4jbE2dR","version":2},{"content":{"summary":{"value":"The paper presents a unified model for zero-shot composed image and video retrieval named Slerp+, aiming to simplify compositional retrieval without relying on supervised triplet annotations. The paper extends Slerp to generate joint embeddings that integrate visual and textual inputs effectively, aiming to enhance retrieval across image and video modalities. Slerp+ achieves improved performance across existing composed image and video retrieval tasks. The author also introduce a new video benchmark."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Please see the points under Weaknesses above."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. Unified Framework: This work  provides a simple yet effective solution that integrates image and video representations with text, improving both image and video composed retrieval performances. \n2. Efficient Training Design: Slerp+ utilizes existing vision-language pre-trained models, with efficient adaptation via low-rank adaptation (LoRA) to avoid high computational costs.\n3. Experimental validation: The authors conduct extensive experiments on multiple benchmarks, offering robust evidence of Slerp+’s effectiveness."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The paper contribution builds on existing interpolation methods, and the paper’s novelty lies primarily in applying Slerp in a unified framework. It would be beneficial to see additional model adjustments that directly address the unique challenges of video-text retrieval. \n\nWhile Slerp+ performs well across modalities, there may be challenges in scenarios where one modality (e.g., video-caption pairs) is limited or unbalanced in the dataset, potentially impacting the performance.\n\nThe current ablation study covers frame selection, but it would be helpful to explore how different interpolation parameters in Slerp influence performance, particularly for more varied or longer video content. It would be good to see other fusion strategies beyond average and cross -attention."}},"nonreaders":[],"tmdate":1731427508303,"tcdate":1730800936905,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission1907/Reviewer_tr5q"],"signatures":["ICLR.cc/2025/Conference/Submission1907/Reviewer_tr5q"],"forum":"YCOVTlMFIG","number":6,"license":"CC BY 4.0","cdate":1730800936905,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission1907/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427508303,"domain":"ICLR.cc/2025/Conference","replyto":"YCOVTlMFIG","id":"u1BnJxvdNE","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["multi-modal representation learning","composed retrieval"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Zero-shot composed image/video retrieval is a challenging task that involves using a combination of a reference visual input and a relative caption as a query to search for target visual data. Earlier studies have treated composed image retrieval and composed video retrieval methods separately, potentially neglecting the benefits of integrating image-video-text representation learning.  In this paper, we consolidate these tasks into a single Composed \\emph{Visual} Retrieval (CVR) task, which requires the composition of image and video samples with textual modifications using a unified retrieval model. Our principal insight is that the video modality can be effectively added to existing vision-language pretrained models. When integrated with the Spherical Linear Interpolation (Slerp) method previously proposed for Composed Image Retrieval (CoIR), we found that it results in an effective approach for solving the CVR task, which we called $\\text{Slerp}^{+}$. Extensive experiments demonstrate $\\text{Slerp}^{+}$'s superiority across various composed image and video retrieval benchmarks, including our newly proposed video benchmark. Notably, $\\text{Slerp}^{+}$ mutually enhances image and video retrieval performance over single-modality models, underscoring its potential to transform the field of compositional visual retrieval."},"_bibtex":{"value":"@misc{\njang2024textslerp,\ntitle={\\${\\textbackslash}text\\{Slerp\\}{\\textasciicircum}\\{+\\}\\$: Spherical Linear Interpolation for Unified Compositional Retrieval},\nauthor={Young Kyun Jang and Donghyun Kim and Bo He and Zihang Meng and Ser-Nam Lim},\nyear={2024},\nurl={https://openreview.net/forum?id=YCOVTlMFIG}\n}"},"title":{"value":"$\\text{Slerp}^{+}$: Spherical Linear Interpolation for Unified Compositional Retrieval"},"pdf":{"value":"/pdf/25ed87289cb3c609d20358983e2e562ad83f3404.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"jang|\\textslerp^_spherical_linear_interpolation_for_unified_compositional_retrieval"},"authorids":{"value":["~Young_Kyun_Jang1","~Donghyun_Kim2","~Bo_He1","~Zihang_Meng1","~Ser-Nam_Lim3"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Young Kyun Jang","Donghyun Kim","Bo He","Zihang Meng","Ser-Nam Lim"]}},"version":2},{"content":{"TLDR":{"value":"We introduce SenseMath, a benchmark that evaluates LLMs' ability to exploit numerical shortcuts, showing that improving number-sense enhances mathematical reasoning efficiency and accuracy."},"venue":{"value":"ICLR 2027 Conference Submission"},"keywords":{"value":["Number Sense","Reasoning","LLMs","Education"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Large language models often default to step-by-step computation even when efficient numerical shortcuts are available. This raises a basic question: do they exhibit number sense in a human-like behavioral sense, i.e., the ability to\n  recognize numerical structure, apply shortcuts when appropriate, and avoid them when they are not? We introduce SenseMath, a controlled benchmark for evaluating structure-sensitive numerical reasoning in LLMs. SenseMath contains\n  4,800 items spanning eight shortcut categories and four digit scales, with matched strong-shortcut, weak-shortcut, and control variants. Guided by Cognitive Load Theory, we manipulate intrinsic load through digit scale and shortcut\n  availability while controlling extraneous load through a consistent item type and format. Following Bloom's Taxonomy, SenseMath supports three evaluation settings of increasing cognitive complexity: Shortcut Use, Applicability\n  Judgment, and Problem Generation. Our evaluation across six model configurations reveals a consistent gap between executing shortcuts and understanding their applicability. Explicit number-sense prompting yields accuracy gains of up\n  to 23.0 percentage points on shortcut-amenable items, yet under standard chain-of-thought prompting, shortcut usage on four-digit strong-shortcut items ranges from 26.3% to 58.8%. Models systematically overgeneralize shortcut\n  applicability, with judgment accuracy ranging from 52.5% to 55.9%, and struggle to generate valid strong/control problem pairs, with only 0–21.9% passing all six automatic checks. Together, these results suggest that current LLMs\n  exhibit procedural shortcut fluency without consistently demonstrating the structural understanding of when and why shortcuts work that underlies human number sense."},"_bibtex":{"value":"@inproceedings{\nanonymous2026sensemath,\ntitle={SenseMath: Do {LLM}s Have Number Sense? Evaluating Shortcut Use, Judgment, and Generation},\nauthor={Anonymous},\nbooktitle={Submitted to The Fifteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=fORSAsyOix},\nnote={under review}\n}"},"title":{"value":"SenseMath: Do LLMs Have Number Sense? Evaluating Shortcut Use, Judgment, and Generation"},"pdf":{"value":"/pdf/712da44a2a289adfba16d26c7d62973feed3450b.pdf"},"venueid":{"value":"ICLR.cc/2027/Conference/Submission"}},"tmdate":1791231966889,"tcdate":1789694653394,"writers":["ICLR.cc/2027/Conference","ICLR.cc/2027/Conference/Submission35516/Authors"],"signatures":["ICLR.cc/2027/Conference/Submission35516/Authors"],"forum":"fORSAsyOix","license":"CC BY 4.0","number":35516,"cdate":1789694653394,"readers":["everyone"],"invitations":["ICLR.cc/2027/Conference/-/Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Bidding","ICLR.cc/2027/Conference/Submission35516/-/Full_Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Reviewing","ICLR.cc/2027/Conference/-/Edit"],"mdate":1791231966889,"odate":1791031328286,"domain":"ICLR.cc/2027/Conference","id":"fORSAsyOix","version":2},{"content":{"summary":{"value":"For multi-shot videos, current video benchmarks tend to annotate captions in a coarse-grained manner, either by providing a holistic caption for the entire video or by allowing annotators to subjectively define the boundaries of each event. But multi-shot videos often include rich narrations that correspond to distinct events occurring within the video. Given this context, this work introduces a new benchmark, *Shot2Story*, aimed at enhancing audio-visual understanding of multi-shot videos through the detailed textual annotation of each individual shot. The authors conduct extensive experiments from three perspectives: single-shot video captioning, multi-shot video summarization, and video question-answering with video summaries. These experiments demonstrate significant limitations in current models’ abilities to understand multi-shot content. Additionally, the authors design two baseline models based on MiniGPT-4 and Video-Chat2. Experimental results reveal that the multi-shot instruction data constructed by the authors substantially improves the performance of multi-shot video understanding."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"See weaknesses."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"This paper addressed here underscores a critical issue in current Video LLM work, where shot-level temporal understanding in video instruction data has largely been neglected. Most existing video instruction datasets lack the fine-grained, shot-level annotations necessary to capture nuanced transitions between events within videos.\n\nTo address and evaluate this issue, this work introduces a challenging multi-shot video understanding instruction dataset and benchmark. The dataset construction process involves substantial human oversight, which significantly reduces the likelihood of inaccuracies and biases often introduced by external models (e.g., GPT).\n\nIn the experiments, the authors analyze and compare several existing Video LLM approaches, highlighting the benchmark’s applicability and challenges. Additionally, they propose two suitalbe baseline models ,which achieve considerable improvements in the understanding of multi-shot video content."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The methods compared across the three tasks—video shot captioning, video summarization, and multi-shot video question answering—lack consistency, making it difficult to observe a clear, uniform improvement in performance across these tasks with the proposed dataset. To improve consistency, I suggest that the authors consider using the same set of baseline models across all tasks. \n\nThe constructed multi-shot dataset appears to resemble a series of single-shot descriptions loosely connected by simple transition words, such as \"begin with,\" \"then,\" and \"next.\" The annotation process seems more akin to first splitting multi-shot videos into individual shots, generating isolated descriptions for each, and then using an external tool (e.g., GPT-4) to concatenate these descriptions. This method does not accurately represent the interconnections between shots, which may pose challenges in teaching large language models (LLMs) multi-shot temporal reasoning skills using the constructed instructional data. Although the actual content contained in different shots is inconsistent, there is an obvious time sequence and event relationship between them. I suggest that the authors improve the annotation process by having annotators explicitly describe relationships or continuity between shots, rather than relying on simple transition words. This approach would better capture the temporal flow and coherence between shots, which is essential for developing multi-shot temporal reasoning capability in LLMs.\n\n\nMoreover, it raises the question of whether multi-shot understanding is at odds with simple video understanding and long-video comprehension. I noticed that the average length of videos in the dataset is only 17.1 seconds, and shots with minimal variation between them were deliberately excluded during construction. Although the videos are relatively short, they contain numerous individual shots with significant differences between them. This could lead models to develop a bias toward segmenting any input video into multiple shots for interpretation, which may pose challenges for understanding both simpler and longer videos. I recommend that the authors conduct additional experiments to evaluate whether models trained on this dataset struggle with simple single-shot videos or longer videos, as this could help clarify any potential biases introduced by the dataset construction."}},"nonreaders":[],"tmdate":1731428921880,"tcdate":1730554078447,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission7580/Reviewer_n4p7"],"signatures":["ICLR.cc/2025/Conference/Submission7580/Reviewer_n4p7"],"forum":"FZv3kPHTtB","number":2,"license":"CC BY 4.0","cdate":1730554078447,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission7580/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428921880,"domain":"ICLR.cc/2025/Conference","replyto":"FZv3kPHTtB","id":"kYkNOy3Eux","forumContent":{"venue":{"value":"ICLR 2025 Poster"},"TLDR":{"value":"Shot2Story presents a large-scale dataset with 43,000 multi-shot videos and 188,000 manually annotated shots, offering detailed visual/audio captions, summaries, and QA pairs to advance multi-shot video understanding."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["vision language model","video question answering","video captioning","multi-shot videos"]},"supplementary_material":{"value":"/attachment/19e43f5040b7d5e255d15b50136dc016857cf8ad.pdf"},"primary_area":{"value":"datasets and benchmarks"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"A short clip of video may contain progression of multiple events and an interesting story line. A human need to capture both the event in every shot and associate them together to understand the story behind it. In this work, we present a new multi-shot video understanding benchmark \\dataset with detailed shot-level captions, comprehensive video summaries and question-answering pairs. To facilitate better semantic understanding of videos, we provide captions for both visual signals and human narrations. We design several distinct tasks including single-shot video captioning, multi-shot video summarization, and multi-shot video question answering. Preliminary experiments show some challenges to generate a long and comprehensive video summary for multi-shot videos. Nevertheless, the generated imperfect summaries can already achieve competitive performance on existing video understanding tasks such as video question-answering, promoting an under-explored setting of video understanding with detailed summaries."},"_bibtex":{"value":"@inproceedings{\nhan2025shotstory,\ntitle={Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos},\nauthor={Mingfei Han and Linjie Yang and Xiaojun Chang and Lina Yao and Heng Wang},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=FZv3kPHTtB}\n}"},"title":{"value":"Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos"},"pdf":{"value":"/pdf/316c2e51532fadb7d781d28a19d13e78f4aa4ca0.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"han|shot2story_a_new_benchmark_for_comprehensive_understanding_of_multishot_videos"},"authorids":{"value":["~Mingfei_Han1","~Linjie_Yang4","~Xiaojun_Chang4","~Lina_Yao2","~Heng_Wang2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Mingfei Han","Linjie Yang","Xiaojun Chang","Lina Yao","Heng Wang"]}},"version":2},{"content":{"summary":{"value":"- This paper presents Shot2Story, a large-scale benchmark for comprehensive multi-shot video understanding. \n\n- The dataset contains detailed shot-level captions for both visual signals and human narrations, and comprehensive video summaries based on shot-level captions.\n\n- The authors further design a challenging video question-answering benchmark for multi-shot video understanding."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"I think the main concern is on its contribution of methodology, as it doesn't involve any architecture/method design for tackling the multi-shot video understanding problem."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"- The paper presents a new multi-shot video understanding benchmark (Shot2Story), with detailed shot-level captions, comprehensive video summaries and question-answering pairs. This will benefit the community, as a benchmark for multi-modal, multi-shot video understanding benchmark.\n\n- The paper has made significant efforts for curating the datasets, to maintain high quality standard, getting the data from 2.1M video clips to 42,958 video clips. The descriptions of the filtering stages are clear and reasonable.\n\n- For baseline establishment, the authors have extensively evaluated a series of open-source models."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- My main concern is on its scientific novelty, as the paper does not include new designs on tackling the multi-shot video evaluation problem, only providing a benchmark, it seems to be more suitable for a benchmark track.\n\n- For baseline evaluation, the focus has been mainly on open-source models, what about models, for example, gpt4v, claude3.5 ?\n\n- In Table 1, \\hline is incorrect ?"}},"nonreaders":[],"tmdate":1734313606742,"tcdate":1730688378337,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission7580/Reviewer_G1Gp"],"signatures":["ICLR.cc/2025/Conference/Submission7580/Reviewer_G1Gp"],"forum":"FZv3kPHTtB","number":3,"license":"CC BY 4.0","cdate":1730688378337,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission7580/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1734313606742,"domain":"ICLR.cc/2025/Conference","replyto":"FZv3kPHTtB","id":"UjIVKX8oGb","forumContent":{"venue":{"value":"ICLR 2025 Poster"},"TLDR":{"value":"Shot2Story presents a large-scale dataset with 43,000 multi-shot videos and 188,000 manually annotated shots, offering detailed visual/audio captions, summaries, and QA pairs to advance multi-shot video understanding."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["vision language model","video question answering","video captioning","multi-shot videos"]},"supplementary_material":{"value":"/attachment/19e43f5040b7d5e255d15b50136dc016857cf8ad.pdf"},"primary_area":{"value":"datasets and benchmarks"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"A short clip of video may contain progression of multiple events and an interesting story line. A human need to capture both the event in every shot and associate them together to understand the story behind it. In this work, we present a new multi-shot video understanding benchmark \\dataset with detailed shot-level captions, comprehensive video summaries and question-answering pairs. To facilitate better semantic understanding of videos, we provide captions for both visual signals and human narrations. We design several distinct tasks including single-shot video captioning, multi-shot video summarization, and multi-shot video question answering. Preliminary experiments show some challenges to generate a long and comprehensive video summary for multi-shot videos. Nevertheless, the generated imperfect summaries can already achieve competitive performance on existing video understanding tasks such as video question-answering, promoting an under-explored setting of video understanding with detailed summaries."},"_bibtex":{"value":"@inproceedings{\nhan2025shotstory,\ntitle={Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos},\nauthor={Mingfei Han and Linjie Yang and Xiaojun Chang and Lina Yao and Heng Wang},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=FZv3kPHTtB}\n}"},"title":{"value":"Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos"},"pdf":{"value":"/pdf/316c2e51532fadb7d781d28a19d13e78f4aa4ca0.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"han|shot2story_a_new_benchmark_for_comprehensive_understanding_of_multishot_videos"},"authorids":{"value":["~Mingfei_Han1","~Linjie_Yang4","~Xiaojun_Chang4","~Lina_Yao2","~Heng_Wang2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Mingfei Han","Linjie Yang","Xiaojun Chang","Lina Yao","Heng Wang"]}},"version":2},{"content":{"summary":{"value":"The paper addresses challenges in co-speech gesture video generation through a proposed method called Realistic-Gesture. This method incorporates three main components: a speech-aware gesture motion representation, a masked gesture motion generator, and a pixel-level refinement module. Experimental results show that the proposed approach generates realistic co-speech gesture videos, while also enabling long-sequence generation and video editing capabilities."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"See weaknesses part."},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. The authors identify three key challenges in co-speech gesture video generation and propose solutions to address each one.\n\n2. Amount of ablation studies are conducted to verify the effectiveness of the proposed modules."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Generally, some challenges identified in this paper have been addressed in prior research; however, the authors failed to cite these studies or compare their findings with the experiments conducted in this work. Notable examples include:\n\n1. **\"Rhythmic Gesticulator: Rhythm-Aware Co-Speech Gesture Synthesis with Hierarchical Neural Embeddings.\"** This study proposed a shared hierarchical embedding for both speech content and motion, which closely resembles the approach taken by the authors.\n  \n2. **\"Make-Your-Anchor: A Diffusion-based 2D Avatar Generation Framework.\"** This research adopted SMPL-X parameters as motion representations for co-speech video generation, enhancing the clarity of hand movements.\n\nAdditional weaknesses include:\n\n1. While the proposed method can perform gesture inpainting, the authors inaccurately claim that it supports video gesture editing. This assertion is misleading, as the inpainted gestures are not controllable.\n\n2. In some of the \"Gesture Pattern Transfer\" videos, the character appears distorted, likely due to differences in body proportions between the source and target characters."}},"nonreaders":[],"tmdate":1731427566941,"tcdate":1730640203888,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission2259/Reviewer_bDof"],"signatures":["ICLR.cc/2025/Conference/Submission2259/Reviewer_bDof"],"forum":"EXsiGFkwV6","number":5,"license":"CC BY 4.0","cdate":1730640203888,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission2259/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427566941,"domain":"ICLR.cc/2025/Conference","replyto":"EXsiGFkwV6","id":"26SgEfnPPY","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["gesture generation; motion representation; video generation"]},"supplementary_material":{"value":"/attachment/cfe92844e4bbcfa27796d460cd200dd923536aa3.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Co-speech gesture generation is crucial for creating lifelike avatars and enhancing human-computer interactions by synchronizing gestures with speech in computer vision. Despite recent advancements, existing methods often struggle with accurately aligning gesture motions with speech signals and achieving pixel-level realism. To address these challenges, we introduce Realistic-Gesture, a groundbreaking framework that transforms co-speech gesture video generation through three innovative components: (1) a speech-aware gesture tokenization that incorporate speech context into motion pattern representation, (2) a mask gesture generator that learns to map audio signals to gestures by predicting masked motion tokens, enabling bidirectional contextually relevant gesture synthesis and editing, and (3) a structure-aware refinement module that employs differentiable edge connection to link gesture keypoints to improve video generation. Our extensive experiments demonstrate that Realistic-Gesture not only produces highly realistic and speech-aligned gesture videos but also supports long-sequence generation and video gesture editing applications."},"_bibtex":{"value":"@misc{\nliu2025realisticgesture,\ntitle={Realistic-Gesture: Co-Speech Gesture Video Generation through Semantic-aware Gesture Representation},\nauthor={Pinxin Liu and Pengfei Zhang and Hyeongwoo Kim and Pablo Garrido and Ari Shapiro and Kyle Olszewski},\nyear={2025},\nurl={https://openreview.net/forum?id=EXsiGFkwV6}\n}"},"title":{"value":"Realistic-Gesture: Co-Speech Gesture Video Generation through Semantic-aware Gesture Representation"},"pdf":{"value":"/pdf/c55f7b536099cd47de6a891a58165127933b1943.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"liu|realisticgesture_cospeech_gesture_video_generation_through_semanticaware_gesture_representation"},"authorids":{"value":["~Pinxin_Liu1","~Pengfei_Zhang12","~Hyeongwoo_Kim3","~Pablo_Garrido1","~Ari_Shapiro3","~Kyle_Olszewski1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Pinxin Liu","Pengfei Zhang","Hyeongwoo Kim","Pablo Garrido","Ari Shapiro","Kyle Olszewski"]}},"version":2},{"content":{"summary":{"value":"The paper presents a vision-language model, CLIP-style, for echocardiography. VLMs for echocardiography suffer from internal challenges in accurate measurement predictions, and sparse and unfocused data of image-text pairs. The proposed model combines image and text encoders trained on the novel EchoGround-MIMIC dataset (19,065 measurement-grounded image-text pairs). The dataset comes from preprocessing and organizing existing public repos (MIMIC-ECHO), and could be valuable for further research—so far there is no similar open-source data, therefore the data is useful. The model uses two specialized contrastive losses: a view-informed contrastive loss (same-view positives, different-view negatives) and a negation-aware contrastive loss for distinguishing negative vs. positive clinical findings in text. However, these losses appear to offer limited technical novelty. For downstream tasks like segmentation and landmark detection, task-specific heads are added to the pre-trained encoder and fine-tuned on benchmark datasets."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Given that EchoPrime is publicly available while EchoApex is unreleased, why not use EchoPrime as the primary baseline? This would enable reproducible comparisons and address potential selection bias concerns."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"-- Data: EchoGround-MIMIC (~20K measurement-grounded image-text pairs) - first open-source dataset of its kind for echocardiography. Data Processing Innovation -- Successfully integrates MIMIC-IV-ECHO imaging with MIMIC-IV-Note reports. \n\n\n-- Clinical Relevance: Addresses critical gap between free-text narratives and quantitative measurements essential for guideline-based echo diagnosis.\n\n-- Comprehensive Evaluation Framework: 36 tasks across 5 clinical application types (classification, retrieval, segmentation, landmark detection) - would be valuable if released as a benchmark.\n\n-- Community Value: Fills significant resource gap for medical AI research in echocardiography."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"-- Limited Technical Novelty: View-informed loss is just constrained negative sampling; negation-aware loss potentially similar to existing work (e.g., MICCAI 2025, \"EchoViewCLIP: Advancing Video Quality Control through High-performance View Recognition of Echocardiography\")\n\n-- Evaluation Methodology Issues: Primary comparison against unreleased EchoApex (weights/data are not released, based on reported results) instead of available EchoPrime (weights are open to download) raises reproducibility concerns, and if all models were trained in the same manner. \n\n-- Technical Details: Frame vs. video level processing unclear; mathematical formulation of negation-aware loss may lack sufficient innovation\n\n-- Algorithmic Contributions Questionable: Technical contributions may not meet novelty bar for top-tier venues - relies heavily on dataset contribution rather than methodological innovation"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762915930291,"tcdate":1761833867382,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission1886/Reviewer_zREj"],"signatures":["ICLR.cc/2026/Conference/Submission1886/Reviewer_zREj"],"forum":"DKOIADzbtM","number":4,"license":"CC BY 4.0","cdate":1761833867382,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission1886/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762915930291,"domain":"ICLR.cc/2026/Conference","replyto":"DKOIADzbtM","id":"uqTActq3Nq","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Echocardiography","vision-language model","ultrasound"]},"primary_area":{"value":"applications to physical sciences (physics, chemistry, biology, etc.)"},"abstract":{"value":"Echocardiography is the most widely used imaging modality in cardiology, yet its interpretation remains labor-intensive and inherently multimodal, which requires view recognition, quantitative measurements, qualitative assessments, and guideline-based reasoning. While recent vision–language models (VLMs) have achieved broad success in natural images and certain medical domains, their potential in echocardiography has been limited by the lack of large-scale, clinically grounded image–text datasets and the absence of measurement-based reasoning central to echo interpretation. We introduce EchoGround-MIMIC, the first measurement-grounded multimodal echocardiography dataset, comprising 19,065 image–text pairs from 1,572 patients with standardized views, structured measurements, measurement-grounded captions, and guideline-derived disease labels. Building on this resource, we propose EchoVLM, a vision–language model that incorporates two novel pretraining objectives: (i) a view-informed contrastive loss that encodes the view-dependent structure of echocardiographic imaging, and (ii) a negation-aware contrastive loss that distinguishes clinically critical negative from positive findings. Across five types of clinical applications with 36 tasks spanning multimodal disease classification, image–text retrieval, view classification, chamber segmentation, and landmark detection, EchoVLM achieves state-of-the-art performance (86.5\\% AUC in zero-shot disease classification and 95.1\\% accuracy in view classification). We demonstrate that clinically grounded multimodal pretraining yields transferable visual representations and establish EchoVLM as foundation model for end-to-end echocardiography interpretation. We will release EchoGround-MIMIC and data curation code, enabling reproducibility and further research in multimodal echocardiography interpretation."},"_bibtex":{"value":"@misc{\nli2026echovlm,\ntitle={Echo{VLM}: Measurement-Grounded Multimodal Learning for Echocardiography},\nauthor={Yuheng Li and Yue Zhang and Abdoul Aziz Amadou and Yuxiang Lai and Jike Zhong and Tiziano Passerini and Dorin Comaniciu and Puneet Sharma},\nyear={2026},\nurl={https://openreview.net/forum?id=DKOIADzbtM}\n}"},"title":{"value":"EchoVLM: Measurement-Grounded Multimodal Learning for Echocardiography"},"pdf":{"value":"/pdf/cb33697f5979a8bae4c5a16ee5f471480af12fd4.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"li|echovlm_measurementgrounded_multimodal_learning_for_echocardiography"},"authorids":{"value":["~Yuheng_Li4","~Yue_Zhang18","~Abdoul_Aziz_Amadou1","~Yuxiang_Lai1","~Jike_Zhong1","~Tiziano_Passerini2","~Dorin_Comaniciu1","~Puneet_Sharma5"]},"authors":{"value":["Yuheng Li","Yue Zhang","Abdoul Aziz Amadou","Yuxiang Lai","Jike Zhong","Tiziano Passerini","Dorin Comaniciu","Puneet Sharma"]}},"version":2},{"content":{"summary":{"value":"- This paper introduces a web-scale video language dataset InternVid comprised of 234M clips and 760K hours.\n- It introduces ViCLIP, a video-language model based on ViT-L, pretrained with contrastive learning (like CLIP) and masked autoencoder (like VideoMAE, MAE).\n- This work expands the utility of their dataset to video understanding tasks like recognition and retrieval, and video generation."},"presentation":{"value":"3 good"},"contribution":{"value":"3 good"},"soundness":{"value":"3 good"},"strengths":{"value":"- It introduces a web-scale video-language dataset bridging the gap of lack of such datasets, unlike image-text domains.\n- It evaluates on multiple benchmarks in both finetuned and zero-shot setups.\n- The dataset curation is fairly detailed and I also like the hierarchical video caption generation strategy."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The majority of the comparisons are with CLIP which was pretrained with image-text pairs and not video-text; I would like to see how ViCLIP performs compared to other video or video-language models (e.g., pretrained on HD-VILA, HT100M, YT100M), I think that would be a fair comparison than comparing with image-text models.\n- I would encourage authors to dig deeper and find a more concrete argument/explanation: why does pretraining with InternVid-10M-FLT or InternVid-10M-DIV perform better in **zero-shot** than InternVid-200M, but not when **finetuned**. \n- In fig. 7 and 8, the experiments are done only on scaling the dataset size, it would be interesting to see the effect of model scaling in addition to the dataset scaling. I would encourage you to add such experiments in the final version.\n- I would be interested to see linear evaluation (i.e., a single FC layer, or use linear SVM) performance on the downstream benchmarks."},"confidence":{"value":"5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully."},"questions":{"value":"- Will you share the pretrained and finetuned models with the supporting code base (e.g., data processing, pretraining, finetuning, generation)? \n- Could you please share the processed clips, even processing the data by individuals (typically for academic researchers) would be a difficult task considering its massive size. Additionally, we all know the unavailability of videos due to location constraints, permission issues, etc., so even if the full data is not possible to share at least share the 3 10M versions.\n- I suggest releasing the fixed embeddings of the datasets from the trained (e.g., pretrained, finetuned) models.\n- Did you investigate if the InternVid has any sort of bias in the curated clips, could you please share a report with such details? Bias could be of many forms e.g., location/race/gender per action category. A suggested reference: https://arxiv.org/abs/1505.01257\n- Did you investigate, if ViCLIP is robust against some of the OOD setups, some of the popular benchmarks are Mimetics, RareAct etc. For more details please see: https://arxiv.org/abs/2306.02014"},"rating":{"value":"8: accept, good paper"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636477339,"tcdate":1698680389582,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission4917/Reviewer_6e9v"],"signatures":["ICLR.cc/2024/Conference/Submission4917/Reviewer_6e9v"],"forum":"MLBdiWu4Fw","number":2,"license":"CC BY 4.0","cdate":1698680389582,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission4917/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636477339,"domain":"ICLR.cc/2024/Conference","replyto":"MLBdiWu4Fw","id":"tT8nT4UOl8","forumContent":{"venue":{"value":"ICLR 2024 spotlight"},"TLDR":{"value":"This paper introduces InternVid, a large-scale video-centric multimodal dataset that enables learning powerful and transferable video-text representations for multimodal understanding and generation."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["video-language dataset","video understanding","video generation","multimodal understanding","action recognition","video retrieval"]},"supplementary_material":{"value":"/attachment/ab6262d86bfdbdf3844255ccbd7ff6bdca41c17b.zip"},"primary_area":{"value":"datasets and benchmarks"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"This paper introduces InternVid, a large-scale video-centric multimodal dataset that enables learning powerful and transferable video-text representations for multimodal understanding and generation. InternVid contains over 7 million videos lasting nearly 760K hours, yielding 234M video clips accompanied by detailed descriptions of total 4.1B words. Our core contribution is to develop a scalable approach to autonomously build a high-quality video-text dataset with large language models (LLM), thereby showcasing its efficacy in learning video-language representation at scale. Specifically, we utilize a multi-scale approach to generate video-related descriptions. Furthermore, we introduce ViCLIP, a video-text representation learning model based on ViT-L. Learned on InternVid via contrastive learning, this model demonstrates leading zero-shot action recognition and competitive video retrieval performance. Beyond basic video understanding tasks like recognition and retrieval, our dataset and model have broad applications. They are particularly beneficial for generating interleaved video-text data for learning a video-centric dialogue system, advancing video-to-text and text-to-video generation research. These proposed resources provide a tool for researchers and practitioners interested in multimodal video understanding and generation."},"_bibtex":{"value":"@inproceedings{\nwang2024internvid,\ntitle={InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation},\nauthor={Yi Wang and Yinan He and Yizhuo Li and Kunchang Li and Jiashuo Yu and Xin Ma and Xinhao Li and Guo Chen and Xinyuan Chen and Yaohui Wang and Ping Luo and Ziwei Liu and Yali Wang and Limin Wang and Yu Qiao},\nbooktitle={The Twelfth International Conference on Learning Representations},\nyear={2024},\nurl={https://openreview.net/forum?id=MLBdiWu4Fw}\n}"},"title":{"value":"InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation"},"pdf":{"value":"/pdf/5355ce2fec3ff26dca65a969b767fd7b1102bb05.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference"},"paperhash":{"value":"wang|internvid_a_largescale_videotext_dataset_for_multimodal_understanding_and_generation"},"authorids":{"value":["~Yi_Wang19","~Yinan_He1","~Yizhuo_Li1","~Kunchang_Li1","~Jiashuo_Yu1","~Xin_Ma3","~Xinhao_Li1","~Guo_Chen2","~Xinyuan_Chen1","~Yaohui_Wang1","~Ping_Luo2","~Ziwei_Liu1","~Yali_Wang1","~Limin_Wang1","~Yu_Qiao1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Yi Wang","Yinan He","Yizhuo Li","Kunchang Li","Jiashuo Yu","Xin Ma","Xinhao Li","Guo Chen","Xinyuan Chen","Yaohui Wang","Ping Luo","Ziwei Liu","Yali Wang","Limin Wang","Yu Qiao"]}},"version":2},{"content":{"venue":{"value":"Video-Langauge Models Oral"},"pdf":{"value":"/pdf/8f671fd31d8ccdd231b7d43d858aaa1a319fc0f5.pdf"},"keywords":{"value":["Video Captioning","Video Understanding","Multimodal Learning"]},"venueid":{"value":"NeurIPS.cc/2024/Workshop/Video-Langauge_Models"},"paperhash":{"value":"li|wolf_captioning_everything_with_a_world_summarization_framework"},"authorids":{"value":["~Boyi_Li1","~Ligeng_Zhu1","~Ran_Tian2","~Shuhan_Tan2","~Yuxiao_Chen3","~Yao_Lu13","~Yin_Cui1","~Sushant_Veer1","~Max_Ehrlich1","~Jonah_Philion1","~Xinshuo_Weng3","~Fuzhao_Xue1","~Andrew_Tao1","~Ming-Yu_Liu1","~Sanja_Fidler1","~Boris_Ivanovic1","~Trevor_Darrell2","~Jitendra_Malik2","~Song_Han5","~Marco_Pavone1"]},"abstract":{"value":"We propose Wolf, a WOrLd summarization Framework for accurate video captioning. Wolf is an automated captioning framework that adopts a mixture-of-experts approach, leveraging complementary strengths of Vision Language Models (VLMs). By utilizing both image and video models, our framework captures different levels of information and summarizes them efficiently. Our approach can be applied to enhance video understanding, auto-labeling, and captioning. To evaluate caption quality, we introduce CapScore, an LLM-based metric to assess the similarity and quality of generated captions compared to the ground truth captions. We further build four human-annotated datasets in three domains: autonomous driving, general scenes, and robotics, to facilitate comprehensive comparisons. We show that Wolf achieves superior captioning performance compared to state-of-the-art approaches from the research community (VILA1.5, CogAgent) and commercial solutions (Gemini-Pro-1.5, GPT-4V). For instance, in comparison with GPT-4V, Wolf improves CapScore (caption quality) by 55.6% and CapScore (caption similarity) by 77.4% on challenging driving videos. Finally, we establish a benchmark for video captioning and introduce a leaderboard, aiming to accelerate advancements in video understanding, captioning, and data alignment."},"_bibtex":{"value":"@inproceedings{\nli2025wolf,\ntitle={Wolf: Captioning Everything with a World Summarization Framework},\nauthor={Boyi Li and Ligeng Zhu and Ran Tian and Shuhan Tan and Yuxiao Chen and Yao Lu and Yin Cui and Sushant Veer and Max Ehrlich and Jonah Philion and Xinshuo Weng and Fuzhao Xue and Andrew Tao and Ming-Yu Liu and Sanja Fidler and Boris Ivanovic and Trevor Darrell and Jitendra Malik and Song Han and Marco Pavone},\nbooktitle={Workshop on Video-Language Models @ NeurIPS 2024},\nyear={2025},\nurl={https://openreview.net/forum?id=YtLubsMyU1}\n}"},"title":{"value":"Wolf: Captioning Everything with a World Summarization Framework"},"track":{"value":"Long Paper Track (up to 9 pages)"},"authors":{"value":["Boyi Li","Ligeng Zhu","Ran Tian","Shuhan Tan","Yuxiao Chen","Yao Lu","Yin Cui","Sushant Veer","Max Ehrlich","Jonah Philion","Xinshuo Weng","Fuzhao Xue","Andrew Tao","Ming-Yu Liu","Sanja Fidler","Boris Ivanovic","Trevor Darrell","Jitendra Malik","Song Han","Marco Pavone"]}},"tmdate":1736861079760,"pdate":1730081751760,"tcdate":1724476950509,"writers":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission1/Authors"],"signatures":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission1/Authors"],"forum":"YtLubsMyU1","license":"CC BY 4.0","number":1,"cdate":1724476950509,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Post_Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Edit","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission1/-/Camera-Ready_Revision"],"mdate":1736861079760,"odate":1736861079744,"domain":"NeurIPS.cc/2024/Workshop/Video-Langauge_Models","id":"YtLubsMyU1","version":2},{"content":{"summary":{"value":"The paper proposes SPLAVU, a lightweight anonymizing adapter placed on top of a frozen video encoder to remove spatially private information from video features while preserving temporal information for downstream tasks. The method is trained with three objectives: a clip-level self-supervised privacy objective to reduce mutual information between static clips, task-specific utility losses, and a latent consistency term to preserve transferability. Experiments on action recognition, temporal detection, and anomaly detection show about a 35% drop in private-attribute predictability with minimal utility loss. Privacy is evaluated on VISPR and CASIA-B, and a bias protocol is included on NTU RGB+D and Toyota Smarthome. The approach is simple and does not require retraining the video backbone."},"soundness":{"value":2},"confidence":{"value":5},"questions":{"value":"My questions are based on the weakness above. Specifically:\n1. Can the authors please clarify the implementation details of the privacy classifier and provide more information on how privacy was evaluated?\n2. Do the authors evaluate the robustness of their method against feature inversion or reconstruction attacks? Can the anonymized features $h'_t$ be used to approximate or reconstruct the original frames $x_t$?\n3. Does the proposed method actually generalize to unseen tasks such as video retrieval, video object segmentation, captioning, or re-identification?\n4. If there is no adversarial classifier or adversarial optimization during training, how do the authors ensure that the optimization process is actually guided to remove privacy-related attributes rather than destroying useful features? Simply reducing mutual information between frame features does not seem to be a strong proxy for privacy removal.\n5. How does the input of the privacy classifier look like?\n6. How stable is the proposed approach? Since the privacy and utility objectives are inherently competing, I wonder wheter the joint optimization of $f_A$ and $f_T^{*}$, considering $L_B$ and $L_{LC}$, remains stable during training.\n7. Why is the ablation on the latent-consistency loss reported only on HMDB51? Could you extend it to another dataset shown in Table 1?\n8. Did the authors analyze how clip length or sampling strategy affects privacy and utility? Would longer or shorter clips change the results?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- The paper is well motivated. Privacy preservation in video representations is an important problem, and the community should work more on this topic.\n- The paper introduces the novel idea of anonymizing latent features instead of input pixels or frames.\n- The proposed anonymization adapter (AAM) is lightweight, plug-and-play, and works on top of a frozen video encoder, making the approach simple and efficient to train.\n- The visual explanation in Figure 2 helps readers understand the method.\n- The evaluation includes privacy, fairness, and bias analysis, which broadens the impact and shows awareness of social implications."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The privacy evaluation is not clear. It appears that the authors train a separate classifier on anonymized features f_A(f_E(x)) for VISPR, but the implementation details (e.g., classifier type, training schedule, multi-label setting, etc) are missing. Clarifying this would help assess reproducibility and fairness when comparing to prior privacy-preserving frameworks that integrate the classifier during training.\n\n- The paper enforces a latent-consistency loss to keep anonymized features close to the encoder's latent space, aiming for generalization to unseen tasks. However, this design could also make the anonymized representations vulnerable to feature inversion or reconstruction attacks, since $h'_t=f_A(f_E(x_t))$ may remain close to $h_t=f_E(x_t)$. An attacker with access to the frozen encoder or a decoder trained on its features could potentially approximate the original frames. Specifically, an adversary could learn an approximate inverse $f_E^{-1}$ or a regression model $g: h'_t \\rightarrow x_t$ trained on pairs $(h_t, x_t)$ or $(h'_t, x_t)$, enabling recovery of frames from anonymized features. A discussion of this risk would strengthen the paper's privacy claims.\n\n- The paper claims generalization to \"unseen tasks,\" yet the evaluated ones (action recognition, temporal action detection, and anomaly detection) are all closely related motion-centric problems. Assessing the method on more diverse downstream video understanding tasks, such as video retrieval, video object segmentation, video captioning, or person re-identification, would provide stronger and more convincing evidence of true task-level generalization.\n\n- Unlike prior works such as SPAct, this paper does not include an explicit privacy classifier or adversarial optimization loop. While this design choice simplifies training, it also weakens the link between the training objective and actual privacy protection. The self-supervised \"privacy budget\" loss acts only as an indirect proxy for removing spatial identity information and may not guarantee robustness. It would also be helpful to visualize the input to the privacy classifier model.\n\n- Since the privacy and utility objectives are inherently competing, I wonder if the joint optimization of $f_A$ and $f_T^{*}$, considering $L_B$ and $L_{LC}$, remains stable.\n\n- The ablation on the latent-consistency (cycle-consistency) loss is evaluated only on HMDB51, whereas the main results (Table 1) include multiple datasets and tasks. It would be more convincing to show the same ablation across others.\n\n- The paper fixes the clip sampling and length configuration but does not analyze its impact on privacy or utility. Since privacy leakage and temporal consistency can depend strongly on clip duration and sampling strategy, an ablation on clip length or selection would help clarify whether the proposed method's gains generalize across different temporal windows.\n\n- The structure could be improved for clarity. While $f_A$ appears in the equations of Section 3.1, its definition is only given in Section 3.2. Introducing it earlier would make the methodological flow easier to follow.\n\n- The notation is not clearly defined and is sometimes hard to follow. For example, $T^n$ is used in the loss formulation but never explicitly defined; it's unclear whether n indexes tasks, samples, or temporal segments."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762918505230,"tcdate":1761728982079,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission6153/Reviewer_dNWy"],"signatures":["ICLR.cc/2026/Conference/Submission6153/Reviewer_dNWy"],"forum":"ncA3UUL0Ri","number":1,"license":"CC BY 4.0","cdate":1761728982079,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission6153/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762918505230,"domain":"ICLR.cc/2026/Conference","replyto":"ncA3UUL0Ri","id":"G4TjVQyeF6","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"TLDR":{"value":"We propose a plug-and-play latent anonymization adapter for video foundation models that substantially reduces private-attribute leakage while preserving performance across multiple downstream video tasks and also mitigates gender bias."},"keywords":{"value":["Privacy Preservation","Video Understanding"]},"primary_area":{"value":"alignment, fairness, safety, privacy, and societal considerations"},"abstract":{"value":"We introduce a novel formulation of visual privacy preservation for video foundation models that operates entirely in the latent space. While spatio-temporal features learned by foundation models have deepened general understanding of video content, sharing or storing these extracted visual features for downstream tasks inadvertently reveals sensitive personal information like skin color, gender, or clothing. Current privacy preservation methods focus on input-pixel-level anonymization, which requires retraining the entire utility video model and results in task-specific anonymization, making them unsuitable for recent video foundational models. To address these challenges, we introduce a lightweight Anonymizing Adapter Module (AAM) that removes private information from video features while retaining general task utility. AAM can be applied in a plug-and-play fashion to frozen video encoders, minimizing the computational burden of finetuning and re-extracting features. Our framework employs three newly designed training objectives: (1) a clip-level self-supervised privacy objective to reduce mutual information between static clips, (2) a co-training objective to retain utility across seen tasks, and (3) a latent consistency loss for generalization on unseen tasks. Our extensive evaluations demonstrate a significant 35% reduction in privacy leakage while maintaining near-baseline utility performance across various downstream tasks: Action Recognition (Kinetics400, UCF101, HMDB51), Temporal Action Detection (THUMOS14), and Anomaly Detection (UCF-Crime). We also provide an analysis on anonymization for sensitive temporal attribute recognition. Additionally, we propose new protocols for assessing gender bias in action recognition models, showing that our method effectively mitigates such biases and promotes more equitable video understanding."},"_bibtex":{"value":"@inproceedings{\nfioresi2026privacy,\ntitle={Privacy Beyond Pixels: Latent Anonymization for Privacy-Preserving Video Understanding},\nauthor={Joseph Fioresi and Ishan Rajendrakumar Dave and Mubarak Shah},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=ncA3UUL0Ri}\n}"},"title":{"value":"Privacy Beyond Pixels: Latent Anonymization for Privacy-Preserving Video Understanding"},"pdf":{"value":"/pdf/6f2d5b7b27685bc8d0446000471f0bfcdbdc01ad.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"fioresi|privacy_beyond_pixels_latent_anonymization_for_privacypreserving_video_understanding"},"authorids":{"value":["~Joseph_Fioresi1","~Ishan_Rajendrakumar_Dave1","~Mubarak_Shah3"]},"authors":{"value":["Joseph Fioresi","Ishan Rajendrakumar Dave","Mubarak Shah"]}},"version":2},{"content":{"summary":{"value":"This paper introduces TVBench, a new benchmark for testing video understanding capability of multimodal models. Flaws in widely used existing benchmark (MVBench) are demonstrated, namely, spatial biases, textual biases and reliance on world knowledge. In addition, it's also shown that open-ended benchmarks can contain similar biases. TVBench is constructed from pre-defined templates in order to mitigate these biases and test temporal reasoning capabilities. Supporting experimental results demonstrate that state-of-the-art models struggle on this benchmark. Similarly, text-only or image-text foundation models struggle to beat random chance signifying the difficulty of this benchmark compared with existing benchmark."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"- Can we use the same model and ablate text-only, image, video, shuffle, reverse, etc. in Table 4? Ideally Gemini 1.5 pro as it performs the best on this benchmark?\n- MVBench is presented as not a great benchmark. However, its performance is also not saturated (best model achieves 67.7 in Table 4). Do the remaining QA pairs satisfy the criteria set in the paper? What is the size of the data? Can we remove the bad examples from MVBench and get a bigger and better dataset than TVBench?\n\nEDIT: Updated score based on the rebuttal."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":3},"strengths":{"value":"The paper is well-written and easy to understand. It tackles an important area of video understanding, i.e. the lack of strong benchmarks that test temporal reasoning in videos. The presentation clearly analyzes drawbacks of existing benchmarks and proposes a new benchmark.\n- The QA pairs don't use LLMs in the loop, and thus can avoid many hallucination related issues.\n- The performance of SOTA models is very low (Table 4). This indicates the benchmark is indeed difficult.\n- Clear contrast with MVBench is demonstrated, especially using text-only and image-only models. This justifies most of the claims in the paper. \n- A significant, and often overlooked issue in open-ended evaluations is pointed out in Section 4. Using closed-source proprietary models whose back-ends may change arbitrarily to score open-ended responses and track our progress on video understanding can be misleading."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The main weakness of this work is around experimentation. \n- Human baseline performance is not presented. This is important to judge the quality of the benchmark and the presented results.\n- Different models are used in Table 2 to make the claim that MVBench has textual bias. Ideally, the same model (ideally the best model) needs to be presented with text-only and video as inputs to justify the claim.\n- Similarly, in Table 4, different models are used to compare different biases (text, image, video) of the model. \n\nFurther Limitations:\n- Using standard template QA pairs may limit the range of video understanding being assessed.\n- In Figure 2 and the associated text in the paper, it's presented as if detecting the absence of something is an easy task. However, by definition, one must watch the entire video to make sure what we're detecting is indeed absent."}},"nonreaders":[],"tmdate":1733274920891,"tcdate":1730700263228,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission11521/Reviewer_TPmi"],"signatures":["ICLR.cc/2025/Conference/Submission11521/Reviewer_TPmi"],"forum":"DrNN5qx66Z","number":4,"license":"CC BY 4.0","cdate":1730700263228,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission11521/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1733274920891,"domain":"ICLR.cc/2025/Conference","replyto":"DrNN5qx66Z","id":"B6qVAoJIL5","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"TLDR":{"value":"We reveal the limitations of existing video-language benchmarks and present TVBench, a benchmark designed to truly assess temporal video-language understanding."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video-Language evaluation","Video-Language benchmark"]},"primary_area":{"value":"datasets and benchmarks"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Large language models have demonstrated impressive performance when integrated with vision models even enabling video understanding. However, evaluating these video models presents its own unique challenges, for which several benchmarks have been proposed. In this paper, we show that the currently most used video-language benchmarks can be solved without requiring much temporal reasoning. We identified three main issues in existing datasets: (i) static information from single frames is often sufficient to solve the tasks (ii) the text of the questions and candidate answers is overly informative, allowing models to answer correctly without relying on any visual input (iii) world knowledge alone can answer many of the questions, making the benchmarks a test of knowledge replication rather than visual reasoning. In addition, we found that open-ended question-answering benchmarks for video understanding suffer from similar issues while the automatic evaluation process with LLMs is unreliable, making it an unsuitable alternative. As a solution, we propose TVBench, a novel open-source video multiple-choice question-answering benchmark, and demonstrate through extensive evaluations that it requires a high level of temporal understanding. Surprisingly, we find that most recent state-of-the-art video-language models perform similarly to random performance on TVBench, with only a few models such as Qwen2-VL, and Tarsier clearly surpassing this baseline."},"_bibtex":{"value":"@misc{\ncores2025tvbench,\ntitle={{TVB}ench: Redesigning Video-Language Evaluation},\nauthor={Daniel Cores and Michael Dorkenwald and Manuel Mucientes and Cees G. M. Snoek and Yuki M Asano},\nyear={2025},\nurl={https://openreview.net/forum?id=DrNN5qx66Z}\n}"},"title":{"value":"TVBench: Redesigning Video-Language Evaluation"},"pdf":{"value":"/pdf/d10e747940b697aa51aeb790b853e347a10f151d.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"cores|tvbench_redesigning_videolanguage_evaluation"},"authorids":{"value":["~Daniel_Cores1","~Michael_Dorkenwald1","~Manuel_Mucientes1","~Cees_G._M._Snoek1","~Yuki_M_Asano1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Daniel Cores","Michael Dorkenwald","Manuel Mucientes","Cees G. M. Snoek","Yuki M Asano"]}},"version":2},{"content":{"venue":{"value":"IEEE Trans. Circuits Syst. Video Technol. 2025"},"pdf":{"value":"https://ieeexplore.ieee.org/iel8/76/10916540/10741539.pdf"},"venueid":{"value":"dblp.org/journals/TCSV/2025"},"paperhash":{"value":"wu|collaborative_aware_bidirectional_semantic_reasoning_for_video_question_answering"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Xize_Wu:","https://dblp.org/search/pid/api?q=author:Jiasong_Wu:","~Lei_Zhu8","~Lotfi_Senhadji1","https://dblp.org/search/pid/api?q=author:Huazhong_Shu:"]},"html":{"value":"https://doi.org/10.1109/TCSVT.2024.3490665"},"_bibtex":{"value":"@article{DBLP:journals/tcsv/WuWZSS25,\n  author={Xize Wu and Jiasong Wu and Lei Zhu and Lotfi Senhadji and Huazhong Shu},\n  title={Collaborative Aware Bidirectional Semantic Reasoning for Video Question Answering},\n  year={2025},\n  month={March},\n  cdate={1740787200000},\n  journal={IEEE Trans. Circuits Syst. Video Technol.},\n  volume={35},\n  number={3},\n  pages={2074-2086},\n  url={https://doi.org/10.1109/TCSVT.2024.3490665}\n}\n"},"abstract":{"value":"Video question answering (VideoQA) is the challenging task of accurately responding to natural language questions based on a given video. Most previous methods focus on designing complex cross-modal interactions to perform question-oriented video scene mining and semantic reasoning, and utilize straightforward classification and matching strategies with different decoders to forcibly associate the predicted representation with ground-truth answer. However, the limitations of question-oriented reasoning and the overlapping semantic co-occurrences between questions and candidates may cause them to fall into spurious correlation reasoning. In this paper, we propose a Collaborative aware Bidirectional Semantic Reasoning (CBSR) model to alleviate this challenging problem. Specifically, we first propose a collaborative aware adaptive correlation reasoning module to collaboratively mine multi-granularity text-aware critical video scenes and reason about the complex intrinsic correlations between them via bottom-up cross-granularity adaptive aggregation. By progressively performing video reasoning from object-level to frame-level, we can obtain a set of semantically rich critical video representations. Then, we collaboratively decode it together with question and knowledge semantics into an implicit representation through the proposed unified answer semantic collaborated decoding module. Finally, a novel bidirectional semantic reasoning learning strategy is proposed to bridge and strengthen the unique positive semantic correlation between the learned implicit representation and the ground-truth answer, and explicitly alleviate the challenge of overlapping semantic co-occurrence. Benefiting from the same model structure and learning strategy, our method can achieve seamless transfer between Open-Ended and Multi-Choice tasks. Extensive experimental results on seven commonly tested datasets (i.e. MSVD-QA, MSRVTT-QA, NExT-QA, Causal-VidQA, NExT-OOD, ActivityNet-QA and EgoSchema) verify the superior performance of our method and the effectiveness of each reasoning module. We provide our source codes and experimental datasets at https://github.com/XizeWu/CBSR."},"title":{"value":"Collaborative Aware Bidirectional Semantic Reasoning for Video Question Answering"},"authors":{"value":["Xize Wu","Jiasong Wu","Lei Zhu","Lotfi Senhadji","Huazhong Shu"]}},"tmdate":1768971745847,"pdate":1735689600000,"externalIds":["dblp:journals/tcsv/WuWZSS25"],"tcdate":1767874824107,"writers":["~"],"signatures":["~Lotfi_Senhadji1"],"forum":"LTRAG1YyUt","license":"CC BY-SA 4.0","number":734858,"cdate":1740787200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1768971745847,"domain":"DBLP.org","id":"LTRAG1YyUt","version":2},{"content":{"summary":{"value":"This paper presents SALSA-V, a shortcut-augmented latent flow matching model for video-to-audio generation that achieves high-fidelity and temporally synchronized audio in just a few sampling steps. By combining masked training for audio conditioning and outpainting with contrastively trained synchronization features, SALSA-V enables both efficient short-form and stable long-form audio generation without distillation."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"**1. Clarification on Feature Roles and Figure References**\n\nIn the Method section, the paper sequentially introduces the VAE encoder, visual-text representation, and synchronization feature.\nHowever, it is unclear how each of these corresponds to the components shown in Figure 1, and what specific roles they play in the generation pipeline. A clearer mapping between the described features and their positions in Figure 1 would be helpful.\nIn addition, Figure 5 is never explicitly referenced in the main text—please consider mentioning and explaining it within the corresponding experimental section.\n\n**2. Inconsistency between Figure 5 and Table 1 Results**\n\nIn Figure 5, the proposed model shows lower performance than MMAudio for 10-second generation, whereas Table 1 indicates a different trend. Could you clarify whether these results are based on different experimental setups or evaluation protocols, and explain what accounts for the discrepancy?\n\n**3. Quantitative Analysis of Sampling Efficiency**\n\nThe paper highlights sampling efficiency as a major advantage of SALSA-V, but does not provide concrete runtime comparisons.\nCould the authors include or discuss actual inference speed measurements (e.g., seconds per 10-second clip, GPU type, batch size), and how they compare to other models?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":1},"contribution":{"value":2},"strengths":{"value":"**1. Enables Long-Form Audio Generation**\n\nThe model supports audio-conditioned outpainting, allowing stable and coherent generation of extended audio sequences (30 s+) from video inputs.\n\n**2. Improved Audio-Visual Synchronization**\n\nA contrastively trained synchronization encoder provides precise temporal alignment between visual motion and audio events, achieving SOTA synchronization performance.\n\n**3. Efficient Few-Step Sampling**\n\nThe shortcut-augmented flow matching formulation reduces sampling steps to ≤ 8 without quality degradation."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"**1. Low Readability and Clarity**\n\nThe paper’s presentation is occasionally hard to follow.\n\n**2. Inconsistency between Quantitative Results (Fig. 5 vs. Table 1)**\n\nThere appears to be a mismatch between the 10s performance trend in Figure 5 and the metrics in Table 1.\n\n**3. Insufficient Analysis of Inference Efficiency**\n\nWhile the paper emphasizes few-step generation, there is no thorough empirical analysis of inference speed. More concrete runtime results would strengthen the efficiency claims."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762922765960,"tcdate":1761652106377,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission11727/Reviewer_7VCU"],"signatures":["ICLR.cc/2026/Conference/Submission11727/Reviewer_7VCU"],"forum":"3FcjKYmNY5","number":2,"license":"CC BY 4.0","cdate":1761652106377,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission11727/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762922765960,"domain":"ICLR.cc/2026/Conference","replyto":"3FcjKYmNY5","id":"93rH6jCBIM","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["video-to-audio","audio generation","diffusion"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"We propose SALSA-V, a multimodal video-to-audio generation model capable of synthesizing highly synchronized, high-fidelity long-form audio from silent video content. Our approach introduces a masked diffusion objective, enabling audio-conditioned generation and the seamless synthesis of audio sequences of unconstrained length. Additionally, by integrating a shortcut loss into our training process, we achieve rapid generation of high-quality audio samples in as few as eight sampling steps, paving the way for near-real-time applications without requiring dedicated fine-tuning or retraining. We demonstrate that SALSA-V significantly outperforms existing state-of-the-art methods in both audiovisual alignment and synchronization with video content in quantiative evaluation and a human listening study. Furthermore, our use of random masking during training enables our model to match spectral characteristics of reference audio samples, broadening its applicability to professional audio synthesis tasks such as Foley generation and sound design."},"_bibtex":{"value":"@misc{\ndellali2026salsav,\ntitle={{SALSA}-V: Shortcut-Augmented Long-form Synchronized Audio from Videos},\nauthor={Amir Dellali and Luca A Lanzend{\\\"o}rfer and Florian Gr{\\\"o}tschla and Roger Wattenhofer},\nyear={2026},\nurl={https://openreview.net/forum?id=3FcjKYmNY5}\n}"},"title":{"value":"SALSA-V: Shortcut-Augmented Long-form Synchronized Audio from Videos"},"pdf":{"value":"/pdf/780501365dd6cdc6dc1c38131015c6ed7be54b52.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"dellali|salsav_shortcutaugmented_longform_synchronized_audio_from_videos"},"authorids":{"value":["~Amir_Dellali1","~Luca_A_Lanzendörfer1","~Florian_Grötschla1","~Roger_Wattenhofer1"]},"authors":{"value":["Amir Dellali","Luca A Lanzendörfer","Florian Grötschla","Roger Wattenhofer"]}},"version":2},{"content":{"summary":{"value":"The paper investigates the potential of general-purpose vision-language models to be directed towards effective biomedical video understanding by training them on publicly available educational content, primarily sourced from YouTube. The authors have constructed a substantial instruction-tuning corpus, named OpenBiomedVid, which comprises approximately 1,031 hours of clinician-guided biomedical videos, meticulously cleaned captions, and GPT-assisted yet human-verified question-answer pairs. Additionally, they introduce two more challenging expert-curated benchmarks—MIMICEchoQA (echo) and SurgeryVideoQA (surgery)—to ensure that evaluations are not conducted on the same noisy distribution as the training data. Fine-tuning Qwen2-VL (2B/7B) using this dataset results in significant relative improvements in biomedical video question answering (QA) and notable advancements in image visual question answering (VQA). In some instances, performance approaches or even surpasses that of stronger general models when applied to echo-style videos. This finding suggests that \"videos made for humans\" can still serve as an effective supervisory signal for medical vision-language models. However, it is important to note that results on the surgical benchmark remain distinctly lower. This indicates that long and heterogeneous procedural videos continue to pose challenges and highlights concerns regarding stylistic inconsistencies and potential hallucination risks introduced by LLM-in-the-loop curation—a concern the authors attempt to address through human verification. In summary, this work presents a well-motivated dataset and benchmark package along with compelling empirical evidence supporting the notion that public educational videos represent a viable resource for domain adaptation of open vision-language models."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"1. You cite several domain-specific multimodal/medical VLMs in Related Work (e.g., LLaVA-Med, Med-Flamingo, MedCLIP), but they do not appear in the main comparison tables. Can you clarify which of these models you actually attempted to run, and what prevented you from including their results?\n\n2. Right now the improvements are mostly against Qwen2-VL/InternVL-style backbones. How can we be sure the gains come from your video-centric data/pipeline rather than simply from using a stronger base model?\n\n3. The evaluation of SurgeryVideoQA was conducted using GPT-4o as the automatic grading system. Given that GPT-4o was also employed for the refinement of captions and question-and-answer pairs, what measures were taken to mitigate potential style bias towards outputs resembling those generated by GPT-4o?\n\n4. You mention using Gemini-2.0-Flash as a second judge and finding similar rankings. Could you provide the agreement numbers between GPT-4o and Gemini on this benchmark?\n\n5. You report ~95% agreement on frame filtering, but on how many samples, with how many annotators, and at which stage of the pipeline was this measured (video-level vs clip-level vs QA-level)?\n\n6. Is human verification applied to every GPT-generated Q/A pair, or only to a sampled subset? If sampled, what was the sampling strategy and coverage?\n\n7. Would it be feasible to run them in a constrained setting (e.g., fixed frame sampling, short 8–16 frame clips) just to position your benchmarks relative to the broader video-VLM ecosystem?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"- Clear Motivation & Well-Defined Problem: The authors identify an important gap: most biomedical video datasets are either small, narrowly focused, or unsuitable for large-scale multimodal learning. They correctly observe that public educational videos are a massive, untapped source for such efforts.\n- Novel Dataset and Benchmark Creation: The curation of a 1031-hour biomedical video-text dataset from public sources is commendable. The pipeline is systematic, leveraging both expert clinical input and multi-stage LLM filtering. The inclusion of 79,367 Q/A pairs and structured metadata adds substantial utility and diversity.\n- Expert-Curated Evaluation: The introduction of the MIMICEchoQA and SurgeryVideoQA benchmarks is a valuable contribution, offering much-needed standardized, clinically relevant tests for biomedical VLMs. The expert review and curation of questions are a particular highlight."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The paper creates a same-model, same-style loop: GPT-4o is used to clean captions and generate the training/eval Q&A and metadata, and then the very same GPT-4o is used as the automatic judge to score open-ended answers (binary 0/1). The authors even note the risk of “stylistic bias” and only add a one-off check with Gemini-2.0-Flash, but GPT-4o remains the primary scorer. This tightly couples data generation and evaluation to one model’s style, making gains plausibly evaluation-protocol-driven rather than true capability. \n\n- Lack of evaluation against domain-specific medical VLMs. The paper compares mainly with general-purpose multimodal backbones (Qwen2-VL at different scales, InternVL3-8B, GPT-4o, Gemini-2.0-Flash, etc.) on the proposed biomedical video QA benchmarks, but does not report results for the most directly relevant medical instruction-tuned VLMs such as LLaVA-Med, Med-Flamingo (medical adaptation of Flamingo), or widely used medical multi-image models (e.g., MedCLIP), even though many of them are cited in Related Work. Because these models target clinical/biomedical vision-language understanding on VQA-RAD, PathVQA, SLAKE and related tasks, including at least a frame-based or image-only adaptation of them on the authors’ benchmarks would make the claimed gains over “prior medical VLMs” easier to judge. The current tables therefore make it hard to tell whether the improvement comes from the proposed video-centric data and pipeline or simply from using stronger general backbones. (We acknowledge that some models, e.g. MedCLIP or MedGemini, are image-centric or not fully open, but even a best-effort comparison or a discussion of feasibility would strengthen the evaluation.)\n- Reliance on LLM-as-a-judge without human grounding. For the open-ended SurgeryVideoQA benchmark, the paper evaluates all models with GPT-4o as the automatic grader, assigning binary correctness against the reference answer. Since GPT-4o was also involved in earlier stages of caption refinement and QA generation, this creates a potential style / phrasing bias toward GPT-4o-like answers. The authors do run a second pass with Gemini-2.0-Flash and report that the ranking is broadly consistent, which is helpful, but both judges are frontier LLMs with unknown overlap in training data and alignment objectives, and no human adjudication, inter-rater agreement, or adversarial stress tests are provided. As a result, part of the reported gains on SurgeryVideoQA could still be influenced by judge-specific preferences rather than purely by better video understanding.\n- Potential residual data-quality and provenance issues. Although the paper presents a multi-stage, human-in-the-loop curation pipeline (YouTube retrieval → GPT-labeled frame filtering with a fine-tuned SigLIP → GPT-4o caption refinement → GPT-4o Q/A generation → human verification) and even reports ~95% agreement on the frame-filtering stage, many of the quality controls are described only at a high level. In particular, the paper does not spell out how deep the human verification went (per-clip vs per-QA, random spot checks vs full passes), how many annotators were involved, or what the inter-annotator agreement was beyond the small sample quoted. Moreover, because GPT-4o is used both to clean captions and to synthesize Q/A pairs, typical LLM failure modes—overly generic answers, temporal misalignment with the actual video segment, or medical hallucination—are plausible but not systematically audited or reported. This matters especially because the paper itself shows that models trained on the noisy YouTube-derived corpus still perform noticeably worse on the cleaner, expert-curated benchmarks (MIMICEchoQA, SurgeryVideoQA), suggesting that remaining noise in the large-scale training split may be a limiting factor. A more explicit error analysis (e.g., failed frame segmentation, ambiguous captions, low-quality or hallucinated Q/A) would make the dataset contribution stronger.\n- Underdeveloped video baselines on the proposed benchmarks. Because the core claim of the paper is about video-centric biomedical understanding, the experimental setup on MIMICEchoQA and SurgeryVideoQA would be more convincing if it included stronger, publicly available video–language models beyond the authors’own Qwen2-VL fine-tunes. At minimum, prior general-purpose video chat models (e.g., Video-ChatGPT, Video-LLaVA, PALIGemma-style video variants) could be run in a frame-sampling or short-clip regime to provide external points of reference, even if they are not domain-tuned. Their absence makes the reported gains look partly relative to the authors’ chosen baselines rather than to the broader video–language landscape, and it also hides whether the new benchmarks are genuinely “hard” for off-the-shelf video models or only for image-first VLMs."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762942037191,"tcdate":1761844922411,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission22059/Reviewer_fKBP"],"signatures":["ICLR.cc/2026/Conference/Submission22059/Reviewer_fKBP"],"forum":"u4PmZOmtko","number":2,"license":"CC BY 4.0","cdate":1761844922411,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission22059/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762942037191,"domain":"ICLR.cc/2026/Conference","replyto":"u4PmZOmtko","id":"QuPCWhrInP","forumContent":{"TLDR":{"value":"Instruction-tuning Qwen-2-VL on 1,031 hours of pedagogical biomedical videos dramatically boosts video and image understanding and includes new expert-curated benchmarks, with all data and code released."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["vision-language models","biomedicine","datasets","evaluations"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Publicly available biomedical videos, such as those on YouTube, serve as valuable educational resources for medical students. Unlike standard machine learning datasets, these videos are designed for human learners, often mixing medical imagery with narration, explanatory diagrams, and contextual framing. In this work, we investigate whether such pedagogically rich, yet non-standardized and heterogeneous videos can effectively teach general-domain vision-language models biomedical knowledge. To this end, we introduce OpenBiomedVid, a biomedical video instruction tuning dataset comprising 1031 hours of video-caption and Q/A pairs, curated through a multi-step human-in-the-loop pipeline. Diverse biomedical video datasets are rare, and OpenBiomedVid fills an important gap by providing instruction-style supervision grounded in real-world educational content. Surprisingly, despite the informal and heterogeneous nature of these videos, the fine-tuned Qwen-2-VL models exhibit substantial performance improvements across most benchmarks. The 2B model achieves gains of 98.7% on video tasks, 71.2% on image tasks, and 0.2% on text tasks. The 7B model shows improvements of 37.09% on video and 11.2% on image tasks, with a slight degradation of 2.7% on text tasks compared to their respective base models. To address the lack of standardized biomedical video evaluation datasets, we also introduce two new expert curated benchmarks, MIMICEchoQA and SurgeryVideoQA. On these benchmarks, the 2B model achieves gains of 99.1% and 98.1%, while the 7B model shows gains of 22.5% and 52.1%, respectively, demonstrating the models' ability to generalize and perform biomedical video understanding on cleaner and more standardized datasets than those seen during training. These results suggest that educational videos created for human learning offer a surprisingly effective training signal for biomedical VLMs. We release OpenBiomedVid, MIMICEchoQA, SurgeryVideoQA, the fine-tuned models, and the complete codebase to support future research."},"_bibtex":{"value":"@misc{\nthapa2026how,\ntitle={How Well Can General Vision-Language Models Learn Medicine By Watching Public Educational Videos?},\nauthor={Rahul Thapa and Andrew Li and Qingyang Wu and Bryan He and Yuki Sahashi and Christina Binder and Angela Zhang and Ben Athiwaratkun and Shuaiwen Leon Song and David Ouyang and James Zou},\nyear={2026},\nurl={https://openreview.net/forum?id=u4PmZOmtko}\n}"},"title":{"value":"How Well Can General Vision-Language Models Learn Medicine By Watching Public Educational Videos?"},"pdf":{"value":"/pdf/fcf643dd9df9746824583d9bb992dfdc21d7b0d9.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"thapa|how_well_can_general_visionlanguage_models_learn_medicine_by_watching_public_educational_videos"},"authorids":{"value":["~Rahul_Thapa1","~Andrew_Li4","~Qingyang_Wu1","~Bryan_He1","~Yuki_Sahashi1","~Christina_Binder1","~Angela_Zhang1","~Ben_Athiwaratkun1","~Shuaiwen_Leon_Song1","~David_Ouyang1","~James_Zou1"]},"authors":{"value":["Rahul Thapa","Andrew Li","Qingyang Wu","Bryan He","Yuki Sahashi","Christina Binder","Angela Zhang","Ben Athiwaratkun","Shuaiwen Leon Song","David Ouyang","James Zou"]}},"version":2},{"content":{"summary":{"value":"This paper presents a nested adaptor and a LLM-driven semantic data generation pipeline to improve video moment retrieval. The pipeline generates semantically similar queries to enrich query diversity, while the nested adapter uses both augmented and human-annotated queries for coarse-tuning and fine-tuning."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"See Weaknesses."},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":1},"strengths":{"value":"1. Using LLM-generated captions to enhance video moment retrieval is reasonable.\n2. Experiments demonstrate some effectiveness."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The writing and figures require improvement.\n2. The contribution is limited, with minimal improvements. Using LLMs to enhance textual semantics has been explored with more substantial gains (e.g. [1-2]). The nested adaptor is also trivial.\n\n   [1]  ChatVTG: Video Temporal Grounding via Chat with Video Dialogue Large Language Models.\n\n   [2]  Context-Enhanced Video Moment Retrieval with Large Language Models.\n\n3. The experiments are insufficient, lacking results on popular datasets like Charades-STA and ActivityNet Captions.\n4. Minor errors, such as \"Figure 1: Figure 1:\"."}},"nonreaders":[],"tmdate":1731429315568,"tcdate":1729949576895,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission9631/Reviewer_PPHd"],"signatures":["ICLR.cc/2025/Conference/Submission9631/Reviewer_PPHd"],"forum":"8Ds99sdp3U","number":1,"license":"CC BY 4.0","cdate":1729949576895,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission9631/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731429315568,"domain":"ICLR.cc/2025/Conference","replyto":"8Ds99sdp3U","id":"rE2I6YydKT","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Moment Retrieval","Highlight Detection","Adapter","Data Augmentation"]},"primary_area":{"value":"unsupervised, self-supervised, semi-supervised, and supervised representation learning"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Existing transformer-based video-moment retrieval models achieve sub-optimal\nperformance when using the pretrain-finetuning learning paradigm – a pretrained\nmultimodal encoder is finetuned using the target training data. While current work\nhas explored different model architectures and training paradigms to explore this\nproblem, the problem of data dilemma has been under addressed. Specifically,\nthere exists high diversity of how semantic is captured in textual query and the\ntraining dataset only consist of limited moment-query pairs for the highly diverse moments. This work addresses this problem with a novel nested adaptor\nand a LLM-driven semantic data generation pipeline. First, a LLM-driven data\naugmentation generates queries that are semantically similar to the ground truth,\nwhich enrich the semantic boundary captured by textual query. We empirically\nanalyze the effectiveness of data augmentation, and proposed a simple yet effective quality measure to retain high quality samples. Second, we propose a novel\nnested adapter that utilises both augmented queries and human annotated queries\nfor model coarse-tuning and fine-tuning, respectively. By combining semantic\nperturbation with domain adaptation, our approach addresses the variability in\nvideo content while capturing nuanced features more effectively. Experimental\nresults on various baseline models show the efficacy of our proposed approach."},"_bibtex":{"value":"@misc{\nbhandari2024a,\ntitle={A Semantic Data Augmentation driven Nested Adapter for Video Moment Retrieval},\nauthor={Arkaprabha Bhandari and KAJAL KANSAL and Yongkang Wong and Jianquan Liu and Mohan Kankanhalli},\nyear={2024},\nurl={https://openreview.net/forum?id=8Ds99sdp3U}\n}"},"title":{"value":"A Semantic Data Augmentation driven Nested Adapter for Video Moment Retrieval"},"pdf":{"value":"/pdf/304e4d94e9fa018116dae27e89f956d0af08ec41.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"bhandari|a_semantic_data_augmentation_driven_nested_adapter_for_video_moment_retrieval"},"authorids":{"value":["~Arkaprabha_Bhandari1","~KAJAL_KANSAL1","~Yongkang_Wong1","~Jianquan_Liu1","~Mohan_Kankanhalli1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Arkaprabha Bhandari","KAJAL KANSAL","Yongkang Wong","Jianquan Liu","Mohan Kankanhalli"]}},"version":2},{"content":{"summary":{"value":"This paper introduces InfiniBench, an innovative and comprehensive benchmark focused on evaluating large multimodal models' performance in understanding very long videos. InfiniBench is notable for its ultra-long video duration (averaging 52.59 minutes per video) and massive question-answer pairs (108.2K), covering nine different skills including multiple-choice and open-ended questions. These questions are designed to be both diverse and human-centric, with videos primarily sourced from movies and TV shows. Experimental results show that even leading AI models like GPT-4V and Gemini 1.5 Flash face significant challenges in long video understanding, achieving average accuracies of only 49.16% and 42.72%, with mean scores of 3.22 and 2.71 (out of 5) respectively. This indicates that while these models perform relatively well on local skills, they still have limitations in skills requiring global reasoning and deep contextual understanding, such as scene transitions and movie spoiler questions. Open-source models generally perform below random chance on multiple-choice questions, highlighting long-sequence global reasoning as a major challenge for existing models. Additionally, models relying on both video and text information perform poorly without caption input, emphasizing the importance of processing both visual and textual information for long video understanding. The introduction of InfiniBench aims to fill the gap in long video understanding benchmarks, drive the development of open-source large language models, and motivate multimodal large models toward more human-like long video understanding and reasoning capabilities, despite current limitations such as video source restrictions and script dependency."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1. How is GPT-4V's scoring aligned with human evaluation?\n2. Why weren't the latest models tested, and why wasn't there comparison and discussion of the latest benchmarks?"},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":1},"strengths":{"value":"1. InfiniBench provides a comprehensive evaluation of large multimodal models' capabilities in long video understanding through including the longest video duration and a large number of question-answer pairs, as well as designing diverse question types (multiple-choice and open-ended questions) covering nine different skills, thus thoroughly examining models' performance across multiple dimensions of long video understanding.\n\n2. By evaluating various models including both commercial and open-source models, InfiniBench reveals the challenges and limitations of existing models in long video understanding, especially in tasks requiring deep contextual understanding and critical thinking. This in-depth assessment helps identify model deficiencies and provides clear directions for future research and model improvements.\n\n3. InfiniBench's design not only tests models' technical capabilities but also drives models toward more human-like understanding and reasoning abilities. Through proposing human-centric questions, such as movie spoiler questions, it promotes model performance improvement in long video understanding tasks, which is significant for achieving more advanced AI applications and advancing the field of artificial intelligence."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The benchmark only uses movies and TV shows for testing, which is too limited. It should include more types of videos that show different parts of real life, like nature documentaries or home videos. The problem is that movies and TV shows follow certain storytelling patterns, so AI models might just learn these patterns instead of truly understanding the videos. They should add more casual videos like vlogs and livestreams to make the testing more realistic.\n\n2. The benchmark needs written scripts to create its questions and answers. This is a big problem because most real-world videos don't come with scripts. Without scripts or captions, the benchmark can't test how well AI models understand regular videos that people actually watch and share online.\n\n3. InfiniBench's testing does not cover current mainstream open-source models such as Qwen2VL, LLaVA-Onevision, and InternVL2. This makes it difficult to obtain a more comprehensive and in-depth comparison between open-source and closed-source models.\n\n4. In Table 1, the benchmark comparison is insufficient, especially regarding some recent video benchmarks such as Video-MME and LongVideoBench. Additionally, the authors' definition of \"very long\" is problematic - MLVU and MovieChat have only a 3-minute gap, yet MLVU is defined as very long. This is not reasonable."}},"nonreaders":[],"tmdate":1731427565473,"tcdate":1730613870276,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission2250/Reviewer_147D"],"signatures":["ICLR.cc/2025/Conference/Submission2250/Reviewer_147D"],"forum":"2D0uXQbntW","number":4,"license":"CC BY 4.0","cdate":1730613870276,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission2250/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427565473,"domain":"ICLR.cc/2025/Conference","replyto":"2D0uXQbntW","id":"APDAwAVRTS","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["video understanding","benchmark","long video benchmark","long video understanding"]},"supplementary_material":{"value":"/attachment/e4286990af860ccb3f16a8fce15b1da8635ef0b5.zip"},"primary_area":{"value":"datasets and benchmarks"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Understanding long videos, ranging from tens of minutes to several hours, presents unique challenges in video comprehension. Despite the increasing importance of long-form video content, existing benchmarks primarily focus on shorter clips. To address this gap, we introduce InfiniBench a comprehensive benchmark for very long video understanding which presents 1)very long video duration, averaging 52.59 minutes per video 2)The largest number of question-answer pairs, 108.2K 3) Diversity in questions that examine nine different skills and include both multiple-choice questions and open-ended questions 4) Memory questions, such as Global Appearance that require remembering and tracking the visual aspects through the video. Using InfiniBench, we comprehensively evaluate existing Large Multi-Modality Models (LMMs) on each skill, including the commercial models such as GPT-4o and Gemini 1.5 Flash and the recent open-source models. \nThe evaluation shows significant challenges in our benchmark.\nOur findings reveal that even leading AI models like GPT-4o and Gemini 1.5 Flash face challenges in achieving high performance in long video understanding, with average accuracies of just 56.01 % and 43.32 %, and average scores of 3.25 and 2.79 out of 5, respectively.\nQwen2-VL matches Gemini's performance in the MCQ skills but lags significantly in open-ended question tasks.\nWe hope this benchmark will stimulate the LMMs community towards long video and human-level understanding."},"_bibtex":{"value":"@misc{\nataallah2025infinibench,\ntitle={InfiniBench: A Comprehensive Benchmark for Large Multimodal Models in Very Long Video Understanding},\nauthor={Kirolos Ataallah and Chenhui Gou and Eslam Mohamed BAKR and Khushbu Pahwa and Jian Ding and Mohamed Elhoseiny},\nyear={2025},\nurl={https://openreview.net/forum?id=2D0uXQbntW}\n}"},"title":{"value":"InfiniBench: A Comprehensive Benchmark for Large Multimodal Models in Very Long Video Understanding"},"pdf":{"value":"/pdf/032fb6c70b723929924e625fabf7cebb80b48405.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"ataallah|infinibench_a_comprehensive_benchmark_for_large_multimodal_models_in_very_long_video_understanding"},"authorids":{"value":["~Kirolos_Ataallah1","~Chenhui_Gou1","~Eslam_Mohamed_BAKR1","~Khushbu_Pahwa1","~Jian_Ding3","~Mohamed_Elhoseiny1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Kirolos Ataallah","Chenhui Gou","Eslam Mohamed BAKR","Khushbu Pahwa","Jian Ding","Mohamed Elhoseiny"]}},"version":2},{"content":{"summary":{"value":"This paper presents Adversarial Object Hallucination (AOH), a novel white-box attack that compels Video-Language Models (Vid-LLMs) to perceive and describe non-existent objects in videos via intermediate feature alignment. Rather than perturbing raw pixels or final outputs, AOH manipulates the Connector features of Vid-LLMs by maximizing cosine similarity between the adversarial and target video representations. The authors also curate a benchmark of 535 clean/target video pairs with VQA annotations to systematically evaluate object-hallucination robustness. Experiments across multiple Vid-LLMs (VideoLLaMA3, LLaVA-OneVision, InternVL 2.5, Video-ChatGPT) show AOH induces strong and transferable hallucinations, particularly from smaller to larger model scales, while maintaining imperceptible perturbations. Grad-CAM analyses indicate the attacks remain visually covert."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"- Could the authors extend evaluations to black-box or transfer-based settings to demonstrate broader applicability?\n- Could the author provide some comparison with the suggested related works? It would be good to discuss them in the related works sections as well.\n- Why is selecting Connector features particularly effective for inducing object hallucination? Some theoretical justification or ablation study would strengthen the argument."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"- This work introduces an intermediate feature alignment strategy that injects target semantics into Connector embeddings rather than logits, offering an interesting and underexplored direction for adversarial research in Vid-LLMs.\n- Curates 535 clean/target video pairs by integrating HQVI, ROVI, Video-Sham, and BDD100K sources, with rigorous human filtering and LLM-verified VQA annotations. This benchmark enables quantitative evaluation of hallucination attacks, a contribution not seen in prior work.\n- Demonstrates consistent attack success across eight Vid-LLMs (0.5B–8B parameters). Grad-CAM visualizations indicate that model attention remains focused on natural regions, suggesting stealthy internal manipulation by the proposed method."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The study evaluates only under a white-box threat model. The black-box practicality and real-world feasibility remain untested, limiting external validity.\n- Limited comparison with related works, for example, CAVALRY-V [1], and its baselines [2–3].\n- Benchmarks focus solely on VQA; no experiments are conducted on open-ended captioning, reasoning, or safety-critical tasks where hallucinations could be particularly harmful.\n- While effective, cosine-similarity-based feature alignment is well established. The key novelty lies in applying it to Connector features, but no theoretical analysis or empirical evidence is provided to justify why this layer choice is optimal.\n- The paper identifies vulnerabilities but offers no mitigation strategies, defense baselines, or robustness analyses.\n\n[1] Zhang, J., Hu, R., Guo, Q., & Lim, W. Y. B. (2025). CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs. arXiv preprint arXiv:2507.00817.\\\n[2] Huang, H., Erfani, S. M., Li, Y., Ma, X., & Bailey, J. (2025). X-Transfer attacks: Towards super transferable adversarial attacks on CLIP. In Proceedings of the 42nd International Conference on Machine Learning (Vol. 267, pp. 25204–25234).\\\n[3] Zhang, J., Ye, J., Ma, X., Li, Y., Yang, Y., Chen, Y., ... & Yeung, D. Y. (2025). AnyAttack: Towards Large-scale Self-supervised Adversarial Attacks on Vision-language Models. In Proceedings of the Computer Vision and Pattern Recognition Conference (pp. 19900-19909)."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762931565496,"tcdate":1761698819874,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission19729/Reviewer_xvVX"],"signatures":["ICLR.cc/2026/Conference/Submission19729/Reviewer_xvVX"],"forum":"cdhk58Z761","number":2,"license":"CC BY 4.0","cdate":1761698819874,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission19729/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762931565496,"domain":"ICLR.cc/2026/Conference","replyto":"cdhk58Z761","id":"51TALlTqY3","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Video Large Language Models","Adversarial Attack","Object Hallucination"]},"primary_area":{"value":"foundation or frontier models, including LLMs"},"abstract":{"value":"Video Large Language Models (Vid-LLMs) have rapidly advanced video understanding, yet their robustness against semantic adversarial manipulation, especially object hallucination, remains largely unexplored. We introduce Adversarial Object Hallucination (AOH), a novel attack that compels Vid-LLMs to ``see\" non-existent objects in videos by injecting visually imperceptible perturbations. Unlike prior attacks limited to inputs or outputs of videos, AOH directly manipulates intermediate connector features, aligning them with representations from a target video to induce controllable hallucinations. To systematically assess this threat, we curate a benchmark of 535 clean/target video pairs with high-quality VQA annotations. Extensive experiments show that AOH poses a severe threat to state-of-the-art Vid-LLMs, achieving highly effective attacks with alarming \\emph{cross-scale transferability}: adversarial examples optimized on smaller models transfer even more strongly to larger counterparts of the same architecture, amplifying attack impact while reducing adversarial cost. Further analyses reveal that perturbations encode semantic object contours, while Grad-CAM highlights their covert influence. These findings expose a severe and previously overlooked vulnerability in Vid-LLMs, raising urgent concerns about their secure deployment and providing a foundation for future adversarial research in video-language modeling."},"_bibtex":{"value":"@misc{\nzhang2025adversarial,\ntitle={Adversarial Object Hallucination Attacks in Video-Language Models via Intermediate Feature Alignment},\nauthor={Lu Zhang and Liang Zeng},\nyear={2025},\nurl={https://openreview.net/forum?id=cdhk58Z761}\n}"},"title":{"value":"Adversarial Object Hallucination Attacks in Video-Language Models via Intermediate Feature Alignment"},"pdf":{"value":"/pdf/504513cce9144d84fafcf5759b866e7e55223945.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"zhang|adversarial_object_hallucination_attacks_in_videolanguage_models_via_intermediate_feature_alignment"},"authorids":{"value":["~Lu_Zhang19","~Liang_Zeng1"]},"authors":{"value":["Lu Zhang","Liang Zeng"]}},"version":2},{"content":{"venue":{"value":"ICLR 2024 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["graph rationalization","shortcut learning"]},"primary_area":{"value":"representation learning for computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"The remarkable success in graph neural networks (GNNs) promotes the Graph Rationalization methods that aim to provide explanations to support the prediction results by identifying a small subset of the original graph (i.e., rationale). Although existing methods have achieved promising results, recent studies have proved that these methods still suffer from  exploiting shortcuts in the data to yield task results and compose rationales. Different from previous methods plagued by shortcuts, in this paper, we propose a Shortcut-guided Graph Rationalization (SGR) method, which identifies rationales by learning from shortcuts. Specifically, SGR consists of two training stages. In the first stage, we train a shortcut guider with an early stop strategy to obtain shortcut information. During the second stage, SGR separates the graph into the rationale and non-rationale subgraphs and lets them learn from the shortcut information generated by the frozen shortcut guider to identify which information belongs to shortcuts and which does not. Finally, we employ the non-rationale subgraphs as environments and identify the invariant rationales which filter out the shortcuts under environment shifts. Extensive experimental results on both synthetic and real-world datasets clearly validate the effectiveness of our proposed method."},"_bibtex":{"value":"@misc{\nyue2024learning,\ntitle={Learning from Shortcut: A Shortcut-guided Approach for Graph Rationalization},\nauthor={Linan Yue and Qi Liu and Ye Liu and Weibo Gao and Chao Song},\nyear={2024},\nurl={https://openreview.net/forum?id=XcwHDoKvVg}\n}"},"title":{"value":"Learning from Shortcut: A Shortcut-guided Approach for Graph Rationalization"},"pdf":{"value":"/pdf/40f49261208b6613b6cba47b97c7d708165772cf.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Withdrawn_Submission"},"paperhash":{"value":"yue|learning_from_shortcut_a_shortcutguided_approach_for_graph_rationalization"},"authorids":{"value":["~Linan_Yue1","~Qi_Liu3","~Ye_Liu10","~Weibo_Gao1","~Chao_Song2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Linan Yue","Qi Liu","Ye Liu","Weibo Gao","Chao Song"]}},"tmdate":1711385471329,"tcdate":1695448601738,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission6956/Authors"],"signatures":["ICLR.cc/2024/Conference/Submission6956/Authors"],"forum":"XcwHDoKvVg","number":6956,"cdate":1695448601738,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/-/Submission","ICLR.cc/2024/Conference/-/Post_Submission","ICLR.cc/2024/Conference/Submission6956/-/Revision","ICLR.cc/2024/Conference/-/Withdrawn_Submission","ICLR.cc/2024/Conference/-/Edit"],"mdate":1711385471329,"odate":1697213872796,"domain":"ICLR.cc/2024/Conference","id":"XcwHDoKvVg","version":2},{"content":{"summary":{"value":"The paper proposes a dual mutual information framework to disentangle embodiement representation and task representation. This disentangled representation learning is then used to edit human demonstration video into a robot demonstration video. Specifically, the two representations are encouraged to be maximally informative of each other within the modality, whereas they are encouraged to have minimal information across the modality. The authors used two different mutual information (MI) estimators. For intramodal MI maximization, InfoNCE estimator is used. For the intermodal MI bottlenecking, CLUB estimator is used. It should be noted that the this cross-embodiment generalization is an emergent result of this disentanglement, not a result of explicit supervised training. The training itself is done in an autoencoding manner without ground-truth cross-embodiment pair."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"Please find the questions in the weakness section."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"The proposed approach is well motivated and the information-theoretic approach is technically sound. While the idea dual contrastive learning is not new in the so-called multiview representation learning field, its adapation to embodiment-transferrable video robot policy is an important and timely contribution."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"While the methodology itself makes a lot of sense, empirical results are insufficient to support its advantage over baseline methods or applicability to wide variety of videos and embodiments. Specifically, I have two major concerns.\n\n### **[Concern 1]** Metrics used measure video quality, not the transferrability.\nI understand that it is incredibly difficult to objectively measure the performance for this kind of problem. That said, the evaluation protocols seems inappropriate or hard to understand. \n\nFor instance, the metrics in the Video Fidelity category requires target video sample (PSNR, LPIPS, SSIM) or distribution (FVD). However, there cannot be a ground truth target sample/distribution for this kind of cross-embodiment editing. As such, I believe the evaluation is done in an autoencoding fashion rather than with the unavailable GT cross-embodiment target. If the former, it contradicts the claim that the general purpose VACE performs better in-distribution but does not transfer to cross-embodiment editing. If the latter, it is important to explain how the GT pairs are obtained.\n\nThe VBench result is also counterintuitive. If the cross-embodiment transferrability is the key, then the metrics like Subject Consistency (SC) should be higher than VACE. However, the result is the opposite. The proposed method instead excels at more detailed consistency like temporal flickering or motion smoothness. By only looking at these results, I would be more inclined to believe that VACE actually transfers better. It would be helpful if the authors can provide video generated by VACE that corresponds to the one shown in Figure 3.\n\n### **[Concern 2]** Diversity of evaluation settings is limited.\nThe provided qualitative results (supp. video and figures) are limited in diversity in terms of embodiment and background. These qualitative results are only shown for a single robot and a background. It is also unclear how diverse the quantitative evaludation settings are. Necessary details such as the number of embodiments, backgrounds, and tasks are missing.\n\n\nAs the methodology itself makes sense and the figures/videos looks convincing (although they are limited in diversity), I am willing to adjust my rating to more positive ones if these concerns are addressed."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762921291655,"tcdate":1761853706798,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission9805/Reviewer_7n3w"],"signatures":["ICLR.cc/2026/Conference/Submission9805/Reviewer_7n3w"],"forum":"hGcb46DWQD","number":2,"license":"CC BY 4.0","cdate":1761853706798,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission9805/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762921291655,"domain":"ICLR.cc/2026/Conference","replyto":"hGcb46DWQD","id":"wFu8aiSjUR","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"TLDR":{"value":"We achieve egocentric cross-embodiment video editing by using a dual contrastive objective to disentangle a human demonstration and generate a coherent, robot-centric video."},"keywords":{"value":["video generation","diffusion transformer","egocentric videos","cross-embodiment gap","contrastive learning"]},"supplementary_material":{"value":"/attachment/5f821f0dec5ca9c147549137c8588a2a9f5d8324.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Learning robotic manipulation from human videos is a promising solution to the data bottleneck in robotics, but the distribution shift between humans and robots remains a critical challenge. Existing approaches often produce entangled representations, where task-relevant information is coupled with human-specific kinematics, limiting their adaptability. We propose a generative framework for cross-embodiment video editing that directly addresses this by learning explicitly disentangled task and embodiment representations. Our method factorizes a demonstration video into two orthogonal latent spaces by enforcing a dual contrastive objective: it minimizes mutual information between the spaces to ensure independence while maximizing intra-space consistency to create stable representations. A parameter-efficient adapter injects these latent codes into a frozen video diffusion model, enabling the synthesis of a coherent robot execution video from a single human demonstration, without requiring paired cross-embodiment data. Experiments show our approach generates temporally consistent and morphologically accurate robot demonstrations, offering a scalable solution to leverage internet-scale human video for robot learning."},"_bibtex":{"value":"@misc{\nli2026egocentric,\ntitle={Egocentric Cross-Embodiment Video Editing via Dual Contrastive Representation Learning},\nauthor={Zhiyuan Li and Wenyan Yang and Wenshuai Zhao and Yue Ma and Yuanpeng Tu and Pekka Marttinen and Joni Pajarinen},\nyear={2026},\nurl={https://openreview.net/forum?id=hGcb46DWQD}\n}"},"title":{"value":"Egocentric Cross-Embodiment Video Editing via Dual Contrastive Representation Learning"},"pdf":{"value":"/pdf/64d034ce4330f0c5e658b1527cd1dc6b83be9a83.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"li|egocentric_crossembodiment_video_editing_via_dual_contrastive_representation_learning"},"authorids":{"value":["~Zhiyuan_Li9","~Wenyan_Yang1","~Wenshuai_Zhao1","~Yue_Ma2","~Yuanpeng_Tu1","~Pekka_Marttinen1","~Joni_Pajarinen2"]},"authors":{"value":["Zhiyuan Li","Wenyan Yang","Wenshuai Zhao","Yue Ma","Yuanpeng Tu","Pekka Marttinen","Joni Pajarinen"]}},"version":2},{"content":{"venue":{"value":"CoRR 2017"},"pdf":{"value":"https://arxiv.org/pdf/1701.00576v1"},"venueid":{"value":"dblp.org/journals/CORR/2017"},"paperhash":{"value":"wu|shortcut_sequence_tagging"},"authorids":{"value":["~Huijia_Wu1","",""]},"html":{"value":"http://arxiv.org/abs/1701.00576"},"_bibtex":{"value":"@article{DBLP:journals/corr/WuZZ17,\n  publtype={informal},\n  author={Huijia Wu and Jiajun Zhang and Chengqing Zong},\n  title={Shortcut Sequence Tagging},\n  year={2017},\n  cdate={1483228800000},\n  journal={CoRR},\n  volume={abs/1701.00576},\n  url={http://arxiv.org/abs/1701.00576}\n}\n"},"abstract":{"value":"Deep stacked RNNs are usually hard to train. Adding shortcut connections across different layers is a common way to ease the training of stacked networks. However, extra shortcuts make the recurrent step more complicated. To simply the stacked architecture, we propose a framework called shortcut block, which is a marriage of the gating mechanism and shortcuts, while discarding the self-connected part in LSTM cell. We present extensive empirical experiments showing that this design makes training easy and improves generalization. We propose various shortcut block topologies and compositions to explore its effectiveness. Based on this architecture, we obtain a 6% relatively improvement over the state-of-the-art on CCGbank supertagging dataset. We also get comparable results on POS tagging task."},"title":{"value":"Shortcut Sequence Tagging"},"authors":{"value":["Huijia Wu","Jiajun Zhang","Chengqing Zong"]}},"tmdate":1773322973161,"pdate":1514678400000,"externalIds":["dblp:journals/corr/WuZZ17"],"tcdate":1773322963743,"writers":["~"],"signatures":["~Huijia_Wu1"],"forum":"ZgumnUzsUm","license":"CC BY-SA 4.0","number":851095,"cdate":1483228800000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1773322973161,"domain":"DBLP.org","id":"ZgumnUzsUm","version":2},{"content":{"summary":{"value":"This paper proposes WorldSense, a benchmark designed for omni-modal video understanding. Specifically, to answer questions in WorldSense, an MLLM’s response must rely on both video and audio information. Experiments on the proposed WorldSense benchmark reveal the limitations of current MLLMs in omni-modal reasoning."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. #82–83: There’s a typo ,“THe”.\n2. #129–130: Pay attention to the spacing between the image title and the main text. It currently looks a bit confusing."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"1. Requiring both video and audio modalities for accurate responses to each question in WorldSense facilitates a more comprehensive evaluation of current MLLMs in omni-modal reasoning.\n2. Experimentally, the performance drop of current video-audio MLLMs indicates that the fusion between modalities is ineffective or even detrimental."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. While WorldSense emphasizes real-world omni-modal perception, understanding, and reasoning, the benchmark primarily consists of QA pairs. I believe that interactive question answering would be more practical for real-world scenarios. Moreover, isn’t the term \"omni-modal\" somewhat overstated, given that the benchmark only includes video, audio, and text modalities?\n2. Lack of analysis on why the fusion of open-source audio and video models failed. While I wouldn’t tend to reject the paper for this reason, providing such an analysis would offer the community deeper insights than merely presenting the conclusion."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762915850032,"tcdate":1761979962645,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission1665/Reviewer_zBjg"],"signatures":["ICLR.cc/2026/Conference/Submission1665/Reviewer_zBjg"],"forum":"YxsfxAvJv4","number":2,"license":"CC BY 4.0","cdate":1761979962645,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission1665/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762915850032,"domain":"ICLR.cc/2026/Conference","replyto":"YxsfxAvJv4","id":"lZgMmlJ2Qz","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"TLDR":{"value":"We introduce WorldSense, the first benchmark to assess models' omni-modal understanding ability."},"keywords":{"value":["OmniModality","Multimodal LLMs","Benchmark","Real-World Understanding"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"We introduce WorldSense, the first benchmark to assess the multi-modal video understanding, that simultaneously encompasses visual, audio, and text inputs. In contrast to existing benchmarks, our WorldSense has several features: (i) collaboration of omni-modality, we design the evaluation tasks to feature a strong coupling of audio and video, requiring models to effectively utilize the synergistic perception of omni-modality; (ii) diversity of videos and tasks, WorldSense encompasses a diverse collection of 1,662 audio-visual synchronised videos, systematically categorized into 8 primary domains and 67 fine-grained subcategories to cover the broad scenarios, and 3,172 multi-choice QA pairs across 26 distinct tasks to enable the comprehensive evaluation; (iii) high-quality annotations, all the QA pairs are manually labeled by 80 expert annotators with multiple rounds of correction to ensure quality. Based on our WorldSense, we extensively evaluate various state-of-the-art models. The experimental results indicate that existing models face significant challenges in understanding real-world scenarios (65.1% best accuracy). By analyzing the limitations of current models, we aim to provide valuable insight to guide development of real-world understanding. We hope our WorldSense can provide a platform for evaluating the ability in constructing and understanding coherent contexts from omni-modality."},"_bibtex":{"value":"@inproceedings{\nhong2026worldsense,\ntitle={WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal {LLM}s},\nauthor={Jack Hong and Shilin Yan and Jiayin Cai and Xiaolong Jiang and Yao Hu and Weidi Xie},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=YxsfxAvJv4}\n}"},"title":{"value":"WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs"},"pdf":{"value":"/pdf/6717503caabf743a8006bc51dc8aec712e400345.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"hong|worldsense_evaluating_realworld_omnimodal_understanding_for_multimodal_llms"},"authorids":{"value":["~Jack_Hong2","~Shilin_Yan1","~Jiayin_Cai1","~Xiaolong_Jiang2","~Yao_Hu1","~Weidi_Xie3"]},"authors":{"value":["Jack Hong","Shilin Yan","Jiayin Cai","Xiaolong Jiang","Yao Hu","Weidi Xie"]}},"version":2},{"content":{"venue":{"value":"CVPR 2026"},"abstract":{"value":"Referring video object segmentation (RVOS) aims to segment objects in a video described by a natural language expression. However, most existing approaches focus only on the referred object (typically the actor), even when the expression clearly describes an interaction involving multiple objects with distinct roles. In this paper, we introduce Interaction-Aware Referring Video Object Segmentation (InterRVOS), a novel task that focuses on explicit interaction modeling by requiring separate segmentation of actor and target objects.This formulation enables fine-grained understanding of object relationships, as many video events are defined by such interactions rather than individual objects. We present InterRVOS-127K, a large-scale dataset of over 127K automatically annotated expressions with distinct actor-target mask pairs, and propose ReVIOSa, a MLLM-based architecture that introduces interaction-aware special tokens and attention mask loss (AML) to enhance interaction-aware segmentation. We also propose a new evaluation protocol that separately evaluates actor and target segmentation for more accurate role distinction. Comprehensive experiments demonstrate that ReVIOSa outperforms existing baselines on the proposed InterRVOS-127K benchmark, with further analyses validating the necessity and effectiveness of both ReVIOSa and InterRVOS-127K."},"_bibtex":{"value":"@inproceedings{\njin2026interrvos,\ntitle={Inter{RVOS}: Interaction-Aware Referring Video Object Segmentation},\nauthor={Woojeong Jin and Seongchan Kim and Jaeho Lee and Seungryong Kim},\nbooktitle={Conference on Computer Vision and Pattern Recognition 2026},\nyear={2026},\nurl={https://openreview.net/forum?id=FWyiqGtKbr}\n}"},"title":{"value":"InterRVOS: Interaction-Aware Referring Video Object Segmentation"},"pdf":{"value":"https://openaccess.thecvf.com/content/CVPR2026/papers/Jin_InterRVOS_Interaction-Aware_Referring_Video_Object_Segmentation_CVPR_2026_paper.pdf"},"venueid":{"value":"thecvf.com/CVPR/2026/Conference"},"paperhash":{"value":"jin|interrvos_interactionaware_referring_video_object_segmentation"},"authorids":{"value":["~Woojeong_Jin2","~Seongchan_Kim2","~Jae_Ho_Lee1","~Seungryong_Kim1"]},"authors":{"value":["Woojeong Jin","Seongchan Kim","Jaeho Lee","Seungryong Kim"]}},"tmdate":1789656516844,"pdate":1789656437446,"tcdate":1765221986289,"writers":["thecvf.com/CVPR/2026/Conference","thecvf.com/CVPR/2026/Conference/Submission38339/Authors"],"signatures":["thecvf.com/CVPR/2026/Conference/Submission38339/Authors"],"forum":"FWyiqGtKbr","license":"CC BY 4.0","number":38339,"cdate":1765221986289,"readers":["everyone"],"invitations":["thecvf.com/CVPR/2026/Conference/-/Submission","thecvf.com/CVPR/2026/Conference/Submission38339/-/Full_Submission","thecvf.com/CVPR/2026/Conference/-/Post_Submission","thecvf.com/CVPR/2026/Conference/Submission38339/-/Supplementary_Material","thecvf.com/CVPR/2026/Conference/-/Edit","thecvf.com/CVPR/2026/Conference/-/Compute_Flag"],"mdate":1789656516844,"odate":1789656437446,"domain":"thecvf.com/CVPR/2026/Conference","id":"FWyiqGtKbr","version":2},{"content":{"venue":{"value":"IEEE Trans. Circuits Syst. Video Technol. 2023"},"pdf":{"value":"https://ieeexplore.ieee.org/iel7/76/10241245/10056330.pdf"},"venueid":{"value":"dblp.org/journals/TCSV/2023"},"paperhash":{"value":"sun|video_moment_retrieval_via_comprehensive_relationaware_network"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Xin_Sun:","~Jialin_Gao1","","https://dblp.org/search/pid/api?q=author:Xuan_Wang:","https://dblp.org/search/pid/api?q=author:Xi_Zhou:"]},"html":{"value":"https://doi.org/10.1109/TCSVT.2023.3250518"},"_bibtex":{"value":"@article{DBLP:journals/tcsv/SunGZWZ23,\n  author={Xin Sun and Jialin Gao and Yizhe Zhu and Xuan Wang and Xi Zhou},\n  title={Video Moment Retrieval via Comprehensive Relation-Aware Network},\n  year={2023},\n  month={September},\n  cdate={1693526400000},\n  journal={IEEE Trans. Circuits Syst. Video Technol.},\n  volume={33},\n  number={9},\n  pages={5281-5295},\n  url={https://doi.org/10.1109/TCSVT.2023.3250518}\n}\n"},"abstract":{"value":"Video moment retrieval aims to retrieve a target moment from an untrimmed video that semantically corresponds to the given language query. Existing methods commonly treat it as a regression task or a ranking task from the perspective of computer vision. Most of these works neglect comprehensive relations between video content and language context at a multi-granularity level and fail to efficiently model temporal relations among different video moments. In this paper, we formulate video moment retrieval into video reading comprehension by treating the input video as a text passage and language query as a question. To tackle the above impediments, we propose a Comprehensive Relation-aware Network (CRNet) to perceive comprehensive relations from extensive aspects. Specifically, we unite visual and textual features simultaneously at both clip-level and moment-level to thoroughly exploit inter-modality information, leading to a coarse-and-fine cross-modal interaction. Moreover, a background suppression module is introduced to restrain irrelevant background clips, meanwhile, a novel IoU attention mechanism and graph attention layer are efficiently devised to focus on the dependencies among highly-correlated video moments for the best choice selection. In-depth experiments on three public datasets TACoS, ActivityNet Captions, and Charades-STA demonstrate the superiority of our solution."},"title":{"value":"Video Moment Retrieval via Comprehensive Relation-Aware Network"},"authors":{"value":["Xin Sun","Jialin Gao","Yizhe Zhu","Xuan Wang","Xi Zhou"]}},"tmdate":1774394057967,"pdate":1672531200000,"externalIds":["dblp:journals/tcsv/SunGZWZ23"],"tcdate":1762417671691,"writers":["~"],"signatures":["~Jialin_Gao1"],"forum":"hXoGhdlzz5","license":"CC BY-SA 4.0","number":664945,"cdate":1693526400000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1774394057967,"domain":"DBLP.org","id":"hXoGhdlzz5","version":2},{"content":{"summary":{"value":"This work investigates potential shortcuts in existing ToM datasets, and train LMs via RL on the shortcut-free datasets.\n\nThe main findings are as follows:\n\n(1) several datasets are prone to shortcut exploitation and are therefore unsuitable for training;\n\n(2) when trained on shortcut-free datasets, the performance follows the order thinking-RFT > no-thinking-RFT > SFT;\n\n(3) thinking-RFT exhibits strong generalization to unseen domains and higher-order ToM questions; and\n\n(4) RFT improves performance by learning to ground its reasoning in cues that correspond to causal factors.\n\nOverall, the paper attempts to address an interesting question, but it lacks important methodological details and sufficient evidence to support its claims, especially (1) and (4). The writing is generally clear, though there are some typos. I will outline my specific concerns below."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"- Regarding generalization, could you test on existing evaluation datasets instead of constructing new ones derived from the training data? This would make the generalization claim more convincing.\n\n- In Section 5.2, you mention that 'we manually mark the minimal cues in each narrative that establish the causal hinge between agents’ intentions and outcomes.' This appears to suggest a quantitative evaluation across multiple data, yet the paper only provides a single qualitative example. Please include quantitative results or clarify how the experiments were truly performed to better support the crucial claim of 'RFT improves performance by learning to ground its reasoning in cues that correspond to causal factors'."},"rating":{"value":4},"details_of_ethics_concerns":{"value":"No ethics concerns."},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"- The writing is mostly clear and easy to follow.\n\n- A major contribution of this work lies in its attempt to identify and analyze shortcuts in existing ToM datasets. Training on such datasets may encourage models to exploit superficial correlations rather than develop genuine ToM capabilities.\n\n- The finding that RFT training outperforms SFT is not entirely surprising, since similar trends have been observed in other domains. However, this work reaches different conclusions from Lu et al. (2025) that explored similar questions.\n\n- I appreciate the use of mechanistic analyses to investigate where RFT demonstrates its advantages, although the presented evidence is not fully convincing.\n\nReference:\nLu, Y. L., Zhang, C., Song, J., Fan, L., & Wang, W. (2025). Do Theory of Mind Benchmarks Need Explicit Human-like Reasoning in Language Models?. arXiv preprint arXiv:2504.01698."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The paper reaches a different conclusion from Lu et al. (2025) but does not discuss the differences in detail. It seems the authors suggest that Lu et al. (2025) trained on shortcut-prone datasets, which may explain their opposite findings. Please provide a more detailed comparison and discussion in the related work section to clarify the source of these discrepancies.\n\n- Some terms and methods are not clearly illustrated.\n  - In Section 2.1, what exactly is the “stratified seed set”? For example, does it include (x, y) pairs where x is a question and y is the answer? How is the set “stratified”? \n  - How are the heuristics implemented and combined? These parts should be described more clearly for reproducibility and clarity, since it's suppose to support a crucial claim of this work.\n\n- Not enough evidence to support the claims:\n\n  - The paper claims that training on shortcut datasets can harm ToM abilities, but there isn’t enough evidence to support this. It would be more convincing if the authors trained the same model on these shortcut datasets using the same method and compared the results, instead of only showing a qualitative figure with reasoning-trace errors (Figure 2).\n\n  - Using procedurally generated data is not very convincing for testing generalization (e.g., generating data from 'apartment' to unseen 'outer space' for MMToM), since such data still follow similar logic as the training data. It would be stronger to test on truly out-of-domain datasets, e.g., the shortcut-prone datasets mentioned earlier, to see if RL really improves generalizable ToM abilities.\n\n- Typos for revising the manuscript:\n  - Line 118: casual -> causal\n  - Line 152: there should be '.' before 'On'\n  - Line 155: mentioned section 2 -> mentioned in Section 2\n  - Line 159: Table 2 -> Figure 2"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1764360073712,"tcdate":1761452185249,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission1862/Reviewer_c2tS"],"signatures":["ICLR.cc/2026/Conference/Submission1862/Reviewer_c2tS"],"forum":"BsEYEXjkVO","number":1,"license":"CC BY 4.0","cdate":1761452185249,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission1862/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1764360073712,"domain":"ICLR.cc/2026/Conference","replyto":"BsEYEXjkVO","id":"opkQj4L53R","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["theory of mind","reasoning","reinforcement finetuning","large language model"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Theory of Mind (ToM) is a must-acquire skill for modern foundation model systems to operate effectively and safely in the real world. Recent works have explored honing ToM via post-training; however, we show that such progress is confounded by a pervasive “shortcut” issue: tasks can reach up to 99% accuracy by simply exploiting spurious causal correlations, leading to a false sense of ToM. Motivated by this, we first develop a framework to systematically examine ToM datasets for shortcuts and provide guidance for future development. We find that questions reducible to pure state tracking (e.g., “belief”) are especially shortcut-prone compared to mind questions (e.g., “intention”) where reasoning beyond tracking is required. Using four shortcut-free datasets across three ToM contexts, we then comprehensively study whether reinforcement-learning fine-tuning with verifiable rewards and explicit reasoning (Thinking-RFT) elevates ToM beyond supervised fine-tuning (SFT). Our key findings are: 1) Thinking-RFT effectively improves ToM in all scenarios (+6% vs. SFT), particularly in complex higher-order reasoning (+10% vs. SFT) and multimodal cases (+7% vs. SFT), and generalizes notably better to unseen domains and higher-order queries while being more robust to counterfactuals. 2) ToM benefits specifically from the joint effect of reasoning and RL: Thinking-RFT outperforms No-Thinking-RFT by 7% on average. 3) RFT works by learning to ground its reasoning on anchor cues (keywords/state changes) that correspond to causal factors. We believe our study is useful for developing effective and robust ToM post-training datasets and advancing critical ToM capabilities in foundation models."},"_bibtex":{"value":"@misc{\nzhong2026from,\ntitle={From Shortcuts to Reasoning: Robust Post-Training of Theory of Mind with Reinforcement Learning},\nauthor={Jike Zhong and Yuxiang Lai and Ming Li and Yuheng Li and Wuao Liu and Behzad Dariush and Konstantinos Psounis and Shao-Yuan Lo},\nyear={2026},\nurl={https://openreview.net/forum?id=BsEYEXjkVO}\n}"},"title":{"value":"From Shortcuts to Reasoning: Robust Post-Training of Theory of Mind with Reinforcement Learning"},"pdf":{"value":"/pdf/81759a9364a3e76b0e9cc8d4b9ce2a7b855c88a7.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"zhong|from_shortcuts_to_reasoning_robust_posttraining_of_theory_of_mind_with_reinforcement_learning"},"authorids":{"value":["~Jike_Zhong1","~Yuxiang_Lai1","~Ming_Li28","~Yuheng_Li4","~Wuao_Liu1","~Behzad_Dariush2","~Konstantinos_Psounis1","~Shao-Yuan_Lo1"]},"authors":{"value":["Jike Zhong","Yuxiang Lai","Ming Li","Yuheng Li","Wuao Liu","Behzad Dariush","Konstantinos Psounis","Shao-Yuan Lo"]}},"version":2},{"content":{"summary":{"value":"This paper presents JenBridge, a framework for generating long-form video soundtracks. It segments video, generates music per segment using a video-aware MMDIT, and introduces an adaptive transition mechanism. This mechanism uses an LLM Agent to select the optimal transition style to connect segments across scene changes. It introduce a comprehensive benchmark with rich annotations and a holistic evaluation protocol for long-form video soundtracking."},"soundness":{"value":2},"confidence":{"value":5},"questions":{"value":"1. Considering the variability inherent in generative music models (dependence on random seeds), how does the segment-wise generation and stitching methodology maintain global musical coherence across potentially disparate, independently generated segments, not just ensuring smooth local transitions?\n2. Who performed the manual selection for the LVS benchmark curation ? Did this involve professionals with audio-visual production expertise?\n3. Are all conditioning inputs (sequence text, global text, visual features) necessary ? Could ablations demonstrate their individual contributions?\n4. Why were some metrics omitted from the ablation study results in Table 2 ?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The work tackles video-to-music temporal alignment from a novel perspective: segmenting the video, generating music per scene, and then adaptively transitioning between these segments to form a complete soundtrack .\n2. The proposed JenBridge framework offers a modular and interpretable approach to the complex task of long-form video soundtracking.\n3. The paper introduces the LVS Benchmark, a comprehensive resource featuring rich annotations and a holistic evaluation protocol specifically designed for long-form video soundtracking ."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The supplementary materials appear incomplete, lacking full comparative results for all examples and providing insufficient evidence to substantiate claims of coherence and fidelity across diverse transitions (only 3 examples).\n2. Applying transitions only at segment boundaries may disrupt overall musical integrity, with examples showing potentially unnatural abrupt shifts.\n3. Comparisons with recent SOTA methods like Diff-BGM [4], MuMu-LLaMA [1], GVMGen [2], and CCCG [3] are missing.\n4. Evaluation is confined to the proposed LVS benchmark, lacking validation on other public datasets (e.g., BGM909 [4], SymMV [5]).\n5. Key quantitative metrics such as FAD and KL divergence are absent from the objective evaluation.\n\n[1]  MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models\n\n[2] GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions\n\n[3] Customized Condition Controllable Generation for Video Soundtrack\n\n[4] Diff-BGM: A Diffusion Model for Video Background Music Generation\n\n[5] Video Background Music Generation: Dataset, Method and Evaluation"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762916844068,"tcdate":1761312847030,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission3583/Reviewer_ikiE"],"signatures":["ICLR.cc/2026/Conference/Submission3583/Reviewer_ikiE"],"forum":"wpQCA4yMkq","number":1,"license":"CC BY 4.0","cdate":1761312847030,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission3583/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762916844068,"domain":"ICLR.cc/2026/Conference","replyto":"wpQCA4yMkq","id":"C4WNoJVvJt","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"TLDR":{"value":"A video-to-music diffusion-based framework to generate arbitrary-length high-fidelity music waveforms with smooth transition."},"keywords":{"value":["video-to-music generation; video soundtracking; music generation; diffusion model; generative models; transition"]},"supplementary_material":{"value":"/attachment/41b97d45353a9809b5277903e87aabe45e2c26d4.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"We address the challenge of generating high-fidelity, long-form soundtracks that remain coherent across scene transitions. Existing AI music systems are mainly designed for short, isolated clips and lack mechanisms to ensure narrative continuity. We present \\texttt{JenBridge}, a modular and interpretable framework for adaptive long-form video soundtracking that ensures both high-fidelity audio generation and transition naturalness. The core architecture is a Transformer-based generative model trained with a flow-matching objective, following a two-stage paradigm: pretraining on large-scale text–audio corpora to establish robust musical priors, then adapting to the video domain with dual text–visual conditioning for precise cross-modal alignment. Crucially, to achieve long-form coherence across diverse scene changes, \\texttt{JenBridge} incorporates a novel adaptive transition mechanism. This system features a versatile toolkit of transition styles, including a generative transition method, and uniquely employs a Large Language Model (LLM) Agent that acts as a director to select the most appropriate transition for each narrative shift intelligently. To rigorously assess this task, we propose the LVS Benchmark, a new benchmark that includes a curated dataset and novel evaluation metrics focusing on holistic and transition-aware assessment. Extensive experiments on the proposed benchmark demonstrate that \\texttt{JenBridge} significantly outperforms existing methods in both objective and subjective metrics, particularly in terms of transition naturalness and overall narrative coherence. JenBridge represents a significant step towards fully automated, professional-quality video soundtracking. The codes and benchmark will be made publicly available."},"_bibtex":{"value":"@misc{\nyu2025jenbridge,\ntitle={JenBridge: Adaptive Long-Form Video Soundtracking across Scene Transition},\nauthor={Jiashuo Yu and Yao Yao and Boyu Chen and Alex Wang},\nyear={2025},\nurl={https://openreview.net/forum?id=wpQCA4yMkq}\n}"},"title":{"value":"JenBridge: Adaptive Long-Form Video Soundtracking across Scene Transition"},"pdf":{"value":"/pdf/6f3b80b651b61790240d82df9e0432ec3f286512.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"yu|jenbridge_adaptive_longform_video_soundtracking_across_scene_transition"},"authorids":{"value":["~Jiashuo_Yu1","~Yao_Yao5","~Boyu_Chen3","~Alex_Wang3"]},"authors":{"value":["Jiashuo Yu","Yao Yao","Boyu Chen","Alex Wang"]}},"version":2},{"content":{"summary":{"value":"This paper focuses on online video super-resolution which aims to reconstruct a high-resolution video from its low-resolution counterpart using only past frames, i.e., without access to information from future frames. This research has high industrial values, especially for real-time and streaming applications. In this paper, authors propose a variant of Mamba models, called TS-Mamba, to model the long-term temporal information bewtween frames. Specifically, to effectively leverage similar features from previous frames, the proposed methods aims to select similar features along the trajectory. This token selection strategy keeps spatial consistency and avoids artifacts in reconstruction. The original Mamba models efficiently process 1D sequential data, but the Raster scan mechanism cannot handle image data, especially in boundary areas. To this end, TS-Mamba introduces a variant of Hilbert scan in the spatial domain, which is restricted within and between local windows. By incoorporating the shifting operator, the proposed model can effectively process image content from adjacent windows for restoration. To better supervise model during training, it additionally propsoes trajectory loss. The experiment results have demonstrated promising results in REDS4 and Vimeo-90K datasets with BI degradation and BD degradation. Overall, this is a good paper. Unlike previous methods, it introduce Mamba structure for online video super-resolution and has demonstrated a better trade-off between restoration quality and speed. However, I still have several suggestions, questions and confusions about this paper listed as below."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"I include the questions and concerns on this paper in Weakness section. Please authors prepare their rebuttal reference to the questions listed in Section 6."},"rating":{"value":8},"details_of_ethics_concerns":{"value":"N/A"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The paper comprehensively discusses the related works in online video super-resolution and points out the existing limitations of the existing methods.\n2. TS-Mamba solely utilizes past frames for online video super-resolution and aggregates long-range information through trajectory-aware token selection, which motion paths across multiple previous frames.\n3. The proposed method is efficient without sacrificing quality, leading to a better trade-off between reconstruction quality and processing time."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. what is $$v_{\\tau_{i}^{h_{j}}}? Are they math typos in Eq.(7) and Eq.(8)?\n2. The method proposes a trajectory-aware method and define a temporal trajectory among video frames in Eq.(3). However, I am confused on this definition. It is better to elaborate more on what positions among video frames belong to the same trajectory.\n3. The method introduces a dual-path block, i.e., intra-window compensation branch and inter-window compensation branch. What is the key difference between two blocks? Why one block can handle intra-window content and another one can hanld inter-window content?\n3. The method proposes variant scanning strategies. What scanning strategy it adopts in experiment? And why using this scanning strategy instead of other strategies?\n4. It is better elaborate on the descriptions of shifted SSMs block, especially for math notations. It seems that elimintation value is a hyper-parameter. Does it matter for the restoration performance?\n5. In the paper, it proposes trajectory-aware loss. How to compute this loss? Does it compute L2 loss between the trajectories in LR and the corresponding counter in HR images?\n6. It is better to demonstrate visual results associated with Table 2. However, due to page limited, it is still acceptable."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762919222453,"tcdate":1761890857114,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission7012/Reviewer_BFtp"],"signatures":["ICLR.cc/2026/Conference/Submission7012/Reviewer_BFtp"],"forum":"RygnSGcV49","number":3,"license":"CC BY 4.0","cdate":1761890857114,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission7012/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762919222453,"domain":"ICLR.cc/2026/Conference","replyto":"RygnSGcV49","id":"DgR6Yrd8UJ","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Video Super-resolution","Online","Mamba","Trajectory"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Online video super-resolution (VSR) is an important technique for many real-world video processing applications, which aims to restore the current high-resolution video frame based on temporally previous frames. Most of the existing online VSR methods solely employ one neighboring previous frame to achieve temporal alignment, which limits long-range temporal modeling of videos. Recently, state space models (SSMs) have been proposed with linear computational complexity and a global receptive field, which significantly improve computational efficiency and performance. In this context, this paper presents a novel online VSR method based on Trajectory-aware Shifted SSMs (TS-Mamba), leveraging both long-term trajectory modeling and low-complexity Mamba to achieve efficient spatio-temporal information aggregation. Specifically, TS-Mamba first constructs the trajectories within a video to select the most similar tokens from the previous frames. Then, a Trajectory-aware Shifted Mamba Aggregation (TSMA) module consisting of proposed shifted SSMs blocks is employed to aggregate the selected tokens. The shifted SSMs blocks are designed based on Hilbert scannings and corresponding shift operations to compensate for scanning losses and strengthen the spatial continuity of Mamba. Additionally, we propose a trajectory-aware loss function to supervise the trajectory generation, ensuring the accuracy of token selection when training our model. Extensive experiments on three widely used VSR test datasets demonstrate that compared with six online VSR benchmark models, our TS-Mamba achieves state-of-the-art performance in most cases and over 22.7% complexity reduction (in MACs). The source code for TS-Mamba is available at https://github.com/QZ1-boy/TS-Mamba."},"_bibtex":{"value":"@inproceedings{\nzhu2026trajectoryaware,\ntitle={Trajectory-aware Shifted State Space Models for Online Video Super-Resolution},\nauthor={Qiang Zhu and Xiandong MENG and Yuxuan Jiang and Fan Zhang and David Bull and Shuyuan Zhu and Bing Zeng and Ronggang Wang},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=RygnSGcV49}\n}"},"title":{"value":"Trajectory-aware Shifted State Space Models for Online Video Super-Resolution"},"pdf":{"value":"/pdf/63207ba1dfa93f1f2594d2cd3d80db2064dc0722.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"zhu|trajectoryaware_shifted_state_space_models_for_online_video_superresolution"},"authorids":{"value":["~Qiang_Zhu6","~Xiandong_MENG1","~Yuxuan_Jiang8","~Fan_Zhang6","~David_Bull1","~Shuyuan_Zhu1","~Bing_Zeng1","~Ronggang_Wang1"]},"authors":{"value":["Qiang Zhu","Xiandong MENG","Yuxuan Jiang","Fan Zhang","David Bull","Shuyuan Zhu","Bing Zeng","Ronggang Wang"]}},"version":2},{"content":{"summary":{"value":"The paper proposes Amodal SAM, a framework that extends SAM to predict both visible and occluded regions of objects, aiming to improve generalization in open-world amodal segmentation. The authors introduce a lightweight Spatial Completion Adapter to infer hidden regions, a Target-Aware Occlusion Synthesis pipeline that generates synthetic occlusions from the SA-1B dataset without manual labeling, and new loss terms promoting regional consistency and topological coherence. Evaluations on six benchmarks spanning both image and video amodal segmentation (KINS, COCOA, COCOA-cls, MP3D-Amodal, FISHBOWL, and MOViD-A) show consistent gains over several prior methods in both closed- and cross-domain settings.\n\nWhile the method demonstrates solid empirical performance, its core ideas substantially overlap with prior work. The proposed TAOS pipeline closely parallels existing synthetic occlusion generation strategies such as Amodal-LVIS in SAMEO and the mixed real–synthetic approach in SAMBA, both of which already integrate similar data synthesis procedures into SAM-based models (neither is cited). Moreover, the paper omits direct comparisons to these closely related and publicly available baselines — pix2gestalt, SAMEO, and SAMBA — which weakens the empirical validation and makes it difficult to substantiate claims of state-of-the-art performance or novel methodological contribution."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Please discuss your work's relationship to the state-of-the-art methods for the problem you are trying to address. Update novelty claims accordingly. \n\nCompare to these methods on the datasets used in their paper using the same metrics."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":1},"strengths":{"value":"The paper is relatively well written and easy to follow. \n\nThe proposed approach is sound.\n\nThe proposed design offers a natural extension from image to video amodal segmntation. \n\nA minimal ablation study is reported."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The authors completely omit the most relevant works in the literature on open-world amodal segmntation. Specifically, they do not cite, discuss or compare to pix2gestalt [1], SAMEO [2], and SAMBA [3]. \n\nMoreover the dataset and methodological contributions are minimal compared to SAMEO and SAMBA, which also fine-tune SAM for amodal segmntation using synthetically \"generated\" occlusions. \n\n[1] Ozguroglu, E., Liu, R., Surís, D., Chen, D., Dave, A., Tokmakov, P., and Vondrick, C. “pix2gestalt: Amodal Segmentation by Synthesizing Wholes.”, CVPR'24\n\n[2] Tai, W.-E., Shih, Y.-L., Sun, C., Wang, Y.-C. F., and Chen, H.-T. “Segment Anything, Even Occluded.”, CVPR'25\n\n[3] Liu, Z., Qiao, L., Chu, X., Ma, L., and Jiang, T. “Towards Efficient Foundation Model for Zero-shot Amodal Segmentation.”, CVPR'25"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762920486501,"tcdate":1761504541906,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission8671/Reviewer_QvEt"],"signatures":["ICLR.cc/2026/Conference/Submission8671/Reviewer_QvEt"],"forum":"YJHuiCMkHS","number":2,"license":"CC BY 4.0","cdate":1761504541906,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission8671/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762920486501,"domain":"ICLR.cc/2026/Conference","replyto":"YJHuiCMkHS","id":"uSm9suBRhr","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Amodal segmentation","SAM","Open world"]},"supplementary_material":{"value":"/attachment/cda11fbd3c125040f18eeaff662e0947870e3ea4.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Amodal segmentation, which aims to predict complete object shapes including occluded regions, remains challenging in open-world scenarios where models must generalize to novel objects and contexts. While the Segment Anything Model (SAM) has demonstrated remarkable zero-shot generalization capabilities, it is fundamentally limited to visible region segmentation. This paper presents Amodal SAM, a framework that extends SAM's capabilities to amodal segmentation while preserving its powerful generalization ability. The improvements lie in three aspects: (1) a lightweight Spatial Completion Adapter that enables occluded region reconstruction, (2) a Target-Aware Occlusion Synthesis (TAOS) pipeline that addresses the scarcity of amodal annotations by generating diverse synthetic training data, and (3) novel learning objectives that enforce regional consistency and topological regularization. Extensive experiments demonstrate that Amodal SAM achieves state-of-the-art performance on standard benchmarks while exhibiting strong generalization to novel scenarios. Furthermore, our framework seamlessly extends to video sequences, as the first attempt to tackle the open-world video amodal segmentation. We hope our research can advance the field toward practical amodal segmentation systems that can operate effectively in unconstrained real-world environments. Code and models will be made publicly available."},"_bibtex":{"value":"@misc{\nzhang2025amodal,\ntitle={Amodal {SAM}: Open-World Amodal Segmentation},\nauthor={Bo Zhang and Zhuotao Tian and Xin Tao and Songlin Tang and Guangming Lu and Jun Yu and Wenjie Pei},\nyear={2025},\nurl={https://openreview.net/forum?id=YJHuiCMkHS}\n}"},"title":{"value":"Amodal SAM: Open-World Amodal Segmentation"},"pdf":{"value":"/pdf/86c0f3223bb299c5fe7b0344183c1068ee37483d.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"zhang|amodal_sam_openworld_amodal_segmentation"},"authorids":{"value":["~Bo_Zhang43","~Zhuotao_Tian1","~Xin_Tao3","~Songlin_Tang2","~Guangming_Lu2","~Jun_Yu1","~Wenjie_Pei1"]},"authors":{"value":["Bo Zhang","Zhuotao Tian","Xin Tao","Songlin Tang","Guangming Lu","Jun Yu","Wenjie Pei"]}},"version":2},{"content":{"summary":{"value":"The paper proposes InfiniBench, a novel benchmark for long video understanding based on movies and TV shows. The benchmark has 108.2k question-answer pairs on 1,219 videos that average 52.59 minutes in length. The benchmark tests 9 different reasoning abilities including visual, long-context and local reasoning. This makes InfiniBench the largest-scale long video understanding benchmark to date. InfiniBench was constructed by combining and augmenting from two existing video benchmarks, TVQA and MovieNet. Most question types were generated by prompting GPT-4 with the transcript of the video while a custom pipeline was used to generate questions on changes in character appearance. The paper presents benchmark results of 8 long video understanding models, including 6 open source ones and 2 commercial ones, and discusses insights into their performance across various tasks."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Given the concerns listed above, I have doubts that this paper is suitable for publication at ICLR. I hope that authors can provide evidence to address my concerns as well as answers to the following questions.\n\n* Could authors provide evidence of transcript quality? How accurate and complete are they? How much focus do they have on vision? Could authors provide examples?  \n* Why are multiple-choice questions evaluated by asking the model to generate an answer and then using GPT to match this answer to the options? Authors state in the appendix that the reason is that models often do not follow the prescribed answer format, but from my experience at least the larger VLMs are good at following instructions about the answer format.  \n* I am worried that using GPT for option matching introduces additional bias. I believe this could be measured by evaluating GPT or Gemini again by giving it the answer options in the prompt and asking it to respond with only the answer letter. Results could then be compared against the GPT-matched results.  \n  * Also to the above point, did authors verify that event ordering type questions get matched correctly with GPT? These answers only differ in their ordering of options, so I am wondering whether GPT matches them correctly.  \n* The benchmark was constructed using GPT, and GPT is the best performing model across all tasks. It would be interesting to quantify if there is bias towards GPT, e.g. by generating part of the data with Gemini and checking if relative model performance is consistent with the original benchmark.  \n* How are copyright concerns handled? Did authors obtain permission from the copyright owners to use the video material for this purpose and to reproduce this content in a publication? If the dataset will be publicly released, how are copyright concerns handled?  \n  * l. 198: “To address this limitation, we transformed the TVQA dataset from a collection of short clips into a long video dataset by gathering and sequencing the clips corresponding to each episode thereby reconstructing the full episode frames.“ How was this done and what data source was used?  \n* Appendix l. 12: “The remaining two skills, i.e., local visual questions and summarizing, do not need human verification, as the first one is adopted from the TVQA dataset, and the latter is scrapped from human responses on the web.” I do not fully agree with this statement since existing benchmarks and the humans writing the summaries that were pulled from the web could still contain errors. Do authors have evidence to the quality of TVQA annotations and summaries obtained from the web?  \n* How does the number of video frames provided affect the model accuracy?  \n* Appendix B is quite important to understand the evaluation results presented, so I think it would be better suited to be in the main text.  \n* Appendix B mentions that the benchmark videos have no audio, so video and subtitles are provided to the model separately. Does this mean that alignment between frames and subtitles is missing? Did authors measure the effect of this?  \n* Could authors explain how spoiler questions are generated and provide the prompt used?  \n* How does the “I don’t know” option affect results? How accurately does GPT match model answers to this option?  \n* Fig. 5 (left) is redundant with Tab. 3, so one of them should be removed.  \n* l. 363: The explanation of local vision and text questions is not clear. It is not explained what these questions are nor how they were generated.  \n* It would be good to have random accuracy in Tab. 5 for direct comparability. Then, Tab. 4 could be omitted.  \n* l. 482: “As shown in the table 5, MiniGPT4-video and LLaVA-NeXT-Interleave match lower than the random performance” What random performance is being compared to here? It would help to add this to the table as suggested above.  \n* l. 482, l. 505: How can a model’s performance be lower than random?  \n* l. 488: “One reason may be that eliminating the noisy information and focus on only the related information helps more in answering the questions“ How does the Goldfish model eliminate noisy information?  \n* For the human verification, how were human responses on open-ended questions evaluated?\n\nMinor points\n\n* Tab. 1: I would not agree with the “human” checkmark for InfiniBench since questions were generated fully automatically.  \n* Tab. 2 is never referenced.  \n* Appendix B: It would be helpful to express this in tabular form so readers can see at a glance how many frames and what modalities were used in each model.  \n* Tab. 5.: I would suggest to organize this into one big table with one column per task type. Also would be nice to visualize as a radar chart.  \n* It would be helpful to annotate question types in Sec 3.2.2 and Fig. 1 with whether they are MCQ or OE.  \n* It would be helpful to see a listing of modalities (vision, summary, transcript) used to generate each question.  \n* Please use \\\\citep for citations to place citations in parentheses.  \n* In tables, please right-justify numerical columns and use a consistent number of digits after the decimal point.  \n* Fig. 4: The font size in these charts is very small in print. I suggest increasing it. Also I would suggest to change the pie chart into a bar chart for easier readability.  \n* Fig. 5: Same concern as above about the font size.  \n* l. 373: Here, the reference to Fig. 4 is repeated, but Fig. 5 is wrongly referenced. Suggest correcting this sentence to refer to Fig. 3\\.  \n* l. 406: Broken reference.  \n* l. 413: The reference should point to Sec. B in the supplementary material."},"rating":{"value":8},"details_of_ethics_concerns":{"value":"The following are my original concerns which have been mitigated:\n\nI have a concern about potential copyright infringement in this work. The proposed dataset is based on copyrighted content (video frames and subtitles of movies and TV shows) that authors have downloaded and used for experiments. The paper also includes figures of frames from TV shows. It is unclear whether the authors obtained permission from copyright owners for their use of the data. Authors do not mention whether they intend to release the dataset publicly, but if they do, this would raise further concerns."},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"* The presented benchmark has an impressive scale with 108.2k questions on 1,219 videos that average 52.59 minutes in length.  \n* There are 9 different question types that test long video understanding models across a variety of skills.  \n* The paper presents results of 8 long video models and draws interesting conclusions on their performance.  \n* There is a large gap between human performance and model performance, suggesting the benchmark has ample room for improvement.  \n* The paper has a good in-depth discussion of related work."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"* The question-answer pairs in the benchmark were generated fully automatically without any human intervention. This raises questions about soundness and of the questions and potential bias. A human evaluation is performed on a subset of the data, but good human performance is no proof that questions are well-formed and free of hallucinations.  \n* Most questions are generated from transcripts that authors obtained online, but it is unclear what information these transcripts contain, whether they are complete and error-free. It is also unclear how much visual information the transcripts contain and therefore it is unclear to what degree this is a multimodal benchmark.  \n* The use of movies and TV shows raises questions about generalizability. Most MLLMs likely know the plots of popular movies and shows because their summaries or transcripts were part of their training data. So, they may be able to answer the questions in the dataset without any context, which is not the case for most videos from the web. The effect of this is not examined.  \n* It is unclear how much the benchmark relies on multimodal reasoning. Questions about movies and TV shows could often be answerable from subtitles alone, which are provided as context in the evaluation. It would be interesting to see an ablation that uses (1) No context, only the question itself (2) Only the question and subtitles (3) the question, subtitles and video frames.  \n* The copyright implications of using movies and TV shows and possibly releasing the dataset are not discussed and raise ethical concerns.  \n* Since the dataset has \\~100 questions per video, it is likely that there are (near) duplicate questions. However there is no analysis of this and no mention of a filtering stage to remove duplicates.  \n* There are several issues with the presentation such as redundant figures, tables that are not referenced, and wrong references. The limitations section also exceeds the 10-page limit."}},"nonreaders":[],"tmdate":1733184911698,"tcdate":1730567275033,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission2250/Reviewer_oZkF"],"signatures":["ICLR.cc/2025/Conference/Submission2250/Reviewer_oZkF"],"forum":"2D0uXQbntW","number":3,"license":"CC BY 4.0","cdate":1730567275033,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission2250/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1733184911698,"domain":"ICLR.cc/2025/Conference","replyto":"2D0uXQbntW","id":"MBr0wRlWgD","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["video understanding","benchmark","long video benchmark","long video understanding"]},"supplementary_material":{"value":"/attachment/e4286990af860ccb3f16a8fce15b1da8635ef0b5.zip"},"primary_area":{"value":"datasets and benchmarks"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Understanding long videos, ranging from tens of minutes to several hours, presents unique challenges in video comprehension. Despite the increasing importance of long-form video content, existing benchmarks primarily focus on shorter clips. To address this gap, we introduce InfiniBench a comprehensive benchmark for very long video understanding which presents 1)very long video duration, averaging 52.59 minutes per video 2)The largest number of question-answer pairs, 108.2K 3) Diversity in questions that examine nine different skills and include both multiple-choice questions and open-ended questions 4) Memory questions, such as Global Appearance that require remembering and tracking the visual aspects through the video. Using InfiniBench, we comprehensively evaluate existing Large Multi-Modality Models (LMMs) on each skill, including the commercial models such as GPT-4o and Gemini 1.5 Flash and the recent open-source models. \nThe evaluation shows significant challenges in our benchmark.\nOur findings reveal that even leading AI models like GPT-4o and Gemini 1.5 Flash face challenges in achieving high performance in long video understanding, with average accuracies of just 56.01 % and 43.32 %, and average scores of 3.25 and 2.79 out of 5, respectively.\nQwen2-VL matches Gemini's performance in the MCQ skills but lags significantly in open-ended question tasks.\nWe hope this benchmark will stimulate the LMMs community towards long video and human-level understanding."},"_bibtex":{"value":"@misc{\nataallah2025infinibench,\ntitle={InfiniBench: A Comprehensive Benchmark for Large Multimodal Models in Very Long Video Understanding},\nauthor={Kirolos Ataallah and Chenhui Gou and Eslam Mohamed BAKR and Khushbu Pahwa and Jian Ding and Mohamed Elhoseiny},\nyear={2025},\nurl={https://openreview.net/forum?id=2D0uXQbntW}\n}"},"title":{"value":"InfiniBench: A Comprehensive Benchmark for Large Multimodal Models in Very Long Video Understanding"},"pdf":{"value":"/pdf/032fb6c70b723929924e625fabf7cebb80b48405.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"ataallah|infinibench_a_comprehensive_benchmark_for_large_multimodal_models_in_very_long_video_understanding"},"authorids":{"value":["~Kirolos_Ataallah1","~Chenhui_Gou1","~Eslam_Mohamed_BAKR1","~Khushbu_Pahwa1","~Jian_Ding3","~Mohamed_Elhoseiny1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Kirolos Ataallah","Chenhui Gou","Eslam Mohamed BAKR","Khushbu Pahwa","Jian Ding","Mohamed Elhoseiny"]}},"version":2},{"content":{"summary":{"value":"The paper proposes a novel model for video moment retrieval, DTAM, which introduces dual temporal adjacent maps to enhance retrieval accuracy. It designs a semantic-aware contrastive loss that clusters features for the same query while distancing those for different queries, and incorporates a moment-aware mechanism to further strengthen the temporal adjacent maps. Through systematic experiments on three benchmark datasets and ablation studies, the paper thoroughly analyzes and validates the contributions of each module."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Clarification on Branch Specialization: The paper states that one branch focuses on visual appearance while the other emphasizes semantic information. However, there is no explicit supervisory signal in the model to ensure that each branch indeed specializes in its respective area. Could the authors provide clarification on how they ensure this specialization during training?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"This article has advantages in the following ways:\n1.  The paper introduces the Dual Temporal Adjacent Maps (DTAM) model, which decouples visual and semantic information for video moment retrieval. This design allows the model to effectively distinguish between moments that are visually similar but semantically different, addressing the limitations of traditional methods that couple visual and semantic features. \n2.  Innovative Structure Design: The paper introduces the Dual Temporal Adjacent Maps (DTAM) model, which decouples visual and semantic information for video moment retrieval. This design allows the model to effectively distinguish between moments that are visually similar but semantically different, addressing the limitations of traditional methods that couple visual and semantic features. \n3.  The introduction of the moment-aware mechanism allows the model to dynamically adjust the importance of video segments. This mechanism strengthens the representational power of the temporal adjacent maps in the video moment retrieval task, enabling better capture and modeling of temporal relationships in videos."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"I think this is a convincing paper. The research questions are all reasonable. However, I believe that some improvements can be made.\n1. The formatting of the tables and images on pages 6 and 7 of the paper is not aesthetically pleasing. Could you consider realigning and reformatting them?  For example, the interval between table 1 and table 2 should be widened.\n2. The Introduction section in Chapter 1 states that \"To this end, we propose a semantic-enhanced Dual Temporal Adjacent Maps (DTAM) for effective video grounding...\" uses “video grounding”, but “video moment retrieval” used in subsequent articles. It is recommended to unify the entire text.  Video Moment Retrieval is recommended, as this is also the usage of most articles.\n3.  It is unclear what knowledge the two branches have actually learned. The paper suggests that one branch focuses on appearance and the other on semantics, but this seems to be a subjective interpretation.  The paper lacks explicit supervisory signals to ensure that each branch focuses on either visual or semantic features, and there is no interpretability analysis or ablation studies to demonstrate that the branches have indeed learned their claimed distinct features.  I recommend that the authors address this issue in the manuscript. You can visualize the feature map and illustrate its results"}},"nonreaders":[],"tmdate":1731428021208,"tcdate":1730205153701,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission4047/Reviewer_4XrJ"],"signatures":["ICLR.cc/2025/Conference/Submission4047/Reviewer_4XrJ"],"forum":"l3CSCOnGPB","number":2,"license":"CC BY 4.0","cdate":1730205153701,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission4047/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428021208,"domain":"ICLR.cc/2025/Conference","replyto":"l3CSCOnGPB","id":"IlEVsT3uFQ","forumContent":{"TLDR":{"value":"a novel semantic-enhanced dual temporal adjacent maps (DTAM) for effective video moment retrieval, which models temporal dependencies between moments in an appearance-semantic decoupled fashion."},"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Computer Vision","Muylti-modal Understanding","Video Grounding"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Retrieving a specific moment from an untrimmed video via a text description is a central problem in vision-language learning. It is a challenging task due to the sophisticated temporal dependency among moments. Existing methods fail to deal with this issue well since they establish temporal relations of moments in a way that visual content and semantics are coupled. This paper studies temporal dependence schemes that decouple content and semantic information, establishing semantic-enhanced Dual Temporal Adjacent Maps for video moment retrieval, conferred as DTAM. Specifically, DTAM designs two branches to encode visual appearance and semantic knowledge from video clips respectively, where knowledge from the appearance branch is distilled into the semantic branch to help DTAM distinguish features with the same visual content but different semantics with a well-designed semantic-aware contrastive loss. Besides, we also develop a moment-aware mechanism to assist temporal adjacent maps' learning for better video grounding. Finally, extensive experimental results and analysis demonstrate the superiority of the proposed DTAM over existing state-of-the-art approaches on three challenging video moment retrieval benchmarks, i.e., TACoS, Charades-STA, and ActivityNet Captions."},"_bibtex":{"value":"@misc{\nwang2025learning,\ntitle={Learning Semantic-Enhanced Dual Temporal Adjacent Maps for Video Moment Retrieval},\nauthor={Yu Wang and Shengjie Zhao and Shiwei Chen},\nyear={2025},\nurl={https://openreview.net/forum?id=l3CSCOnGPB}\n}"},"title":{"value":"Learning Semantic-Enhanced Dual Temporal Adjacent Maps for Video Moment Retrieval"},"pdf":{"value":"/pdf/83e3b8d8b0191f0c4f665fc641a4a682d0f67e99.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"wang|learning_semanticenhanced_dual_temporal_adjacent_maps_for_video_moment_retrieval"},"authorids":{"value":["~Yu_Wang32","~Shengjie_Zhao1","~Shiwei_Chen3"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Yu Wang","Shengjie Zhao","Shiwei Chen"]}},"version":2},{"content":{"venue":{"value":"IEEE Trans. Circuits Syst. Video Technol. 2025"},"pdf":{"value":"https://ieeexplore.ieee.org/iel8/76/10857816/10666730.pdf"},"venueid":{"value":"dblp.org/journals/TCSV/2025"},"paperhash":{"value":"hou|bidirectional_erroraware_fusion_network_for_video_inpainting"},"authorids":{"value":["","~Zhong_Ji1","https://dblp.org/search/pid/api?q=author:Jinyu_Yang:","https://dblp.org/search/pid/api?q=author:Feng_Zheng:"]},"html":{"value":"https://doi.org/10.1109/TCSVT.2024.3454641"},"_bibtex":{"value":"@article{DBLP:journals/tcsv/HouJYZ25,\n  author={Jiacheng Hou and Zhong Ji and Jinyu Yang and Feng Zheng},\n  title={Bidirectional Error-Aware Fusion Network for Video Inpainting},\n  year={2025},\n  month={January},\n  cdate={1735689600000},\n  journal={IEEE Trans. Circuits Syst. Video Technol.},\n  volume={35},\n  number={1},\n  pages={577-588},\n  url={https://doi.org/10.1109/TCSVT.2024.3454641}\n}\n"},"abstract":{"value":"Existing video inpainting approaches tend to adopt vision transformers with rare customized designs, which poses two limitations. Firstly, the conventional self-attention mechanism treats tokens from invalid and valid regions equally and mingles them, which may incur blurriness. Secondly, these approaches merely employ forward frames as references, while ignoring the past inpainted frames, which are also valuable in enhancing temporal consistency and offering more available information. In this paper, we propose a new video inpainting network, called Bidirectional Error-Aware Fusion Network (BEAF-Net). Concretely, on one hand, we propose a tailored Error-Aware Transformer (EAT) that discerns different tokens by assigning dynamic weights to bridle the use of erroneous tokens. Meanwhile, each EAT is equipped with a Spatial Feature Enhancement (SFE) layer to synthesize features with multi-scales. On the other hand, we apply a pair of EATs to utilize forward reference frames and past inpainted frames simultaneously, and a proposed Bidirectional Fusion (BiF) layer is exerted to blend the aggregation results adaptively. By coupling these novel designs, our proposed BEAF-Net completely leverages the location priors, multi-scale perception, and past predictions to produce more faithful and consistent inpainting results. We corroborate our BEAF-Net on two commonly-used video inpainting datasets: DAVIS and Youtube-VOS, where the experimental results demonstrate BEAF-Net compares favorably with state-of-the-art solutions. Video examples can be found at https://github.com/JCATCV/BEAF-Net."},"title":{"value":"Bidirectional Error-Aware Fusion Network for Video Inpainting"},"authors":{"value":["Jiacheng Hou","Zhong Ji","Jinyu Yang","Feng Zheng"]}},"tmdate":1768272322051,"pdate":1735689600000,"tcdate":1747296608358,"writers":["~"],"signatures":["~Zhong_Ji2"],"forum":"FxjRUBMKOk","license":"CC BY-SA 4.0","number":478461,"cdate":1735689600000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1768272322051,"domain":"DBLP.org","id":"FxjRUBMKOk","version":2},{"content":{"summary":{"value":"This paper proposed a strong foundation model for promptable visual segmentation in images and videos. It proposed a new data engine  that enhances model and data through user interaction, creating the largest video segmentation dataset to date.  A streaming memory augmented transformer is proposed  for real-time video processing."},"soundness":{"value":4},"confidence":{"value":4},"questions":{"value":"1. How is the annotation quality of the SA-V dataset, and how do you ensure the quality of the annotations? What is the difficulty level of this dataset (such as the movement of objects in the video, occlusions, etc.) compared to previous datasets?\n\n2. Has SAM 2 attempted stability testing for results on ultra-long videos? Is object tracking in long videos more prone to errors?\n\n\n3. If the memory bank is of fixed size, will it lead to forgetting when dealing with long videos?"},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":4},"strengths":{"value":"1. This paper proposed a strong foundation model for the video and image segmentation. The data, model, and insights will serve\nas a significant milestone for video segmentation.\n\n2. The writing of the paper is good and the paper is easy to understand."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. More experiments should be conducted. For example, more interactive VOS methods should be compared.\n[1*] Modular interactive video object segmentation: Interaction-to-mask, propagation and difference-aware fusion. CVPR 2021\n[2*] Memory aggregation networks for efficient interactive video object segmentation. CVPR 2020\n\n2. More VOS datasets (e.g., VIPOSeg[4*]) should be included in this paper. \n[4*] Video Object Segmentation in Panoptic Wild Scenes. IJCAI 2023"}},"nonreaders":[],"tmdate":1731427195152,"tcdate":1730468091756,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission92/Reviewer_DQTY"],"signatures":["ICLR.cc/2025/Conference/Submission92/Reviewer_DQTY"],"forum":"Ha6RTeWMd0","number":3,"license":"CC BY 4.0","cdate":1730468091756,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission92/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427195152,"domain":"ICLR.cc/2025/Conference","replyto":"Ha6RTeWMd0","id":"M5CFKoUyTa","forumContent":{"venue":{"value":"ICLR 2025 Oral"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["computer vision","video segmentation","image segmentation"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"We present Segment Anything Model 2 (SAM 2), a foundation model towards solving promptable visual segmentation in images and videos. We build a data engine, which improves model and data via user interaction, to collect the largest video segmentation dataset to date. Our model is a simple transformer architecture with streaming memory for real-time video processing. SAM 2 trained on our data provides strong performance across a wide range of tasks. In video segmentation, we observe better accuracy, using 3x fewer interactions than prior approaches. In image segmentation, our model is more accurate and 6x faster than the Segment Anything Model (SAM). We believe that our data, model, and insights will serve as a significant milestone for video segmentation and related perception tasks. We are releasing our main model, the dataset, an interactive demo and code."},"_bibtex":{"value":"@inproceedings{\nravi2025sam,\ntitle={{SAM} 2: Segment Anything in Images and Videos},\nauthor={Nikhila Ravi and Valentin Gabeur and Yuan-Ting Hu and Ronghang Hu and Chaitanya Ryali and Tengyu Ma and Haitham Khedr and Roman R{\\\"a}dle and Chloe Rolland and Laura Gustafson and Eric Mintun and Junting Pan and Kalyan Vasudev Alwala and Nicolas Carion and Chao-Yuan Wu and Ross Girshick and Piotr Dollar and Christoph Feichtenhofer},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=Ha6RTeWMd0}\n}"},"title":{"value":"SAM 2: Segment Anything in Images and Videos"},"pdf":{"value":"/pdf/7c41968163abe4e3700e3e3a15174a9d679fcd52.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"ravi|sam_2_segment_anything_in_images_and_videos"},"authorids":{"value":["~Nikhila_Ravi1","~Valentin_Gabeur1","~Yuan-Ting_Hu1","~Ronghang_Hu1","~Chaitanya_Ryali1","~Tengyu_Ma3","~Haitham_Khedr1","~Roman_Rädle1","~Chloe_Rolland1","~Laura_Gustafson1","~Eric_Mintun1","~Junting_Pan2","~Kalyan_Vasudev_Alwala1","~Nicolas_Carion1","~Chao-Yuan_Wu1","~Ross_Girshick1","~Piotr_Dollar1","~Christoph_Feichtenhofer4"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Nikhila Ravi","Valentin Gabeur","Yuan-Ting Hu","Ronghang Hu","Chaitanya Ryali","Tengyu Ma","Haitham Khedr","Roman Rädle","Chloe Rolland","Laura Gustafson","Eric Mintun","Junting Pan","Kalyan Vasudev Alwala","Nicolas Carion","Chao-Yuan Wu","Ross Girshick","Piotr Dollar","Christoph Feichtenhofer"]}},"version":2},{"content":{"summary":{"value":"This paper proposed PPLLAVA, a video-LLM based on a prompt-guide pooling strategy. The core idea is to conduct text query/prompt-dependent pooling for video features before putting them into LLM. The authors conduct extensive experiments to show the effectiveness of the proposed method. The proposed PPLLAVA outperforms other training-free/training-based video-LLM at a similar model scale."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Please see the weaknesses."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"(1) this paper conducts extensive experiments to show the effectiveness of the proposed methods.\n\n(2) pooling redundant tokens based on visual-text similarity is an efficient solution to capture query-related information in a video with large redundancy.\n\n(3) The design of extending the CLIP content window is smart enough to adapt CLIP to a longer context with minimal modification."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"(1) one thing I'm a bit confused about is the video captioning setting when adapting this prompt-guided pooling strategy. As the caption prompts are not as diverse as that question answering, holistic captioning should be affected unless the user is querying a specific object-related caption. I would suggest some experiments on video caption benchmarks, e.g. activitynet caption or newer DREAM-1K, to further support the strong video understanding ability claim.\n\n(2) it seems the paper missed some related baseline methods, see LongVA [1], and Kangaroo [2]. It should be helpful to include a more comprehensive analysis and comparison with those methods, even if some of them achieve better performance on some datasets. \n\n(3) My other big concern is the novelty side. While I understand this paper did a good job of validating their idea/design with extensive experiments and ablations, comparing some zero-shot baseline like PLLAVA/LLaVA-NeXT-Video, the big boost of the performance might come from fine-tuning on 1.3M multimodal data and the prompt-guided pooling seems a bit incremental from my personal perspective. So I am a bit worried that if we conduct the same level of training with a stronger LLM backbone that has a long content window, the model can learn this query-related attention itself, so the proposed core idea, prompt-guided pooling, seems sub-optimal. \n\n(4) A potential solution might be to fine-tune zero methods that conduct their token merging idea with the same data and show the performance difference.\n\n[1] Long Context Transfer from Language to Vision\n[2] Kangaroo: A powerful video-language model supporting long-context video input"}},"nonreaders":[],"tmdate":1732557342697,"tcdate":1730703412997,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission1865/Reviewer_vFqm"],"signatures":["ICLR.cc/2025/Conference/Submission1865/Reviewer_vFqm"],"forum":"qUZY7ymDPr","number":2,"license":"CC BY 4.0","cdate":1730703412997,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission1865/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732557342697,"domain":"ICLR.cc/2025/Conference","replyto":"qUZY7ymDPr","id":"z97QpYzXlO","forumContent":{"TLDR":{"value":"visual token pooling with instruction guidance, 1/8 throughput, better video&image performance."},"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video LLM","Prompt-guided Pooling","PPLLaVA"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"The past year has witnessed the significant advancement of video-based large language models. However, the challenge of developing a unified model for both short and long video understanding remains unresolved. Most existing video LLMs cannot handle hour-long videos, while methods custom for long videos tend to be ineffective for shorter videos and images. In this paper, we identify the key issue as the redundant content in videos. To address this, we propose a novel pooling strategy that simultaneously achieves token compression and instruction-aware visual feature aggregation. Our model is termed Prompt-guided Pooling LLaVA, or PPLLaVA for short. Specifically, PPLLaVA consists of three core components: the CLIP-based visual-prompt alignment that extracts visual information relevant to the user's instructions, the prompt-guided pooling that compresses the visual sequence to arbitrary scales using convolution-style pooling, and the clip context extension designed for lengthy prompt common in visual dialogue. Moreover, our codebase also integrates the most advanced video Direct Preference Optimization (DPO) and visual interleave training. Extensive experiments have validated the performance of our model. With superior throughput, PPLLaVA achieves better results on image benchmarks as a video LLM, while achieving state-of-the-art performance across various video benchmarks, excelling in tasks ranging from caption generation to multiple-choice questions, and handling video lengths from seconds to hours. The codes are promised to be made public."},"_bibtex":{"value":"@misc{\nliu2025ppllava,\ntitle={{PPLL}a{VA}: Varied Video Sequence Understanding With Prompt Guidance},\nauthor={Ruyang Liu and Chen Li and Haoran Tang and Yixiao Ge and Ying Shan and Haibo Lu and Jiankun Yang},\nyear={2025},\nurl={https://openreview.net/forum?id=qUZY7ymDPr}\n}"},"title":{"value":"PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance"},"pdf":{"value":"/pdf/bb6fc0273484f207f47d34813d971987179b27a5.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"liu|ppllava_varied_video_sequence_understanding_with_prompt_guidance"},"authorids":{"value":["~Ruyang_Liu1","~Chen_Li34","~Haoran_Tang4","~Yixiao_Ge2","~Ying_Shan2","~Haibo_Lu1","~Jiankun_Yang1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Ruyang Liu","Chen Li","Haoran Tang","Yixiao Ge","Ying Shan","Haibo Lu","Jiankun Yang"]}},"version":2},{"content":{"summary":{"value":"This paper introduces fast flow joint distillation(F2D2), which jointly learns sampling trajectory and divergence using a joint self-distillation  objective loss and a single model to enable few-step sampling and fast log-likelihood evaluation in flow-based models. It shows that the method is compatible with Shortcut and Meanflow Models and can be further distilled. The proposed method performs on image-based dataset and shows optimistic results."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"Following the weaknesses:\n\n1. In table 1, I found it is hard to interpret the results as no method performs consistently well than the others across different sampling steps, either it is existing method or the proposed methods. The results from shortcut distilling method outperforms the shortcut model sometimes. I am wondering if there is actually no pattern and the good performance happens by chance.\n2. In Shortcur-Distilll-F2D2, it only talks about the vector field training, how's divergence trained here?\n3. How are invalid NLLs computed in table 1?\n4. I am not pretty sure why we should learn the divergence term from the beginning. Presumably we can learn instantaneous/average vector fields well, and this will induce the marginal density distribution for each $\\rho_t(x)$, then it will automatically the continuity function, which means we know the divergence in the meanwhile. If this is correct, learning divergence is redundant. Please correct me if I am wrong."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. It's interesting to learn vector field and divergence by sharing the same backbone with separate prediction heads since they are related to each other.\n2. The proposed method can be extended to shortcut model and meanflow model."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The objective loss of shortcut F2D2 is composed of 4 components. It is not quite clear if any one of the components might dominate the whole training and make the performance better. An ablation study is required here. Also, are the 4 components equally weighted? It appears equally weighted from the equation but what if the weights are tuned.\n2. Though the idea to learn velocity field and divergence is interesting, I am wondering if it might lead to instability during training since velocity field learns the directional vector and divergence learns the expectation of the trace through Hutchinson estimator, which involves a gradient of vector field. What if they do not share the same numerical range? Again, it is better to have an ablation study here.\n3. For the experiments, no computation cost and runtime is provided."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762924285142,"tcdate":1761817502039,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission13749/Reviewer_5Bkt"],"signatures":["ICLR.cc/2026/Conference/Submission13749/Reviewer_5Bkt"],"forum":"8uZ5UdIul2","number":3,"license":"CC BY 4.0","cdate":1761817502039,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission13749/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762924285142,"domain":"ICLR.cc/2026/Conference","replyto":"8uZ5UdIul2","id":"zvz5XHggQO","forumContent":{"TLDR":{"value":"We propose a modular distillation framework for simultaneous fast likelihood evaluation and fast sampling for flow matching models"},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["flow matching","distillation models","fast likelihood evaluation","fast sampling","generative models"]},"primary_area":{"value":"generative models"},"abstract":{"value":"Log-likelihood evaluation enables important capabilities in generative models, including model comparison, certain fine-tuning objectives, and many downstream applications. Yet paradoxically, some of today's best generative models -- diffusion and flow-based models -- still require hundreds to thousands of neural function evaluations (NFEs) to compute a single likelihood. While recent distillation methods have successfully accelerated sampling to just a few steps, they achieve this at the cost of likelihood tractability: existing approaches either abandon likelihood computation entirely or still require expensive integration over full trajectories. We present fast flow joint distillation (F2D2), a framework that simultaneously reduces the number of NFEs required for both sampling and likelihood evaluation by two orders of magnitude. Our key insight is that in continuous normalizing flows, the coupled ODEs for sampling and likelihood are computed from a shared underlying velocity field, allowing us to jointly distill both the sampling trajectory and cumulative divergence using a single flow map.   F2D2 is modular, compatible with existing flow-based few-step sampling models, and requires only an additional divergence prediction head. Experiments demonstrate F2D2's capability of achieving accurate log-likelihood with few-step evaluations while maintaining high sample quality, solving a long-standing computational bottleneck in flow-based generative models. As an application of our approach, we propose a lightweight self-guidance method that enables a 2-step MeanFlow to outperform a 1024 step flow matching model with only a single additional backward NFE."},"_bibtex":{"value":"@inproceedings{\nai2026joint,\ntitle={Joint Distillation for Fast Likelihood Evaluation and Sampling in Flow-based Models},\nauthor={Xinyue Ai and Yutong He and Albert Gu and Ruslan Salakhutdinov and J Zico Kolter and Nicholas Matthew Boffi and Max Simchowitz},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=8uZ5UdIul2}\n}"},"title":{"value":"Joint Distillation for Fast Likelihood Evaluation and Sampling in Flow-based Models"},"pdf":{"value":"/pdf/68520d640fe4e747b1255302b54d5ae03df166b3.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"ai|joint_distillation_for_fast_likelihood_evaluation_and_sampling_in_flowbased_models"},"authorids":{"value":["~Xinyue_Ai1","~Yutong_He1","~Albert_Gu1","~Ruslan_Salakhutdinov1","~J_Zico_Kolter1","~Nicholas_Matthew_Boffi1","~Max_Simchowitz1"]},"authors":{"value":["Xinyue Ai","Yutong He","Albert Gu","Ruslan Salakhutdinov","J Zico Kolter","Nicholas Matthew Boffi","Max Simchowitz"]}},"version":2},{"content":{"summary":{"value":"The authors propose a novel approach to 2D video compression using 3D Gaussian Splatting (3DGS). They represent 2D video sequences as a volumetric distribution in 3D space, using X-Y-T coordinates. To reduce temporal redundancy, they introduce a Toast-like Sliding Window (TSW) projection method that considers Gaussian points within a temporal window centered on the rendering plane. The framework incorporates optical flow information and a FiLM-based deformation field to better capture temporal dynamics. An entropy coding scheme is implemented to optimize model compression and enhance streaming efficiency.\n\nThe paper presents three main contributions:\n* Introduction of a novel paradigm that represents 2D video using 3D Gaussian models, supported by the TSW projection technique\n* Development of GSVC, a comprehensive compression pipeline that combines HAC, deformable fields, and opacity flow regularization\n* Experimental validation showing GSVC achieves comparable quality to NeRV while providing 30% faster rendering and flexible bitrate-availability tradeoffs"},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"- In Figure 5, there remains a gap in frame availability even after complete bit decoding. What factors contribute to this limitation?\n- Another concern is about the depth sorting when it comes to rendering. there could be a scenario when two gaussians (lets say Ga and Gb), in the frame t-1, fall into Vf, while in the next frame t, they fall into Vb, due to the time plane moving forward. Yet in this case the relative depths of Ga and Gb to the image plane are reversed \n  - for example, in t-1, Ga, Gb in Vf, and Ga is close to the plane, and Gb is further to the plane, then in t, Ga, Gb in Vb, it must be that Ga is further to the plane and Gb looks closer to the plane\n  - Such kind of reversing might lead to totally different coloring of the same pixel after alpha blending. has such inconsistency in consecutive frames ever been observed in your experiments?\n  - Also how the two images projected from Vb and Vf get mixed into a single image? simply averaging them?\n  - It would be interesting if you could visualize the xyt space as a normal 3D space and in the meantime, visualize the rendering of one of the frames from nearby Gaussians. That will be cool."},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"* Pioneering application of 3DGS to video compression\n* Minimal assumptions about video content, relying primarily on temporal continuity, which enables broad applicability\n* High adaptability with various 3DGS models, opening promising research directions\n* Demonstrated practical utility through flexible bitrate-availability tradeoffs, particularly valuable for streaming applications"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"I do really like the idea of the paper yet the evaluation section could be further polished.\n* Limited comparison baseline, with NeRV as the only reference. The evaluation would benefit from comparisons with other neural video compression methods\n* Evaluation only focus on low bitrate/bpp scenarios, while contemporary neural compression methods (e.g., VCT, DCVC) typically demonstrate performance at higher bitrates with PSNR values exceeding 34 on the UVG dataset."}},"nonreaders":[],"tmdate":1732865275193,"tcdate":1730581494254,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission10656/Reviewer_JGLW"],"signatures":["ICLR.cc/2025/Conference/Submission10656/Reviewer_JGLW"],"forum":"JbRM5QKRDd","number":3,"license":"CC BY 4.0","cdate":1730581494254,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission10656/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732865275193,"domain":"ICLR.cc/2025/Conference","replyto":"JbRM5QKRDd","id":"1NgibnLnCb","forumContent":{"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video Compression","3D Gaussian Splatting","Entropy Coding"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"3D Gaussian Splatting (3DGS) has witnessed its rapid development in novel view synthesis, which attains high quality reconstruction and real-time rendering. At the same time, there is still a gap before implicit neural representation (INR)  can become a practical compressor due to the lack of stream decoding and real-time frame reconstruction on consumer-grade hardware. It remains a question whether the fast rendering and partial parameter decoding characteristics of 3DGS are applicable to video compression. To address these challenges, we propose a Toast-like Sliding Window (TSW) orthographic projection for converting any 3D Gaussian model into a video representation model. This method efficiently represents video by leveraging temporal redundancy through a sliding window approach. Additionally, the converted model is inherently stream-decodable and offers a higher rendering frame rate compared to INR methods. Building on TSW, we introduce an end-to-end trainable video compression method, GSVC, which employs deformable Gaussian representation and optical flow guidance to capture dynamic content in videos. Experimental results demonstrate that our method effectively transforms a 3D Gaussian model into a practical video compressor.  GSVC further achieves better rate-distortion performance than NeRV on the UVG dataset, while achieving higher frame reconstruction speed (+30%~40% fps) and stream decoding. Code is available at [Github](https://github.com/actcwlf/GSVC)"},"_bibtex":{"value":"@inproceedings{\nliu2025an,\ntitle={An Exploration with Entropy Constrained 3D Gaussians for 2D Video Compression},\nauthor={Xiang Liu and Bin Chen and Zimo Liu and Yaowei Wang and Shu-Tao Xia},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=JbRM5QKRDd}\n}"},"title":{"value":"An Exploration with Entropy Constrained 3D Gaussians for 2D Video Compression"},"pdf":{"value":"/pdf/dd2ad454120d5409842b2c4c893de22844fdec79.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"liu|an_exploration_with_entropy_constrained_3d_gaussians_for_2d_video_compression"},"authorids":{"value":["~Xiang_Liu11","~Bin_Chen4","~Zimo_Liu1","~Yaowei_Wang1","~Shu-Tao_Xia1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Xiang Liu","Bin Chen","Zimo Liu","Yaowei Wang","Shu-Tao Xia"]}},"version":2},{"content":{"summary":{"value":"This paper proposes Vidmuse, a video-to-music generation framework that generates high-fidelity music in sync with visual content. The authors also propose a large-scale video-to-music dataset containing 360k video-music pairs and a new benchmark V2M-bench. The proposed framework outperforms several previous methods both on subjective and objective metrics."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"My major concerns are listed in the weaknesses part mentioned above, and I only have minor questions here.\n\n1. For the dataset composition, there are 400K videos derived from YouTube and IMDB, what is the proportion? What kind of query set is adopted to retrieve the videos?   \n\n2. Why does the model perform worse when using MusicGen-large as the decoder? In the manuscript, it says 'this discrepancy can be partly attributed to limited GPU resources', can the model be trained using some parameter-efficient training strategy such as LoRA?  \n\n3. For the model architecture, why the music token decoder is involved in training considering that the vanilla MusicGen is able to generate high-fidelity music? Maybe adopting a trainable linear projection layer to the decoder could significantly reduce the model parameter and solve the training difficulty of MusicGen-large.\n\n4. Table 4 is overlapped with Table 5, please consider adjusting the table spacing."},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. Video-music paired datasets are scarce. The proposed high-quality, large-scale dataset benefits the community. The authors also design a reasonable and effective coarse-to-fine filtering pipeline to ensure data quality. The proposed benchmark also helps the validation of video-to-music models.\n\n2. The proposed framework is intuitive and easy to understand. Incorporating several pretrained models (Clip, Encodec, and MusicGen transformer), the proposed method achieves state-of-the-art performance on several metrics.\n\n3. The writing is clear and easy to follow."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. I am curious about the role of the video-to-music generation task. Though some previous advances tackle the task of video-to-music generation, additional constraints are attached to these models to make them more applicable. For example, some previous works [1-6] explore the rhythm synchronization of music and video, which can generate musical soundtracks with high audio-visual rhythm correspondence [1-5], and some other previous advances generate background music with corresponding emotional responses [7] or combine the music with additional audio effects [8]. However, the proposed model seems to only be able to generate semantic-matched music, which can be easily fulfilled in a training-free way, especially considering the proposed method directly leverages the pre-trained MusicGen as the music generator. There are at least three ways to achieve a similar goal: 1). Use some video-music model (such as M2UGen [9]) to generate musical captions and then leverage MusicGen to generate semantic-matched background music. 2) Use a video captioner to generate video captions and transform them into musical captions based on its semantic information using LLM, and then leverage MusicGen to generate semantic-matched background music. 3) Use Imagebind-av, the very same model that the authors use to construct the dataset, to retrieve music with the same semantics as the visual contents, and use music captioner to generate music captions, then leverage MusicGen to generate semantic-matched background music. In other words, generating semantic-matched music, especially leveraging several existing modules, seems to be an unnecessary need, which can be solved in a training-free manner, using almost the same pretrained models. From another perspective, a good soundtrack for a given video should respond timely to the semantic change in the visual contents, yet I cannot find any explicit control module in the model architecture, nor the musical rhythm change in the provided demos. What will the music be like when the video's rhythm of the former part is rapid and enthusiastic, yet suddenly becomes slow and sad in the latter part? Consequently, the restricted applicability of the proposed model significantly diminishes the paper's contribution.  \n\n2. The model architecture is trivial. Clip is used for visual encoding, Encodec is utilized for audio codec, and MusicGen is used for music generation. That is to say, only the long-short-term visual module is the newly proposed module, while it is constructed by several attention-based integration and fusion blocks. The entire framework is more likely to be a successful industrial product rather than a highlighted research finding.  \n\n3. The experiments are insufficient. The authors only conduct experiments on some weak baseline methods. For example, VM-Net and CMT are works published 7 and 3 years ago, and M2UGen is a music-centric multi-task model that is not specifically designed for video-to-music generation. On the contrary, some newly proposed video-to-music generation methods [1-6] are not compared. Besides, have the authors tested the model's performance on other existing benchmarks, such as BGM909 [5], LORIS [3], or SymMV [6]? Experiments on more available benchmarks and comparisons with more recent advances are needed to support the authors' claim.  \n\nReference:  \n\n[1]: Zhu, Ye, et al. \"Quantized gan for complex music generation from dance videos.\" European Conference on Computer Vision. Cham: Springer Nature Switzerland, 2022.  \n[2]: Zhu, Ye, et al. \"Discrete contrastive diffusion for cross-modal music and image generation.\" arXiv preprint arXiv:2206.07771 (2022).  \n[3]: Yu, Jiashuo, et al. \"Long-term rhythmic video soundtracker.\" International Conference on Machine Learning. PMLR, 2023.  \n[4]: Su, Kun, et al. \"V2Meow: Meowing to the Visual Beat via Video-to-Music Generation.\" Proceedings of the AAAI Conference on Artificial Intelligence. Vol. 38. No. 5. 2024.  \n[5]: Li, Sizhe, et al. \"Diff-BGM: A Diffusion Model for Video Background Music Generation.\" Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2024.  \n[6] Zhuo, Le, et al. \"Video background music generation: Dataset, method and evaluation.\" Proceedings of the IEEE/CVF International Conference on Computer Vision. 2023.  \n[7] Kang, Jaeyong, Soujanya Poria, and Dorien Herremans. \"Video2music: Suitable music generation from videos using an affective multimodal transformer model.\" Expert Systems with Applications 249 (2024): 123640.  \n[8] Movie Gen: A Cast of Media Foundation Models, meta, 2024  \n[9] Liu, Shansong, et al. \"M $^{2} $ UGen: Multi-modal Music Understanding and Generation with the Power of Large Language Models.\" arXiv preprint arXiv:2311.11255 (2023)."}},"nonreaders":[],"tmdate":1731427364164,"tcdate":1730710709991,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission1045/Reviewer_uyZr"],"signatures":["ICLR.cc/2025/Conference/Submission1045/Reviewer_uyZr"],"forum":"6cGKi7FqJS","number":4,"license":"CC BY 4.0","cdate":1730710709991,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission1045/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427364164,"domain":"ICLR.cc/2025/Conference","replyto":"6cGKi7FqJS","id":"v2sc1ZdLXb","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video-to-Music Generation","Transformer"]},"supplementary_material":{"value":"/attachment/9beeb773a63ce531e2fd1badaa250607f286a9c4.zip"},"primary_area":{"value":"generative models"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"In this work, we systematically study music generation conditioned solely on the video. First, we present a large-scale dataset by collecting 360K video-music pairs, including various genres such as movie trailers, advertisements, and documentaries. Furthermore, we propose VidMuse, a simple framework for generating music aligned with video inputs. VidMuse stands out by producing high-fidelity music that is both acoustically and semantically aligned with the video. By incorporating local and global visual cues, VidMuse enables the creation of coherent music tracks that consistently match the video content through Long-Short-Term modeling. Through extensive experiments, VidMuse outperforms existing models in terms of audio quality, diversity, and audio-visual alignment."},"_bibtex":{"value":"@misc{\ntian2024vidmuse,\ntitle={VidMuse: A Simple Video-to-Music Generation Framework with Long-Short-Term Modeling},\nauthor={Zeyue Tian and Zhaoyang Liu and Ruibin Yuan and Jiahao Pan and Qifeng Liu and Xu Tan and Qifeng Chen and Wei Xue and Yike Guo},\nyear={2024},\nurl={https://openreview.net/forum?id=6cGKi7FqJS}\n}"},"title":{"value":"VidMuse: A Simple Video-to-Music Generation Framework with Long-Short-Term Modeling"},"pdf":{"value":"/pdf/16417a40922983f089cc73c28dba6f7dbb0713da.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"tian|vidmuse_a_simple_videotomusic_generation_framework_with_longshortterm_modeling"},"authorids":{"value":["~Zeyue_Tian2","~Zhaoyang_Liu1","~Ruibin_Yuan1","~Jiahao_Pan1","~Qifeng_Liu1","~Xu_Tan1","~Qifeng_Chen1","~Wei_Xue5","~Yike_Guo1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Zeyue Tian","Zhaoyang Liu","Ruibin Yuan","Jiahao Pan","Qifeng Liu","Xu Tan","Qifeng Chen","Wei Xue","Yike Guo"]}},"version":2},{"content":{"venue":{"value":"CVPR 2024"},"pdf":{"value":"https://ieeexplore.ieee.org/iel8/10654794/10654797/10655313.pdf"},"venueid":{"value":"dblp.org/conf/CVPR/2024"},"paperhash":{"value":"zou|languageaware_visual_semantic_distillation_for_video_question_answering"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Bo_Zou:","~Chao_Yang3","~Yu_Qiao1","https://dblp.org/search/pid/api?q=author:Chengbin_Quan:","~Youjian_Zhao1"]},"html":{"value":"https://doi.org/10.1109/CVPR52733.2024.02560"},"_bibtex":{"value":"@inproceedings{DBLP:conf/cvpr/Zou00QZ24,\n  author={Bo Zou and Chao Yang and Yu Qiao and Chengbin Quan and Youjian Zhao},\n  title={Language-aware Visual Semantic Distillation for Video Question Answering},\n  year={2024},\n  cdate={1704067200000},\n  pages={27103-27113},\n  url={https://doi.org/10.1109/CVPR52733.2024.02560},\n  booktitle={CVPR},\n  crossref={conf/cvpr/2024}\n}\n"},"abstract":{"value":"Significant progress in video question answering (VideoQA) have been made thanks to thriving large image-language pretraining frameworks. Although image-language models can efficiently represent both video and language branches, they typically employ goal-free vision perception and do not interact vision with language well during the answer generation, thus omitting crucial visual cues. In this paper, we are inspired by the human recognition and learning pattern and propose VideoDistill, a framework with language-aware (i.e., goal-driven) behavior in both vision perception and answer generation. VideoDistill generates answers only from question-related visual embeddings and follows a thinking-observing-answering approach that closely resembles human behavior, distinguishing it from previous research. Specifically, we develop a language-aware gating mechanism to replace the standard cross-attention, avoiding language's direct fusion into visual representations. We incorporate this mechanism into two key components of the entire framework. The first component is a differentiable sparse sampling module, which selects frames containing the necessary dynamics and semantics relevant to the questions. The second component is a vision refinement module that merges existing spatial-temporal attention layers to ensure extracting multi-grained visual semantics associated with the questions. We conduct evaluations on various challenging video question-answering benchmarks, and VideoDistill achieves state-of-the-art performance in both general and long-form VideoQA datasets. In Addition, we verify that VideoDistill can effectively alleviate the utilization of language shortcut solutions in the EgoTaskQA dataset."},"title":{"value":"Language-aware Visual Semantic Distillation for Video Question Answering"},"authors":{"value":["Bo Zou","Chao Yang","Yu Qiao","Chengbin Quan","Youjian Zhao"]}},"tmdate":1762521058205,"pdate":1704067200000,"tcdate":1731488166184,"writers":["~"],"signatures":["~Yu_Qiao1"],"forum":"XD7VqVdNiW","license":"CC BY-SA 4.0","number":218622,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1762521058205,"domain":"DBLP.org","id":"XD7VqVdNiW","version":2},{"content":{"venue":{"value":"CoRR 2024"},"pdf":{"value":"http://arxiv.org/pdf/2404.00973v1"},"venueid":{"value":"dblp.org/journals/CORR/2024"},"paperhash":{"value":"zou|videodistill_languageaware_vision_distillation_for_video_question_answering"},"authorids":{"value":["~Bo_Zou3","~Chao_Yang3","~Yu_Qiao1","https://dblp.org/search/pid/api?q=author:Chengbin_Quan:","https://dblp.org/search/pid/api?q=author:Youjian_Zhao:"]},"html":{"value":"https://doi.org/10.48550/arXiv.2404.00973"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2404-00973,\n  publtype={informal},\n  author={Bo Zou and Chao Yang and Yu Qiao and Chengbin Quan and Youjian Zhao},\n  title={VideoDistill: Language-aware Vision Distillation for Video Question Answering},\n  year={2024},\n  cdate={1704067200000},\n  journal={CoRR},\n  volume={abs/2404.00973},\n  url={https://doi.org/10.48550/arXiv.2404.00973}\n}\n"},"abstract":{"value":"Significant advancements in video question answering (VideoQA) have been made thanks to thriving large image-language pretraining frameworks. Although these image-language models can efficiently represent both video and language branches, they typically employ a goal-free vision perception process and do not interact vision with language well during the answer generation, thus omitting crucial visual cues. In this paper, we are inspired by the human recognition and learning pattern and propose VideoDistill, a framework with language-aware (i.e., goal-driven) behavior in both vision perception and answer generation process. VideoDistill generates answers only from question-related visual embeddings and follows a thinking-observing-answering approach that closely resembles human behavior, distinguishing it from previous research. Specifically, we develop a language-aware gating mechanism to replace the standard cross-attention, avoiding language's direct fusion into visual representations. We incorporate this mechanism into two key components of the entire framework. The first component is a differentiable sparse sampling module, which selects frames containing the necessary dynamics and semantics relevant to the questions. The second component is a vision refinement module that merges existing spatial-temporal attention layers to ensure the extraction of multi-grained visual semantics associated with the questions. We conduct experimental evaluations on various challenging video question-answering benchmarks, and VideoDistill achieves state-of-the-art performance in both general and long-form VideoQA datasets. In Addition, we verify that VideoDistill can effectively alleviate the utilization of language shortcut solutions in the EgoTaskQA dataset."},"title":{"value":"VideoDistill: Language-aware Vision Distillation for Video Question Answering"},"authors":{"value":["Bo Zou","Chao Yang","Yu Qiao","Chengbin Quan","Youjian Zhao"]}},"tmdate":1739962066715,"pdate":1704067200000,"tcdate":1727614024125,"writers":["~"],"signatures":["~Bo_Zou3"],"forum":"tcJrxrglv4","license":"CC BY-SA 4.0","number":106319,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1739962066715,"domain":"DBLP.org","id":"tcJrxrglv4","version":2},{"content":{"summary":{"value":"This paper proposes an Egocentric Video QA benchmark. The authors say that many questions in current datasets don’t really test whether a model understands the video from the recorder’s point of view. Instead, models can often guess the answer using easy “shortcuts,” like recognizing the general scene or common actions. This gives a false impression that Vision-Language Models (VLMs) are better at first-person understanding than they actually are.\n\nTo fix this, the authors make three main contributions: Defining good egocentric questions, creating a question-generation pipeline, and building the EgoQuestions Benchmark. Using this pipeline, they produce EgoQuestions, a benchmark containing 2,500 carefully made question–answer pairs.\n\nThey then test several Vision-Language Models on both the new and existing datasets. The results show that current models perform much worse on the new, stricter questions. There’s about a 10% performance drop compared to the older, easier benchmarks. This confirms that existing benchmarks overestimate how well current models truly understand first-person videos."},"soundness":{"value":1},"confidence":{"value":4},"questions":{"value":"- The definitions of the data subsets—“SPO (because we provide the Subject and Predicate in the question while inquiring about the Object), SOAP, and SPAO”—are unclear. The distinctions among these subsets are not well explained, and it is not evident why they differ. This section requires significant clarification and improvement in writing quality.\n\n- Did you perform any human validation on a subset of the LLM-as-judge's scores? How can we be confident that the judge itself isn't systematically scoring the \"flawed\" F2 questions more leniently, thereby contributing to the inflated gap? \n\n- The current question templates (SPO, SOAP, SPAO) are focused on grounding actions and objects. Do you believe your three principles and crafting pipeline could be extended to generate more complex, abstract egocentric questions, such as those reasoning about the recorder's intentions (\"Why am I lifting the pan?\") or beliefs?\n\n- Could you provide significantly more qualitative examples of your benchmark data?\n\n- How do the SOTA models fail on your T1 questions? Could you analyze some examples? For instance, when asked an SPO question (\"What am I picking up?\"), do models tend to name a plausible but incorrect object from the scene, hallucinate an object, or simply state they don't know?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":1},"contribution":{"value":2},"strengths":{"value":"- The paper provides insights about how to better build egocentric video understanding benchmarks by eliminating the shortcut in questions."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Limited Question Scope: The benchmark does not explicitly assess more complex reasoning abilities, such as those involving long-term intentions, social dynamics, or the causal consequences of the recorder’s actions.\n\n- Reliance on LLM-as-Judge: While using an LLM to evaluate open-ended responses is a common practice, it remains imperfect. The observed 10% $\\Delta_{acc}$ performance gap might partly result from the judge model itself finding the F2 questions (which include shortcuts) easier to assess. Incorporating more reliable multiple-choice questions into the benchmark is recommended.\n\n- Writing Quality: The paper’s writing requires improvement, as several key concepts—such as the SPO, SOAP, and SPAO subsets—are insufficiently defined.\n\n- Insufficient Model Coverage: The number of evaluated models is too limited for a benchmark paper, which reduces the generality and robustness of the findings.\n\n- Missing qualitative study: The qualitative analysis of benchmark data and wrong answers by different models is missing."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762916262835,"tcdate":1761627997729,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission2515/Reviewer_3o3B"],"signatures":["ICLR.cc/2026/Conference/Submission2515/Reviewer_3o3B"],"forum":"ym7L1by6iO","number":3,"license":"CC BY 4.0","cdate":1761627997729,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission2515/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762916262835,"domain":"ICLR.cc/2026/Conference","replyto":"ym7L1by6iO","id":"7rNkgreBAe","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Egocentric Vision"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"A thorough understanding of models’ egocentric capabilities is crucial for robotics, autonomous driving, smart glasses, etc.\nEgocentric VideoQA aims to assess models’ understanding of first‑person videos, \nbut existing benchmarks often include questions that do not reliably probe recorder‑centric reasoning. \nUsing these datasets to train and evaluate models can obscure true model capabilities and reduce the value of curated egocentric data.\nTo address this, we define egocentric questions and propose three clear principles: a question should focus on the video recorder and their activities;\nit must avoid shortcut cues that allow answers via generic scene or action recognition (e.g., simultaneously naming an action and its object); while intentions and attributes may serve as shortcuts for actions and objects, those that require understanding of the recorder’s perspective will not. \nGuided by these principles, we build a checking pipeline to filter existing QA pairs and a crafting pipeline to generate valid egocentric questions. We release EgoQuestions, a benchmark of 2,500 curated egocentric QA instances created with our pipeline, and evaluate several proprietary and open‑source VLMs. \nResults reveal substantial room for improvement in current models’ egocentric capabilities and a clear performance gap  (about 10%) between egocentric questions that adhere to our principles and flawed alternatives, \ndemonstrating existing egocentric benchmarks tend to overrate models’ first-person capabilities,\nand the need for rigorously designed egocentric benchmarks\nto more accurately assess models’ first-person vision capabilities."},"_bibtex":{"value":"@misc{\ndong2025egoquestions,\ntitle={EgoQuestions: Crafting Egocentric Questions for Egocentric Video Question Answering},\nauthor={Xinzhi Dong and Fang-Lue Zhang and Meng-Hao Guo},\nyear={2025},\nurl={https://openreview.net/forum?id=ym7L1by6iO}\n}"},"title":{"value":"EgoQuestions: Crafting Egocentric Questions for Egocentric Video Question Answering"},"pdf":{"value":"/pdf/da1642200236a767ca2d63afab12b01a87b546b9.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"dong|egoquestions_crafting_egocentric_questions_for_egocentric_video_question_answering"},"authorids":{"value":["~Xinzhi_Dong1","~Fang-Lue_Zhang1","~Meng-Hao_Guo1"]},"authors":{"value":["Xinzhi Dong","Fang-Lue Zhang","Meng-Hao Guo"]}},"version":2},{"content":{"TLDR":{"value":"We investigate the data distributional properties that drive in-context learning in vision-language models for videos."},"venue":{"value":"Video-Langauge Models Poster"},"pdf":{"value":"/pdf/f800aeac2c8a7e485527c6c23dfdd6b2bcc8ef20.pdf"},"keywords":{"value":["cross-modal pretraining","video processing","multimodality"]},"venueid":{"value":"NeurIPS.cc/2024/Workshop/Video-Langauge_Models"},"paperhash":{"value":"yu|eliciting_incontext_learning_in_visionlanguage_models_for_videos_through_curated_data_distributional_properties"},"authorids":{"value":["~Keunwoo_Peter_Yu1","~Zheyuan_Zhang4","~Fengyuan_Hu2","~Shane_Storks1","~Joyce_Chai2"]},"abstract":{"value":"A major reason behind the recent success of large language models (LLMs) is their $\\textit{in-context learning}$ capability, which makes it possible to rapidly adapt them to downstream text-based tasks by prompting them with a small number of relevant demonstrations. While large vision-language models (VLMs) have recently been developed for tasks requiring both text and images, they largely lack in-context learning over visual information, especially in understanding and generating text about videos. In this work, we implement $\\textbf{E}$mergent $\\textbf{I}$n-context $\\textbf{Le}$arning on $\\textbf{V}$ideos ($\\textbf{EILeV}$), a novel training paradigm that induces in-context learning over video and text by capturing key properties of pre-training data found by prior work to be essential for in-context learning in transformers. In our experiments, we show that $\\textbf{EILeV}$-trained models outperform other off-the-shelf VLMs in few-shot video narration for novel, rare actions. Furthermore, we demonstrate that these key properties of bursty distributions, skewed marginal distributions, and dynamic meaning each contribute to varying degrees to VLMs' in-context learning capability in narrating procedural videos. Our results, analysis, and $\\textbf{EILeV}$-trained models yield numerous insights about the emergence of in-context learning over video and text, creating a foundation for future work to optimize and scale VLMs for open-domain video understanding and reasoning."},"_bibtex":{"value":"@inproceedings{\nyu2025eliciting,\ntitle={Eliciting In-Context Learning in Vision-Language Models for Videos Through Curated Data Distributional Properties},\nauthor={Keunwoo Peter Yu and Zheyuan Zhang and Fengyuan Hu and Shane Storks and Joyce Chai},\nbooktitle={Workshop on Video-Language Models @ NeurIPS 2024},\nyear={2025},\nurl={https://openreview.net/forum?id=wLFmFPLlo1}\n}"},"title":{"value":"Eliciting In-Context Learning in Vision-Language Models for Videos Through Curated Data Distributional Properties"},"track":{"value":"Long Paper Track (up to 9 pages)"},"authors":{"value":["Keunwoo Peter Yu","Zheyuan Zhang","Fengyuan Hu","Shane Storks","Joyce Chai"]}},"tmdate":1736861079856,"pdate":1730081751851,"tcdate":1725042877720,"writers":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission4/Authors"],"signatures":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission4/Authors"],"forum":"wLFmFPLlo1","license":"CC BY 4.0","number":4,"cdate":1725042877720,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Post_Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Edit","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission4/-/Camera-Ready_Revision"],"mdate":1736861079856,"odate":1736861079747,"domain":"NeurIPS.cc/2024/Workshop/Video-Langauge_Models","id":"wLFmFPLlo1","version":2},{"content":{"venue":{"value":"CoRR 2024"},"pdf":{"value":"http://arxiv.org/pdf/2411.06096v1"},"venueid":{"value":"dblp.org/journals/CORR/2024"},"paperhash":{"value":"liu|zhoblimp_a_systematic_assessment_of_language_models_with_linguistic_minimal_pairs_in_chinese"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Yikang_Liu:","https://dblp.org/search/pid/api?q=author:Yeting_Shen:","https://dblp.org/search/pid/api?q=author:Hongao_Zhu:","https://dblp.org/search/pid/api?q=author:Lilong_Xu:","https://dblp.org/search/pid/api?q=author:Zhiheng_Qian:","~Siyuan_Song1","https://dblp.org/search/pid/api?q=author:Kejia_Zhang:","https://dblp.org/search/pid/api?q=author:Jialong_Tang:","https://dblp.org/search/pid/api?q=author:Pei_Zhang_0011:","https://dblp.org/search/pid/api?q=author:Baosong_Yang:","https://dblp.org/search/pid/api?q=author:Rui_Wang_0015:","~Hai_Hu1"]},"html":{"value":"https://doi.org/10.48550/arXiv.2411.06096"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2411-06096,\n  publtype={informal},\n  author={Yikang Liu and Yeting Shen and Hongao Zhu and Lilong Xu and Zhiheng Qian and Siyuan Song and Kejia Zhang and Jialong Tang and Pei Zhang and Baosong Yang and Rui Wang and Hai Hu},\n  title={ZhoBLiMP: a Systematic Assessment of Language Models with Linguistic Minimal Pairs in Chinese},\n  year={2024},\n  cdate={1704067200000},\n  journal={CoRR},\n  volume={abs/2411.06096},\n  url={https://doi.org/10.48550/arXiv.2411.06096}\n}\n"},"abstract":{"value":"Whether and how language models (LMs) acquire the syntax of natural languages has been widely evaluated under the minimal pair paradigm. However, a lack of wide-coverage benchmarks in languages other than English has constrained systematic investigations into the issue. Addressing it, we first introduce ZhoBLiMP, the most comprehensive benchmark of linguistic minimal pairs for Chinese to date, with 118 paradigms, covering 15 linguistic phenomena. We then train 20 LMs of different sizes (14M to 1.4B) on Chinese corpora of various volumes (100M to 3B tokens) and evaluate them along with 14 off-the-shelf LLMs on ZhoBLiMP. The overall results indicate that Chinese grammar can be mostly learned by models with around 500M parameters, trained on 1B tokens with one epoch, showing limited benefits for further scaling. Most (N=95) linguistic paradigms are of easy or medium difficulty for LMs, while there are still 13 paradigms that remain challenging even for models with up to 32B parameters. In regard to how LMs acquire Chinese grammar, we observe a U-shaped learning pattern in several phenomena, similar to those observed in child language acquisition."},"title":{"value":"ZhoBLiMP: a Systematic Assessment of Language Models with Linguistic Minimal Pairs in Chinese"},"authors":{"value":["Yikang Liu","Yeting Shen","Hongao Zhu","Lilong Xu","Zhiheng Qian","Siyuan Song","Kejia Zhang","Jialong Tang","Pei Zhang","Baosong Yang","Rui Wang","Hai Hu"]}},"tmdate":1764337386468,"pdate":1704067200000,"tcdate":1749488273535,"writers":["~"],"signatures":["~Siyuan_Song1"],"forum":"PrByOBxKVi","license":"CC BY-SA 4.0","number":558633,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1764337386468,"domain":"DBLP.org","id":"PrByOBxKVi","version":2},{"content":{"summary":{"value":"The paper proposes a 4D generation method that aims to generate 4D content from a monocular video. A video-to-multi-view-video diffusion model is presented to create multi-view videos given a monocular video, a text prompt, and a sequence of camera poses. The trained multi-view-video diffusion model is leveraged to optimize 4D representation, i.e., dynamic NeRF.  In addition, 4D-aware SDS loss and an anchor loss are introduced to train dynamic NeRF. Experimental results show the proposed method achieves the best performance, compared with state-of-the-art methods."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"Please refer to my comments above"},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The paper is well-written and easy to follow.\n\n2. A multi-view-video diffusion model is presented to generate multi-view videos from a monocular video\n\n3. The paper addresses an interesting problem, and  4D generation significantly impacts various applications."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Some technical details are unclear. The paper builds a multi-view-video diffusion model by inserting a learnable motion module into ImageDream. The learnable motion module is critical to the proposed method. However, the paper does not provide detailed information about the motion module, such as the architecture and the layer information. Without this information, it is difficult to reproduce the proposed method.\n\n2.  The paper only uses 996 training data to train the multi-view-video diffusion model or the motion module, while the training takes two days on 16 NVIDIA Tesla A100 GPUs. Does such a small data set and such an extensive training cost lead to significant overfitting? How many parameters are in the motion modules?\n\n3.  Instead of the input monocular video, the anchor loss chooses a monocular video generated by the presented multi-view-video diffusion model as an anchor, due to the difficulty in estimating the camera pose of the input video. Why not use all videos generated by the multi-view-video diffusion? Would this operation degrade the 4D generation performance?  In addition, the input video typically has better quality than the generated one.\n\n4. Table 2 shows that using ImageDream achieves better CLIP-I than using the present multi-view-video diffusion model. Could the authors provide more explanations?"},"limitations":{"value":"The paper provides the limitations and societal impact of the proposed work"}},"nonreaders":[],"tmdate":1730879347928,"tcdate":1721811663863,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission9911/Reviewer_ioTR"],"signatures":["NeurIPS.cc/2024/Conference/Submission9911/Reviewer_ioTR"],"forum":"SFk7AMpyhx","number":4,"license":"CC BY 4.0","cdate":1721811663863,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission9911/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730879347928,"domain":"NeurIPS.cc/2024/Conference","replyto":"SFk7AMpyhx","id":"iuWlyHssyl","forumContent":{"TLDR":{"value":"we present a novel 4D generation pipeline, 4Diffusion, to create high-quality spatial-temporally consistent 4D content from a monocular video."},"venue":{"value":"NeurIPS 2024 poster"},"keywords":{"value":["Diffusion Model","4D Generation","NeRF"]},"supplementary_material":{"value":"/attachment/93d614b0ed825bfaa519b19418cdb02188b26762.zip"},"primary_area":{"value":"generative_models"},"abstract":{"value":"Current 4D generation methods have achieved noteworthy efficacy with the aid of advanced diffusion generative models. However, these methods lack multi-view spatial-temporal modeling and encounter challenges in integrating diverse prior knowledge from multiple diffusion models, resulting in inconsistent temporal appearance and flickers. In this paper, we propose a novel 4D generation pipeline, namely $\\textbf{4Diffusion}$, aimed at generating spatial-temporally consistent 4D content from a monocular video. We first design a unified diffusion model tailored for multi-view video generation by incorporating a learnable motion module into a frozen 3D-aware diffusion model to capture multi-view spatial-temporal correlations. After training on a curated dataset, our diffusion model acquires reasonable temporal consistency and inherently preserves the generalizability and spatial consistency of the 3D-aware diffusion model. Subsequently, we propose 4D-aware Score Distillation Sampling loss, which is based on our multi-view video diffusion model, to optimize 4D representation parameterized by dynamic NeRF. This aims to eliminate discrepancies arising from multiple diffusion models, allowing for generating spatial-temporally consistent 4D content. Moreover, we devise an anchor loss to enhance the appearance details and facilitate the learning of dynamic NeRF. Extensive qualitative and quantitative experiments demonstrate that our method achieves superior performance compared to previous methods."},"_bibtex":{"value":"@inproceedings{\nzhang2024diffusion,\ntitle={4Diffusion: Multi-view Video Diffusion Model for 4D Generation},\nauthor={Haiyu Zhang and Xinyuan Chen and Yaohui Wang and Xihui Liu and Yunhong Wang and Yu Qiao},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=SFk7AMpyhx}\n}"},"title":{"value":"4Diffusion: Multi-view Video Diffusion Model for 4D Generation"},"pdf":{"value":"/pdf/c82cad9ac993b7e7299eb3787716e0166d8f972c.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"zhang|4diffusion_multiview_video_diffusion_model_for_4d_generation"},"authorids":{"value":["~Haiyu_Zhang1","~Xinyuan_Chen1","~Yaohui_Wang1","~Xihui_Liu1","~Yunhong_Wang1","~Yu_Qiao1"]},"authors":{"value":["Haiyu Zhang","Xinyuan Chen","Yaohui Wang","Xihui Liu","Yunhong Wang","Yu Qiao"]}},"version":2},{"content":{"summary":{"value":"The paper presents multi-modal pretraining approach with modalities N=5 (video, text, audio, depth, infrared) by using language as bind across different modalities. A frozen text encoder from a pretrained VL model is used as the feature extractor for the text modality and aligned with other modalities (pair-wise) using contrastive loss. It also introduces a dataset called VIDAL-10M with 10M data pairs from VL, DL, IL,and AL. The dataset and method is evaluated on standard retrieval benchmarks to show the effectiveness of the pretraining data as well as the technique."},"presentation":{"value":"3 good"},"contribution":{"value":"3 good"},"soundness":{"value":"3 good"},"strengths":{"value":"1) The paper attempts to learn a unified embedding space for 5 modalities where the modalities are guided by language during pre-training. Such an embedding space can be very useful for tasks involving: i) multi-modal data for ex: video containing audio, ii) tasks where paired data is not available, for ex: Video-Infrared, Video-depth etc \n\n2) The paper introduces a dataset with 10M paired data from AL, VL, IL and DL which is important for driving research in the multimodal learning area as many of the real-world applications contain multimodal data. It follows a careful approach by leveraging existing vision and language models (OFA, mPLUG-owl, chatgpt etc) to collect a balanced (in topic) and diverse (in semantic) data-pairs.\n\n3) The introduced dataset and the pretrained model is shown to be useful for:\na) cross-modality video retrieval task where it outperforms its counterparts (ImageBind, CLIP-straight, CLIP4clip).\nb) AL, DL, IL zero-shot classification tasks.\nThis shows that the model has learned good representations in the joint embedding space."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1) It is not clear from the text or Table2, the size of the pretraining data used for MSR-VTT and MSVD datasets. For a fair comparison, all methods should be pretrained with same amount of data but here CLIP-Straight is trained with WIT400M only (initialized from CLIP but no fine-tuning), CLIP4clip is trained with WIT400M+HT100M-380k whereas the proposed technique (although CLIP4clip technique is used) is pretrained with WIT400M+VIDAL-10M. It would be a fair comparison if all methods use similar sized data, i.e what would be the performance of other technique like CLIP4clip if additional data (not VIDAL-10M) is used for training.\n\n2) One of the goals of learning a model from multimodal data is  that the data can use all available modalities to learn stronger representations but there are no experiments to demonstrate this, for ex: instead of using just video -> text retrieval, it would interesting to show that video+audio -> text retrieval has better performance. \n\n3) There are few other advantages of multimodal learning in situations where:\na) one of the modalities is corrupted \nb) one of the modalities has some weaknesses (videos taken in the dark, OR audio from multiple sources) \nc) one of the modality undergoes a domain change while the other doesn't (eg: videos under weather changes etc) \nbut none of these has been addressed in this paper. It would be interesting to see results on at least one of the above scenarios.\n\n4) It would also be interesting to see an experiment where the model is evaluated on retrieval task where the modalities doesn't contain text. For ex (video<->audio, video<->infrared). This will evaluate the quality of learned representations."},"confidence":{"value":"4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work."},"questions":{"value":"I would like authors to discuss all the points described above.\n\nFinal Rating: After reading the rebuttal and comments from other reviewers, I have decided not to change the score."},"rating":{"value":"6: marginally above the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1701158299031,"tcdate":1699292484551,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission4730/Reviewer_Bowz"],"signatures":["ICLR.cc/2024/Conference/Submission4730/Reviewer_Bowz"],"forum":"QmZKc7UZCy","number":2,"license":"CC BY 4.0","cdate":1699292484551,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission4730/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1701158299031,"domain":"ICLR.cc/2024/Conference","replyto":"QmZKc7UZCy","id":"3ylHEwF1ZG","forumContent":{"venue":{"value":"ICLR 2024 poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["LanguageBind","Multi-modal Pretraining","Multi-modal Dataset"]},"primary_area":{"value":"representation learning for computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"The video-language (VL) pretraining has achieved remarkable improvement in multiple downstream tasks. However, the current VL pretraining framework is hard to extend to multiple modalities (N modalities, N ≥ 3) beyond vision and language. We thus propose LanguageBind, taking the language as the bind across different modalities because the language modality is well-explored and contains rich semantics. Specifically, we freeze the language encoder acquired by VL pretraining and then train encoders for other modalities with contrastive learning. As a result, all modalities are mapped to a shared feature space, implementing multi-modal semantic alignment. While LanguageBind ensures that we can extend VL modalities to N modalities, we also need a high-quality dataset with alignment data pairs centered on language. We thus propose VIDAL-10M with 10 Million data with Video, Infrared, Depth, Audio and their corresponding Language. In our VIDAL-10M, all videos are from short video platforms with complete semantics rather than truncated segments from long videos, and all the video, depth, infrared, and audio modalities are aligned to their textual descriptions. LanguageBind has achieved superior performance on a wide range of 15 benchmarks covering video, audio, depth, and infrared. Moreover, multiple experiments have provided evidence for the effectiveness of LanguageBind in achieving indirect alignment and complementarity among diverse modalities."},"_bibtex":{"value":"@inproceedings{\nzhu2024languagebind,\ntitle={LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment},\nauthor={Bin Zhu and Bin Lin and Munan Ning and Yang Yan and Jiaxi Cui and WANG HongFa and Yatian Pang and Wenhao Jiang and Junwu Zhang and Zongwei Li and Cai Wan Zhang and Zhifeng Li and Wei Liu and Li Yuan},\nbooktitle={The Twelfth International Conference on Learning Representations},\nyear={2024},\nurl={https://openreview.net/forum?id=QmZKc7UZCy}\n}"},"title":{"value":"LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment"},"pdf":{"value":"/pdf/598284cab1807220a6ff084736e65e7ba5baea3b.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference"},"paperhash":{"value":"zhu|languagebind_extending_videolanguage_pretraining_to_nmodality_by_languagebased_semantic_alignment"},"authorids":{"value":["~Bin_Zhu6","~Bin_Lin1","~Munan_Ning1","~Yang_Yan3","~Jiaxi_Cui1","~WANG_HongFa1","~Yatian_Pang1","~Wenhao_Jiang1","~Junwu_Zhang2","~Zongwei_Li1","~Cai_Wan_Zhang1","~Zhifeng_Li5","~Wei_Liu3","~Li_Yuan2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Bin Zhu","Bin Lin","Munan Ning","Yang Yan","Jiaxi Cui","WANG HongFa","Yatian Pang","Wenhao Jiang","Junwu Zhang","Zongwei Li","Cai Wan Zhang","Zhifeng Li","Wei Liu","Li Yuan"]}},"version":2},{"content":{"TLDR":{"value":"Shortcut World Models predict environment dynamics across multiple horizons in a single forward pass, achieving 33-64x speedup with up to 50% lower error than autoregressive rollout on discontinuous dynamics."},"venue":{"value":"ICLR 2026 Workshop World Models"},"pdf":{"value":"/pdf/36b75971e783993697d7a4e5e3c2b79472c6a7be.pdf"},"keywords":{"value":["World Models","Multi-Horizon Prediction","Discontinuous Dynamics","Model-Based Reinforcement Learning","Learned Simulation"]},"venueid":{"value":"ICLR.cc/2026/Workshop/World_Models"},"paperhash":{"value":"lakshmanan|[tiny_paper]_shortcut_world_models_learning_to_leap_not_step"},"authorids":{"value":["~Pranav_Lakshmanan1","~Paras_Chopra1"]},"abstract":{"value":"Autoregressive world models chain single-step predictions, requiring N forward\npasses for N steps into the future. We introduce Shortcut World Models, trained\nto predict environment dynamics across multiple horizons, enabling direct leaping\nto any learned step-size in a single pass rather than iteratively stepping through intermediate states. Beyond speed, skipping intermediate predictions also improves\naccuracy: errors compound through state discontinuities in autoregressive rollout,\nbut shortcuts sidestep this accumulation entirely. At inference, adaptive chaining\ndecomposes arbitrary horizons into learned sub-steps, handling step-sizes beyond\ntraining while maximizing accuracy with minimal sacrifice in speed. On discontinuous particle dynamics, Shortcut World Models achieve 33–64× fewer forward\npasses with up to 50% lower error, demonstrating a path toward learned simulators\nand model-based planning that are both faster and more accurate."},"_bibtex":{"value":"@inproceedings{\nlakshmanan2026tiny,\ntitle={[Tiny Paper] Shortcut World Models: Learning to Leap, Not Step},\nauthor={Pranav Lakshmanan and Paras Chopra},\nbooktitle={ICLR 2026 the 2nd Workshop on World Models: Understanding, Modelling and Scaling},\nyear={2026},\nurl={https://openreview.net/forum?id=O8nGPjY7Tj}\n}"},"title":{"value":"[Tiny Paper] Shortcut World Models: Learning to Leap, Not Step"},"authors":{"value":["Pranav Lakshmanan","Paras Chopra"]}},"tmdate":1776292156543,"pdate":1772419702629,"tcdate":1770115795785,"writers":["ICLR.cc/2026/Workshop/World_Models","ICLR.cc/2026/Workshop/World_Models/Submission92/Authors"],"signatures":["ICLR.cc/2026/Workshop/World_Models/Submission92/Authors"],"forum":"O8nGPjY7Tj","license":"CC BY 4.0","number":92,"cdate":1770115795785,"readers":["everyone"],"invitations":["ICLR.cc/2026/Workshop/World_Models/-/Submission","ICLR.cc/2026/Workshop/World_Models/-/Post_Submission","ICLR.cc/2026/Workshop/World_Models/-/Edit"],"mdate":1776292156543,"odate":1772673286401,"domain":"ICLR.cc/2026/Workshop/World_Models","id":"O8nGPjY7Tj","version":2},{"content":{"summary":{"value":"The paper introduces EVER, an EEG and video emotion recognition framework that aims to make physiological signals and visual cues work in a single structured model. The authors argue that emotion recognition with only video or only EEG gives a partial view and is brittle across subjects and recording conditions, so they build a multimodal pipeline that extracts high level features from video with AdaMAE and from EEG spectrograms with AudioMamba, then fuses them through a two stage Brain anatomy aware Inter modal Hierarchical GCN. The first stage aggregates channel features into anatomically grounded regions through masked graph propagation and attention pooling. The second stage places both global modality embeddings and the region nodes in one graph and performs structured message passing so that video and EEG can exchange information while remaining connected to brain regions. A correlation based distribution alignment loss is added to reduce statistical mismatch between the two modalities. The authors further provide what is essentially a benchmark for EEG and video pair emotion recognition on three public datasets, namely MDMER, Emognition and EAV, and compare twelve representative unimodal and multimodal baselines. EVER gives the best overall numbers on the three datasets and the ablations show that both the hierarchical graph and the correlation loss are necessary."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1-The model forms the second stage graph with two global nodes and K region nodes and then directly projects every node to the label space. Why was this direct projection chosen over first producing a unified feature and then applying a task head. \n2-It would be helpful to know whether the gains come from graph reasoning or from letting regions vote directly for the label. The correlation alignment is applied to batches of unimodal embeddings. What is the effect of the batch size on this loss and did the authors try instance wise alignment or moving average statistics to make it less sensitive to the composition of a batch. \n3-In the Emognition dataset there are only four EEG channels. In this case the first stage graph is almost trivial. Can the authors report how much of the improvement on Emognition comes from the alignment loss alone. \n4-In the benchmark comparison the existing multimodal models are audio and video systems that were re purposed because audio is structurally similar to EEG. Have the authors tried to re implement these models with the same EEG spectrogram encoder used in EVER in order to separate architectural gains from encoder gains. For subject independent evaluation it is important to consider domain shift across subjects. Did the authors observe subject specific clustering of the embeddings before and after alignment.\n5-Figure 2 shows a nicer shared space, yet a quantitative measure such as class conditional Fréchet distance across subjects would make the claim stronger."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"The work targets a setting that is clearly under serviced. Existing affective computing studies often focus on video alone or on EEG alone, while audio and video fusion is better studied. Positioning the model around paired EEG and video recordings is therefore meaningful for the ICLR audience that is interested in multimodal learning and in models that can work with physiological signals. The model design is careful. The first graph stage does not simply connect electrodes according to similarity. It enforces a mask derived from standard scalp regions and then performs attention pooling inside each region, which makes the subsequent fusion less sensitive to the specific channel layout of a dataset. This is a sound way to bring domain priors into a modern backbone. The second graph stage is the part that actually unifies the modalities. By treating the video embedding, the EEG embedding and the region representations as nodes in the same graph, the model creates an explicit route for inter modality reasoning instead of the usual concatenation that appears in audio and video models. This is consistent with the motivation stated in the introduction. The distribution alignment objective is a small but useful addition. It operates on correlations rather than raw covariance and therefore avoids scale problems. The ablations confirm that alignment alone does not solve the task but in combination with the anatomy aware graph it brings the model over both unimodal baselines and over naive fusion. The experimental section is richer than what is typical for a first paper on a new multimodal pairing. The authors collect three public datasets that have time aligned EEG and video, they impose subject independent splits, and they report accuracy, unweighted average recall and weighted F1, so the reader can interpret results under both class balanced and class imbalanced regimes. On all three datasets EVER improves over the strongest single modality model, which is the right standard to use in multimodal work. The paper is technically self contained. The construction of the two adjacency masks, the attention pooling, the fusion of the two global nodes into logits and the formulation of the alignment loss are all given in sufficient detail so that an experienced reader can implement the model. The writing is clear and connects the graph construction to actual neuroscience practice on five scalp regions."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Although the paper claims a unified framework, in practice the current system is anchored to a very specific choice of backbones. Video features come from AdaMAE and EEG features come from a spectrogram that is processed by AudioMamba. The paper does not test the framework with lighter or more conventional EEG encoders such as shallow CNNs or temporal transformers trained from scratch, nor does it show whether the hierarchical graph can compensate when one modality is considerably weaker. This limits the generality of the claim. The benchmark is valuable but the size of the paired data is still modest. Emognition has slightly more than four hundred pairs. MDMER has about two point three thousand pairs. EAV is larger but its labels are discrete and balanced. Generalization to truly large scale recordings, to unlabeled synchronized streams, or to field recordings with missing frames is not studied. The model relies on discrete brain regions derived from the 10 20 system and maps channels to these regions through a fixed mask. This is reasonable for the three datasets at hand, yet cross dataset variation in channel count is already visible. The paper partially alleviates this through attention pooling, yet it does not quantify how sensitive the model is to errors in the region assignment or to setups with very few channels such as six or fewer electrodes which are common in wearable EEG devices. The correlation based alignment is motivated as a way to reduce modality discrepancy, yet there is no comparison with other simple alignment objectives such as maximum mean discrepancy or contrastive matching of the global embeddings. The current ablation only shows with and without. A stronger analysis would include parameter free alternatives to confirm that choosing correlations is important. The interpretation part could go further. Figure A1 shows that different emotions activate different regions, which is interesting, but the main paper does not link these observations to classification outcomes or to failure cases. A reviewer would expect at least one analysis of wrong predictions where the video backbone is correct and the EEG branch is not, and the other way around, to justify the structured fusion. Finally, the paper is long and dense and some implementation choices appear late in the appendix. A more compact main text would help readers who want to re implement the method quickly."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762927226092,"tcdate":1762093641383,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission17283/Reviewer_yFBv"],"signatures":["ICLR.cc/2026/Conference/Submission17283/Reviewer_yFBv"],"forum":"hga7TjP3uD","number":4,"license":"CC BY 4.0","cdate":1762093641383,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission17283/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762927226092,"domain":"ICLR.cc/2026/Conference","replyto":"hga7TjP3uD","id":"YOsnhX2Soi","forumContent":{"TLDR":{"value":"We introduce a comprehensive benchmark and a novel EEG–video fusion framework that leverages brain anatomy-aware inter-modal hierarchical GCN for robust emotion recognition."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Emotion Recognition","EEG-Video","Benchmark","Multi-modal Fusion","Graph Convolutional Network"]},"supplementary_material":{"value":"/attachment/08b75f4068f93fb715afe361aa29745e4bb8805f.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Recent studies in video- and EEG-based emotion recognition have shown notable progress. However, multi-modal emotion recognition remains largely unexplored, particularly the integration of physiological signals with video. This integration is crucial, as EEG–video fusion combines observable behavioral cues with internal neural dynamics and enables a more comprehensive and robust characterization of human emotion. To this end, we propose EVER, a novel EEG–Video Emotion Recognition framework that effectively integrates complementary information from both modalities. Specifically, EVER employs a Brain anatomy-aware Inter-modal Hierarchical Graph Convolution Network (BIH-GCN), which aggregates EEG channel features into region-level representations guided by anatomical priors. These region-level features are combined with global EEG and video embeddings to form a unified representation for emotion classification. Furthermore, we introduce a correlation-based distribution alignment loss to reconcile modality-specific embeddings and reduce cross-modal discrepancies. To provide a comprehensive evaluation, we conduct comprehensive benchmark across three public EEG-video paired datasets---Emognition, MDMER, and EAV. We evaluate 12 representative models, consisting of 5 EEG-only, 5 video-only, and 2 audio-video models, and report their performance under EEG, video, and EEG–video settings. Our benchmark highlights the strengths and limitations of both unimodal and multi-modal approaches across diverse environments. Extensive experiments demonstrate that the proposed EVER achieves state-of-the-art performance by jointly modeling behavioral cues from video and physiological responses from EEG, thereby enabling the recognition of emotional patterns unattainable by either modality alone."},"_bibtex":{"value":"@misc{\nkim2026a,\ntitle={A Unified Framework for {EEG}{\\textendash}Video Emotion Recognition with Brain Anatomy Guidance},\nauthor={JangHyun Kim and Seongro Yoon and Temo Saghinadze and Aowen Shi and Mingyun Jeong and Donghyeon Cho and Jinsun Park and Francois Bremond},\nyear={2026},\nurl={https://openreview.net/forum?id=hga7TjP3uD}\n}"},"title":{"value":"A Unified Framework for EEG–Video Emotion Recognition with Brain Anatomy Guidance"},"pdf":{"value":"/pdf/ed06f3177534209f9a3997c6268680bd74fd86ef.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"kim|a_unified_framework_for_eegvideo_emotion_recognition_with_brain_anatomy_guidance"},"authorids":{"value":["~JangHyun_Kim3","~Seongro_Yoon1","~Temo_Saghinadze1","~Aowen_Shi1","~Mingyun_Jeong1","~Donghyeon_Cho1","~Jinsun_Park1","~Francois_Bremond1"]},"authors":{"value":["JangHyun Kim","Seongro Yoon","Temo Saghinadze","Aowen Shi","Mingyun Jeong","Donghyeon Cho","Jinsun Park","Francois Bremond"]}},"version":2},{"content":{"summary":{"value":"This paper proposes VideoTetris, a novel framework for compositional text-to-video generation. It addresses the limitations of existing methods in handling complex scenes with multiple objects and dynamic changes. VideoTetris achieves this through several key innovations: I) Spatio-Temporal Compositional Diffusion: Manipulates the cross-attention of denoising networks to synthesize videos that follow complex instructions. II) Dynamic-Aware Video Data Processing: Filters and recaptions video-text pairs to enhance consistency in auto-regressive long video generation. III) Consistency Regularization with Reference Frame Attention: Maintains coherence in multi-object generation by aligning object features across frames. Extensive experiments demonstrate that VideoTetris significantly outperforms state-of-the-art methods in both short and long video generation tasks, showcasing its ability to generate high-quality, coherent, and compositional videos."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"N/A. See weakness for more details."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1.\tThe paper is easy to follow.\n2.\tThis paper introduces a new approach for compositional text-to-video generation, addressing the limitations of existing methods in handling complex scenes and dynamic changes.\n3.\tExtensive experiments demonstrating the superior performance of VideoTetris compared to state-of-the-art methods in both short and long video generation tasks."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1.\tThis paradigm introduces more computational cost by the decomposition of input text using LLMs and the computation of cross-attention for multiple sub-objects and frames. These operations require more computation, especially in scenarios with numerous objects or a high number of frames, potentially impacting the efficiency of the overall video generation process. Besides, the use of ControlNet for auto-regressive generation leads to a high computational cost, which may limit the practical applicability of the method. \n2.\tThe Reference Frame Attention module relies on the assumption that object features remain consistent across frames. However, this may not always hold true, especially for dynamic scenes with fast movements or significant changes in lighting. Investigating more robust consistency regularization methods that can handle these challenges would be beneficial.\n3.\tThe generated videos exhibit relatively subtle variations in content, which could result in static or less dynamic scenes. This may limit their practical applicability."},"limitations":{"value":"The authors have discussed both limitations and potential negative social impacts."}},"nonreaders":[],"tmdate":1730879408741,"tcdate":1720625605976,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission10727/Reviewer_9Z5k"],"signatures":["NeurIPS.cc/2024/Conference/Submission10727/Reviewer_9Z5k"],"forum":"RPM7STrnVz","number":1,"license":"CC BY 4.0","cdate":1720625605976,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission10727/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730879408741,"domain":"NeurIPS.cc/2024/Conference","replyto":"RPM7STrnVz","id":"3OUKxD4vXa","forumContent":{"venue":{"value":"NeurIPS 2024 poster"},"keywords":{"value":["Text-to-Video Generation","Video Diffusion Models"]},"supplementary_material":{"value":"/attachment/f72c15c52e21ddf70f17cb82f2289107037072b9.zip"},"primary_area":{"value":"generative_models"},"abstract":{"value":"Diffusion models have demonstrated great success in text-to-video (T2V) generation. However, existing methods may face challenges when handling complex (long) video generation scenarios that involve multiple objects or dynamic changes in object numbers. To address these limitations, we propose VideoTetris, a novel framework that enables compositional T2V generation. Specifically, we propose spatio-temporal compositional diffusion to precisely follow complex textual semantics by manipulating and composing the attention maps of denoising networks spatially and temporally. Moreover, we propose a new dynamic-aware data processing pipeline and a consistency regularization method to enhance the consistency of auto-regressive video generation. Extensive experiments demonstrate that our VideoTetris achieves impressive qualitative and quantitative results in compositional T2V generation. Code is available at: https://github.com/YangLing0818/VideoTetris"},"_bibtex":{"value":"@inproceedings{\ntian2024videotetris,\ntitle={VideoTetris: Towards Compositional Text-to-Video Generation},\nauthor={Ye Tian and Ling Yang and Haotian Yang and Yuan Gao and Yufan Deng and Xintao Wang and Zhaochen Yu and Xin Tao and Pengfei Wan and Di ZHANG and Bin CUI},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=RPM7STrnVz}\n}"},"title":{"value":"VideoTetris: Towards Compositional Text-to-Video Generation"},"pdf":{"value":"/pdf/bb68997ead13efc218660269a1fa4d189f0588a2.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"tian|videotetris_towards_compositional_texttovideo_generation"},"authorids":{"value":["~Ye_Tian15","~Ling_Yang1","~Haotian_Yang1","~Yuan_Gao32","~Yufan_Deng2","~Xintao_Wang1","~Zhaochen_Yu2","~Xin_Tao3","~Pengfei_Wan1","~Di_ZHANG3","~Bin_CUI2"]},"authors":{"value":["Ye Tian","Ling Yang","Haotian Yang","Yuan Gao","Yufan Deng","Xintao Wang","Zhaochen Yu","Xin Tao","Pengfei Wan","Di ZHANG","Bin CUI"]}},"version":2},{"content":{"venue":{"value":"Vis. Intell. 2024"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/s44267-023-00034-7.pdf"},"venueid":{"value":"dblp.org/journals/VISINTELLIGENCE/2024"},"paperhash":{"value":"qian|controllable_augmentations_for_video_representation_learning"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Rui_Qian:","https://dblp.org/search/pid/api?q=author:Weiyao_Lin:","~John_See1","~Dian_Li1"]},"html":{"value":"https://doi.org/10.1007/s44267-023-00034-7"},"_bibtex":{"value":"@article{DBLP:journals/visintelligence/QianLSL24,\n  author={Rui Qian and Weiyao Lin and John See and Dian Li},\n  title={Controllable augmentations for video representation learning},\n  year={2024},\n  cdate={1704067200000},\n  journal={Vis. Intell.},\n  volume={2},\n  number={1},\n  url={https://doi.org/10.1007/s44267-023-00034-7}\n}\n"},"abstract":{"value":"This paper focuses on self-supervised video representation learning. Most existing approaches follow the contrastive learning pipeline to construct positive and negative pairs by sampling different clips. However, this formulation tends to bias the static background and has difficulty establishing global temporal structures. The major reason is that the positive pairs, i.e., different clips sampled from the same video, have limited temporal receptive fields, and usually share similar backgrounds but differ in motions. To address these problems, we propose a framework to jointly utilize local clips and global videos to learn from detailed region-level correspondence as well as general long-term temporal relations. Based on a set of designed controllable augmentations, we implement accurate appearance and motion pattern alignment through soft spatio-temporal region contrast. Our formulation avoids the low-level redundancy shortcut with an adversarial mutual information minimization objective to improve the generalization ability. Moreover, we introduce local-global temporal order dependency to further bridge the gap between clip-level and video-level representations for robust temporal modeling. Extensive experiments demonstrate that our framework is superior on three video benchmarks in action recognition and video retrieval, and captures more accurate temporal dynamics."},"title":{"value":"Controllable augmentations for video representation learning"},"authors":{"value":["Rui Qian","Weiyao Lin","John See","Dian Li"]}},"tmdate":1744175582045,"pdate":1704067200000,"tcdate":1727539348571,"writers":["~"],"signatures":["~Dian_Li1"],"forum":"i97F2qrqai","license":"CC BY-SA 4.0","number":102617,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1744175582045,"domain":"DBLP.org","id":"i97F2qrqai","version":2},{"content":{"summary":{"value":"This paper introduces Phantom-Data, a cross-pair subject-to-video consistency dataset for video personalization, comprising one billion identity-consistent pairs across diverse categories. Experiments indicate the high quality and potential impact of the dataset."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"n/a"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"- The paper presents a well-designed pipeline for curating subject-video paired datasets, leveraging Visual Language Models (VLMs) to achieve robust subject-attribute pairing.\n- The dataset is large-scale, providing valuable resources for advancing research in video personalization.\n- The curated data effectively addresses the prevalent copy-paste issue encountered in video personalization tasks."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- It remains unclear how much the VLM-based pipeline improves over simpler approaches, such as using GPT-4o for generating image variations. As shown in Figure 7, though having drawbacks, generative models might sometimes even offer advantages in achieving controllable variations. Additionally, potential artifacts from generative approaches could be mitigated through VLM-driven verification and filtering, which the paper does not explore.\n- The evaluation is somewhat limited, as the dataset is tested exclusively with Wan2.1. Given that many established methods in video personalization predate the Wan series, broader evaluation across multiple models would strengthen the paper’s claims and demonstrate the wider utility of the dataset."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762917940415,"tcdate":1761934872489,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission5200/Reviewer_EDtH"],"signatures":["ICLR.cc/2026/Conference/Submission5200/Reviewer_EDtH"],"forum":"IjqKXnzUXx","number":3,"license":"CC BY 4.0","cdate":1761934872489,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission5200/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762917940415,"domain":"ICLR.cc/2026/Conference","replyto":"IjqKXnzUXx","id":"1Z3DZKRv2u","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Video Generation","Multimodal Generation"]},"primary_area":{"value":"generative models"},"abstract":{"value":"Subject-to-video generation has witnessed substantial progress in recent years. However, existing models still face significant challenges in faithfully following textual instructions. This limitation, commonly known as the copy-paste problem, arises from the widely used in-pair training paradigm. This approach inherently entangles subject identity with background and contextual attributes by sampling reference images from the same scene as the target video. To address this issue, we introduce \\textbf{Phantom-Data, the first general-purpose cross-pair subject-to-video consistency dataset}, containing approximately one million identity-consistent pairs across diverse categories. Our dataset is constructed via a three-stage pipeline: (1) a general and input-aligned subject detection module, (2) large-scale cross-context subject retrieval from more than 53 million videos and 3 billion images, and (3) prior-guided identity verification to ensure visual consistency under contextual variation. Comprehensive experiments show that training with Phantom-Data significantly improves prompt alignment and visual quality while preserving identity consistency on par with in-pair baselines."},"_bibtex":{"value":"@inproceedings{\nchen2026phantomdata,\ntitle={Phantom-Data:  Towards a General Subject-Consistent Video Generation Dataset},\nauthor={Zhuowei Chen and Bingchuan Li and Tianxiang Ma and Lijie Liu and Mingcong Liu and Yunsheng Jiang and Gen Li and Xinghui Li and Liyang Chen and SiYu Zhou and Qian HE and Xinglong Wu},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=IjqKXnzUXx}\n}"},"title":{"value":"Phantom-Data:  Towards a General Subject-Consistent Video Generation Dataset"},"pdf":{"value":"/pdf/2bd728a9d3ab000d579dcb848f86efb31635b881.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"chen|phantomdata_towards_a_general_subjectconsistent_video_generation_dataset"},"authorids":{"value":["~Zhuowei_Chen1","~Bingchuan_Li2","~Tianxiang_Ma1","~Lijie_Liu1","~Mingcong_Liu1","~Yunsheng_Jiang1","~Gen_Li8","~Xinghui_Li2","~Liyang_Chen1","~SiYu_Zhou3","~Qian_HE3","~Xinglong_Wu2"]},"authors":{"value":["Zhuowei Chen","Bingchuan Li","Tianxiang Ma","Lijie Liu","Mingcong Liu","Yunsheng Jiang","Gen Li","Xinghui Li","Liyang Chen","SiYu Zhou","Qian HE","Xinglong Wu"]}},"version":2},{"content":{"summary":{"value":"This paper targets the gap between classic REC benchmarks (RefCOCO series) and real visual–language reasoning by constructing a hard, shortcut-resistant benchmark of thousands of instances. \nReferring expressions are produced via a two-stage LLM pipeline that first extracts attributes and then composes a minimal-sufficient description, followed by tri-annotator human verification for the correctness.\nEvaluation uses Acc@IoU at multiple thresholds and tests a broad slate of MLLMs (open and closed, with/without CoT). \nResults show strong models on traditional datasets RefCOCO drop substantially on Ref-Adv. Anti-shortcut ablations demonstrate that Ref-Adv requires order-sensitive, compositional grounding rather than keyword matching, which is an potential issue in the traditional benchmarks."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Could the author please add category/attribute/relation/occlusion histograms/stats for the benchmark as a comparison vs RefCOCO(/+/g). \nIt will be better if the author can clarify the CoT setup and explain the different conclusions with ARGUS on RefCOCO/+/g. Please consider address other evaluation pipeline & baselines."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"It proposes anti-shortcut design including bag-of-words shuffling and descriptor-deletion both cause larger drops than on legacy benchmarks. This is helpful for the need for compositional, order-aware grounding.\nIts data construction is considerable with strong filter pipeline and covering negation. The final benchmark is processed with strict 3-human-annotator agreement. \nFrom experiments the benchmark exposes failure modes that legacy REC underestimates, creating a clear diagnostic “stress test” where many strong MLLMs collapse. Reporting across multiple IoU thresholds and conditional slices yields informative analysis.\nModel coverage spans major open/closed families and CoT settings."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Coverage analysis of this paper is limited. There is no thorough breakdown of the 2,833 images / 5,000 instances (categories/attributes/relations/occlusion, long tails) or side-by-side coverage vs RefCOCO/+/g traiditional widely used benchmarks.\n\nThe CoT conclusion on RefCOCO conflicts with prior work in top venue. ARGUS[1] reports grounded CoT improves MLLM performance on RefCOCO/+/g, but this paper finds CoT can hurt the performance on RefCOCO, making the conlcusion not convincing. \n\nThere are gaps and missing baselines in the evaluation. SoM + Semantic-SAM may propagate segmentation errors to IoU in evaluation with no sensitivity analysis(quantitative or qualitative) in this work. Also, this paper does not analyze the answer–grounding consistency, which is about the inconsistency between the text prediction and the actual grounding. Key baseline like LLaVA-OneVision does not been tested.\n\n[1] ARGUS: Vision-Centric Reasoning with Grounded Chain-of-Thought, CVPR 2025"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762926388885,"tcdate":1761989000878,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission16232/Reviewer_txuD"],"signatures":["ICLR.cc/2026/Conference/Submission16232/Reviewer_txuD"],"forum":"iEBgrepR9i","number":3,"license":"CC BY 4.0","cdate":1761989000878,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission16232/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762926388885,"domain":"ICLR.cc/2026/Conference","replyto":"iEBgrepR9i","id":"XvebRf7zoG","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["MLLM","Referring Expression Comprehensions"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Referring Expression Comprehension (REC) links language to region level visual\nperception. Standard benchmarks (RefCOCO, RefCOCO+, RefCOCOg) have\nprogressed rapidly with multimodal LLMs but remain weak tests of visual reasoning and grounding: (i) many expressions are very short, leaving little reason-\ning demand; (ii) images often contain few distractors, making the target easy to\nfind; and (iii) redundant descriptors enable shortcut solutions that bypass genuine\ntext understanding and visual reasoning. We introduce Ref-Adv, a modern REC\nbenchmark that suppresses shortcuts by pairing linguistically nontrivial expressions with only the information necessary to uniquely identify the target. The\ndataset contains various expressions on real images, curated with hard distractors and annotated with reasoning facets including negation. We conduct comprehensive ablations (word order perturbations and\ndescriptor deletion sufficiency) to show that solving Ref-Adv requires reasoning\nbeyond simple cues, and we evaluate a broad suite of contemporary multimodal\nLLMs on Ref-Adv. Despite strong results on RefCOCO, RefCOCO+, and RefCOCOg, models drop markedly on Ref-Adv, revealing reliance on shortcuts and\ngaps in visual reasoning and grounding. We provide an in depth failure analysis\nand aim for Ref-Adv to guide future work on visual reasoning and grounding in\nMLLMs.\nThe dataset is available at \\url{https://ref-adv.github.io/}."},"_bibtex":{"value":"@inproceedings{\ndong2026refadv,\ntitle={Ref-Adv: Exploring {MLLM} Visual Reasoning in Referring Expression Tasks},\nauthor={Qihua Dong and Kuo Yang and Lin Ju and Handong Zhao and Yitian Zhang and Yizhou Wang and Huimin Zeng and Jianglin Lu and Yun Fu},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=iEBgrepR9i}\n}"},"title":{"value":"Ref-Adv: Exploring MLLM Visual Reasoning in Referring Expression Tasks"},"pdf":{"value":"/pdf/9be59c3a7f7bd55caf7a10db58375ffb778a8649.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"dong|refadv_exploring_mllm_visual_reasoning_in_referring_expression_tasks"},"authorids":{"value":["~Qihua_Dong2","~Kuo_Yang6","~Lin_Ju2","~Handong_Zhao3","~Yitian_Zhang1","~Yizhou_Wang3","~Huimin_Zeng2","~Jianglin_Lu2","~Yun_Fu1"]},"authors":{"value":["Qihua Dong","Kuo Yang","Lin Ju","Handong Zhao","Yitian Zhang","Yizhou Wang","Huimin Zeng","Jianglin Lu","Yun Fu"]}},"version":2},{"content":{"summary":{"value":"This paper introduces Video Needle-In-a-Haystack (VideoNIAH), a framework for synthetic video generation that utilizes intra-frame and inter-frame edits to create specific modifications within video segments. These modifications, referred to as \"needles,\" are used to assess the capabilities of video MLLMs in handling challenging video content. Additionally, the authors present VNBench, a benchmark compiled from these synthetic videos, which evaluates video MLLMs across three video understanding tasks: retrieval, ordering, and counting. The benchmark is tested on 12 MLLMs, both proprietary and open-source."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"- **Counting E2 Task:** What factors might explain the poor performance of Gemini and GPT-4 on the counting E2 task (Figure 5)?\n- **Consistency Check:** Could the authors verify the data in Figure 5 and Table 2? For instance, Video-LLaVa and VideoChat2 seem to perform better than Gemini on the counting E2 task in Figure 5, but Table 2 suggests otherwise.\n- **Haystack Length Variation:** In Section 5, where the \"Effect of Needle Position\" is discussed, the authors mention that the haystack is fixed. Could the authors clarify how the length of the haystack is adjusted in this scenario?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"- The paper presents a framework to create a synthetic (i.e.: edited video) video that can be automated based on certain predetermined rules and tasks. These edits are posed as “needles” in a “haystack” (the original video segment), and the query for the MLLM is based on the inserted needles. The proposed method is easily scalable given the minimal labeling/definition process that needs to be set manually.\n- The curated VNBench benchmark can be used for evaluation of MLLM in video understanding, primarily spatial and temporal understanding. This is presented through tasks of retrieval, counting and ordering based queries from the curated video benchmark.\n- The paper evaluates 12 MLLMs on the proposed benchmark with extensive evaluation on the “needle” placements and difficulty."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The rule-based automation process for dataset curation is unclear. For example, Figure 1(b) lacks clarity in illustrating this process, which is essential to understanding the scalability benefits of this approach, which can be improved.\n- While the method is scalable and captures the temporal dependency well, the scope of video understanding through this framework seems limited. For instance, the ability of the model in understanding actions, fine-grained action recognition cannot be ascertained, especially since MLLMs are used in this context extensively.\n- Given the nature of the benchmark, the sampling strategy of the video MLLM I believe plays a significant role in its performance. For instance, if any of the frames with a “needle” is not sampled, the model fails to correctly answer that query. Clarification on how this issue is addressed, or suggested solutions, would strengthen the framework’s reliability.\n- Can the authors explain if each cell in Figure 8 corresponds to a single test sample? It would be better to enhance this image with a heat map to better understand how the depth of the needle impacts the MLLMs performance. In addition, is there a reason that a depth higher than 10% was not considered?\n- Section 6: The mention of the \"training recipe\" lacks context on what is being trained, creating some confusion. A clearer introduction to this section would aid comprehension."}},"nonreaders":[],"tmdate":1731428629976,"tcdate":1730694009343,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission8848/Reviewer_FQGX"],"signatures":["ICLR.cc/2025/Conference/Submission8848/Reviewer_FQGX"],"forum":"ZJo6Radbqq","number":3,"license":"CC BY 4.0","cdate":1730694009343,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission8848/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428629976,"domain":"ICLR.cc/2025/Conference","replyto":"ZJo6Radbqq","id":"FWj96kZdja","forumContent":{"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["video MLLM"]},"primary_area":{"value":"datasets and benchmarks"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Video understanding is a crucial next step for multimodal large language models (MLLMs).\nVarious benchmarks are introduced for better evaluating the MLLMs.\nNevertheless, current video benchmarks are still inefficient for evaluating video models during iterative development due to the high cost of constructing datasets and the difficulty in isolating specific skills.\nIn this paper, we propose VideoNIAH (Video Needle in A Haystack), a benchmark construction framework through synthetic video generation. \nVideoNIAH decouples video content from their query-responses by inserting unrelated visual 'needles' into original videos. \nThe framework automates the generation of query-response pairs using predefined rules, minimizing manual labor.  The queries focus on specific aspects of video understanding, enabling more skill-specific evaluations. The separation between video content and the queries also allow for increased video variety and evaluations across different lengths.\nUtilizing VideoNIAH, we compile a video benchmark, VNBench, which includes tasks such as retrieval, ordering, and counting to evaluate three key aspects of video understanding: temporal perception, chronological ordering, and spatio-temporal coherence. We conduct a comprehensive evaluation of both proprietary and open-source models, uncovering significant differences in their video understanding capabilities across various tasks. Additionally, we perform an in-depth analysis of the test results and model configurations. Based on these findings, we provide some advice for improving video MLLM training, offering valuable insights to guide future research and model development."},"_bibtex":{"value":"@inproceedings{\nzhao2025needle,\ntitle={Needle In A Video Haystack: A Scalable  Synthetic Evaluator for Video {MLLM}s},\nauthor={Zijia Zhao and Haoyu Lu and Yuqi Huo and Yifan Du and Tongtian Yue and Longteng Guo and Bingning Wang and weipeng chen and Jing Liu},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=ZJo6Radbqq}\n}"},"title":{"value":"Needle In A Video Haystack: A Scalable  Synthetic Evaluator for Video MLLMs"},"pdf":{"value":"/pdf/875d019fb5070873ce0564e8a175d24d3df4dc3f.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"zhao|needle_in_a_video_haystack_a_scalable_synthetic_evaluator_for_video_mllms"},"authorids":{"value":["~Zijia_Zhao1","~Haoyu_Lu1","~Yuqi_Huo1","~Yifan_Du1","~Tongtian_Yue1","~Longteng_Guo1","~Bingning_Wang3","~weipeng_chen2","~Jing_Liu1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Zijia Zhao","Haoyu Lu","Yuqi Huo","Yifan Du","Tongtian Yue","Longteng Guo","Bingning Wang","weipeng chen","Jing Liu"]}},"version":2},{"content":{"summary":{"value":"This paper proposes MIMIC, a two-stage diffusion framework to generate manipulation videos by leveraging a reference video for semantic and motion cues. The method first generates a sequence of interaction masks (Stage I) and then renders the final video (Stage II), aiming to provide a scalable data source for embodied AI. The authors demonstrate strong quantitative and qualitative performance over existing methods."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"My overall feeling is that the paper presents a strong method, but the presentation is unclear in several critical areas, as detailed in the comments. I hope the authors can clarify these points in their response. Happy to increase my score once these are clarified."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"- The paper addresses the important and challenging problem of data scarcity for training embodied intelligence systems.\n- The proposed two-stage framework is logical. Decoupling the generation process into (1) motion/interaction understanding (mask generation) and (2) high-fidelity rendering (video generation) is a sensible approach to decompose a complex problem.\n- The qualitative comparisons provided in the Supp. video are compelling.\n- The ablation studies are informative and effectively validate the contributions of the proposed components, i.e., the IMA Attention and the Pair Prompt Control mechanisms ."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Relation to Video Editing: The related work section focuses on \"Video Motion Customization\" and \"Interaction Video Generation\" but does not explicitly discuss the field of generative video editing. The task in this paper shares conceptual overlap with reference-based video editing. It would strengthen the paper to include this discussion, clarify the key distinctions, and situate the work relative to recent highly relevant papers in the HOI image/video editing domain, such as [1][2].\n- The paper does not clearly explain how the (reference, target) video pairs are constructed for training. Is a video simply used as its own reference? Are pairs formed in some other way (based on action label)? This is a critical methodological detail that is currently missing and needs to be explicitly stated.\n- The paper's description of Stage I is confusing. It's trained to \"identify\" the manipulated object , but Grounding-SAM2 is already used for annotation. Why train a generative model for a recognition task? Why Grounding-SAM2 isn't used at test time to provide the initial object mask, allowing the diffusion model to focus on its core task: generating the new motion pattern (i.e., the mask trajectory)?\n- It is unclear what the scope of the tested actions and objects is. For instance, do the experiments demonstrate generalization to unseen object categories, or new instances of seen categories? The experimental setup and data splits should be described more explicitly to clarify what level of generalization is being evaluated.\n- About the use of language: Appendix A.2 says that the method relies on structured language templates (e.g., \"move [something] down\") and an NLP pipeline to \"extract structured action-object pairs\". I feel this should be explicitly stated in the main paper, as it clarifies that the model is not conditioned on free-form, natural language prompts, right?\n- About the masks: the text repeatedly describes them as \"binary masks\". However, numerous figures visualize multi-class masks, with hands/grippers and manipulated objects colored differently. The authors need to clarify whether the masks are binary (e.g., interaction vs. background) or multi-class (e.g., hand vs. object vs. background). Also Ln204 “soft binary mask”, what does “soft” mean?\n\n[1] Affordance Diffusion: Synthesizing Hand-Object Interactions\n\n[2] HOI-Swap: Swapping Objects in Videos with Hand-Object Interaction Awareness"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762919749789,"tcdate":1761863130990,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission7691/Reviewer_Kzh4"],"signatures":["ICLR.cc/2026/Conference/Submission7691/Reviewer_Kzh4"],"forum":"COrUdVuInH","number":2,"license":"CC BY 4.0","cdate":1761863130990,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission7691/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762919749789,"domain":"ICLR.cc/2026/Conference","replyto":"COrUdVuInH","id":"IyLbOj3kDJ","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["video diffusion model","manipulation video"]},"supplementary_material":{"value":"/attachment/b6d9b8cb3acbab21f3cbce260e6148e91bc5581b.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Embodied intelligence faces a fundamental bottleneck from limited large-scale interaction data. Video generation offers a scalable alternative, but manipulation videos remain particularly challenging, as they require capturing subtle, contact-rich dynamics. Despite recent advances, video diffusion models still struggle to balance semantic understanding with fine-grained visual details, restricting their effectiveness in manipulation scenarios. Our key insight is that reference videos provide rich semantic and motion cues that can effectively drive manipulation video generation. Building on this, we propose MIMIC, a two-stage image-to-video diffusion framework. (1) We first introduce an Interaction-Motion-Aware (IMA) module to fuse visual features from the reference video, producing coherent semantic masks that correspond to the target image. (2) then utilize these masks as semantic control signals to guide the video generation process. Moreover, considering the ambiguity of the motion attribution,  we introduce a Pair Prompt Control mechanism to disentangle object and camera motion by adding the reference video as an additional input. Extensive experiments demonstrate that MIMIC significantly outperforms existing methods, effectively preserves manipulation intent and motion details, even when handling diverse and deformable objects. Our findings underscore the effectiveness of reference-driven semantics for controllable and realistic manipulation video generation."},"_bibtex":{"value":"@inproceedings{\nchen2026mimic,\ntitle={{MIMIC}: Mask-Injected Manipulation Video Generation with Interaction Control},\nauthor={Tianxiao Chen and Jintao Rong and Huajin Chen and Jingya Wang and Tao Zhou and Jiming Chen and Qi Ye},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=COrUdVuInH}\n}"},"title":{"value":"MIMIC: Mask-Injected Manipulation Video Generation with Interaction Control"},"pdf":{"value":"/pdf/84bc45b7fa57b5fa895a891bd9cf268ab48a7972.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"chen|mimic_maskinjected_manipulation_video_generation_with_interaction_control"},"authorids":{"value":["~Tianxiao_Chen1","~Jintao_Rong2","~Huajin_Chen1","~Jingya_Wang3","~Tao_Zhou8","~Jiming_Chen1","~Qi_Ye2"]},"authors":{"value":["Tianxiao Chen","Jintao Rong","Huajin Chen","Jingya Wang","Tao Zhou","Jiming Chen","Qi Ye"]}},"version":2},{"content":{"summary":{"value":"The paper introduces AuroraCap, a video captioning model based on LMM, along with VDC, a detailed captioning benchmark, and VDCScore, an evaluation metric designed for long captions. AuroraCap handles video sequences with lower visual token numbers by the token merge strategy, outperforming state-of-the-art models, even including GPT-4V and Gemini-1.5 Pro.  VDC addresses the limitations of existing video captioning benchmarks by providing over one thousand videos with detailed, structured captions. since existing metrics are insufficient to evaluate detailed captioning performance, they develop a new LLM-assisted metric VDC-SCORE, that breaks down long captions into short question-answer pairs, ensuring better alignment with human judgments."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. **Training Strategy Clarity:** Was the token merging strategy used during AuroraCap’s training phase? How does varying the token merge ratio during training impact inference performance?\n2. **Performance in Video Question Answering:** What factors limit AuroraCap’s performance in video question answering compared to captioning tasks?\n3. **Token Merge Curve on VDC** As shown in Figure 6, why does retaining all visual tokens not lead to the highest VDCScore?\n\nI will consider raising my score if the authors can further address my concerns."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. Despite adopting a simple architecture design consistent with LLaVA, AuroraCap achieves remarkable results on captioning tasks with its three-stage training recipe.\n2. AuroraCap combines the Token Merging strategy, significantly reducing the number of visual tokens input and thereby improving inference efficiency.\n3. This paper also presents VDC, a detailed video description benchmark, which includes over one thousand structured descriptions, providing detailed annotations for videos from multiple perspectives.\n4. To accurately evaluate the model’s detailed captioning capabilities, the paper introduces VDCScore. This metric uses a divide-and-conquer strategy to break long captions into short QA pairs and employs LLM assistance for scoring. This approach avoids the challenges traditional metrics face with long captions and mitigates potential hallucination issues associated with using LLM evaluations.\n5. Compared to existing models, AuroraCap achieves excellent performance in both image captioning and video captioning tasks.\n6. The paper also explores the impact of the visual token kept ratio on model performance, demonstrating that with the token merging strategy, fewer visual tokens can still maintain high performance."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. **Spatial Token Merging:** AuroraCap only performs spatial token merging within frames and does not implement temporal token merging between frames, which could further enhance inference efficiency.\n2. **Static Token Merging:** It seems that for different visual inputs, the number of visual tokens remains consistent by setting a visual token kept ratio. However, this ‘static’ token merging strategy is unreasonable for tokens with varying levels of informativeness.\n3. **Unclear Training Strategy:** It is not clear whether the token merging strategy was used during training. Additionally, the impact of using different token merge ratios during training on inference performance is not clear.\n4. **Performance in Video Question Answering:** Although AuroraCap demonstrates leading performance in captioning tasks, its advantages in video question answering tasks are not apparent, which is perplexing."}},"nonreaders":[],"tmdate":1731427292511,"tcdate":1730276014941,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission640/Reviewer_ZSxM"],"signatures":["ICLR.cc/2025/Conference/Submission640/Reviewer_ZSxM"],"forum":"tTDUrseRRU","number":1,"license":"CC BY 4.0","cdate":1730276014941,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission640/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427292511,"domain":"ICLR.cc/2025/Conference","replyto":"tTDUrseRRU","id":"JjUcC49rez","forumContent":{"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video Captioning","Benchmark","Multimodel Large Language Model"]},"primary_area":{"value":"datasets and benchmarks"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Video detailed captioning is a key task which aims to generate comprehensive and coherent textual descriptions of video content, benefiting both video understanding and generation. In this paper, we propose AuroraCap, a video captioner based on a large multimodal model. We follow the simplest architecture design without additional parameters for temporal modeling. To address the overhead caused by lengthy video sequences, we implement the token merging strategy, reducing the number of input visual tokens. Surprisingly, we found that this strategy results in little performance loss. AuroraCap shows superior performance on various video and image captioning benchmarks, for example, obtaining a CIDEr of 88.9 on Flickr30k, beating GPT-4V (55.3) and Gemini-1.5 Pro (82.2). However, existing video caption benchmarks only include simple descriptions, consisting of a few dozen words, which limits research in this field. Therefore, we develop VDC, a video detailed captioning benchmark with over one thousand carefully annotated structured captions. In addition, we propose a new LLM-assisted metric VDCscore for bettering evaluation, which adopts a divide-and-conquer strategy to transform long caption evaluation into multiple short question-answer pairs. With the help of human Elo ranking, our experiments show that this benchmark better correlates with human judgments of video detailed captioning quality."},"_bibtex":{"value":"@inproceedings{\nchai2025auroracap,\ntitle={AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark},\nauthor={Wenhao Chai and Enxin Song and Yilun Du and Chenlin Meng and Vashisht Madhavan and Omer Bar-Tal and Jenq-Neng Hwang and Saining Xie and Christopher D Manning},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=tTDUrseRRU}\n}"},"title":{"value":"AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark"},"pdf":{"value":"/pdf/54378994740e36257bfa53e767f50f28a4e36483.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"chai|auroracap_efficient_performant_video_detailed_captioning_and_a_new_benchmark"},"authorids":{"value":["~Wenhao_Chai1","~Enxin_Song2","~Yilun_Du1","~Chenlin_Meng1","~Vashisht_Madhavan3","~Omer_Bar-Tal2","~Jenq-Neng_Hwang1","~Saining_Xie2","~Christopher_D_Manning1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Wenhao Chai","Enxin Song","Yilun Du","Chenlin Meng","Vashisht Madhavan","Omer Bar-Tal","Jenq-Neng Hwang","Saining Xie","Christopher D Manning"]}},"version":2},{"content":{"summary":{"value":"This paper introduces **PreciseCache**, a training-free acceleration framework for video diffusion models that reuses redundant computations both across denoising steps (LFCache) and within transformer blocks (BlockCache). The method identifies truly redundant features based on low-frequency difference (LFD) analysis, achieving up to 2.6× speedup without noticeable quality degradation. The paper evaluates performance across several open-source video generation backbones."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"**Evaluation on compact video generation models:**\n\n Please evaluate PreciseCache on a smaller, efficient video generation model (e.g., with reduced transformer depth or parameter count) to demonstrate whether the proposed caching mechanism provides benefits beyond over-parameterized models. Please **specify the exact training datasets** and report **metric scores per model** (LPIPS, SSIM, PSNR, VBench) and training time for at least one video geneartion model.  I am willing to increase score if the authors can help reduce some concerns about the video generation models."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"The paper identifies redundancy in transformer-based video generation models, distinguishing between *pivotal* and *non-pivotal* blocks, which aligns with observations of over-parameterization in large diffusion models. Training-free and plug-and-play: The proposed method requires no fine-tuning or retraining, making it applicable to existing models with minimal integration cost."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Limited generalization evidence. The proposed method relies on the redundancy of existing open-weight video generation models. Its effectiveness may diminish when applied to smaller or more efficient architectures with less redundancy. The paper does not test PreciseCache on small-scale, parameter-efficient video generation models, leaving uncertainty about its general applicability."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1764359985741,"tcdate":1761611362860,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission7215/Reviewer_j8AN"],"signatures":["ICLR.cc/2026/Conference/Submission7215/Reviewer_j8AN"],"forum":"DjfRkr82jn","number":1,"license":"CC BY 4.0","cdate":1761611362860,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission7215/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1764359985741,"domain":"ICLR.cc/2026/Conference","replyto":"DjfRkr82jn","id":"K58jhiecJR","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Diffusion Model"]},"primary_area":{"value":"generative models"},"abstract":{"value":"High computational costs and slow inference hinder the practical application of video generation models. While prior works accelerate the generation process through feature caching, they often suffer from notable quality degradation. In this work, we reveal that this issue arises from their inability to distinguish truly redundant features, which leads to the unintended skipping of computations on important features. To address this, we propose \\textbf{PreciseCache}, a plug-and-play framework that precisely detects and skips truly redundant computations, thereby accelerating inference without sacrificing quality. Specifically, PreciseCache contains two components: LFCache for step-wise caching and BlockCache for block-wise caching. For LFCache, we compute the Low-Frequency Difference (LFD) between the prediction features of the current step and those from the previous cached step. Empirically, we observe that LFD serves as an effective measure of step-wise redundancy, accurately detecting highly redundant steps whose computation can be skipped through reusing cached features. To further accelerate generation within each non-skipped step, we propose BlockCache, which precisely detects and skips redundant computations at the block level within the network. Extensive experiments on various backbones demonstrate the effectiveness of our PreciseCache, which achieves an average of $2.6\\times$ speedup without noticeable quality loss. Source code will be released."},"_bibtex":{"value":"@inproceedings{\nwang2026precisecache,\ntitle={PreciseCache: Precise Feature Caching for Efficient and High-fidelity Video Generation},\nauthor={Jiangshan Wang and Kang Zhao and Jiayi Guo and Jiayu Wang and Hang Guo and Chenyang Zhu and Xiu Li and Xiangyu Yue},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=DjfRkr82jn}\n}"},"title":{"value":"PreciseCache: Precise Feature Caching for Efficient and High-fidelity Video Generation"},"pdf":{"value":"/pdf/d3d9a11482cae5d091718222592b236b720787a6.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"wang|precisecache_precise_feature_caching_for_efficient_and_highfidelity_video_generation"},"authorids":{"value":["~Jiangshan_Wang2","~Kang_Zhao7","~Jiayi_Guo2","~Jiayu_Wang2","~Hang_Guo3","~Chenyang_Zhu5","~Xiu_Li1","~Xiangyu_Yue1"]},"authors":{"value":["Jiangshan Wang","Kang Zhao","Jiayi Guo","Jiayu Wang","Hang Guo","Chenyang Zhu","Xiu Li","Xiangyu Yue"]}},"version":2},{"content":{"summary":{"value":"The paper introduces Neural Concept Verifier (NCV), a method that combines Concept Bottleneck Models (CBMs) with Prover-Verifier Games (PVGs) to improve robustness against shortcut learning while maintaining interpretability. NCV uses a concept encoder to map inputs to a concept space and employs a prover-verifier framework to ensure that predictions are based on meaningful concepts rather than spurious correlations. The method is evaluated on several image benchmarks, demonstrating improved robustness and interpretability compared to some existing approaches."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"1. Can you provide an ablation study on the concept mask size? You report the results for 32 concepts, but it would be interesting to see the relationship between performance and the concept count. \n2. Why do you think Pixel-MAC performs poorly on CIFAR-100, ImageNet-1k, and COCOLogic compared to CLEVR-Hans? \n3. How was the CLIP-Sim vocabulary chosen? Was it selected based on relevance to the datasets, or was it a generic set of concepts? \n4. Can you report the validation–test gap for Pixel-MAC? \n5. How does the performance change if you only optimize Merlin's loss (i.e., $\\gamma=0$)? Can you report the performance and loss curves for different values of $\\gamma$? \n6. Can you evaluate your method against input-space adversarial attacks (e.g., FGSM) to see if the robustness gains in concept space translate to input space robustness? \n7. What are some real-world image-based settings where adversarial provers would be relevant? Real-world settings for pixel-level attacks are fairly established, but I'm not aware of ones where a concept selector may be compromised in the image domain."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. Framing prediction as a PVG on a concept bottleneck is an interesting approach to the problem of shortcut learning. \n2. PVGs offer a different and more stable spin on adversarial training by shifting the adversarial intervention to the concept space instead of the input space. \n3. The proposed method provides good semantic explanations without compromising predictive performance, which is a key challenge in interpretable ML. \n4. The paper is easy to follow."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The claim that CBMs are \"limited by their reliance on low-capacity linear predictors\" is not quite fair; CBMs can easily be combined with non-linear predictors, so this is not really a serious limitation that this paper resolves. \n2. The first 3 benchmarks in Table 1 are missing a key baseline, namely, CBM on the CLIP-Sim feature space with a nonlinear probe. This is to test whether the PVG approach has a distinct performance edge over simply applying CBMs on a rich concept vocabulary with nonlinear probes. \n3. Table 2 is missing the Pixel-MAC baseline, which is important to see how NCV compares to other PVG methods on shortcut learning. \n4. It is unclear how significant the robustness gains are since the paper assumes a safe (non-adversarial) concept encoder, which ignores an important risk/attack vector of shortcut learning where the concept encoder itself may learn spurious correlations.\n5. (Minor) The last sentence of the Table 1 caption, \"NCV consistently matches or outperforms baselines in completeness,\" is not totally accurate since the CBM baseline outperforms NCV on ImageNet-1k in terms of completeness."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762928162976,"tcdate":1761663818015,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission18469/Reviewer_ACh6"],"signatures":["ICLR.cc/2026/Conference/Submission18469/Reviewer_ACh6"],"forum":"zSEThawime","number":2,"license":"CC BY 4.0","cdate":1761663818015,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission18469/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762928162976,"domain":"ICLR.cc/2026/Conference","replyto":"zSEThawime","id":"agaSeBxqU3","forumContent":{"TLDR":{"value":"We propose the Neural Concept Verifier (NCV), a new framework combining Prover-Verifier Games and concept-level encodings, enabling interpretable and nonlinear classification for high-dimensional data."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Interpretability","Prover-Verifier Games","Concept Bottleneck Models","Concept Explanation","XAI"]},"supplementary_material":{"value":"/attachment/ce0bf6ee154dd8f72b1132ccd4e82aab3fb1b9df.zip"},"primary_area":{"value":"interpretability and explainable AI"},"abstract":{"value":"While *Prover-Verifier Games* (PVGs) offer a promising and much needed path toward verifiability in nonlinear classification models, they have not yet been applied to complex inputs such as high-dimensional images. Conversely, *Concept Bottleneck Models* (CBMs) effectively translate such data into interpretable concepts but are limited by their reliance on low-capacity linear predictors.  \nIn this work, we push towards real-world verifiability by combining the strengths of both approaches. We introduce *Neural Concept Verifier (NCV)*, a unified framework combining PVGs for formal verifiability with concept encodings to handle complex, high-dimensional inputs in an interpretable way. NCV achieves this by utilizing recent minimally supervised concept discovery models to extract structured concept encodings from raw inputs. A *prover* then selects a subset of these encodings, which a *verifier*, implemented as a nonlinear predictor, uses exclusively for decision-making.  \nOur evaluations show that NCV outperforms CBM and pixel-based PVG classifier baselines on high-dimensional, logically complex datasets and also helps mitigate shortcut behavior. Overall, we demonstrate NCV as a promising step toward performative, verifiable AI."},"_bibtex":{"value":"@misc{\nturan2026neural,\ntitle={Neural Concept Verifier: Scaling Prover-Verifier Games via Concept Encodings},\nauthor={Berkant Turan and Suhrab Asadulla and David Steinmann and Kristian Kersting and Wolfgang Stammer and Sebastian Pokutta},\nyear={2026},\nurl={https://openreview.net/forum?id=zSEThawime}\n}"},"title":{"value":"Neural Concept Verifier: Scaling Prover-Verifier Games via Concept Encodings"},"pdf":{"value":"/pdf/5f47ee462a2f77a584de66d006417e17ef2b5b67.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"turan|neural_concept_verifier_scaling_proververifier_games_via_concept_encodings"},"authorids":{"value":["~Berkant_Turan1","~Suhrab_Asadulla1","~David_Steinmann1","~Kristian_Kersting1","~Wolfgang_Stammer1","~Sebastian_Pokutta1"]},"authors":{"value":["Berkant Turan","Suhrab Asadulla","David Steinmann","Kristian Kersting","Wolfgang Stammer","Sebastian Pokutta"]}},"version":2},{"content":{"summary":{"value":"The paper presents MDSGen, an efficient framework for vision-guided sound generation that minimizes model size, memory usage, and inference time. Key innovations include a temporal-aware masking strategy to enhance alignment accuracy and a redundant feature removal module to filter unnecessary video information. Using a lightweight masked diffusion transformer, MDSGen outperforms larger Unet-based models on VGGSound and Flickr-SoundNet, achieving high synchronization and alignment with significantly reduced computational costs."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"Please see Weaknesses."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"The paper introduces a novel framework for video-to-audio sound generation that effectively combines a temporal-aware masking strategy with a redundant feature removal module. \n\nMDSGen demonstrates significant improvements in model efficiency by using a smaller masked diffusion transformer architecture. The framework achieves high alignment accuracy on benchmark datasets with a fraction of the parameters, memory usage, and inference time compared to baselines. \n\nThe paper provides a structured explanation of MDSGen’s architecture and mechanisms, including the Temporal-Awareness Masking (TAM) and the Reducer module for filtering out redundant features. Extensive experimental results on VGGSound and Flickr-SoundNet datasets clearly validate the method’s effectiveness, with MDSGen achieving superior performance across alignment accuracy and efficiency metrics."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Lack of Novelty and Contribution: The paper presents the primary contributions are the Temporal-Awareness Masking (TAM) strategy and the visual Reducer module. However, masking strategies have been widely explored in audio generation research, as seen in works like [1, 2], and the specific concept of Temporal-Awareness Masking has been studied in [3, 4]. The visual Reducer module, primarily a 1x1 convolutional layer (line 181), lacks detailed design innovations, which limits its distinctiveness and impact.\n\n2. Insufficient Exploration of Design Choices: For video-to-audio generation, the choice of video encoder plays a crucial role in understanding the video content. Clarification on the selection of CAVP as the video encoder would add valuable insight. Additionally, the paper could explore using more video encoders, such as CLIP [5], VideoMAE [6], ViVit [7], and TAM [8], which could enrich the technical depth of the proposed method.\n\n3. Presentation and Writing:\nSome claims in the paper lack supporting evidence, such as the statements in lines 183-185 that the proposed method “minimizes redundant features that could lead to overfitting” and in line 224 that setting N_2 = 4 “gives better performance for audio data.” These points would benefit from empirical support to substantiate their validity.\n\n4. Supplementary Material: The quality of generated audio samples in the supplementary material raises concerns regarding the overall quality of results produced by the proposed method, which may affect its effectiveness and appeal.\n\n[1]Pascual S, Yeh C, Tsiamas I, et al. Masked Generative Video-to-Audio Transformers with Enhanced Synchronicity[J]. arXiv preprint arXiv:2407.10387, 2024.\n\n[2]Borsos Z, Sharifi M, Vincent D, et al. Soundstorm: Efficient parallel audio generation[J]. arXiv preprint arXiv:2305.09636, 2023.\n\n[3]Bai H, Zheng R, Chen J, et al. A $^ 3$ T: Alignment-Aware Acoustic and Text Pretraining for Speech Synthesis and Editing[C]//International Conference on Machine Learning. PMLR, 2022: 1399-1411.\n\n[4]Garcia H F, Seetharaman P, Kumar R, et al. Vampnet: Music generation via masked acoustic token modeling[J]. arXiv preprint arXiv:2307.04686, 2023.\n\n[5]Radford A, Kim J W, Hallacy C, et al. Learning transferable visual models from natural language supervision[C]//International conference on machine learning. PMLR, 2021: 8748-8763.\n\n[6]Tong Z, Song Y, Wang J, et al. Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training[J]. Advances in neural information processing systems, 2022, 35: 10078-10093.\n\n[7]Arnab A, Dehghani M, Heigold G, et al. Vivit: A video vision transformer[C]//Proceedings of the IEEE/CVF international conference on computer vision. 2021: 6836-6846."}},"nonreaders":[],"tmdate":1733609107421,"tcdate":1730674463549,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission265/Reviewer_mkFe"],"signatures":["ICLR.cc/2025/Conference/Submission265/Reviewer_mkFe"],"forum":"yFEqYwgttJ","number":3,"license":"CC BY 4.0","cdate":1730674463549,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission265/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1733609107421,"domain":"ICLR.cc/2025/Conference","replyto":"yFEqYwgttJ","id":"uPpLzmUOfa","forumContent":{"TLDR":{"value":"A novel approach is presented for highly efficient vision-guided sound synthesis using masked diffusion models"},"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["vision-guided audio generation","fast inference","open-domain sound synthesis","masked diffusion models","temporal learning","visual sound source localization","generative AI"]},"supplementary_material":{"value":"/attachment/540764b582c3bd4ee182a390db04bfe032eb2087.zip"},"primary_area":{"value":"generative models"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"We introduce MDSGen, a novel framework for vision-guided open-domain sound generation optimized for model parameter size, memory consumption, and inference speed. This framework incorporates two key innovations: (1) a redundant video feature removal module that filters out unnecessary visual information, and (2) a temporal-aware masking strategy that leverages temporal context for enhanced audio generation accuracy. In contrast to existing resource-heavy Unet-based models, MDSGen employs denoising masked diffusion transformers,  facilitating efficient generation without reliance on pre-trained diffusion models. Evaluated on the benchmark VGGSound dataset, our smallest model (5M parameters) achieves 97.9% alignment accuracy, using 172x fewer parameters, 371% less memory, and offering 36x faster inference than the current 860M-parameter state-of-the-art model (93.9% accuracy). The larger model (131M parameters) reaches nearly 99% accuracy while requiring 6.5x fewer parameters. These results highlight the scalability and effectiveness of our approach. The code is available at https://bit.ly/mdsgen."},"_bibtex":{"value":"@inproceedings{\npham2025mdsgen,\ntitle={{MDSG}en: Fast and Efficient Masked Diffusion Temporal-Aware Transformers for Open-Domain Sound Generation},\nauthor={Trung X. Pham and Tri Ton and Chang D. Yoo},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=yFEqYwgttJ}\n}"},"title":{"value":"MDSGen: Fast and Efficient Masked Diffusion Temporal-Aware Transformers for Open-Domain Sound Generation"},"pdf":{"value":"/pdf/ac9a020becd94dd25f980950f2ac0b59550f2139.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"pham|mdsgen_fast_and_efficient_masked_diffusion_temporalaware_transformers_for_opendomain_sound_generation"},"authorids":{"value":["~Trung_X._Pham1","~Tri_Ton1","~Chang_D._Yoo1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Trung X. Pham","Tri Ton","Chang D. Yoo"]}},"version":2},{"content":{"summary":{"value":"This Manuscript introduces `video-gpt`, a generative self-supervised solution to represent videos. The main idea is to train video embeddings in a GPT style, where each clip is acting as a token. The training idea is to have interleaved noisy and clean clips, and the training objective is to de-noise the noisy clips."},"soundness":{"value":4},"confidence":{"value":5},"questions":{"value":"1- Authors propose `Clips as tokens` but later they propose frame-level and patch-level masking. It reads to me as `partial-token` masking. IMO, it would be better not to name Clips as tokens.\n\n2- What is the intuition behind having interleaved noisy and clean clips during training? Why not going with a classic next frame prediction formulation and have `k clean clips` and diffuse the `k+1th noisy clip`? What is the advantage of interleaved modeling? \n\n3- Why `from scratch` model performs poorly in `Tab 6`?\n\n4- In Section 3.3, are all previous K clips (some of them being generated diffused clips) being used to predict the k+1? or there is a limit on K to keep the context window capped?"},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":3},"strengths":{"value":"State of the results on multiple tasks is the main strength of this paper.\n\nThe idea of interleaved clip level noise/de-noise, although intuitive, is novel IMO.\n\nVideo level self-supervision has been overlooked, IMO. Research like this can bring more attention and opens the road for future works in video domain."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Some design choices for training is not trivial to me and I need more clarification. (will ask in question section).\n\nI believe such a method works only on single camera and continious (or single scene) videos. If there is a POV change in a video, like Movies or TV shows, I believe that it will break the whole network.\n\nMotion is not modeled very well in this work. I am curious to know how this model can predict videos where there is partly stationary clips (minimal motion) and partly abrupt motion."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762915836122,"tcdate":1762031245689,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission1629/Reviewer_sVcH"],"signatures":["ICLR.cc/2026/Conference/Submission1629/Reviewer_sVcH"],"forum":"E0ZAcqy9TB","number":4,"license":"CC BY 4.0","cdate":1762031245689,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission1629/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762915836122,"domain":"ICLR.cc/2026/Conference","replyto":"E0ZAcqy9TB","id":"EBBnYjrMIt","forumContent":{"TLDR":{"value":"Video-GPT treats video as new language for visual world modeling."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Video; Diffusion; LLM"]},"primary_area":{"value":"foundation or frontier models, including LLMs"},"abstract":{"value":"GPT has shown its remarkable success in natural language processing. However, the language sequence is not sufficient to describe spatial-temporal details in the visual world. Alternatively, the video sequence is good at capturing such details. Motivated by this fact, we propose a concise Video-GPT in this paper by treating video as new language for visual world modeling. By analogy to next token prediction in GPT, we introduce a novel next clip diffusion paradigm for pretraining Video-GPT. Different from the previous works, this distinct paradigm allows Video-GPT to tackle both short-term generation and long-term prediction, by autoregressively denoising the noisy clip according to the clean clips in the history. Extensive experiments show our Video-GPT achieves the state-of-the-art performance on video prediction, which is the key factor towards world modeling (Physics-IQ Benchmark: Video-GPT 34.97 vs. Kling 23.64 vs. Wan 20.89). Moreover, it can be well adapted on 6 mainstream video tasks in both video generation and understanding, showing its great generalization capacity in downstream."},"_bibtex":{"value":"@inproceedings{\nzhuang2026videogpt,\ntitle={Video-{GPT} via Next Clip Diffusion},\nauthor={Shaobin Zhuang and Zhipeng Huang and Ying Zhang and Fangyikang Wang and Canmiao Fu and Binxin Yang and Chong Sun and Chen Li and Yali Wang},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=E0ZAcqy9TB}\n}"},"title":{"value":"Video-GPT via Next Clip Diffusion"},"pdf":{"value":"/pdf/717aa186ca36c0bc98678fe5f11097effe05f8bf.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"zhuang|videogpt_via_next_clip_diffusion"},"authorids":{"value":["~Shaobin_Zhuang1","~Zhipeng_Huang3","~Ying_Zhang9","~Fangyikang_Wang1","~Canmiao_Fu1","~Binxin_Yang1","~Chong_Sun1","~Chen_Li11","~Yali_Wang1"]},"authors":{"value":["Shaobin Zhuang","Zhipeng Huang","Ying Zhang","Fangyikang Wang","Canmiao Fu","Binxin Yang","Chong Sun","Chen Li","Yali Wang"]}},"version":2},{"content":{"venue":{"value":"ICLR 2027 Conference Submission"},"TLDR":{"value":"We audit 115 video benchmarks with five levels of shortcut attacks, find that attackers without any frame nearly match full-video accuracy on 35, and release Video-Index, 840 verified questions hardest under these attacks."},"keywords":{"value":["video understanding","video benchmarks","benchmark auditing","shortcut learning","multimodal large language models","evaluation"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"A video benchmark should reward the capability it claims to measure, yet models can exploit answer options, question text, or partial visual evidence. We introduce the attack pyramid, five levels of shortcut attacks with increasing access to each item, and audit 115 video benchmarks with it. On 35 benchmarks, attackers that never see a frame approach full-video accuracy. On 51 benchmarks with temporal probes, shuffled frames keep a median 96% of full-video accuracy. Near-duplicate questions make up at least half the items in 63 benchmarks. We screen 505,518 question–answer pairs from 112 of them into an audited pool. Agents turn evaluation requests into specifications, and a deterministic selector with a red-team gate composes reproducible benchmarks. We release Video-Index, the 210 hardest verified items under these attacks in each of four capability groups, 840 items from 76 sources. With the same fixed input, Claude Opus 5 outscores every open-source model by over 37 percentage points, and agent tools add about 20 more, yet all systems leave room to improve efficiency and accuracy."},"_bibtex":{"value":"@inproceedings{\nanonymous2026videoindex,\ntitle={Video-Index: A Curated Meta-Benchmark for Video Understanding Evaluation},\nauthor={Anonymous},\nbooktitle={Submitted to The Fifteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=g0vDUvCHlP},\nnote={under review}\n}"},"title":{"value":"Video-Index: A Curated Meta-Benchmark for Video Understanding Evaluation"},"pdf":{"value":"/pdf/d697af61fed1d5996eab97906d8a286a45df6307.pdf"},"venueid":{"value":"ICLR.cc/2027/Conference/Submission"}},"tmdate":1791219581935,"tcdate":1787237077990,"writers":["ICLR.cc/2027/Conference","ICLR.cc/2027/Conference/Submission1991/Authors"],"signatures":["ICLR.cc/2027/Conference/Submission1991/Authors"],"forum":"g0vDUvCHlP","license":"CC BY 4.0","number":1991,"cdate":1787237077990,"readers":["everyone"],"invitations":["ICLR.cc/2027/Conference/-/Submission","ICLR.cc/2027/Conference/Submission1991/-/Full_Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Bidding","ICLR.cc/2027/Conference/-/Submission_Change_Before_Reviewing","ICLR.cc/2027/Conference/-/Edit"],"mdate":1791219581935,"odate":1791031328286,"domain":"ICLR.cc/2027/Conference","id":"g0vDUvCHlP","version":2},{"content":{"summary":{"value":"The paper highlights the shortcomings of existing LMM benchmarks in evaluating perceptual video quality discernibility. To address this limitation, the paper proposes a video quality-specific benchmark called Q-Bench-Video composed of real, AI-generated, and computer graphics video content. To test the LMMs, the paper collects a series of Yes-No, What-How, and Open-ended question-answer pairs about the videos. Video-pair comparisons are also considered. A set of 17 LMMs (12 open-source + 5 proprietary) are benchmarked on Q-Bench-Video. The performance of these LMMs is compared with human performance. The comparison leads to the observation that humans > LMMs > random choice in terms of accuracy."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1. Can you please offer more useful insights into your findings? The main result of human > LMM > random and the other observations in Section 4 are not unexpected. For example, let's consider the following observation - \"This disparity underscores a significant proficiency gap between LMMs’ capability in handling straightforward, closed-form questions and their effectiveness in navigating the complexities of real-world problem-solving, particularly in the context of video quality evaluation.\" Can you please provide any reasons for such a gap?\n\n2. How can Q-Bench-Video be used to improve LMM performance? Please provide some guidance on this aspect.\n\n3. What factors led to the chosen dataset size? \n\n4. What is the frame rate and the duration of the videos? What is the impact of sampling these videos uniformly?\n\n5. Why were only three subjects used for the human evaluation? Is this a statistically significant number? Please refer to https://www.itu.int/dms_pubrec/itu-r/rec/bt/R-REC-BT.500-15-202305-I!!PDF-E.pdf. Section 2.5.1 explicitly mentions 15 observers for a subjective evaluation.\n\n6. Were statistical consistency tests conducted on the human evaluations? Please refer to Annex 1 of the above ITU document.\n\n7. Can you please clarify the dev and test subsets?"},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. The problem is motivated well, and the solution is timely given the ever-increasing usage of LMMs in VQA.\n\n2. The choice of real, AIG, and CG videos, along with the QA pairs, is well thought out.\n\n3. The coverage of the LMMs is extensive."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. In my opinion, the work’s full impact is not felt in the current form of the paper. Specifically, the key observation of humans > LMMs > random choice is neither surprising nor unexpected. The same can be said of the observation that proprietary models outperform open-source models. In light of this, the study should have been extended to show how the benchmark can help improve LMM performance to reduce the gap with human performance. Without this, a practitioner would be left wanting for the utility of Q-Bench-Video.\n\n2. Given that the original content sources contained a much larger set of videos, the size of Q-Bench-Video (1800) is too small to qualify as a representative benchmark. For example, the LSVQ dataset has 39000 videos, and the MaxWell dataset has 4500 videos. Further, several user-generated content datasets, such as YT-UGC, Konvik-1K, etc., have in excess of 1500 videos. Given that Q-Bench-Video is proposed as a benchmark composed of not only real videos but also AIG and CG videos, a representative dataset should contain at least 4500 videos.  \n\n3. Some of the key details of the work have been relegated to the appendices. For example, the video selection and evaluation strategy are both detailed in the appendix.  \n\n4. The fact that only three human subjects were used for human performance evaluation does not inspire confidence in the work. There is no information on statistical consistency tests on the subjective experiments.  \n\n5. The presentation quality needs to be improved."}},"nonreaders":[],"tmdate":1731427490509,"tcdate":1730625421457,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission1786/Reviewer_xK9V"],"signatures":["ICLR.cc/2025/Conference/Submission1786/Reviewer_xK9V"],"forum":"VaUy5GZO3f","number":1,"license":"CC BY 4.0","cdate":1730625421457,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission1786/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427490509,"domain":"ICLR.cc/2025/Conference","replyto":"VaUy5GZO3f","id":"mww6GvAWoX","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Large multi-modal model","benchmark","video quality assessment"]},"primary_area":{"value":"datasets and benchmarks"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"With the rising interest in research on Large Multi-modal Models (LMMs) for video understanding, many studies have emphasized general video comprehension capabilities, neglecting the **systematic exploration into video quality understanding**. To address this oversight, we introduce **Q-Bench-Video** in this paper, a new benchmark specifically designed to evaluate LMMs' proficiency in discerning video quality. **a)** To ensure the diversity of video sources, Q-Bench-Video encompasses videos from natural scenes, computer graphics (CG), and AI-generated content (AIGC). **b)** Building on the traditional multiple-choice questions format with the *Yes-or-No* and *What-How* categories, we include *Open-ended* questions to better evaluate complex scenarios. Additionally, we incorporate the **video pair quality comparison** question to enhance comprehensiveness. **c)** Beyond the traditional *Technical*, *Aesthetic*, and *Temporal* distortions, we have expanded our evaluation aspects to include the dimension of *AIGC* distortions, which addresses the increasing demand for video generation. Finally, we collect a total of 2,378 question-answer pairs and test them on 12 open-source & 5 proprietary LMMs. Our findings indicate that while LMMs have a foundational understanding of video quality, their performance remains incomplete and imprecise, with a notable discrepancy compared to human-level performance. Through **Q-Bench-Video**, we seek to catalyze community interest, stimulate further research, and unlock the untapped potential of LMMs to close the gap in video quality understanding."},"_bibtex":{"value":"@misc{\nzhang2024qbenchvideo,\ntitle={Q-Bench-Video: Benchmarking the Video Quality Understanding of {LMM}s},\nauthor={Zicheng Zhang and Ziheng Jia and Haoning Wu and Chunyi Li and Zijian Chen and Yingjie Zhou and Wei Sun and Xiaohong Liu and Xiongkuo Min and Weisi Lin and Guangtao Zhai},\nyear={2024},\nurl={https://openreview.net/forum?id=VaUy5GZO3f}\n}"},"title":{"value":"Q-Bench-Video: Benchmarking the Video Quality Understanding of LMMs"},"pdf":{"value":"/pdf/b6665a67a0e193ca939b3b34f488e5e0b380c738.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"zhang|qbenchvideo_benchmarking_the_video_quality_understanding_of_lmms"},"authorids":{"value":["~Zicheng_Zhang7","~Ziheng_Jia1","~Haoning_Wu1","~Chunyi_Li1","~Zijian_Chen1","~Yingjie_Zhou1","~Wei_Sun12","~Xiaohong_Liu2","~Xiongkuo_Min1","~Weisi_Lin1","~Guangtao_Zhai1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Zicheng Zhang","Ziheng Jia","Haoning Wu","Chunyi Li","Zijian Chen","Yingjie Zhou","Wei Sun","Xiaohong Liu","Xiongkuo Min","Weisi Lin","Guangtao Zhai"]}},"version":2},{"content":{"venue":{"value":"CoRR 2022"},"pdf":{"value":"http://arxiv.org/pdf/2208.02245v1"},"venueid":{"value":"dblp.org/journals/CORR/2022"},"paperhash":{"value":"huang|minvis_a_minimal_video_instance_segmentation_framework_without_videobased_training"},"authorids":{"value":["~De-An_Huang1","~Zhiding_Yu1","https://dblp.org/search/pid/api?q=author:Anima_Anandkumar:"]},"html":{"value":"https://doi.org/10.48550/arXiv.2208.02245"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2208-02245,\n  publtype={informal},\n  author={De-An Huang and Zhiding Yu and Anima Anandkumar},\n  title={MinVIS: A Minimal Video Instance Segmentation Framework without Video-based Training},\n  year={2022},\n  cdate={1640995200000},\n  journal={CoRR},\n  volume={abs/2208.02245},\n  url={https://doi.org/10.48550/arXiv.2208.02245}\n}\n"},"abstract":{"value":"We propose MinVIS, a minimal video instance segmentation (VIS) framework that achieves state-of-the-art VIS performance with neither video-based architectures nor training procedures. By only training a query-based image instance segmentation model, MinVIS outperforms the previous best result on the challenging Occluded VIS dataset by over 10% AP. Since MinVIS treats frames in training videos as independent images, we can drastically sub-sample the annotated frames in training videos without any modifications. With only 1% of labeled frames, MinVIS outperforms or is comparable to fully-supervised state-of-the-art approaches on YouTube-VIS 2019/2021. Our key observation is that queries trained to be discriminative between intra-frame object instances are temporally consistent and can be used to track instances without any manually designed heuristics. MinVIS thus has the following inference pipeline: we first apply the trained query-based image instance segmentation to video frames independently. The segmented instances are then tracked by bipartite matching of the corresponding queries. This inference is done in an online fashion and does not need to process the whole video at once. MinVIS thus has the practical advantages of reducing both the labeling costs and the memory requirements, while not sacrificing the VIS performance. Code is available at: https://github.com/NVlabs/MinVIS"},"title":{"value":"MinVIS: A Minimal Video Instance Segmentation Framework without Video-based Training"},"authors":{"value":["De-An Huang","Zhiding Yu","Anima Anandkumar"]}},"tmdate":1731476973535,"pdate":1640995200000,"tcdate":1727773830662,"writers":["~"],"signatures":["~Zhiding_Yu1"],"forum":"OLPWQohP0e","license":"CC BY-SA 4.0","number":129351,"cdate":1640995200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1731476973535,"domain":"DBLP.org","id":"OLPWQohP0e","version":2},{"content":{"summary":{"value":"The paper proposes a method for generating large scale paired image datasets from real or synthetic video data. The method uses SIFT matching to identify spatially related patches between video frames, which then become pairs that can be used for representation learning via MAE, CroCo, etc. They build a dataset, pre-train a representation, then evaluate the representation on downstream tasks."},"presentation":{"value":"3 good"},"contribution":{"value":"4 excellent"},"soundness":{"value":"4 excellent"},"strengths":{"value":"- Being able to mine large data for pairs is a relevant task in representation learning.\n\n- They outperform multiview habitat on their evaluations.\n\n- The method, being based on classical techniques like SIFT and RANSAC, should scale well.\n\n- The method is simple but effective."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- (minor) some exposition on what the \n\n- This paper might be better suited for a computer vision venue\n\n- The paper targets dense vision tasks, but it would be interesting to see the method used to generate pairs for constrastive learning, as well as evaluations on non-dense tasks such as imagenet finetuning/linear probe accuracy."},"confidence":{"value":"4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work."},"questions":{"value":"Is there any reason this wouldn't be useful for general representation learning beyond dense tasks?"},"rating":{"value":"8: accept, good paper"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699637073199,"tcdate":1698946627193,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission8576/Reviewer_AdG9"],"signatures":["ICLR.cc/2024/Conference/Submission8576/Reviewer_AdG9"],"forum":"uhtQyRrTzY","number":3,"license":"CC BY 4.0","cdate":1698946627193,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission8576/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699637073199,"domain":"ICLR.cc/2024/Conference","replyto":"uhtQyRrTzY","id":"SeuCCNQlFs","forumContent":{"TLDR":{"value":"We propose a pretraining data curation method to mine multi-view image pairs from real videos and synthetic 3D environments for downstream dense vision tasks."},"venue":{"value":"ICLR 2024 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Data curation techniques","Masked Image Modeling","Dense vision tasks","Large scale pretraining","Self supervised learning","Datasets"]},"primary_area":{"value":"datasets and benchmarks"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Dense pixel-specific representation learning at scale has been bottlenecked due to the unavailability of large-scale multi-view datasets. Current methods for building effective pretraining datasets heavily rely on annotated 3D meshes, point clouds, and camera parameters from simulated environments, preventing them from building datasets from real-world data sources where such metadata is lacking. We propose a pretraining dataset-curation approach that does not require any additional annotations. Our method allows us to generate multi-view datasets from both real-world videos and simulated environments at scale. Specifically, we experiment with two scales: MIMIC-1M with 1.3M and MIMIC-3M with 3.1M multi-view image pairs. We train multiple models with different masked image modeling objectives to showcase the following findings: Representations trained on our automatically generated MIMIC-3M outperform those learned from expensive crowdsourced datasets (ImageNet-1K) and those learned from synthetic environments (MULTIVIEW-HABITAT) on two dense geometric tasks: depth estimation on NYUv2 (↑1.7%), and surface normals estimation on Taskonomy (↓2.05%). For dense tasks which also require object understanding, we outperform MULTIVIEW- HABITAT, on semantic segmentation on ADE20K (↑3.89%), pose estimation on MSCOCO (↑9.4%), and reduce the gap with models pre-trained on the object-centric expensive ImageNet-1K. We outperform even when the representations are frozen, and when downstream training data is limited to few-shot. Larger dataset (MIMIC-3M) significantly improves performance, which is promising since our curation method can arbitrarily scale to produce even larger datasets."},"_bibtex":{"value":"@misc{\nmarathe2024mimic,\ntitle={{MIMIC}: Masked Image Modeling with Image Correspondences},\nauthor={Kalyani Marathe and Mahtab Bigverdi and Nishat Anjum Khan and Tuhin Kundu and Aniruddha Kembhavi and Linda Shapiro and Ranjay Krishna},\nyear={2024},\nurl={https://openreview.net/forum?id=uhtQyRrTzY}\n}"},"title":{"value":"MIMIC: Masked Image Modeling with Image Correspondences"},"pdf":{"value":"/pdf/ee8ad62b33eb0278b5645e90eb0b1a621d9fd35a.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Withdrawn_Submission"},"paperhash":{"value":"marathe|mimic_masked_image_modeling_with_image_correspondences"},"authorids":{"value":["~Kalyani_Marathe1","~Mahtab_Bigverdi1","~Nishat_Anjum_Khan2","~Tuhin_Kundu1","~Aniruddha_Kembhavi1","~Linda_Shapiro1","~Ranjay_Krishna1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Kalyani Marathe","Mahtab Bigverdi","Nishat Anjum Khan","Tuhin Kundu","Aniruddha Kembhavi","Linda Shapiro","Ranjay Krishna"]}},"version":2},{"content":{"venue":{"value":"Frontiers Comput. Sci. 2025"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/s11704-024-40452-4.pdf"},"venueid":{"value":"dblp.org/journals/FCSC/2025"},"paperhash":{"value":"yue|learning_from_shortcut_a_shortcutguided_approach_for_explainable_graph_learning"},"authorids":{"value":["~Linan_Yue1","~Qi_Liu3","~Ye_Liu10","~Weibo_Gao1","~Fangzhou_Yao1"]},"html":{"value":"https://doi.org/10.1007/s11704-024-40452-4"},"_bibtex":{"value":"@article{DBLP:journals/fcsc/YueLLGY25,\n  author={Linan Yue and Qi Liu and Ye Liu and Weibo Gao and Fangzhou Yao},\n  title={Learning from shortcut: a shortcut-guided approach for explainable graph learning},\n  year={2025},\n  month={August},\n  cdate={1754006400000},\n  journal={Frontiers Comput. Sci.},\n  volume={19},\n  number={8},\n  pages={198338},\n  url={https://doi.org/10.1007/s11704-024-40452-4}\n}\n"},"abstract":{"value":"The remarkable success in graph neural networks (GNNs) promotes the explainable graph learning methods. Among them, the graph rationalization methods draw significant attentions, which aim to provide explanations to support the prediction results by identifying a small subset of the original graph (i.e., rationale). Although existing methods have achieved promising results, recent studies have proved that these methods still suffer from exploiting shortcuts in the data to yield task results and compose rationales. Different from previous methods plagued by shortcuts, in this paper, we propose a Shortcut-guided Graph Rationalization (SGR) method, which identifies rationales by learning from shortcuts. Specifically, SGR consists of two training stages. In the first stage, we train a shortcut guider with an early stop strategy to obtain shortcut information. During the second stage, SGR separates the graph into the rationale and non-rationale subgraphs. Then SGR lets them learn from the shortcut information generated by the frozen shortcut guider to identify which information belongs to shortcuts and which does not. Finally, we employ the non-rationale subgraphs as environments and identify the invariant rationales which filter out the shortcuts under environment shifts. Extensive experiments conducted on synthetic and real-world datasets provide clear validation of the effectiveness of the proposed SGR method, underscoring its ability to provide faithful explanations."},"title":{"value":"Learning from shortcut: a shortcut-guided approach for explainable graph learning"},"authors":{"value":["Linan Yue","Qi Liu","Ye Liu","Weibo Gao","Fangzhou Yao"]}},"tmdate":1768973500477,"pdate":1735689600000,"tcdate":1747225553257,"writers":["~"],"signatures":["~Fangzhou_Yao1"],"forum":"s33GvUqY2Q","license":"CC BY-SA 4.0","number":457468,"cdate":1754006400000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1768973500477,"domain":"DBLP.org","id":"s33GvUqY2Q","version":2},{"content":{"summary":{"value":"This paper proposes a framework for open-set video-to-audio generation.\nThe proposed method consists of two primary components: Modality-aware Masking Alignment (MMA) for capturing fine-grained synchronization patterns between audio and video, and Modality-aware Flow Generation (MFG), a flow-based generative modeling for audio synthesis.\nThe paper also introduces a new open-set benchmark, derived from AudioSet and Panda70M, to evaluate V2A models across a broader range of scenarios.\nExperiments on VGGSound and the new dataset show that the methods outperform existing methods."},"soundness":{"value":1},"confidence":{"value":4},"questions":{"value":"Please see Weaknesses above."},"rating":{"value":0},"code_of_conduct":{"value":"Yes"},"presentation":{"value":1},"contribution":{"value":1},"strengths":{"value":"- Exploring open-set video-to-audio generation is a meaningful and underexplored direction.\n- The attempt to construct a larger-scale and more diverse dataset is valuable for the community if properly validated and documented."},"flag_for_ethics_review":{"value":["Yes, Research integrity issues (e.g., plagiarism, dual submission)"]},"weaknesses":{"value":"While the paper addresses an interesting problem, the contribution lacks novelty and technical clarity.\nThe proposed approach overlaps substantially with prior work (Foley-Flow, CVPR2025), and several methodological components are insufficiently explained.\nMoreover, important baselines are missing, and the experimental validation is not convincing.\nGiven these issues, the reviewer recommends strong rejection.\n\n**Major overlap with prior work**  \nThe proposed framework appears highly similar to the CVPR2025 paper \"Foley-Flow: Coordinated Video-to-Audio Generation with\nMasked Audio-Visual Alignment and Dynamic Conditional Flows.\" However, this paper is not cited nor compared, raising serious concerns about novelty and scholarly diligence.\n\n**Misrepresentation of prior work**  \nThe claim that \"Most prior work conditions audio synthesis on category labels or text descriptions\" is incorrect. Recent models such as DiffFoley, Frieren[1], MMAudio[2] explicitly leverage pre-trained audiovisual representations and capture fine-grained semantic and temporal visual information. The proposed method should be compared with these recent V2A models.\n\n**Missing references to masked audio-visual modeling**  \nThe paper employs a masked audio-visual modeling strategy but fails to cite or position itself with respect to prior work on cross-modal masked audio-video modeling (e.g., [3-7]). This omission makes it difficult to assess the originality of the proposed alignment module.\n\n**Missing references to flow-based generative models**  \nThe paper employs Flow-based models, but it provides no citations or comparisons with existing flow-based generative models. It is hard to identify novelty in generative modeling without comparison.\n\n**Lack of clarity in model description**  \nThe methodology is insufficiently detailed, making it difficult to follow and reproduce. \nKey definitions and computational steps are ambiguous or missing:\n\n- In Eq(4), $\\tilde{\\mathcal{A}}$ and $\\tilde{\\mathcal{V}}$ are defined as reconstructed audio and video features, but it is unclear how they are obtained or by what model.\n- Section 3.1 defines $z$ as \"it follows a known distribution (e.g., Gaussian)\", whereas Section 3.3 states \"Instead of assuming a standard Gaussian prior\", which contradicts the earlier statement.\n- The Selective Frame Conditioning step lacks details on how key frames are selected.\n- The variable $v_t$ in Eq. (7) is unclear. Its relationship to $z_t$ in the flow model is not described.\n\nOverall, the technical explanation is incomplete and inconsistent, preventing reproducibility or meaningful evaluation.\n\n**Limited and outdated evaluation setup**  \n- The evaluation relies only on KLD, FAD, and Align ACC, which do not reflect current standards for video-to-audio generation. Recent works (e.g., Frieren, MMAudio) employ a broader set of metrics for comprehensive audiovisual evaluation (e.g., Inception Score for audio quality, ImageBind similarity for audiovisual semantic alignment, and Desync for audiovisual temporal alignment).\n- For the proposed open-set benchmark, it appears that the model was trained on this dataset (Section A). However, it is unclear whether baseline models were retrained on the same data or evaluated using pre-trained weights. Without consistent training conditions, the comparison would not be valid.\n\n**Accessibility and reproducibility issues**  \n- The project page is not accessible, and no results can be verified. This significantly limits the ability to assess the claimed improvements. \n\n**References:**   \n[1] Frieren: Efficient video-to-audio generation with rectified flow matching, NeurIPS 2024  \n[2] MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis, CVPR 2025  \n[3] Contrastive Audio-Visual Masked Autoencoder, ICLR 2023  \n[4] Audiovisual Masked Autoencoders, ICCV 2023  \n[5] MAViL: Masked Audio-Video Learners, NeurIPS 2023  \n[6] AV-MaskEnhancer: Enhancing Video Representations through Audio-Visual Masked Autoeocoder, ICTAI 2023  \n[7] CrossMAE: Cross-Modality Masked Autoencoders for Region-Aware Audio-Visual Pre-Training, CVPR 2024"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762916583874,"tcdate":1761902542563,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission3175/Reviewer_Z3XJ"],"signatures":["ICLR.cc/2026/Conference/Submission3175/Reviewer_Z3XJ"],"forum":"l81Q8t59N9","number":2,"license":"CC BY 4.0","cdate":1761902542563,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission3175/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762916583874,"domain":"ICLR.cc/2026/Conference","replyto":"l81Q8t59N9","id":"jywq19r2CK","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["video-to-audio generation","video-audio learning"]},"primary_area":{"value":"generative models"},"abstract":{"value":"Video-to-audio generation has emerged as a promising frontier for enriching multimodal understanding and synthesis. However, most existing approaches operate under closed-set assumptions, restricting training and evaluation to predefined categories and limiting generalization in open-world scenarios. Prior methods primarily rely on pre-trained vision-language or audio-language encoders such as CLIP and CLAP, overlooking the strong inherent video–audio correspondence that can directly guide cross-modal grounding. In this work, we present \\method, a novel framework for open-set video-to-audio generation that enforces semantic fidelity and rhythmic synchronization across modalities. Our approach introduces a modality-aware dynamic masking strategy, where audio segments are reconstructed from masked video frames and vice versa, enabling the model to capture fine-grained temporal alignment without relying solely on external encoders. Furthermore, we design a generalized masked flow-based module that conditions generation on selectively sampled video frames, significantly improving efficiency and fidelity while preserving cross-modal coherence. Comprehensive experiments on VGGSound and a newly curated open-set benchmark demonstrate that \\method consistently outperforms state-of-the-art baselines in both objective and perceptual metrics, achieving superior Fréchet Audio Distance (FAD) and Kullback–Leibler (KL) divergence scores. The project page can be found at: https://openfoley.github.io."},"_bibtex":{"value":"@misc{\nmo2025openfoley,\ntitle={OpenFoley: Open-Set Video-to-Audio Generation with Modality-Aware Masking and Flows},\nauthor={Shentong Mo and Yibing Song and Weihua Chen and Fan Wang},\nyear={2025},\nurl={https://openreview.net/forum?id=l81Q8t59N9}\n}"},"title":{"value":"OpenFoley: Open-Set Video-to-Audio Generation with Modality-Aware Masking and Flows"},"pdf":{"value":"/pdf/e681116a9cc0b10048e71bfe8e0ad093b7bf9152.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"mo|openfoley_openset_videotoaudio_generation_with_modalityaware_masking_and_flows"},"authorids":{"value":["~Shentong_Mo1","~Yibing_Song1","~Weihua_Chen1","~Fan_Wang6"]},"authors":{"value":["Shentong Mo","Yibing Song","Weihua Chen","Fan Wang"]}},"version":2},{"content":{"venue":{"value":"ICLR 2027 Conference Submission"},"keywords":{"value":["Shortcut Learning; Observational Identifiability; Model Auditing; Distribution Shift; Physical Sensing"]},"primary_area":{"value":"interpretability and explainable AI"},"abstract":{"value":"A shortcut can be real yet observationally undiagnosable. Shortcut auditing is therefore an identification problem before it is an estimation problem: a fixed predictor may rely on identity-associated information even when the data lack the same-state cross-identity comparisons needed to reveal that reliance. We define predictor-specific reliance under an explicit target comparison distribution and distinguish full-target reliance, the support-restricted population estimand, and the finite audit estimate. If the target assigns positive mass to unsupported comparisons, full-target reliance need not be point-identified. Under bounded prediction discrepancy, partial support yields an identified interval whose width contracts with target-weighted support coverage. This result motivates a support-aware audit that reports prediction contrast, admissible pair count, and target coverage separately, and returns \\textsc{Abstain} when no admissible requested comparison is available. In a controlled intervention that varies only auditor-visible support while holding the predictor and finite full-target oracle reference fixed, recovery fidelity improves in $39/40$ model--seed endpoints, while all $800/800$ zero-support views abstain. In CWRU vibration sensing, support thinning degrades recovery of the full-support observational reference in all five repetitions. At a fixed pair count, concentrated support incurs $12.4\\times$ the recovery error of broadly distributed support at the median seed. Stanford battery EIS provides a complementary boundary case: substantial deployment degradation does not identify its source. Shortcut audits must therefore establish which predictor comparisons are observationally supported before interpreting estimated reliance."},"_bibtex":{"value":"@inproceedings{\nanonymous2026a,\ntitle={A Shortcut Can Be Real Yet Undiagnosable: Identifiability Limits of Observational Shortcut Auditing},\nauthor={Anonymous},\nbooktitle={Submitted to The Fifteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=PnmACq6Yle},\nnote={under review}\n}"},"title":{"value":"A Shortcut Can Be Real Yet Undiagnosable: Identifiability Limits of Observational Shortcut Auditing"},"pdf":{"value":"/pdf/c7bf4a904915f7584fc9fe5eea9ecfbb85dc1d49.pdf"},"venueid":{"value":"ICLR.cc/2027/Conference/Submission"}},"tmdate":1791228533066,"tcdate":1789389232847,"writers":["ICLR.cc/2027/Conference","ICLR.cc/2027/Conference/Submission18035/Authors"],"signatures":["ICLR.cc/2027/Conference/Submission18035/Authors"],"forum":"PnmACq6Yle","license":"CC BY 4.0","number":18035,"cdate":1789389232847,"readers":["everyone"],"invitations":["ICLR.cc/2027/Conference/-/Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Bidding","ICLR.cc/2027/Conference/Submission18035/-/Full_Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Reviewing","ICLR.cc/2027/Conference/-/Edit"],"mdate":1791228533066,"odate":1791031328286,"domain":"ICLR.cc/2027/Conference","id":"PnmACq6Yle","version":2},{"content":{"venue":{"value":"IEEE Trans. Circuits Syst. Video Technol. 2024"},"pdf":{"value":"https://ieeexplore.ieee.org/iel7/76/10739869/10540307.pdf"},"venueid":{"value":"dblp.org/journals/TCSV/2024"},"paperhash":{"value":"liu|enlarged_motionaware_and_frequencyaware_network_for_compressed_video_artifact_reduction"},"authorids":{"value":["~Wang_Liu1","https://dblp.org/search/pid/api?q=author:Wei_Gao_0003:","https://dblp.org/search/pid/api?q=author:Ge_Li_0002:","https://dblp.org/search/pid/api?q=author:Siwei_Ma:","https://dblp.org/search/pid/api?q=author:Tiesong_Zhao:","~Hui_Yuan1"]},"html":{"value":"https://doi.org/10.1109/TCSVT.2024.3406425"},"_bibtex":{"value":"@article{DBLP:journals/tcsv/LiuGLMZY24,\n  author={Wang Liu and Wei Gao and Ge Li and Siwei Ma and Tiesong Zhao and Hui Yuan},\n  title={Enlarged Motion-Aware and Frequency-Aware Network for Compressed Video Artifact Reduction},\n  year={2024},\n  month={October},\n  cdate={1727740800000},\n  journal={IEEE Trans. Circuits Syst. Video Technol.},\n  volume={34},\n  number={10},\n  pages={10339-10352},\n  url={https://doi.org/10.1109/TCSVT.2024.3406425}\n}\n"},"abstract":{"value":"Making full use of spatial-temporal information is the key factor for removing compressed video artifacts. Recently, many deep learning-based compression artifact reduction methods have emerged. Among them, a series of methods based on deformable convolution have shown excellent capabilities in spatio-temporal feature extraction. However, local deformable offset prediction and pixel-wise inter-frame feature alignment in the unidirectional form limit the full utilization of temporal features in the existing method. Additionally, compressed video shows inconsistent degrees of distortion on different frequency components, and their restoration difficulty is also nonuniform. For the above problems presented by existing methods, we propose an enlarged motion-aware and frequency-aware network (EMAFA) to further extract spatio-temporal information and enhance information of different frequency components. To perceive different degrees of motion artifacts between compressed frames as accurately as possible, we design a bidirectional dense propagation pattern with pixel-wise and patch-wise deformable convolution (PIPA) module in the feature domain. In addition, we propose a multi-scale atrous deformable alignment (MSADA) module to enrich spatio-temporal features in image domain. Moreover, we design a multi-direction frequency enhancement (MDFE) module with multiple direction convolution to enhance the features of different frequency components. The experimental results show that the proposed method performs better than the state-of-the-art methods in both objective evaluation and visual perception experience. Supplementary experiments for Internet Streamed Video with hybrid-distortion demonstrate that our method also exhibits considerable generalizability for quality enhancement."},"title":{"value":"Enlarged Motion-Aware and Frequency-Aware Network for Compressed Video Artifact Reduction"},"authors":{"value":["Wang Liu","Wei Gao","Ge Li","Siwei Ma","Tiesong Zhao","Hui Yuan"]}},"tmdate":1731483331573,"pdate":1704067200000,"tcdate":1731471594339,"writers":["~"],"signatures":["~Wang_Liu1"],"forum":"6CnliAiJwX","license":"CC BY-SA 4.0","number":196987,"cdate":1727740800000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1731483331573,"domain":"DBLP.org","id":"6CnliAiJwX","version":2},{"content":{"summary":{"value":"To make up for the gap of lacking attention in streaming video understanding, this paper introduces an novel designed SVBench which try to assess how well large multi-modal language models handle temporal multi-turn question-answering dialogues over streaming videos. SVBench includes 49,979 QA pairs derived from 1,353 videos and all\nannotations are interconnected through temporal linkages, ensuring that the model needs to consider previous and current video segments to answer questions correctly.\nThe authors developed StreamingChat, which significantly improved performance by incorporating long-context reasoning abilities specific to streaming videos. They also leveraged advanced training techniques like fine-tuning with LoRA (Low-Rank Adaptation) to handle long video contexts efficiently"},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. According to the paper, the length of video data used by SVBench is between 5 and 30 seconds. Has the author considered the maximum video processing time of StreamingChat?\n2. Can the author provide the test results of the latest streaming video understanding model Flash-VStream and Video-online on SVBench? This will further highlight the advantages of the article.\n3. Can the author provide some data on the model's inference speed and resource consumption? For example, the change curve of inference speed and current consumption under different video lengths, which may be what I am more interested in."},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The SVBench proposed fill the gap between video benchmarks and streaming video understanding. In the real world, streaming video is a more challenging data form, so this benchmark has very important practical significance.\n2. The authors proposed a semi-automatic annotation process and integrated multiple types of video data, which not only maintained the diversity of the benchmark, but also the accuracy of the annotation information and prevented hallucinations from affecting the benchmark results.\n3. In response to the observed limitations of current large multi-modal langeuage models, the authors introduce StreamingChat, which significantly improves performance on SVBench.\n4. The paper conducts an extensive evaluation of 14 models, comparing both open-source and closed-source models,  provides valuable insights into the current state-of-the-art models in handling streaming video understanding."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Although the authors provide an intuitive expression of the proposed SVBench in terms of video types and the diversity of annotation information through visualization results, the authors lack a description of the distribution of video lengths, which may be important for a benchmark.\n2. The training method of the StreamingChat architecture proposed by the author is similar to the recently proposed streaming video understanding model Video-online. I hope the author can explain the difference between them in more detail.\n3. The author's experimens are rich, but lacks measurements of the overall system's latency and inference speed like FPS, which are also very important matrics for streaming video understanding."}},"nonreaders":[],"tmdate":1733125475083,"tcdate":1730528460918,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission7487/Reviewer_6qmJ"],"signatures":["ICLR.cc/2025/Conference/Submission7487/Reviewer_6qmJ"],"forum":"Hz4BYVY8YM","number":2,"license":"CC BY 4.0","cdate":1730528460918,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission7487/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1733125475083,"domain":"ICLR.cc/2025/Conference","replyto":"Hz4BYVY8YM","id":"penOkpLdc0","forumContent":{"venue":{"value":"ICLR 2025 Spotlight"},"TLDR":{"value":"A benchmark with temporal multi-turn dialogues specifically designed to thoroughly assess the capabilities of long-context streaming video understanding of current LVLMs."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Multimodal large language model","Streaming video analysis","Video understanding"]},"primary_area":{"value":"datasets and benchmarks"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Despite the significant advancements of Large Vision-Language Models (LVLMs) on established benchmarks, there remains a notable gap in suitable evaluation regarding their applicability in the emerging domain of long-context streaming video understanding. Current benchmarks for video understanding typically emphasize isolated single-instance text inputs and fail to evaluate the capacity to sustain temporal reasoning throughout the entire duration of video streams. To address these limitations, we introduce SVBench, a pioneering benchmark with temporal multi-turn question-answering chains specifically designed to thoroughly assess the capabilities of streaming video understanding of current LVLMs. We design a semi-automated annotation pipeline to obtain 49,979 Question-Answer (QA) pairs of 1,353 streaming videos, which includes generating QA chains that represent a series of consecutive multi-turn dialogues over video segments and constructing temporal linkages between successive QA chains. Our experimental results, obtained from 14 models in dialogue and streaming evaluations, reveal that while the closed-source GPT-4o outperforms others, most open-source LVLMs struggle with long-context streaming video understanding. We also construct a StreamingChat model, which significantly outperforms open-source LVLMs on our SVBench and achieves comparable performance on diverse vision-language benchmarks. We expect SVBench to advance the research of streaming video understanding by providing a comprehensive and in-depth analysis of current LVLMs. Our benchmark and model can be accessed at https://yzy-bupt.github.io/SVBench."},"_bibtex":{"value":"@inproceedings{\nyang2025svbench,\ntitle={{SVB}ench: A Benchmark with Temporal Multi-Turn Dialogues for Streaming Video Understanding},\nauthor={Zhenyu Yang and Yuhang Hu and Zemin Du and Dizhan Xue and Shengsheng Qian and Jiahong Wu and Fan Yang and Weiming Dong and Changsheng Xu},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=Hz4BYVY8YM}\n}"},"title":{"value":"SVBench: A Benchmark with Temporal Multi-Turn Dialogues for Streaming Video Understanding"},"pdf":{"value":"/pdf/570cf446ad6d2a9682beb96af5d08dfb6b98d95a.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"yang|svbench_a_benchmark_with_temporal_multiturn_dialogues_for_streaming_video_understanding"},"authorids":{"value":["~Zhenyu_Yang6","~Yuhang_Hu3","~Zemin_Du2","~Dizhan_Xue1","~Shengsheng_Qian1","~Jiahong_Wu2","~Fan_Yang30","~Weiming_Dong1","~Changsheng_Xu1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Zhenyu Yang","Yuhang Hu","Zemin Du","Dizhan Xue","Shengsheng Qian","Jiahong Wu","Fan Yang","Weiming Dong","Changsheng Xu"]}},"version":2},{"content":{"summary":{"value":"This paper introduces VideoMV for multi-view 3D content generation. The key insight is to fine-tune existing video generation models for multi-view generation, leveraging the inherent frame consistency in video models. Specifically, VideoMV consists of three stages: 1) fine-tuning a video generation model for multi-view generation; 2) training a feed-forward reconstruction module for explicit 3D modeling; 3) proposing a 3D-aware denoising sampling strategy to enhance multi-view consistency. Experimental results show that VideoMV outperforms the selected existing methods in both generation quality and training efficiency."},"soundness":{"value":2},"confidence":{"value":5},"questions":{"value":"Why reconstruction metrics such as PSNR, Lpips can be considered as 3D consistency metrics as stated in the manuscripts?"},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The paper is well-written and structured, with clear presentation of the methodology and results.\n2. The overall pipeline is comprehensive and elegant, integrating both multi-view image generation and 3D Gaussian reconstruction in a unified framework.\n3. The proposed 3D-aware denoising sampling strategy is well-motivated and effective in improving multi-view consistency."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The comparison with baseline methods is incomplete. Specifically, there are several existing works that also fine-tune video diffusion models for multi-view generation, such as SV3D [1] and VFusion3D [2]. The lack of those baselines make me not convinced about the superiority of the proposed methods in the experiment part.\n\n2. Once again, if we take paper like SV3D into consideration, the author need to prove their proposed 3D-aware denoising sampling strategy --- which I believe is the most novel part in this paper is somehow compariable with the fine-tuning stragtegy in SV3D.\n\n3. Given that the reconstruction network is fine-tuned from LGM, a critical baseline comparison with MVDream+LGM is notably absent in Figure 6. \n\n4. The paper's emphasis appears misaligned with its claimed contributions. If the main novelty lies in the fine-tuned video diffusion model rather than the adapted reconstruction network, the experimental results should focus more on comparing multi-view generation capabilities instead of reconstruction results\n\n5. It is also unfair to compare with single-view reconstruction methods like Open-LRM\n\n\n\n[1] Sv3d: Novel multi-view synthesis and 3d generation from a single image using latent video diffusion\n[2] Vfusion3d: Learning scalable 3d generative models from video diffusion models"}},"nonreaders":[],"tmdate":1731427652174,"tcdate":1730610884623,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission2766/Reviewer_BPpx"],"signatures":["ICLR.cc/2025/Conference/Submission2766/Reviewer_BPpx"],"forum":"RtFWWAXIyH","number":4,"license":"CC BY 4.0","cdate":1730610884623,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission2766/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427652174,"domain":"ICLR.cc/2025/Conference","replyto":"RtFWWAXIyH","id":"RDsmfTlWe5","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"TLDR":{"value":"A novel framework that can generate 24 dense views and converges much faster in training than state-of-the-art approaches and outperforms existing state-of-the-art methods in both quantitative metrics and visual effects."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["3D Generation","Novel View Synthesis"]},"supplementary_material":{"value":"/attachment/6d425d7bd0c2b80fa8e5f2df54e92cfa714cad96.zip"},"primary_area":{"value":"generative models"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Generating multi-view images based on text or single-image prompts is a central topic in 3D content creation. Two fundamental questions on this topic are what data we use for training and how to ensure multi-view consistency. This paper introduces a novel framework that makes fundamental contributions to both questions. Unlike leveraging images from 2D diffusion models for training, we propose a dense consistent multi-view generation model that is fine-tuned from off-the-shelf video generative models. Images from video generative models are more suitable for multi-view generation because the underlying network architecture employs a temporal module to enforce frame consistency. Moreover, the video data sets used to train these models are abundant and diverse, leading to a reduced train-finetuning domain gap. To enhance multi-view consistency during generation, we introduce a 3D-Aware Denoising Sampling procedure, which first employs a feed-forward reconstruction module to get an explicit global 3D model, and then adopts a sampling strategy that effectively involves images rendered from the global 3D model into the denoising sampling loop to improve the multi-view consistency of the final images. As a by-product, this module also provides a fast way to create 3D assets represented by 3D Gaussians within a few seconds. Our approach can generate 24 dense views and converges much faster in training than state-of-the-art approaches (4 GPU hours versus many thousand GPU hours) with comparable visual quality and consistency. By further fine-tuning, our approach outperforms existing state-of-the-art methods in both quantitative metrics and visual effects."},"_bibtex":{"value":"@misc{\nzuo2024videomv,\ntitle={Video{MV}: Consistent Multi-View Generation Based on Large Video Generative Model},\nauthor={Qi Zuo and Xiaodong Gu and Lingteng Qiu and Yuan Dong and zhengyi zhao and Weihao Yuan and Rui Peng and Siyu Zhu and Liefeng Bo and Zilong Dong and Qixing Huang},\nyear={2024},\nurl={https://openreview.net/forum?id=RtFWWAXIyH}\n}"},"title":{"value":"VideoMV: Consistent Multi-View Generation Based on Large Video Generative Model"},"pdf":{"value":"/pdf/ec80f7e01f8b5d18b9ccc08fbe9e10ba43679add.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"zuo|videomv_consistent_multiview_generation_based_on_large_video_generative_model"},"authorids":{"value":["~Qi_Zuo1","~Xiaodong_Gu3","~Lingteng_Qiu1","~Yuan_Dong3","~zhengyi_zhao2","~Weihao_Yuan1","~Rui_Peng1","~Siyu_Zhu1","~Liefeng_Bo1","~Zilong_Dong2","~Qixing_Huang1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Qi Zuo","Xiaodong Gu","Lingteng Qiu","Yuan Dong","zhengyi zhao","Weihao Yuan","Rui Peng","Siyu Zhu","Liefeng Bo","Zilong Dong","Qixing Huang"]}},"version":2},{"content":{"summary":{"value":"The paper studies shortcut learning under a controlled synthetic distribution shift with two features: an invariant causal feature $x_1$ that determines the label, and a spurious training shortcut $x_2$ correlated with the label only during training. The authors studied a binary classification task, how shortcut reliance evolves with dataset scaling."},"significance":{"value":2},"workshop_fit":{"value":3},"strengths":{"value":"1. The authors introduce a direct functional measure of shortcut reliance, spurious-feature gradient sensitivity, showing shortcut dependence can grow even when test accuracy is near-saturated with dataset scaling. \n\n2. The synthetic two-feature construction (causal $x_1$, spurious $x_2$) with an explicit train–test decorrelation makes the shortcut mechanism easy to isolate and interpret."},"confidence":{"value":3},"suggestions":{"value":"1. The optimizer result is interesting. Adam/AdamW substantially reduces spurious gradient sensitivity relative to SGD under the authors’ synthetic setup. However, this appears primarily empirical, and the underlying mechanism is not isolated, so additional investigation across settings (architectures, hyperparameter matching, and real datasets) would strengthen the claim. \n\n2. The related-work discussion would be stronger if it positioned the observed shortcut amplification and optimizer effects relative to prior theory on simplicity bias[1] and gradient starvation[2], which are not currently cited or discussed.\n\n[1] The Pitfalls of Simplicity Bias in Neural Networks, Harshay Shah and Kaustav Tamuly and Aditi Raghunathan and Prateek Jain and Praneeth Netrapalli, NIPS'20\n\n[2] Gradient starvation: a learning proclivity in neural networks, Pezeshki, Mohammad and Kaba, S\\'{e}kou-Oumar and Bengio, Yoshua and Courville, Aaron and Precup, Doina and Lajoie, Guillaume, NIPS' 21"}},"parentInvitations":"ICLR.cc/2026/Workshop/Sci4DL/-/Official_Review","nonreaders":[],"tmdate":1772449479697,"tcdate":1772078255219,"writers":["ICLR.cc/2026/Workshop/Sci4DL","ICLR.cc/2026/Workshop/Sci4DL/Submission58/Reviewer_t5g5"],"signatures":["ICLR.cc/2026/Workshop/Sci4DL/Submission58/Reviewer_t5g5"],"forum":"VzCAvoXjeg","number":2,"license":"CC BY 4.0","cdate":1772078255219,"readers":["everyone"],"invitations":["ICLR.cc/2026/Workshop/Sci4DL/Submission58/-/Official_Review","ICLR.cc/2026/Workshop/Sci4DL/-/Edit"],"mdate":1772449479697,"domain":"ICLR.cc/2026/Workshop/Sci4DL","replyto":"VzCAvoXjeg","id":"RaPGNAx6aQ","forumContent":{"TLDR":{"value":"Scaling training data can increase neural networks’ reliance on spurious shortcut features even when accuracy stays high, while Adam/AdamW reduce this shortcut amplification compared to SGD."},"venue":{"value":"Submitted to Sci4DL 2026"},"keywords":{"value":["shortcut learning","spurious correlations","dataset scaling","gradient sensitivity","distribution shift","out-of-distribution generalization","synthetic binary classification","invariant causal feature","spurious feature reinforcement","optimizer implicit bias","SGD","Adam","AdamW","robustness diagnostics","critical onset scaling","beta-scaling phase boundary"]},"abstract":{"value":"Deep neural networks often exploit spurious shortcuts, non-causal correlations that fail under distribution shift. In a controlled synthetic binary classification setting with one invariant causal feature and one label-correlated shortcut, we study how shortcut reliance evolves with dataset scaling. Using gradient sensitivity to the spurious dimension as a direct functional diagnostic, we show a scaling-induced amplification effect: as training set size increases, models become increasingly sensitive to the shortcut feature despite near-saturated test accuracy. We further find that optimizer choice modulates this reinforcement, with Adam and AdamW substantially suppressing spurious gradient growth relative to stochastic gradient descent (SGD)."},"_bibtex":{"value":"@misc{\nanonymous2026when,\ntitle={When Data Amplifies Shortcuts: Gradient-Flow Evidence of Spurious Feature Reinforcement},\nauthor={Anonymous},\nyear={2026},\nurl={https://openreview.net/forum?id=VzCAvoXjeg}\n}"},"title":{"value":"When Data Amplifies Shortcuts: Gradient-Flow Evidence of Spurious Feature Reinforcement"},"Anonymization":{"value":"This submission has been anonymized for double-blind review via the removal of identifying information such as names, affiliations, and identifying URLs."},"style_files":{"value":"I have used the style files."},"pdf":{"value":"/pdf/dc08a6582d5577c7433d2ecce26a468b2dd0d0b2.pdf"},"venueid":{"value":"ICLR.cc/2026/Workshop/Sci4DL/Rejected_Submission"}},"version":2},{"content":{"summary":{"value":"This paper proposes egocentric-exocentric video groups alignment pretraining. The motivation is to align groups of ego and exo videos, in contrast to aligning pairs of ego and exo videos. The authors propose a two-step pretraining strategy and demonstrate improvement on two downstream tasks."},"soundness":{"value":2},"confidence":{"value":5},"questions":{"value":"See weaknesses"},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"+ The motivation is valid. It makes sense to me to explore the idea on utilizing the multi-view videos to form ego and exo video groups (compared with just enforcing ego-exo pair alignment with multi-view videos). \n\n+ The experiments show performance gain of EVGAP. The novel view inference setting (section 4.5) is interesting."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"+ The paper is not very well-written, a few clarifications are needed:\n1. L158: How are the scenes defined and identified? Also, is the scene definition incorporated in stage-I training?\n2. L161: The objective of EVGAP is to \"align and pair\" ego-exo video groups. However, this description is overly simplistic. Please expand on what \"align and pair\" specifically entails in this context. Provide more formal definitions or explanations to clarify the intended meaning.\n\n\n+ Experiments: evaluation setting is weak. The authors only conduct experiments on two datasets (Charades-Ego and Assembly101) and downstream tasks are only on Assembly101. Moreover, while the authors review a few ego-exo view-invariant works (L37-38), none of them is implemented as a baseline for comparison with EVGAP. The experiments are only about evaluating whether EVGAP gives additional performance gain on top of the task-specific approaches. The gain is expected since EVGAP benefits from more data in the pretraining. However, the real comparison should be about how different ego-exo feature learning approaches perform on these downstream tasks. I believe adding those baselines is essential. \n\n\n+ Method novelty is limited. I feel the major claim of the paper is to better utilize the multi-view ego-exo videos, extend a regular contrastive loss to account for groups of ego and exo videos. The contribution is incremental from my understanding. Moreover, the claim of using groups and objective 2 is better than using pairs is not thoroughly evaluated. For example, in Table 4 and 5, the authors should report results of aligning all possible views as pairs, to demonstrate the superiority of the proposed objective. Otherwise, the improvement could be understood from introducing multi-view data in pretraining.\n\nOverall, my concern with the paper is its limited novelty and weak evaluation. Specifically, I feel that the biggest claim made in the paper is not thoroughly evaluated, and there is a noticeable absence of baseline comparisons."}},"nonreaders":[],"tmdate":1731428551128,"tcdate":1730607346445,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission5997/Reviewer_JGNn"],"signatures":["ICLR.cc/2025/Conference/Submission5997/Reviewer_JGNn"],"forum":"F0K0zxi62U","number":2,"license":"CC BY 4.0","cdate":1730607346445,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission5997/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428551128,"domain":"ICLR.cc/2025/Conference","replyto":"F0K0zxi62U","id":"vk1sB864EU","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["multiview video","view-invariant pretraining","view alignment","ego-exo pair alignment"]},"supplementary_material":{"value":"/attachment/0e02a68395c5d5f7c0c30f848f4b30b2db1e9fcd.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Aligning egocentric and exocentric videos facilitates the learning of view-invariant features, which significantly contributes to video understanding. While previous approaches have primarily focused on aligning individual ego-exo video pairs, our method extends this concept by aligning groups of synchronized egocentric and exocentric videos.  This strategy enables the model to capture more comprehensive cross-view relationships across densely captured viewpoints, enhancing its capacity for robust multi-view understanding.\nTherefore, we develop a pipeline based on contrastive learning for \\textbf{E}gocentric-exocentric \\textbf{V}ideo \\textbf{G}roups \\textbf{A}lignment \\textbf{P}re-training (EVGAP). \nOur method introduces several key innovations: 1) a novel video pre-training paradigm that extends alignment from ego-exo video pairs to ego-exo video group alignments; 2) an innovative two-step training process that leverages the abundant ego-exo video pair data to support the learning of ego-exo video group alignments, transitioning from sparse to dense viewpoints; and 3) the application of auxiliary losses to progressively align videos from different perspectives.\nExtensive ablations illustrate the effectiveness of our approach in single-view and multi-view downstream tasks. We also find that our approach facilitates the tasks inluding novel views. The codes will be available upon acceptance."},"_bibtex":{"value":"@misc{\nwang2024evgap,\ntitle={{EVGAP}: Egocentric-Exocentric Video Groups Alignment Pre-training},\nauthor={Peiyao Wang and Haibin Ling},\nyear={2024},\nurl={https://openreview.net/forum?id=F0K0zxi62U}\n}"},"title":{"value":"EVGAP: Egocentric-Exocentric Video Groups Alignment Pre-training"},"pdf":{"value":"/pdf/9dbc7419b37fbcd3e8d90f9d4c03bc38178865b7.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"wang|evgap_egocentricexocentric_video_groups_alignment_pretraining"},"authorids":{"value":["~Peiyao_Wang2","~Haibin_Ling1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Peiyao Wang","Haibin Ling"]}},"version":2},{"content":{"summary":{"value":"The paper presents a new dataset to evaluate model capability for domain generalization in a video comprehension task. Videos are filtered using a multimodal LLM to ensure they fall into certain domains. The videos are then filtered based on the actions again, using a multimodal LLM so that videos in the different domains have similar semantics. QA pairs are automatically generated for each video, which are then reviewed by human experts. The dataset offers both training and test sets. The paper provides evaluation results over multiple models in multi-domain fine-tuning, single-domain fine-tuning, and zero-shot settings. Human evaluation is also provided to show the relevance and correctness of the automatically generated WA pairs."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"I want to see some discussion on the weaknesses."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. I believe the paper offers a useful dataset (once it is published online), which should contribute to the community. \n2. The paper also offers comprehensive details of how the dataset is created."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. This may be a minor concern, but since the dataset only offers limited semantics, the evaluation capability may also be limited. It would be nice if the paper could provide some discussion on this point (e.g., whether the set of actions is sufficient for evaluating general performance, and why). \n2. As domain generalization can involve fine-tuning, it is desirable to evaluate possible bias in QA pairs and semantics. This may be done by training a language model with only QA pairs and some tokens to represent semantics. \n3. It’s hard for me to see what the human evaluation results in Table 10 mean. Human-annotated QAs and some randomization (e.g., randomly pairing up a video and QA in different domains but in the same semantics, or just random video and QA pairs) may help understand it. \n4. I also don’t see why human evaluation uses a 5-point scale. Relevance and correctness sound more like binary."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762925837731,"tcdate":1762001396799,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission15564/Reviewer_RsGD"],"signatures":["ICLR.cc/2026/Conference/Submission15564/Reviewer_RsGD"],"forum":"0mUiXz1TNq","number":2,"license":"CC BY 4.0","cdate":1762001396799,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission15564/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762925837731,"domain":"ICLR.cc/2026/Conference","replyto":"0mUiXz1TNq","id":"Cj7Tq64cpK","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Video Understanding","Dataset","Domain Generalization"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Video understanding has made remarkable progress in recent years, largely driven by advances in deep models and the availability of large-scale annotated datasets.\nHowever, the robustness of these models to domain shifts encountered in real-world video applications remains a critical yet underexplored problem, limiting their practical reliability.\nTo address this problem, we introduce \\textbf{V}ideo \\textbf{U}nderstanding \\textbf{D}omain \\textbf{G}eneralization (\\textbf{VUDG}), the first dataset designed specifically for evaluating domain generalization in video understanding.\nVUDG contains videos from 11 distinct domains that cover three types of domain shifts, and maintains semantic consistency across different domains to ensure fair and meaningful evaluation. We propose a multi-expert progressive annotation framework to efficiently annotate videos with structured question-answer pairs designed for domain generalization.\nExtensive experiments on 9 representative Large Vision-Language Models (LVLMs) and several traditional video question answering methods show that most models (including state-of-the-art LVLMs) suffer performance degradation under domain shifts. \nThese results highlight the challenges posed by VUDG and the difference in the robustness of current models to data distribution shifts. We believe VUDG provides a critical resource to benefit future research in domain generalization for video understanding."},"_bibtex":{"value":"@inproceedings{\nwang2026vudg,\ntitle={{VUDG}: A Dataset for Video Understanding Domain Generalization},\nauthor={Ziyi Wang and Zhi Gao and Boxuan Yu and Zirui Dai and Peiyao Wang and Yuxiang Song and Qingyuan Lu and Jin Chen and Xinxiao Wu},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=0mUiXz1TNq}\n}"},"title":{"value":"VUDG: A Dataset for Video Understanding Domain Generalization"},"pdf":{"value":"/pdf/feec85898ca8b7da1de77de6c74b1c8ebabd4c20.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"wang|vudg_a_dataset_for_video_understanding_domain_generalization"},"authorids":{"value":["~Ziyi_Wang18","~Zhi_Gao5","~Boxuan_Yu1","~Zirui_Dai2","~Peiyao_Wang5","~Yuxiang_Song2","~Qingyuan_Lu2","~Jin_Chen2","~Xinxiao_Wu1"]},"authors":{"value":["Ziyi Wang","Zhi Gao","Boxuan Yu","Zirui Dai","Peiyao Wang","Yuxiang Song","Qingyuan Lu","Jin Chen","Xinxiao Wu"]}},"version":2},{"content":{"venue":{"value":"IEEE Trans. Circuits Syst. Video Technol. 2025"},"pdf":{"value":"https://ieeexplore.ieee.org/iel8/76/10949577/10755967.pdf"},"venueid":{"value":"dblp.org/journals/TCSV/2025"},"paperhash":{"value":"zhao|finegrained_modality_relationaware_network_for_video_moment_retrieval"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Yibo_Zhao_0001:","~Zan_Gao1","https://dblp.org/search/pid/api?q=author:Chunjie_Ma:","~Weili_Guan3","~Riwei_Wang1","~Shengyong_Chen2"]},"html":{"value":"https://doi.org/10.1109/TCSVT.2024.3494744"},"_bibtex":{"value":"@article{DBLP:journals/tcsv/ZhaoGMGWC25,\n  author={Yibo Zhao and Zan Gao and Chunjie Ma and Weili Guan and Riwei Wang and Shengyong Chen},\n  title={Fine-Grained Modality Relation-Aware Network for Video Moment Retrieval},\n  year={2025},\n  month={April},\n  cdate={1743465600000},\n  journal={IEEE Trans. Circuits Syst. Video Technol.},\n  volume={35},\n  number={4},\n  pages={3315-3327},\n  url={https://doi.org/10.1109/TCSVT.2024.3494744}\n}\n"},"abstract":{"value":"Video moment retrieval (VMR) involves localizing video segments semantically aligned with given queries within videos. Despite the development of numerous methods for VMR in recent years, there remains a need to better incorporate fine-grained modality relation-aware information both in intra-modality and cross-modality. To address these challenges, we propose a Fine-grained Modality Relation-Aware Network (FMRN) tailored for the video moment retrieval task. FMRN effectively explores fine-grained modality relation-aware information within text queries, videos, and proposals. Our approach begins with a semantic graph encoder to capture deep semantic relations in intra-modality. Besides, we introduce a novel fine-grained cross-modality interaction module comprising a cross-similarity weighting module, an intra-modality weighting module, and an adaptive fusion module. These components comprehensively exploit fine-grained relation information within intra-modality and cross-modality contexts. Specifically, the cross-similarity weighting module leverages similarities between text queries and video snippets, as well as between videos and query words. The intra-modality weighting module determines the importance of words and snippets, while the adaptive fusion module combines cross-similarity weighting and intra-modality weighting. Additionally, we design a proposal relation module to enhance retrieval by capturing fine-grained proposals-relation information in videos. Extensive experiments demonstrate that the proposed method can outperform all state-of-the-art methods on the TACoS dataset and obtain comparable results on the Charades-STA and ActivityNet-Captions datasets. Compared with MCMN (TCSVT2024) and DPHANet (TMM2024), FMRN can achieve average improvements of 3.61 % and 5.44 % on the TACoS dataset, respectively."},"title":{"value":"Fine-Grained Modality Relation-Aware Network for Video Moment Retrieval"},"authors":{"value":["Yibo Zhao","Zan Gao","Chunjie Ma","Weili Guan","Riwei Wang","Shengyong Chen"]}},"tmdate":1768523279081,"pdate":1735689600000,"tcdate":1744965452377,"writers":["~"],"signatures":["~Wang_RiWei1"],"forum":"pR2mS0KWn3","license":"CC BY-SA 4.0","number":404292,"cdate":1743465600000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1768523279081,"domain":"DBLP.org","id":"pR2mS0KWn3","version":2},{"content":{"summary":{"value":"The paper leverages existing open-source video datasets and utilizes the capabilities of large models to obtain a training dataset, Video-Thinker-10K, which includes question-answer pairs and chain-of-thought annotations. During the training process, the Video-Thinker-7B model was trained using the SFT+GRPO training strategy, outperforming several existing large model approaches on several common video QA datasets."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. During the data generation process, how can the hallucination phenomenon in the automatic annotation of large models be addressed?  \n2. In the training of Video CoT, what are the specific differences compared to GRPO training in Language or Image CoT?  \n3. Could more comprehensive training details be provided, such as the number of T?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"The paper is clearly written, and the specific prompt design for the dataset construction process is also well-explained. \n\nThe chain-of-thought (CoT) data annotation for video reasoning represents a notable contribution.\n\nThe phenomena observed during the chain-of-thought training process provide valuable insights."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The paper's technical contribution is limited. \nThe CoT annotations for video labeling primarily rely on the capabilities of the DeepSeek and Gemini models. \n\nThe training process of Video-Thinker-7B lacks contrution, as it mainly adopts the conventional approach of SFT+GRPO."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762927238865,"tcdate":1762173715385,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission17303/Reviewer_1Y79"],"signatures":["ICLR.cc/2026/Conference/Submission17303/Reviewer_1Y79"],"forum":"ofNbGPV6Ve","number":4,"license":"CC BY 4.0","cdate":1762173715385,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission17303/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762927238865,"domain":"ICLR.cc/2026/Conference","replyto":"ofNbGPV6Ve","id":"NHn7AigiyM","forumContent":{"TLDR":{"value":"We introduce Video-Thinker, a novel approach that empowers MLLMs to think with videos by autonomously leveraging their intrinsic grounding and captioning capabilities to generate reasoning clues throughout the inference process."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Video reasoning","Multimodal large language model","Thinking with videos"]},"primary_area":{"value":"foundation or frontier models, including LLMs"},"abstract":{"value":"Recent advances in image reasoning methods, particularly \"Thinking with Images\", have demonstrated remarkable success in Multimodal Large Language Models (MLLMs); however, this dynamic reasoning paradigm has not yet been extended to video reasoning tasks. In this paper, we propose Video-Thinker, which empowers MLLMs to think with videos by autonomously leveraging their intrinsic \"grounding\" and \"captioning\" capabilities to generate reasoning clues throughout the inference process. To spark this capability, we construct Video-Thinker-10K, a curated dataset featuring autonomous tool usage within chain-of-thought reasoning sequences. Our training strategy begins with Supervised Fine-Tuning (SFT) to learn the reasoning format, followed by Group Relative Policy Optimization (GRPO) to strengthen this reasoning capability. Through this approach, Video-Thinker enables MLLMs to autonomously navigate grounding and captioning tasks for video reasoning, eliminating the need for constructing and calling external tools. Extensive experiments demonstrate that Video-Thinker achieves significant performance gains on both in-domain tasks and challenging out-of-domain video reasoning benchmarks, including Video-Holmes, CG-Bench-Reasoning, and VRBench. Our Video-Thinker-7B substantially outperforms existing baselines such as Video-R1 and establishes state-of-the-art performance among 7B-sized MLLMs."},"_bibtex":{"value":"@misc{\nwang2026videothinker,\ntitle={Video-Thinker: Sparking ''Thinking with Videos'' via Reinforcement Learning},\nauthor={Shijian Wang and Jiarui Jin and Xingjian Wang and Linxin Song and Runhao Fu and Hecheng Wang and Zongyuan Ge and Yuan Lu and Xuelian Cheng},\nyear={2026},\nurl={https://openreview.net/forum?id=ofNbGPV6Ve}\n}"},"title":{"value":"Video-Thinker: Sparking \"Thinking with Videos\" via Reinforcement Learning"},"pdf":{"value":"/pdf/2434c327ac439ca9afeb57c45297d6f58cbecc77.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"wang|videothinker_sparking_thinking_with_videos_via_reinforcement_learning"},"authorids":{"value":["~Shijian_Wang1","~Jiarui_Jin1","~Xingjian_Wang2","~Linxin_Song1","~Runhao_Fu1","~Hecheng_Wang1","~Zongyuan_Ge1","~Yuan_Lu7","~Xuelian_Cheng2"]},"authors":{"value":["Shijian Wang","Jiarui Jin","Xingjian Wang","Linxin Song","Runhao Fu","Hecheng Wang","Zongyuan Ge","Yuan Lu","Xuelian Cheng"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a new approach to neuronal response modeling by predicting/forecasting the volumetric video instead of the per-neuron calcium trace (dF/F) or spike train, which is the norm in neural response prediction. This approach allows the model to take advantage of the inter-cell activity and spatial organization of the population that is typically discarded when deconvolving the volumetric video to individual response traces. The authors evaluated a range of video and trace-based models on ZAPBench and showed that the video-based model outperforms trace-based models in short temporal context length conditions."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"Major\n- It would be nice to test whether or not the trace-base model performed (relatively) poorly due to the lack of inter-cell activities or imperfect masking as suggested in Section 2 and Figure 2. I suggest applying the segmentation masks to the video, i.e. set regions outside of the identified cells as background, and train the video-based model on the masked videos. This should give us a sense of the influence of inter-cell activities and imperfect masking, and isolate the influence of spatial organization of the cells. \n- I suggest the authors include metrics that are commonly used in neural response prediction so that readers can have a sense of how well these models are performing, such as normalized correlation ($CC_\\text{norm}$) [1], or fraction of explainable variance explained (FEV) [2], basically metrics that takes trial-to-trial variability into account.\n- Can the authors comment on the computational cost of the models? The authors stated in the hyperparameter search section and appendix A.3 that 16 A100 40GB GPUs are used to train the video-based model, and ~5k GPU hours (so ~300 hours in wall-time) was used in the loss ablation experiment in Figure 4, which is a considerable amount of computational time and cost. Can the authors share the time it took to train the final (best) video-based and trace-based models? I believe the authors should discuss the trade-off between the two approaches in computation cost if they are indeed substantially different. To clarify, I think it is fine for the method to be more computationally expensive than other methods, but it is important to point it out.\n\nMinor\n- What is the frame rate of the video? \n- Why and how are the two temporal context lengths (4 and 256) selected? Does it make sense to predict the future 32 frames from only 4 frames?\n- In the hyperparameters section and Figure 1, it is stated that the models optimize the trace-based MAE. Does this include the video-based model? Since the video-based model inputs and outputs a video, does it make a difference to optimize the recorded and predicted video MAE?\n- How are the hyperparameters selected? Hand-picked or via some form of hyperparameter search (random search, bayesian search, etc.)\n\n[1] Schoppe, Oliver, et al. \"Measuring the performance of neural models.\" Frontiers in computational neuroscience 10 (2016): 10.\n\n[2] Cadena, Santiago A., et al. \"Deep convolutional models improve predictions of macaque V1 responses to natural images.\" PLoS computational biology 15.4 (2019): e1006897."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"- To my knowledge, neural response prediction/forecasting directly from volumetric video is novel. This allows minimal preprocessing of the data which can be beneficial to deep learning-based methods.\n- A wide range of training and evaluation conditions are compared, including trade-offs of spatial and temporal resolution, pre-training vs direct training, and training set size and combinations. These empirical results can guide future work in modeling neural responses."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Please find my suggestions for the following points in the Question section.\n- A key motivation of this work, as stated by the authors in Section 2, is that the typical deconvolution step to convert volumetric video to dF/F response traces can lead to loss of information, such as cell spatial organization, inter-cell activities, etc. However, while the video-based model appears to outperform the trace-based model in short temporal context-length conditions (though similar performance in longer context length), it is unclear whether or not this is due to the additional information that exists in the raw data and that the video-based model is indeed taking advantage of such information.\n- MAE might not be the most intuitive metric for getting a sense of how the (video and traced-based) models are performing. For instance, I am not sure if an MAE value of 0.02 is good or bad, or how big of a difference is an MAE of  0.02 to 0.04? In particular, I believe ZAPBench is a new dataset and we don’t have any other models to compare against these MAE values, other than the single trace-based model provided.\n- Unclear trade-off in computation cost between video-based and trace-based models."}},"nonreaders":[],"tmdate":1733175638666,"tcdate":1730423308427,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission4792/Reviewer_5NDV"],"signatures":["ICLR.cc/2025/Conference/Submission4792/Reviewer_5NDV"],"forum":"4UXIGATUTj","number":2,"license":"CC BY 4.0","cdate":1730423308427,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission4792/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1733175638666,"domain":"ICLR.cc/2025/Conference","replyto":"4UXIGATUTj","id":"W6MogSaSQU","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"TLDR":{"value":"We propose a new method for forecasting neuronal activity using volumetric videos, leveraging spatial relationships between neurons and outperforming traditional trace-based methods."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["neuroscience","forecasting","video","lightsheet microscopy","zebrafish","calcium imaging","neuron activity"]},"primary_area":{"value":"applications to neuroscience & cognitive science"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Large-scale neuronal activity recordings with fluorescent calcium indicators are increasingly common, yielding high-resolution 2D or 3D videos. Traditional analysis pipelines reduce this data to 1D traces by segmenting regions of interest, leading to inevitable information loss. Inspired by the success of deep learning on minimally processed data in other domains, we investigate the potential of forecasting neuronal activity directly from volumetric videos. To capture long-range dependencies in high-resolution volumetric whole-brain recordings, we design a model with large receptive fields, which allow it to integrate information from distant regions within the brain. We explore effects of pre-training and perform extensive model selection, analyzing spatio-temporal trade-offs for generating accurate forecasts. Our model outperforms trace-based forecasting approaches on ZAPBench, a recently proposed benchmark on whole-brain activity prediction in zebrafish, demonstrating the advantages of preserving the spatial structure of neuronal activity."},"_bibtex":{"value":"@misc{\nimmer2025forecasting,\ntitle={Forecasting Whole-Brain Neural Activity from Volumetric Video},\nauthor={Alexander Immer and Jan-Matthis Lueckmann and Alex Bo-Yuan Chen and Peter H. Li and Mariela D Petkova and Nirmala A Iyer and Aparna Dev and Gudrun Ihrke and Woohyun Park and Alyson Petruncio and Aubrey Weigel and Wyatt Korff and Florian Engert and Jeff Lichtman and Misha Ahrens and Viren Jain and Michal Januszewski},\nyear={2025},\nurl={https://openreview.net/forum?id=4UXIGATUTj}\n}"},"title":{"value":"Forecasting Whole-Brain Neural Activity from Volumetric Video"},"pdf":{"value":"/pdf/8174bd58ef8838d290a0c38318afaf4e36570ef8.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"immer|forecasting_wholebrain_neural_activity_from_volumetric_video"},"authorids":{"value":["~Alexander_Immer1","~Jan-Matthis_Lueckmann2","~Alex_Bo-Yuan_Chen1","~Peter_H._Li1","~Mariela_D_Petkova1","~Nirmala_A_Iyer1","~Aparna_Dev1","~Gudrun_Ihrke1","~Woohyun_Park2","~Alyson_Petruncio1","~Aubrey_Weigel1","~Wyatt_Korff1","~Florian_Engert1","~Jeff_Lichtman1","~Misha_Ahrens1","~Viren_Jain2","~Michal_Januszewski1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Alexander Immer","Jan-Matthis Lueckmann","Alex Bo-Yuan Chen","Peter H. Li","Mariela D Petkova","Nirmala A Iyer","Aparna Dev","Gudrun Ihrke","Woohyun Park","Alyson Petruncio","Aubrey Weigel","Wyatt Korff","Florian Engert","Jeff Lichtman","Misha Ahrens","Viren Jain","Michal Januszewski"]}},"version":2},{"content":{"summary":{"value":"The paper presents PhyMAGIC, a training-free framework that enables physically consistent 3D motion generation and physical property inference from a single static image. Unlike prior diffusion-based or physics-embedded models that require task-specific training, PhyMAGIC integrates three pretrained components—a video diffusion model (CogVideoX), a large language model (GPT-4o) for physics reasoning, and a differentiable Material Point Method (MPM) simulator—into a unified, closed-loop system. The framework iteratively refines motion generation through confidence-guided LLM feedback, where low-confidence physical attributes (e.g., density, elasticity, friction) trigger targeted prompt updates and new motion synthesis. The refined videos and inferred parameters are then validated and corrected via differentiable simulation, forming a reasoning–generation cycle that progressively improves physical plausibility without any fine-tuning or supervision.\n\nComprehensive experiments on benchmark datasets such as PhysGaussian, PhysGen, and various real-world single-image scenes show that PhyMAGIC achieves superior performance in both semantic and physical consistency compared to state-of-the-art video generation and physics-aware baselines. It enhances text–motion alignment (up to 16% CLIP similarity improvement) and maintains high visual fidelity while being more computationally efficient—over ten times faster than trained counterparts. Overall, the paper contributes a novel and generalizable approach that bridges LLM reasoning, diffusion-based video synthesis, and differentiable physics simulation to realize physically grounded, interpretable, and scalable dynamic generation from minimal visual input."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"The author did not include the code in the Supplementary Material. Do you have any plan for publishing the implementation?"},"rating":{"value":4},"details_of_ethics_concerns":{"value":"N/A"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. The paper introduces a conceptually elegant and practically valuable approach that unifies pretrained video diffusion models, large language model reasoning, and differentiable physics simulation into a closed-loop pipeline. This design enables physical property inference and motion synthesis from a single static image without any task-specific fine-tuning or annotated data, which is a notable advance in the direction of scalable physics-aware generation.\n2. The authors provide clear physical intuition for the design choices. The observation that different motion trajectories reveal distinct levels of physical evidence (e.g., a squeeze motion discloses elasticity better than free fall) is both physically grounded and conceptually insightful, forming a solid theoretical backbone for the proposed iterative reasoning framework.\n3. The explicit use of interpretable physical parameters (density, elasticity, yield stress, etc.) also makes the model’s behavior more transparent compared to end-to-end black-box video diffusion systems."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The paper does not include supplementary videos or visual demonstrations, which are crucial for evaluating a method that claims improvements in physical realism and motion plausibility. From the static figures alone, it is difficult to fully assess the perceptual quality and physical consistency of the generated dynamics. In particular, Figure 5 lacks ground-truth visualizations, making it hard to judge whether the proposed approach indeed performs better than the baselines in practice. The authors are strongly encouraged to provide additional qualitative video results for clearer evaluation.\n2. Several figures are visually coarse and fail to clearly convey the intended comparisons or architectural design. For example, Figure 1 and Figure 3 appear somewhat rough and could be redrawn with more precise annotations or higher-quality visuals. Moreover, the paper does not explain clearly how the baselines are re-implemented or compared. A more transparent comparison protocol would strengthen the credibility of the experimental results.\n3. The experimental section compares PhyMAGIC mainly against earlier open-source baselines such as CogVideoX, PhysDreamer, and Physics3D, but omits several state-of-the-art commercial or research-level models, including Wan 2.2, Sora, and Veo. Given the rapid progress in video generation, evaluating against these newer and higher-performing models would provide a fairer and more convincing benchmark for the claimed advances.\n4. The core component of PhyMAGIC relies on the assumption that large language models can accurately infer physical parameters (e.g., density, Young’s modulus, Poisson’s ratio) from limited visual cues. However, the experimental results do not strongly support this assumption. As shown in Table 8, even after multiple refinement iterations, the predicted physical values often deviate from ground truth by one or more orders of magnitude. This raises concerns about the robustness and reliability of LLM-based physical inference.\n5. While the framework leverages pretrained video diffusion models to synthesize motion evidence, current video generation models themselves often suffer from physical inconsistency, such as non-conservation of momentum and unrealistic object interactions. Relying on such imperfect priors may limit the upper bound of PhyMAGIC’s physical realism, especially when the diffusion model introduces visually plausible but physically implausible trajectories."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762921679678,"tcdate":1761449316471,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission10350/Reviewer_z1t3"],"signatures":["ICLR.cc/2026/Conference/Submission10350/Reviewer_z1t3"],"forum":"nruZar3Aaz","number":1,"license":"CC BY 4.0","cdate":1761449316471,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission10350/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762921679678,"domain":"ICLR.cc/2026/Conference","replyto":"nruZar3Aaz","id":"Js7XGF9zXY","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["3D Dynamic Generation","Phyical Priors","Physical Simulation","LLM Reasoning","MPM Simulation"]},"primary_area":{"value":"generative models"},"abstract":{"value":"Recent advances in 3D content generation have amplified demand for dynamic models that are both visually realistic and physically consistent. However, state-of-the-art video diffusion models frequently produce implausible results such as momentum violations and object interpenetrations. Existing physics-aware approaches often rely on task-specific fine-tuning or supervised data, which limits their scalability and applicability. To address the challenge, we present PhyMAGIC, a training-free framework that generates physically consistent motion from a single image. PhyMAGIC integrates a pre-trained image-to-video diffusion model, confidence-guided reasoning via large language models (LLMs), and a differentiable physics simulator to produce 3D assets ready for downstream physical simulation without fine-tuning or manual supervision. By iteratively refining motion prompts using LLM-derived confidence scores and leveraging simulation feedback, PhyMAGIC steers generation toward physically consistent dynamics. Comprehensive experiments demonstrate that PhyMAGIC outperforms state-of-the-art video generators and physics-aware baselines, enhancing physical property inference and motion–text alignment while maintaining visual fidelity."},"_bibtex":{"value":"@misc{\nmeng2025phymagic,\ntitle={Phy{MAGIC}: Physical Motion-Aware Generative Inference with Confidence-guided {LLM}},\nauthor={Siwei Meng and Yawei Luo and Ping Liu},\nyear={2025},\nurl={https://openreview.net/forum?id=nruZar3Aaz}\n}"},"title":{"value":"PhyMAGIC: Physical Motion-Aware Generative Inference with Confidence-guided LLM"},"pdf":{"value":"/pdf/af1594b19fb1335afb9429816fc8f2f5c4f00a74.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"meng|phymagic_physical_motionaware_generative_inference_with_confidenceguided_llm"},"authorids":{"value":["~Siwei_Meng1","~Yawei_Luo3","~Ping_Liu1"]},"authors":{"value":["Siwei Meng","Yawei Luo","Ping Liu"]}},"version":2},{"content":{"summary":{"value":"The paper proposed a two-stage video editing framework, using T2I diffusion model to edit the key frames and then interpolating between those frames. During T2I diffusion process, the paper leveraged controlnet to jointly keep the edge consistency. After that, a Masked generative transformer model called MaskINT is introduced to generate middle frames. The results show that the proposed network can accelerate generate videos compared with baseline pipelines while suffering slightly temporal and prompt consistency decrease."},"presentation":{"value":"4 excellent"},"contribution":{"value":"3 good"},"soundness":{"value":"3 good"},"strengths":{"value":"1. The proposed MaskINT leverage masked generative transformer to interpolate between keyframes.\n2. The inference speed outperformed the proposed video editing pipelines.\n3. MaskINT is trained on unlabeled video datasets using masked token modeling, without needing text-video pairs."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Although the proposed MaskINT can beat other methods in speed, the method still suffers consistency degradation in both prompt and temporal domain. \n2. Noticeable degradation across key frames and interpolated frames.\n3. No related baseline comparison between video interpolation pipeline."},"confidence":{"value":"4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work."},"questions":{"value":"1. By increasing the decoding step and keyframes, the method can increase the performance in Tem-Con and Pro-Con. Can the method reach comparable qualitative results in less time by increasing those hyper parameters?\n2. The videos in supplementary seem to have heavy moiré patterns. Why does this occur?"},"rating":{"value":"5: marginally below the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699637031172,"tcdate":1698745838745,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission8292/Reviewer_hZt7"],"signatures":["ICLR.cc/2024/Conference/Submission8292/Reviewer_hZt7"],"forum":"NRVW8SShFd","number":2,"license":"CC BY 4.0","cdate":1698745838745,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission8292/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699637031172,"domain":"ICLR.cc/2024/Conference","replyto":"NRVW8SShFd","id":"7AMadlqiza","forumContent":{"venue":{"value":"ICLR 2024 Conference Withdrawn Submission"},"TLDR":{"value":"We propose to disentangle the text-based video editing into a two stage pipeline, that involves key frames joint editing using existing image diffusion model and structure-aware frame interpolation with masked generative transformers."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["video editing","masked generative transformers","frame interpolation","diffusion models"]},"supplementary_material":{"value":"/attachment/ad82dacda2c8822c2ee7e2de65f20cbf10ae3425.zip"},"primary_area":{"value":"generative models"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Recent advances in generative AI have significantly enhanced image and video editing, particularly in the context of text prompt control. State-of-the-art approaches predominantly rely on diffusion models to accomplish these tasks. However, the computational demands of diffusion-based methods are substantial, often necessitating large-scale paired datasets for training, causing them challenges to employ in practical applications. This study addresses this challenge by breaking down the text-based video editing process into two stages. In the first stage, we leverage an existing text-to-image diffusion model to simultaneously edit a select few key frames without any additional fine-tuning. In the second stage, we introduce an efficient model called MaskINT, which is built on non-autoregressive masked generative transformers. MaskINT specializes in frame interpolation between the key frames, benefiting from structural guidance provided by intermediate frames. The training of MaskINT incorporates masked token modeling. Our comprehensive set of experiments illustrates the efficacy and efficiency of MaskINT when compared to other diffusion-based methodologies. This research offers a practical solution for text-based video editing and showcases the potential of non-autoregressive masked generative transformers in this domain."},"_bibtex":{"value":"@misc{\nma2024maskint,\ntitle={Mask{INT}: Video Editing via Interpolative Non-autoregressive Masked Transformers},\nauthor={Haoyu Ma and Shahin Mahdizadehaghdam and Bichen Wu and Zhipeng Fan and Yuchao Gu and Wenliang Zhao and Lior Shapira and Xiaohui Xie},\nyear={2024},\nurl={https://openreview.net/forum?id=NRVW8SShFd}\n}"},"title":{"value":"MaskINT: Video Editing via Interpolative Non-autoregressive Masked Transformers"},"pdf":{"value":"/pdf/0d7a9a176c9eb9869f0108ef3c0af2d2dccd3d70.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Withdrawn_Submission"},"paperhash":{"value":"ma|maskint_video_editing_via_interpolative_nonautoregressive_masked_transformers"},"authorids":{"value":["~Haoyu_Ma1","~Shahin_Mahdizadehaghdam2","~Bichen_Wu1","~Zhipeng_Fan1","~Yuchao_Gu1","~Wenliang_Zhao1","~Lior_Shapira1","~Xiaohui_Xie2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Haoyu Ma","Shahin Mahdizadehaghdam","Bichen Wu","Zhipeng Fan","Yuchao Gu","Wenliang Zhao","Lior Shapira","Xiaohui Xie"]}},"version":2},{"content":{"TLDR":{"value":"The \"Timeline Assembler\" is a generative model for instructed visual assembly, using natural language to edit visual timelines. It leverages a multimodal language model for interpreting and processing visual data."},"venue":{"value":"Video-Langauge Models Poster"},"keywords":{"value":["visual timelines","visual language model","natural language instructions","instruction tuning","video editing"]},"supplementary_material":{"value":"/attachment/55c5bb6546851819cf420ddb75e186f43b9156f3.zip"},"abstract":{"value":"The objective of this work is to manipulate visual timelines (e.g. a video) through natural language instructions, making complex timeline editing tasks accessible to non-expert or potentially even disabled users. We call this task Instructed visual assembly. This task is challenging as it requires (i) identifying relevant visual content in the input timeline as well as retrieving relevant visual content in a given input (video) collection, (ii) understanding the input natural language instruction, and (iii) performing the desired edits of the input visual timeline to produce an output timeline. To address these challenges, we propose the Timeline Assembler, a generative model trained to perform instructed visual assembly tasks. The contributions of this work are three-fold. First, we develop a large multimodal language model, which is designed to process visual content, compactly represent timelines and accurately interpret timeline editing instructions. Second, we introduce a novel method for automatically generating datasets for visual assembly tasks, enabling efficient training of our model without the need for human-labeled data. Third, we validate our approach by creating two novel datasets for image and video assembly, demonstrating that the Timeline Assembler substantially outperforms established baseline models, including the recent GPT-4o, in accurately executing complex assembly instructions across various real-world inspired scenarios."},"_bibtex":{"value":"@inproceedings{\npardo2025generative,\ntitle={Generative Timelines for Instructed Visual Assembly},\nauthor={Alejandro Pardo and Jui-Hsien Wang and Bernard Ghanem and Josef Sivic and Bryan Russell and Fabian Caba Heilbron},\nbooktitle={Workshop on Video-Language Models @ NeurIPS 2024},\nyear={2025},\nurl={https://openreview.net/forum?id=UKpQSUbr0D}\n}"},"title":{"value":"Generative Timelines for Instructed Visual Assembly"},"pdf":{"value":"/pdf/772fcaba26c6881fc72942a46ed9df805095e6ff.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Workshop/Video-Langauge_Models"},"paperhash":{"value":"pardo|generative_timelines_for_instructed_visual_assembly"},"authorids":{"value":["~Alejandro_Pardo1","~Jui-Hsien_Wang1","~Bernard_Ghanem1","~Josef_Sivic1","~Bryan_Russell1","~Fabian_Caba_Heilbron3"]},"track":{"value":"Long Paper Track (up to 9 pages)"},"authors":{"value":["Alejandro Pardo","Jui-Hsien Wang","Bernard Ghanem","Josef Sivic","Bryan Russell","Fabian Caba Heilbron"]}},"tmdate":1736861080239,"pdate":1730081752218,"tcdate":1725690909241,"writers":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission15/Authors"],"signatures":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission15/Authors"],"forum":"UKpQSUbr0D","license":"CC BY 4.0","number":15,"cdate":1725690909241,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Post_Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Edit","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission15/-/Camera-Ready_Revision"],"mdate":1736861080239,"odate":1736861080226,"domain":"NeurIPS.cc/2024/Workshop/Video-Langauge_Models","id":"UKpQSUbr0D","version":2},{"content":{"summary":{"value":"This paper presents a graph-based reinforcement learning method for Video Spatio-Temporal Reasoning which models inter-object relationships using relation graphs. It introduces a graph-based reasoning mechanism in the Group Relative Policy Optimization (GRPO) [Shao et al. 2024] for inferring the underlying spatio-temporal topology of scenarios during the thinking process. The paper proposes STV-205k dataset with 205k question-answering pairs for model training and reports SOTA performance on various benchmarks including STIBench."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"Please see weaknesses above."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"The motivation and problem definition are valid. \n\nThe graph based reasoning approach is interesting and seems novel, although I am not aware of all the research in the area of video reasoning. Nevertheless, a graph-based approach appears to be a viable direction for video reasoning. \n\nThe idea of generating a QA dataset (STV-205K) from existing datasets is innovative and useful. \n\nReported experimental results are promising."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"My biggest concerns include limited contribution, lack of clarity and the possible unfair comparison in experiments. Details below. \n\nGraph-based spatio-temporal reasoning is not a new approach and has been done previously. This work appears to be a combination of existing techniques with limited methodological novelty.  \n\nThe very first sentence of the abstract is grammatically incorrect. Overall, the writing and presentation of the paper can certainly be improved as I find it quite confusing. Which “model” is being referred to at each place? And how distances are calculated in images? How do you get Euclidean distances from images/videos? Are they in pixels and if so, how meaningful are pixel based distances in a graph?  \n\nWhat is X in Eqn.(8)? It should be defined after the equation and not only in the Appendix. Going to the proof in the Appendix, X seems to be the object locations and not distances and locations should not be rotation invariant. On the other hand, graph edge weights being Euclidean distances will remain rotation invariant. Also, images/video frames are 2D images, why is the rotation R \\in SO(3)? \n\nAlso, I did not notice the proof of Theorem 2 in the Appendix (maybe I overlooked but there were no numbers with Theorem proof in the Appendix). I am also not sure if these two theorems are worth mentioning at all. \n\nAt this point, I should also highlight the ambiguity in the proposed dataset description. It is not explicitly mentioned which data modality and labels are obtained from the three datasets (TAO, ScanNet and KITTI) e.g. ScanNet and KITTI also have depth images/point clouds. There is no mention of 3D, depth or point cloud in the paper so I assume the paper is about conventional video reasoning but then “Euclidean distance” is mentioned at many places which is confusing. Yes, the Euclidean distances can be in pixel units but that would be meaningless in a scene graph e.g. the chair is two pixels to the right of the table. \n\nReferring to “The model is prompted to generate a graph” on page 5, which model is prompted? The following sentence is also confusing where “model” is used twice. Are they both the same model or two different models? \n\nOn page 2 it is mentioned that “we extend the Group Relative Policy Optimization (GPRO)” but in Section 4.3, it is not clearly mentioned how GPRO is extended in this paper. \n\nThe results in Table 1 look very good but I am wondering if the improvement is mainly due to the additional training data from newly proposed STV-205k dataset. Since the other methods did not use this additional dataset, the comparison is unfair. \n\nPage 7: “SFT achieves localized improvements”. What does “localized improvement” mean? \n\nIn Fig.4(b), the improvement on “Static” sub-task is quite substantial compared to “Temporal”. This makes me wonder how the graph encodes temporal information. \n\nIn Sec. 5.3, “Video-260k”, “SR-91k” are mentioned, but no reference or any further information is given about these datasets. \n\nOverall, the proposed method may have merit but that is obscured by unclear presentation. Moreover, unfair comparison with existing works also undermines its credability."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762919959872,"tcdate":1761870723683,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission7938/Reviewer_hMWm"],"signatures":["ICLR.cc/2026/Conference/Submission7938/Reviewer_hMWm"],"forum":"D6v3B6oTDA","number":2,"license":"CC BY 4.0","cdate":1761870723683,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission7938/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762919959872,"domain":"ICLR.cc/2026/Conference","replyto":"D6v3B6oTDA","id":"w8vrWExkfF","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Multi-modal large language odel","reinforcement learning with verifiable rewards","spatio-temporal reasoning","video understanding"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Recent progress in Multimodal Large Language Models (MLLMs) has demonstrated strong semantic understanding capabilities, but struggles to perform precise spatio-temporal understanding. Existing spatio-temporal methods primarily focus on the video itself, while overlooking the physical information within the video, such as multi-object layouts and motion. Such limitations restrict the use of MLLMs in downstream applications that demand high precision, including embodied intelligence and VR. To address this issue, we present Video-STR, a novel graph-based reinforcement method for precise Video Spatio-Temporal Reasoning. Building upon the capacity of Reinforcement Learning with Verifiable Reward (RLVR) to improve model abilities, we introduce a reasoning mechanism using graph representation based on Group Relative Policy Optimization (GRPO) to guide the model in inferring the underlying spatio-temporal topology of scenarios during the thinking process.To resolve the lack of spatio-temporal training data, we construct the STV-205k dataset with 205k question-answering pairs, covering dynamic multi-object scenes in both indoor and outdoor environments, to support the model training. Experiments show that Video-STR achieves state-of-the-art results on various benchmarks, outperforming the base model by 13% on STI-Bench, and demonstrating the effectiveness of our approach and dataset. Code, model, and data will be released."},"_bibtex":{"value":"@misc{\nwang2026videostr,\ntitle={Video-{STR}: Reinforcing {MLLM}s in Video Spatio-Temporal Reasoning with Relation Graph},\nauthor={Wentao Wang and Heqing Zou and Tianze Luo and Guiyang Xie and Rui Huang and Yutian Zhao and Zhuochen Wang and Hansheng Zhang and Chengwei Qin and Yan Wang and Lin Zhao and Zhang huaijian},\nyear={2026},\nurl={https://openreview.net/forum?id=D6v3B6oTDA}\n}"},"title":{"value":"Video-STR: Reinforcing MLLMs in Video Spatio-Temporal Reasoning with Relation Graph"},"pdf":{"value":"/pdf/44d5d0d35e4beff6b3afc46e66772eed521c69d1.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"wang|videostr_reinforcing_mllms_in_video_spatiotemporal_reasoning_with_relation_graph"},"authorids":{"value":["~Wentao_Wang9","~Heqing_Zou1","~Tianze_Luo1","~Guiyang_Xie1","~Rui_Huang23","~Yutian_Zhao3","~Zhuochen_Wang1","~Hansheng_Zhang1","~Chengwei_Qin1","~Yan_Wang12","~Lin_Zhao3","~Zhang_huaijian1"]},"authors":{"value":["Wentao Wang","Heqing Zou","Tianze Luo","Guiyang Xie","Rui Huang","Yutian Zhao","Zhuochen Wang","Hansheng Zhang","Chengwei Qin","Yan Wang","Lin Zhao","Zhang huaijian"]}},"version":2},{"content":{"summary":{"value":"This paper introduces a new video and language dataset, INTERNVID, which includes 234M video-text pairs lasting 760K hours. In addition to the dataset, this paper also introduce a baseline model, ViCLIP, demonstrating its performance on various downstream applications after pre-training on INTERNVID dataset."},"presentation":{"value":"3 good"},"contribution":{"value":"3 good"},"soundness":{"value":"2 fair"},"strengths":{"value":"1. The paper is well written, and each section is easy to follow. \n\n2. The proposed large-scaled video-text dataset can contribute to the community, scaling up the current model and largely improving the video and language feature representation learning. \n\n3. This paper provides very detailed statistic for the dataset, and also reveal the method for data curation.\n\n4. By leveraging the new dataset, this paper demonstrates extensive experiments on many downstream applications and achieving promising results on many benchmarks."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. There are not many analyses and design justification for the proposed ViCLIP. If ViCLIP is claimed as one of the contributions of this paper. The proposed architecture is not novel and the masking idea needs further justification. For example, the efficiency gain vs. the performance drop, and the ablation over the masking ratio.\n\n2. The video caption is generated by language model from frame-level captions. In this case, will this reduce the number of motion-related words that need to be captured from video-based understanding?"},"confidence":{"value":"4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work."},"questions":{"value":"Given 10M pretrained data, the ViCLIP receives better zero-shot action recognition results from WebVid and better fine-tune action recognition results in Table 2 and 3. And the gain from 50M, 200M INTERNVID pretraining is minor. Does it mean the pretraining data is not the more the better? The performance of vision and language model will be saturated when the pretraining data reach to a certain scale? Could you please provide more insights here?"},"rating":{"value":"6: marginally above the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636477084,"tcdate":1698790159010,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission4917/Reviewer_G48V"],"signatures":["ICLR.cc/2024/Conference/Submission4917/Reviewer_G48V"],"forum":"MLBdiWu4Fw","number":3,"license":"CC BY 4.0","cdate":1698790159010,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission4917/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636477084,"domain":"ICLR.cc/2024/Conference","replyto":"MLBdiWu4Fw","id":"xt7EedelKQ","forumContent":{"venue":{"value":"ICLR 2024 spotlight"},"TLDR":{"value":"This paper introduces InternVid, a large-scale video-centric multimodal dataset that enables learning powerful and transferable video-text representations for multimodal understanding and generation."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["video-language dataset","video understanding","video generation","multimodal understanding","action recognition","video retrieval"]},"supplementary_material":{"value":"/attachment/ab6262d86bfdbdf3844255ccbd7ff6bdca41c17b.zip"},"primary_area":{"value":"datasets and benchmarks"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"This paper introduces InternVid, a large-scale video-centric multimodal dataset that enables learning powerful and transferable video-text representations for multimodal understanding and generation. InternVid contains over 7 million videos lasting nearly 760K hours, yielding 234M video clips accompanied by detailed descriptions of total 4.1B words. Our core contribution is to develop a scalable approach to autonomously build a high-quality video-text dataset with large language models (LLM), thereby showcasing its efficacy in learning video-language representation at scale. Specifically, we utilize a multi-scale approach to generate video-related descriptions. Furthermore, we introduce ViCLIP, a video-text representation learning model based on ViT-L. Learned on InternVid via contrastive learning, this model demonstrates leading zero-shot action recognition and competitive video retrieval performance. Beyond basic video understanding tasks like recognition and retrieval, our dataset and model have broad applications. They are particularly beneficial for generating interleaved video-text data for learning a video-centric dialogue system, advancing video-to-text and text-to-video generation research. These proposed resources provide a tool for researchers and practitioners interested in multimodal video understanding and generation."},"_bibtex":{"value":"@inproceedings{\nwang2024internvid,\ntitle={InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation},\nauthor={Yi Wang and Yinan He and Yizhuo Li and Kunchang Li and Jiashuo Yu and Xin Ma and Xinhao Li and Guo Chen and Xinyuan Chen and Yaohui Wang and Ping Luo and Ziwei Liu and Yali Wang and Limin Wang and Yu Qiao},\nbooktitle={The Twelfth International Conference on Learning Representations},\nyear={2024},\nurl={https://openreview.net/forum?id=MLBdiWu4Fw}\n}"},"title":{"value":"InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation"},"pdf":{"value":"/pdf/5355ce2fec3ff26dca65a969b767fd7b1102bb05.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference"},"paperhash":{"value":"wang|internvid_a_largescale_videotext_dataset_for_multimodal_understanding_and_generation"},"authorids":{"value":["~Yi_Wang19","~Yinan_He1","~Yizhuo_Li1","~Kunchang_Li1","~Jiashuo_Yu1","~Xin_Ma3","~Xinhao_Li1","~Guo_Chen2","~Xinyuan_Chen1","~Yaohui_Wang1","~Ping_Luo2","~Ziwei_Liu1","~Yali_Wang1","~Limin_Wang1","~Yu_Qiao1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Yi Wang","Yinan He","Yizhuo Li","Kunchang Li","Jiashuo Yu","Xin Ma","Xinhao Li","Guo Chen","Xinyuan Chen","Yaohui Wang","Ping Luo","Ziwei Liu","Yali Wang","Limin Wang","Yu Qiao"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a video-to-audio model by investigating vision encoders, auxiliary embeddings, and data augmentation techniques. It shows an SOTA performance. \nIt uses the CLIP4CLIP as a video encoder, LDM, to generate the mel-spectrogram space.\nIt uses VGGSound 200k videos for training. It also filters out the video-audio pairs with low similarity."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1.\tIs the LDM learned from scratch?\n2.\tThe CLIP4CLIP feature is not clear to me. Does it use the spatial-temporal feature, or is it just a mean pooled feature vector, as indicated by the original paper?\n3.\tWhat is the position embedding in L338? Is it the video position embedding or text position embedding? \n4.\tI’d like to see the explanation on why the attention on video is more focused by adding text, as shown in Fig.3. I think this salience map is the UNet activation as query and the video patch feature as key.\n5.\tHow to remove those videos where the background music is not that correlated with the video content?"},"rating":{"value":5},"details_of_ethics_concerns":{"value":"N/A"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1.\tThis paper proposes a baseline of the LDM model that surpasses the other comparison methods in some metrics. \n2.\tIt performed various ablations on video encoder selection, data augmentation, and text embedding, adding up to this paper's empirical value."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1.\tIt shows that the text embedding would help the model generate the audio; I think it makes sense, but I’d like to know how details of the caption are needed. Does dense video caption help, or just a succinct video caption is enough? \n2.\tMinor: Why do Experiment Setup and Experiments in separate sections? I think it’s common to pub experimental setup as a subsection under the experiment section.\n3.\tIn L383, the data augmentation is to randomly combine video and audio segments. But does this reduce the alignment between video and audio?\n4.\tIt seems the author did many kinds of data preprocessing but only evaluated them independently and expected the mixture of them to work the best. However, I think it should be within the scope of this paper to give a final combined setting of those data augmentation, which makes me feel the unreadiness of this paper.\n5.\tI found this paper a baseline with empirical ablation, but it is not clear where the novelty lands."}},"nonreaders":[],"tmdate":1731427855458,"tcdate":1730852719001,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission3487/Reviewer_7Cqr"],"signatures":["ICLR.cc/2025/Conference/Submission3487/Reviewer_7Cqr"],"forum":"YQjdNC0NkW","number":3,"license":"CC BY 4.0","cdate":1730852719001,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission3487/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427855458,"domain":"ICLR.cc/2025/Conference","replyto":"YQjdNC0NkW","id":"esfTKkoxBW","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["audio generation","video-to-audio","diffusion"]},"supplementary_material":{"value":"/attachment/c2a999bd8d2d4d4e02507685d64c8579022ee977.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Generating semantically and temporally aligned audio content in accordance with video input has become a focal point for researchers, particularly following the remarkable breakthrough in text-to-video generation. In this work, we aim to offer insights into the video-to-audio generation paradigm, focusing on three crucial aspects: vision encoders, auxiliary embeddings, and data augmentation techniques.\nBeginning with a foundational model built on a simple yet surprisingly effective intuition, we explore various vision encoders and auxiliary embeddings through ablation studies. Employing a comprehensive evaluation pipeline that emphasizes generation quality and video-audio synchronization alignment, we demonstrate that our model exhibits state-of-the-art video-to-audio generation capabilities. Furthermore, we provide critical insights into the impact of different data augmentation methods on enhancing the generation framework’s overall capacity. We showcase possibilities to advance the challenge of generating synchronized audio from semantic and temporal perspectives. We hope these insights will serve as a stepping stone toward developing more realistic and accurate audio-visual generation models."},"_bibtex":{"value":"@misc{\nxu2024videotoaudio,\ntitle={Video-to-Audio generation with Hidden Alignment},\nauthor={Manjie Xu and Chenxing Li and Yong Ren and Xinyi Tu and Rilin Chen and Yu Gu and Wei Liang and Dong Yu},\nyear={2024},\nurl={https://openreview.net/forum?id=YQjdNC0NkW}\n}"},"title":{"value":"Video-to-Audio generation with Hidden Alignment"},"pdf":{"value":"/pdf/2cea385bd1f6afbe90241fbd54b5d1f14527b03e.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"xu|videotoaudio_generation_with_hidden_alignment"},"authorids":{"value":["~Manjie_Xu1","~Chenxing_Li1","~Yong_Ren5","~Xinyi_Tu1","~Rilin_Chen1","~Yu_Gu15","~Wei_Liang1","~Dong_Yu2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Manjie Xu","Chenxing Li","Yong Ren","Xinyi Tu","Rilin Chen","Yu Gu","Wei Liang","Dong Yu"]}},"version":2},{"content":{"summary":{"value":"The paper proposes Flow Uniqueness Models (FUM), a framework for achieving high-quality one-step generation within flow-matching models. The key idea is to enforce velocity uniqueness across the sampling path by dividing the flow into two sub-paths: (i) a first sub-path trained on strictly one-to-one image pairs to ensure deterministic flow, and (ii) a second sub-path trained with a flow consistency strategy to align its velocity with the first. Two consistency variants (Shortcut-based and MeanFlow-based) are introduced. Experiments on CIFAR-10, FFHQ, and ImageNet demonstrate competitive one-step and few-step generation performance compared to prior diffusion distillation and consistency models."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"- How sensitive is the method to the choice of the split points between sub-paths?\n\n- Does the performance degrade significantly if the initial diffusion model is not well-trained?\n\n- Can FUM be trained entirely from scratch without relying on a pre-trained ODE sampler?\n\n- What is the computational overhead compared to MeanFlow or Shortcut models during training?\n\n- Could the velocity uniqueness concept be extended to text-to-image or multimodal setups?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":1},"contribution":{"value":1},"strengths":{"value":"- The paper tackles an important practical issue—reducing sampling steps in generative models without major loss in quality.\n\n- The notion of “velocity uniqueness” provides an intuitive way to understand and regularize one-step generation.\n\n- FUM shows consistently strong or comparable performance to leading methods (e.g., MeanFlow, Shortcut) across multiple datasets."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Limited novelty: While the idea of enforcing path or velocity consistency is meaningful, the proposed approach mainly combines existing components (pairwise matching + consistency regularization) without introducing a fundamentally new principle.\n\n- FUM relies on pre-trained diffusion models to obtain the one-to-one image pairs, meaning it is not a fully independent or end-to-end training scheme.\n\n- At several places, the authors mention “strong uniqueness”. Could the authors please define to quantify “strong” here?\n\n- The paper contains many typos and inconsistencies (see list below), which makes me doubt the technical correctness of the paper. \n\nList of potential typos/inconsistencies:\n- Line 128: definition of $a_t$ is not correct\n- Equation (1): $\\epsilon$ is missing in the objective function\n- At the beginning, the flow is defined on $t \\in [0, 1]$. Later on, the index is shifted to $\\{0, …, T\\}$. \n- Around equation (2): it is unclear whether $s \\in [0, T]$ or $s \\in [1, T]$. \n- Equation (3) and (4): How could I sample $j \\sim \\mathcal U(s, j)$?\n- Algorithm 1, line 14: the EMA update is incorrect.\n- Algorithm 2, line 6: Do we miss a $\\Delta k$ term?\n- Proposition 1 only holds under some assumptions of the noise distribution $\\pi$. These assumptions have not been made explicit."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762919151124,"tcdate":1761992417553,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission6913/Reviewer_aCia"],"signatures":["ICLR.cc/2026/Conference/Submission6913/Reviewer_aCia"],"forum":"ZMqIgONdJZ","number":4,"license":"CC BY 4.0","cdate":1761992417553,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission6913/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762919151124,"domain":"ICLR.cc/2026/Conference","replyto":"ZMqIgONdJZ","id":"KuHHuxZWrs","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Flow matching","one-step generation","flow consistency"]},"primary_area":{"value":"generative models"},"abstract":{"value":"Recent advances in generative modeling frameworks, such as diffusion models and flow matching, have achieved record-breaking performance.\nNevertheless, these approaches involve iterative sampling procedures across many neural network passes, which severely limits their practical deployment, particularly in domains demanding real-time interaction.\nAlthough considerable effort has been devoted to accelerating sampling, achieving high-quality one-step generation remains an open challenge, motivating research into a new era of generative modeling.\nMotivated by this, we put forward a novel and effective framework, termed \\textit{Flow Uniqueness Models} (\\textbf{FUM}).\nThe core idea of FUM is to construct strictly one-to-one image pairs, thereby enforcing velocity uniqueness along the entire sampling path, which forms as the foundation for few-step sampling.\nBy leveraging this modeling mechanism, FUM not only achieves remarkable one-step generative performance but also provides the flexibility to balance image quality against the number of sampling steps.\nExtensive experiments on three benchmark datasets comprehensively validate the superiority of our proposed FUM."},"_bibtex":{"value":"@misc{\nzhang2026building,\ntitle={Building Flow Uniqueness in One-step Generative Modeling},\nauthor={Junyu Zhang and Daochang Liu and Liu.Liu and Zhizhong Su and Jong Hwan Ko and Shichao Zhang and Eunbyung Park and Chang Xu},\nyear={2026},\nurl={https://openreview.net/forum?id=ZMqIgONdJZ}\n}"},"title":{"value":"Building Flow Uniqueness in One-step Generative Modeling"},"pdf":{"value":"/pdf/e1f62ddaec95e6411b4597cb38cc6a0699d5c444.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"zhang|building_flow_uniqueness_in_onestep_generative_modeling"},"authorids":{"value":["~Junyu_Zhang2","~Daochang_Liu1","~Liu.Liu19","~Zhizhong_Su3","~Jong_Hwan_Ko2","~Shichao_Zhang3","~Eunbyung_Park1","~Chang_Xu4"]},"authors":{"value":["Junyu Zhang","Daochang Liu","Liu.Liu","Zhizhong Su","Jong Hwan Ko","Shichao Zhang","Eunbyung Park","Chang Xu"]}},"version":2},{"content":{"venue":{"value":"IEEE Trans. Multim. 2023"},"pdf":{"value":"https://ieeexplore.ieee.org/iel7/6046/10016790/10057062.pdf"},"venueid":{"value":"dblp.org/journals/TMM/2023"},"paperhash":{"value":"zhu|motionvideogan_a_novel_video_generator_based_on_the_motion_space_learned_from_image_pairs"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Jingyuan_Zhu:","https://dblp.org/search/pid/api?q=author:Huimin_Ma_0001:","~Jiansheng_Chen1","https://dblp.org/search/pid/api?q=author:Jian_Yuan:"]},"html":{"value":"https://doi.org/10.1109/TMM.2023.3251095"},"_bibtex":{"value":"@article{DBLP:journals/tmm/ZhuMCY23,\n  author={Jingyuan Zhu and Huimin Ma and Jiansheng Chen and Jian Yuan},\n  title={MotionVideoGAN: A Novel Video Generator Based on the Motion Space Learned From Image Pairs},\n  year={2023},\n  cdate={1672531200000},\n  journal={IEEE Trans. Multim.},\n  volume={25},\n  pages={9370-9382},\n  url={https://doi.org/10.1109/TMM.2023.3251095}\n}\n"},"abstract":{"value":"Video generation has achieved rapid progress benefiting from high-quality renderings provided by powerful image generators. We regard the video synthesis task as generating a sequence of images sharing the same contents but varying in motions. However, most previous video synthesis frameworks based on pre-trained image generators treat content and motion generation separately, leading to unrealistic generated videos. Therefore, we design a novel framework to build the motion space, aiming to achieve content consistency and fast convergence for video generation. We present MotionVideoGAN, a novel video generator synthesizing videos based on the motion space learned by pre-trained image pair generators. Firstly, we propose an image pair generator named MotionStyleGAN to generate image pairs sharing the same contents and producing various motions. Then we manage to acquire motion codes to edit one image in the generated image pairs and keep the other unchanged. The motion codes help us edit images within the motion space since the edited image shares the same contents with the other unchanged one in image pairs. Finally, we introduce a latent code generator to produce latent code sequences using motion codes for video generation. Our approach achieves state-of-the-art performance on the most complex video dataset ever used for unconditional video generation evaluation, UCF101."},"title":{"value":"MotionVideoGAN: A Novel Video Generator Based on the Motion Space Learned From Image Pairs"},"authors":{"value":["Jingyuan Zhu","Huimin Ma","Jiansheng Chen","Jian Yuan"]}},"tmdate":1754398534863,"pdate":1672531200000,"externalIds":["dblp:journals/tmm/ZhuMCY23"],"tcdate":1754398510233,"writers":["~"],"signatures":["~Jiansheng_Chen3"],"forum":"IDAt2Dn21L","license":"CC BY-SA 4.0","number":614607,"cdate":1672531200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1754398534863,"domain":"DBLP.org","id":"IDAt2Dn21L","version":2},{"content":{"venue":{"value":"IEEE Trans. Multim. 2023"},"pdf":{"value":"https://ieeexplore.ieee.org/iel7/6046/10016790/10057062.pdf"},"venueid":{"value":"dblp.org/journals/TMM/2023"},"paperhash":{"value":"zhu|motionvideogan_a_novel_video_generator_based_on_the_motion_space_learned_from_image_pairs"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Jingyuan_Zhu:","~Huimin_Ma1","~Jiansheng_Chen3","https://dblp.org/search/pid/api?q=author:Jian_Yuan:"]},"html":{"value":"https://doi.org/10.1109/TMM.2023.3251095"},"_bibtex":{"value":"@article{DBLP:journals/tmm/ZhuMCY23,\n  author={Jingyuan Zhu and Huimin Ma and Jiansheng Chen and Jian Yuan},\n  title={MotionVideoGAN: A Novel Video Generator Based on the Motion Space Learned From Image Pairs},\n  year={2023},\n  cdate={1672531200000},\n  journal={IEEE Trans. Multim.},\n  volume={25},\n  pages={9370-9382},\n  url={https://doi.org/10.1109/TMM.2023.3251095}\n}\n"},"abstract":{"value":"Video generation has achieved rapid progress benefiting from high-quality renderings provided by powerful image generators. We regard the video synthesis task as generating a sequence of images sharing the same contents but varying in motions. However, most previous video synthesis frameworks based on pre-trained image generators treat content and motion generation separately, leading to unrealistic generated videos. Therefore, we design a novel framework to build the motion space, aiming to achieve content consistency and fast convergence for video generation. We present MotionVideoGAN, a novel video generator synthesizing videos based on the motion space learned by pre-trained image pair generators. Firstly, we propose an image pair generator named MotionStyleGAN to generate image pairs sharing the same contents and producing various motions. Then we manage to acquire motion codes to edit one image in the generated image pairs and keep the other unchanged. The motion codes help us edit images within the motion space since the edited image shares the same contents with the other unchanged one in image pairs. Finally, we introduce a latent code generator to produce latent code sequences using motion codes for video generation. Our approach achieves state-of-the-art performance on the most complex video dataset ever used for unconditional video generation evaluation, UCF101."},"title":{"value":"MotionVideoGAN: A Novel Video Generator Based on the Motion Space Learned From Image Pairs"},"authors":{"value":["Jingyuan Zhu","Huimin Ma","Jiansheng Chen","Jian Yuan"]}},"tmdate":1728435349008,"pdate":1672531200000,"tcdate":1727773621316,"writers":["~"],"signatures":["~Jiansheng_Chen3"],"forum":"SZG2JvHoxf","license":"CC BY-SA 4.0","number":129175,"cdate":1672531200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1728435349008,"domain":"DBLP.org","id":"SZG2JvHoxf","version":2},{"content":{"summary":{"value":"This paper proposes Temporal Preference Optimization (TPO), a lightweight post-training framework that improves video-LMMs’ temporal grounding and reasoning without human temporal annotations by generating contrastive preference pairs from the same query answered on original (relevant) versus corrupted or incomplete frames, filtering noisy pairs with a small LLM, and then optimizing with Direct Preference Optimization plus a minor auxiliary SFT loss so the model prefers responses aligned with when evidence occurs, not just what appears; across LongVideoBench, MLVU, and Video-MME, TPO consistently outperforms baselines, with LLaVA-Video-TPO achieving state-of-the-art among 7B models, and the recipe is practical (e.g., ~4 hours on 8×A100 with fixed 32 sampled frames). A scalable annotation-free temporal supervision pipeline via input manipulation and LLM post-filtering; a preference-learning objective instantiating DPO (with small SFT) tailored to temporal grounding; and strong empirical gains across three long-video benchmarks and multiple bases"},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"## Temporal grounding definition & evidence.\nCould you specify which aspects of “temporal grounding” are under-defined here and what concrete operational tests (e.g., frame-level localization, counterfactual masking) you would need to see to accept that the method truly improves when-aware reasoning rather than generic accuracy?\n\n## LLM post-filter robustness & bias.\nWhich failure modes of the LLM-based filtering (e.g., label leakage, stylistic bias, prompt sensitivity) most concern you, and what ablations or audits (alternative teachers, temperature sweeps, bias probes) would convincingly address them?\n\n## Evaluation reliability & generalization.\nWhere do you believe the current results are insufficiently reliable (e.g., short-video subsets, high-motion segments), and which additional analyses—stratified metrics with uncertainty (bootstrap CIs), new baselines, or zero-shot datasets—would meaningfully change your assessment?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- Simple, efficient training recipe. The method trains in roughly four hours on 8×A100 (80 GB) with fixed 32 sampled frames shared across data generation and training, indicating practical scalability.\n- Targeted objective that preserves general ability. DPO on temporal preference pairs is positioned to enhance temporal reasoning while retaining pretrained knowledge."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Limited benefit on short-video settings. The authors note performance is only comparable to SFT baselines on the Video-MME-short subset, suggesting gains concentrate on longer temporal contexts.\n- Fixed-frame sampling could bottleneck long-horizon reasoning. The design uses a constant 32 frames for both generation and training; the implications for very long or high-motion videos are not extensively studied.\n- Dependence on an external LLM for curation. The data pipeline requires GPT-4o-mini for question curation and post-filtering, introducing cost/availability/bias considerations that are not thoroughly quantified."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762921893148,"tcdate":1761816631525,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission10638/Reviewer_p9KZ"],"signatures":["ICLR.cc/2026/Conference/Submission10638/Reviewer_p9KZ"],"forum":"vZLZyNxeOa","number":2,"license":"CC BY 4.0","cdate":1761816631525,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission10638/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762921893148,"domain":"ICLR.cc/2026/Conference","replyto":"vZLZyNxeOa","id":"jj8rh6pDIp","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Video Understanding","Vision-Language Model","Preference Learning","Post-Training"]},"supplementary_material":{"value":"/attachment/2b50b9c583c3ba1e6ae199b123b50293246f93ca.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Despite recent advancements in video large multimodal models (video-LMMs), accurate temporal grounding remains a key challenge. In this work, we introduce Temporal Preference Optimization (TPO)—a post-training framework that unlocks superior temporal reasoning in video-LMMs without requiring human annotations. TPO enables preference modeling by manipulating video inputs to generate contrastive responses, ensuring that preferred responses are more temporally grounded than dis-preferred ones. Through preference learning, TPO enhances the model’s capability for more comprehensive video understanding with better temporal reasoning. Extensive experiments on LongVideoBench, MLVU, and Video-MME demonstrate that TPO significantly improves temporal grounding across multiple video-LMMs.   Notably, LLaVA-Video-TPO achieves state-of-the-art performance among 7B models on Video-MME, establishing TPO as a scalable and effective solution for advancing temporal understanding in video analysis."},"_bibtex":{"value":"@misc{\nli2025temporal,\ntitle={Temporal Preference Optimization of Large Multimodal Models},\nauthor={Rui Li and Xiaohan Wang and Yuhui Zhang and Orr Zohar and Zeyu Wang and Serena Yeung-Levy},\nyear={2025},\nurl={https://openreview.net/forum?id=vZLZyNxeOa}\n}"},"title":{"value":"Temporal Preference Optimization of Large Multimodal Models"},"pdf":{"value":"/pdf/b782ce9c318b44c952f15ad9541830af11a97f24.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"li|temporal_preference_optimization_of_large_multimodal_models"},"authorids":{"value":["~Rui_Li26","~Xiaohan_Wang2","~Yuhui_Zhang3","~Orr_Zohar1","~Zeyu_Wang1","~Serena_Yeung-Levy1"]},"authors":{"value":["Rui Li","Xiaohan Wang","Yuhui Zhang","Orr Zohar","Zeyu Wang","Serena Yeung-Levy"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a LLM-based video understanding framework, which uses text autoencoder + text Q-former and designs the corresponding alignment loss to guide the overall framework to accurately understand the input video and ASR text.\nExperimental results show that the proposed framework achieves comparable performance with fewer pre-training samples and performs well in the corresponding downstream tasks."},"presentation":{"value":"3 good"},"contribution":{"value":"2 fair"},"soundness":{"value":"3 good"},"strengths":{"value":"1.  This paper proposes a video understanding architecture that improves the encoding performance of the video encoder by introducing a text autoencoder and corresponding alignment losses.\n2.  The experimental results demonstrate the effectiveness of the proposed framework in low-resource and zero-shot settings.\n3. The overall design is concise and feasible, and the paper is easy to follow."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The reviewer has the most concern about its limited contributions. Fusing the video features with its corresponding audio signals is a common technique, as in [B]. Aligning video with text before feeding them into a joint Transformer (LLM in this paper) is also a well-known technique, as shown in [A]. Generally speaking, the paper combines multiple existing techniques and the only significant difference is that the authors use the LLM framework which has been popular since 2023.\n\n> [A] Align and Prompt: Video-and-Language Pre-training with Entity Prompts. CVPR 2022.\n\n> [B] MERLOT Reserve: Multimodal Neural Script Knowledge through Vision and Language and Sound. CVPR 2022.\n\n2. The second concern is about the incomprehensive experiments.\n\n(1) This paper compares with existing methods on the data resources used for pre-training, but does not mention the difference in model parameters.\n\n(2) Some experimental results are confusing and lack of explanations, such as the performance of contrastive learning loss in different datasets in Table 4.\n\n(3) The baseline methods used in the experiments lack some SOTA methods, and their performance advantages are not significant compared to the training cost. At least, BLIP-2 should be compared on these benchmarks as it is the baseline.\n\n(4) One important ablation study is missed, where only ASR text is used without visual input. There might exist a shortcut that most of the knowledge is carried by ASR text instead of well-aligned visual features.\n\n(5) The scalability of the method on larger datasets or larger backbones has not been studied. Moreover, it achieves lower accuracy compared to previous methods without LLMs as shown in Tab.1. Its practicality is questionable."},"confidence":{"value":"4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work."},"questions":{"value":"1.  Do the advantages in low-resource settings come from the introduction of LLM?\n2.  Can video-teller handle long videos?\n\nMinor: there are some typos, such as the \"black\" in Table 3 and the \"21.9\" of B@4 in Table 4."},"rating":{"value":"3: reject, not good enough"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636251677,"tcdate":1698658367453,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission3062/Reviewer_wY2h"],"signatures":["ICLR.cc/2024/Conference/Submission3062/Reviewer_wY2h"],"forum":"ia5wG0Fp2U","number":1,"license":"CC BY 4.0","cdate":1698658367453,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission3062/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636251677,"domain":"ICLR.cc/2024/Conference","replyto":"ia5wG0Fp2U","id":"ZzDczFEufz","forumContent":{"venue":{"value":"ICLR 2024 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video Understanding;Modality alignment;Representation decoupling;Fine-grained"]},"primary_area":{"value":"representation learning for computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"This paper proposes Video-Teller, a Video-Text foundational model that leverages multi modal fusion and fine-grained modality alignment to significantly enhance the cross-modal generation task. Video-Teller boosts the training efficiency by utilizing frozen pre-trained vision and language modules. Furthermore, it capitalizes on the robust linguistic capabilities of large language model, enabling the generation of more nuanced descriptions for videos (video summaries). To effectively integrate visual and auditory information and improve the model's understanding of videos, Video-Teller employs cascaded Q-Former to fuse information from different frames and modalities. In addition to conventional loss functions, we introduce an additional Text Auto-Encoder to decouple the target text for fine-grained modality alignment, further optimizing the model. Experimental results demonstrate the efficacy of our proposed video foundational model in accurately comprehending videos and generating coherent and precise language descriptions. It is worth noting that the fine-grained alignment enhances the model's capabilities with a relatively minimal increase in cost."},"_bibtex":{"value":"@misc{\nliu2024videoteller,\ntitle={Video-Teller: Enhancing Cross-Modal Generation with Fusion and Decoupling},\nauthor={Haogeng Liu and Qihang Fan and Tingkai Liu and Linjie Yang and Yunzhe Tao and Huaibo Huang and Ran He and Hongxia Yang},\nyear={2024},\nurl={https://openreview.net/forum?id=ia5wG0Fp2U}\n}"},"title":{"value":"Video-Teller: Enhancing Cross-Modal Generation with Fusion and Decoupling"},"pdf":{"value":"/pdf/7f22d5615efd2648bdc809ea533f15f67a9f6668.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Withdrawn_Submission"},"paperhash":{"value":"liu|videoteller_enhancing_crossmodal_generation_with_fusion_and_decoupling"},"authorids":{"value":["~Haogeng_Liu1","~Qihang_Fan1","~Tingkai_Liu1","~Linjie_Yang4","~Yunzhe_Tao2","~Huaibo_Huang1","~Ran_He1","~Hongxia_Yang2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Haogeng Liu","Qihang Fan","Tingkai Liu","Linjie Yang","Yunzhe Tao","Huaibo Huang","Ran He","Hongxia Yang"]}},"version":2},{"content":{"summary":{"value":"This paper introduces OVID, a massive open video dataset sourced from Common Crawl, containing 1.3B video URLs (∼10M hours). From these, the authors generate 300M high-quality frame-caption pairs by extracting scene-change frames and captioning them with a efficient vision-language model (DeepSeek-VL2), followed by video-level summarization.\n\nWhen used to train CLIP models, OVID demonstrates strong and scalable performance: it achieves state-of-the-art results on COCO image-text retrieval, outperforming comparable datasets like DataComp. However, a noted limitation is its weaker performance on zero-shot ImageNet classification, attributed to a domain gap from its synthetic captions.\n\nThe work provides a valuable, large-scale multimodal resource complementary to existing image-text collections."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. The paper notes a performance drop in ImageNet classification due to synthetic captions. It would be helpful to see a quantitative analysis of caption diversity (e.g., vocabulary size, sentence length) compared to human-written alt-text. Furthermore, could the authors comment on whether mixing OVID's data with smaller human-curated subsets might help bridge this domain gap and improve classification accuracy?\n\n2. The reliance on platform-level moderation is noted. Have the authors conducted any evaluation for unsafe or biased content? Even a small-scale audit or analysis of potential demographic biases or NSFW frames would significantly strengthen the claims about dataset reliability and safety.\n\n3. Could the authors elaborate on how they ensure the selected scene-change frames are semantically meaningful and not overly redundant? Would an adaptive, diversity-based sampling strategy be a worthwhile future direction to improve data efficiency?\n\n4. Given the video origins of the data, has there been any consideration to benchmark OVID on temporal or video-language tasks? Demonstrating performance on benchmarks like MSR-VTT could further validate its value for video understanding, beyond image-text retrieval.\n\n5. The dataset uses a research-only license. Could the authors clarify if there will be restrictions on creating derivative datasets or on fine-tuning commercial models? This clarity is important for the community to understand the full implications of the \"open access\" terms.\n\n6. Figure 6 indicates different scaling trends for retrieval and classification. Could the authors provide more insight into why OVID exhibits stronger scaling for retrieval but weaker scaling for classification? Is this primarily attributed to caption style or domain bias?\n\n7. Do the authors plan to extend OVID with audio or multimodal features? Integrating audio captions could make it an even more valuable resource for training unified vision-language-audio models."},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"This paper's main contribution is OVID, an extremely large-scale video-text dataset (1.3B URLs, 10M hours) that significantly dwarfs existing open collections. Its scale and diversity—multilingual and multi-topic—are its key strengths. The method of generating image-text pairs from scene-change frames is both novel and impactful.\n\nThe empirical validation is robust: CLIP models trained on OVID achieve state-of-the-art COCO retrieval performance across scales, proving the data's tangible value. This focus on large-scale, open multimodal data, combined with its commitment to reproducibility and open access, makes it a timely and valuable contribution to the community."},"flag_for_ethics_review":{"value":["Yes, Discrimination / bias / fairness concerns","Yes, Legal compliance (e.g., GDPR, copyright, terms of use, web crawling policies)"]},"weaknesses":{"value":"While the OVID dataset is a significant contribution, several limitations should be noted. \n\n1. Its reliance on synthetic captions introduces a domain gap; while this benefits retrieval tasks, it leads to a notable drop in zero-shot classification accuracy. \n\n2. The light data filtering strategy, while enabling scale, leaves potential concerns about noise, bias, and unsafe content unquantified in the current evaluation. \n\n3. The empirical study is thorough but narrowly focused on image-text tasks, leaving its value for video-language modeling an open question for future work."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762933737272,"tcdate":1761644143819,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission20250/Reviewer_mjz7"],"signatures":["ICLR.cc/2026/Conference/Submission20250/Reviewer_mjz7"],"forum":"etFOgs8vIb","number":2,"license":"CC BY 4.0","cdate":1761644143819,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission20250/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762933737272,"domain":"ICLR.cc/2026/Conference","replyto":"etFOgs8vIb","id":"iSD2oWKopo","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["video dataset","recaptioning","clip","open foundation models","open datasets"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"We present OVid, a large open video dataset comprising _10 million hours_ of diverse content collected from CommonCrawl. To complement the raw data, we generate image captions for scene-changing frames and video-level captions for a 300M frame–caption subset. Using this subset, we train CLIP models at multiple scales and benchmark them against reference CLIP models trained on DataComp, Re-LAION and DataComp recaptioned with the same captioning pipeline. Observed scaling trends for classification and retrieval show evidence that OVid can be another valuable and scalable source of image-text data, in addition to image-text pairs from public webpages. OVid marks a significant step towards democratizing access to large-scale video data and fostering the development of open multimodal foundation models. To this end, all the data will be freely available to research institutions."},"_bibtex":{"value":"@misc{\nhochlehnert2026ovid,\ntitle={{OV}id: Open Large-Scale Video Dataset as a Novel Source for Image-Text Data},\nauthor={Andreas Hochlehnert and Marianna Nezhurina and Thadd{\\\"a}us Wiedemer and Christoph Schuhmann and Mehdi Cherti and Romain Beaumont and Andrii Matiuk and Andrej Radonjic and Bernhard Sch{\\\"o}lkopf and Wieland Brendel and A. Sophia Koepke and Jenia Jitsev and Matthias Bethge},\nyear={2026},\nurl={https://openreview.net/forum?id=etFOgs8vIb}\n}"},"title":{"value":"OVid: Open Large-Scale Video Dataset as a Novel Source for Image-Text Data"},"pdf":{"value":"/pdf/7a5482f7f60923f65052b07107b7cc4fc627e148.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"hochlehnert|ovid_open_largescale_video_dataset_as_a_novel_source_for_imagetext_data"},"authorids":{"value":["~Andreas_Hochlehnert1","~Marianna_Nezhurina1","~Thaddäus_Wiedemer1","~Christoph_Schuhmann1","~Mehdi_Cherti2","~Romain_Beaumont1","~Andrii_Matiuk2","~Andrej_Radonjic1","~Bernhard_Schölkopf1","~Wieland_Brendel1","~A._Sophia_Koepke1","~Jenia_Jitsev1","~Matthias_Bethge1"]},"authors":{"value":["Andreas Hochlehnert","Marianna Nezhurina","Thaddäus Wiedemer","Christoph Schuhmann","Mehdi Cherti","Romain Beaumont","Andrii Matiuk","Andrej Radonjic","Bernhard Schölkopf","Wieland Brendel","A. Sophia Koepke","Jenia Jitsev","Matthias Bethge"]}},"version":2},{"content":{"summary":{"value":"This paper proposes an idea of “Thinking with video.” Analogous to “thinking with images,”  “thinking with video” treats reasoning as a dynamic process of temporally grounded exploration and decomposition. The model iteratively decides what to look for, where to watch, and at what temporal scale, flexibly combining both fine-grained inspection and coarse-grained temporal grounding for long video understanding. The paper also proposes a two-stage training framework with SFT and DPO on accepted temporal grounding trajectories to improve VQA on long video benchmarks and showcase great performance improvement with modest innovation."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"In Section 2.2, regarding the statement: “This paradigm will inevitably lose the rich visual information in original long videos, leading to sub-optimal performance such as in egocentric videos…”\nThe reviewer finds this argument unconvincing, as captioning from image or video frames does not seem to differ significantly between egocentric and exocentric videos."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The paper clearly extends the idea of Thinking with Images to the video domain from a temporal perspective.\n\n2. It constructs a reasoning-centric dataset with multi-step reasoning trajectories, supervised for a two-stage optimization pipeline that combines structured imitation learning (SFT) and trajectory-level preference alignment (DPO for video)\n\n3. The work demonstrates strong token efficiency in long-video understanding, outperforming uniform-sampling-based methods."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The proposed components include question decomposition and agent planning, question-aware temporal grounding, agentic iterative reasoning, and temporally grounded preference optimization—have been explored in many existing works such as DrVideo [1], VideoINSTA [2], Traveler [3], TPO [4], and Video-R1 (T-GRPO) [5]. The authors either missed direct relevant comparisons or need to further justify the novelty and necessity of their proposed multi-step temporal grounding approach.\n\n[1] DrVideo: Document Retrieval-Based Long Video Understanding. CVPR 2025.\n\n[2] VideoINSTA: Zero-shot Long Video Understanding via Informative Spatial-Temporal Reasoning with LLMs. EMNLP 2024.\n\n[3] Traveler: A Modular Multi-LLM Agent Framework for Video Question Answering. EMNLP 2024.\n\n[4] Temporal Preference Optimization for Long-Form Video Understanding. arXiv:2501.13919.\n\n[5] Video-R1: Reinforcing Video Reasoning in MLLMs. arXiv:2503.21776.\n\n2. The annotation procedure for multi-step reasoning in the proposed dataset is unclear and less described. The definitions of step-wise reasoning and the quality assurance of VLM-based grounding annotations require more explanation. \n\n3. In Table 2, the QA accuracy shows significant improvement. However, the evaluation of Temporal Grounding Accuracy compared to non-grounding methods (e.g., Ego-R1, VideoAgent) needs additional validation, preferably against other temporal grounding-focused baselines such as [6]\n\n[6] ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos. CVPR2025.\n\nFormatting Errors:\n\nIn Section 2.1, the references following “…leading to the emergence of Multi-modal Large Language Models (MLLMs)” are improperly formatted.\n\nLine 146 contains a duplicate “video”."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762921937024,"tcdate":1762354976213,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission10694/Reviewer_F2uk"],"signatures":["ICLR.cc/2026/Conference/Submission10694/Reviewer_F2uk"],"forum":"mPxHWZl9bs","number":3,"license":"CC BY 4.0","cdate":1762354976213,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission10694/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762921937024,"domain":"ICLR.cc/2026/Conference","replyto":"mPxHWZl9bs","id":"xP6HauLFjh","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["long video understanding","agentic frameworks"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Long-video understanding (LVU) is a challenging problem in computer vision. \nExisting methods either downsample frames for single-pass reasoning, sacrificing fine-grained details, or depend on textual reasoning over task-agnostic representations, hindering task-specific perception and exploration.\nIn this paper, we propose VideoExplorer, a framework grounded in the principle of ``thinking with video'', which naturally intertwines planning, temporal grounding, and scalable perception into a coherent reasoning process.\nRather than reasoning over a static context, VideoExplorer iteratively formulates sub-questions, locates relevant moments, and performs task-oriented, temporally scalable video understanding until reaching the final answer, enabling faithful, efficient, and interpretable reasoning.\nTo address the lack of LVU training resources, we construct a long-video reasoning dataset using difficulty-adaptive sampling to ensure high-quality trajectories on complex tasks.\nBuilding on this dataset, we design a two-stage training pipeline: supervised trajectory initialization followed by trajectory-level preference optimization, encouraging adaptive temporal grounding and iterative information integration guided by downstream rewards.\nExtensive evaluations on popular long-video understanding and reasoning benchmarks demonstrate VideoExplorer's significant advantage over existing baselines, highlighting its robustness, adaptability, and efficiency. \nOur code is available in this repository."},"_bibtex":{"value":"@misc{\nyuan2026videoexplorer,\ntitle={VideoExplorer: Boosting Long Video Understanding with Dynamic Temporal Grounding},\nauthor={Huaying Yuan and Zheng Liu and Junjie Zhou and Hongjin Qian and Yan Shu and Nicu Sebe and Ji-Rong Wen and Zhicheng Dou},\nyear={2026},\nurl={https://openreview.net/forum?id=mPxHWZl9bs}\n}"},"title":{"value":"VideoExplorer: Boosting Long Video Understanding with Dynamic Temporal Grounding"},"pdf":{"value":"/pdf/f305e3d691ea334d0c5a04faa311f5c9e10af18a.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"yuan|videoexplorer_boosting_long_video_understanding_with_dynamic_temporal_grounding"},"authorids":{"value":["~Huaying_Yuan1","~Zheng_Liu4","~Junjie_Zhou2","~Hongjin_Qian1","~Yan_Shu3","~Nicu_Sebe1","~Ji-Rong_Wen1","~Zhicheng_Dou1"]},"authors":{"value":["Huaying Yuan","Zheng Liu","Junjie Zhou","Hongjin Qian","Yan Shu","Nicu Sebe","Ji-Rong Wen","Zhicheng Dou"]}},"version":2},{"content":{"venue":{"value":"ICTC 2023"},"pdf":{"value":"https://ieeexplore.ieee.org/iel7/10391854/10391860/10393566.pdf"},"venueid":{"value":"dblp.org/conf/ICTC/2023"},"paperhash":{"value":"park|complex_motionaware_splatting_for_video_frame_interpolation"},"authorids":{"value":["~Minho_Park1","https://dblp.org/search/pid/api?q=author:Yuseok_Bae:"]},"html":{"value":"https://doi.org/10.1109/ICTC58733.2023.10393566"},"_bibtex":{"value":"@inproceedings{DBLP:conf/ictc/ParkB23,\n  author={Minho Park and Yuseok Bae},\n  title={Complex Motion-aware Splatting for Video Frame Interpolation},\n  year={2023},\n  cdate={1672531200000},\n  pages={1872-1876},\n  url={https://doi.org/10.1109/ICTC58733.2023.10393566},\n  booktitle={ICTC},\n  crossref={conf/ictc/2023}\n}\n"},"abstract":{"value":"Video frame interpolation, a crucial component of computer vision, synthesizes additional frames to enhance the frame rate of a video, leading to improved performance with minimal additional cost. Despite recent advancements with deep learning and convolutional neural networks (CNNs), it still remains a challenge to generate precise intermediate frames, especially when complex and fast motions are involved. This paper presents a novel deep learning-based framework for video frame interpolation that incorporates a complex motion detection module and proposes a complex motion-aware splatting (CMS) method. We employ a forward warping approach that uses a complex motion map as a weight map in splatting. The framework further leverages a module that embeds temporal and spatial information from the frame sequence to acquire motion information. The effectiveness of our proposed model is demonstrated through qualitative and quantitative results on a public dataset."},"title":{"value":"Complex Motion-aware Splatting for Video Frame Interpolation"},"authors":{"value":["Minho Park","Yuseok Bae"]}},"tmdate":1744097819553,"pdate":1672531200000,"tcdate":1744097817933,"writers":["~"],"signatures":["~Minho_Park2"],"forum":"5z6VcN4ZwD","license":"CC BY-SA 4.0","number":381393,"cdate":1672531200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1744097819553,"domain":"DBLP.org","id":"5z6VcN4ZwD","version":2},{"content":{"summary":{"value":"The paper investigates three issues of MVBench: 1) independence of video or video motion, 2) bias in the generated question-answer pairs, and 3) heavy reliance on world knowledge in questions. A significant part of the paper was written to prove and showcase these problems in the MVBench. A new benchmark called TVBench is proposed to mitigate these issues by redesigning the questions and available choices. The new benchmark attempts to prove that with no visual input or just image input, the models will perform like random guesses. Some strong video models also perform so even with full video inputs. The experiment also presents the results of inputting video frames in reverse order or shuffled order to prove that the benchmark questions requires understanding on the true video motion to be answered correctly."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"Please see the weaknesses."},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"- The paper addresses some critical issues with existing video understanding benchmarks, which is that the correct answers to many questions do not rely on information from the video or video motion. Thus, proposing a new benchmark to resolve these issues are well-motivated.\n- The paper explains in detailed examples and some ablation studies to prove that these problems exist widely in MVBench.\n- Based solely on the reverse & shuffle order experiment results in Table 4, it seems that TVBench indeed improves some questions' reliance on video inputs and the motion contained in those videos."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"While I appreciate the authors' great efforts to prove their statements of the issues, I am confused by many details after reading through the paper and am not fully convinced by the quality of the new benchmark.\n\n- Some of the results in the figures and tables, or the way they are presented, can be confusing. In Fig.1 left, the trend seems linear after the models achieve a certain level of performance (>50) on MVBench. In Fig. 1, right, MVBench shows a performance drop in VIdeoChat2 when the video is reversed. How do these results support the claim that MVBench does not measure temporal understanding? What is Table 1 trying to prove? I cannot compare the results of GPT-4o + image inputs with Gemini 1.5 Pro + video input to the conclusion that a single image is sufficient. You should at least fix other variables and leave the input as the one changing to prove that. Besides, even though the results are close, did you prove that the questions answered correctly are the same ones? It's a similar issue in Table 2 that I cannot understand how text-only rows could be compared to video input rows since they are using different models. \n\n- Since it's a benchmark, it should attempt to document the performance of as many models as possible. A lot of video models are missing, such as Video-LLaVA, mPLUG-Owl, PandaGPT, ImageBind, Video-LLaMa and etc. In addition, the GPT-4 series can accept multiple images, which is essentially the same as video models with video inputs -- they all need to sample a certain number of frames as multiple image inputs. You can also concatenate multiple frames into one image and feed into the GPT-4 series. It doesn't make sense to me to only benchmark GPT-4o with a single frame input. \n\n- Writing is a big issue in this paper. So many details make it hard to understand the paper without being confused. \n  - In Fig. 1, what is the unit of the axes? \n  - Table 1 is presented but never referred to in the text. If I understand correctly, some \"Tab. 2\" should refer to Table 1 instead. Please also choose between \"Tab 2\" and \"Tab. 2\" so that searching is convenient. \n  - I think the paper shows an excessive amount of bad examples from MVBench, which makes some of these figures unnecessary. While it's good to identify and prove the existence of these problems, more efforts should be spent convincing the readers that the \"proposed\" benchmark is high-quality and indeed resolves these issues. \n  - I understand that Sec. 4 is trying to show that open-ended qa and evaluation are not reliable, but how does that matter with the main point of this paper? Multiple-choice-based QA and open-ended QA are different settings used in different benchmarks or evaluations. It doesn't convince me that TVBench is high quality by showing the weaknesses of open-ended QA -- they are different settings. \n  - In line 430, *following the model provided in Tab. 5 for each task*, what is *model* in Table 5? There is no *model* column in Table 5 and the appendix is too short to provide enough context. How many templates are you using? Are the templates in Table 5 showing all you are using? How did you collect these templates? If you have hired annotators, how did you ensure the quality of these templates? These are all important details to be included in the paper to convince readers about the quality of TVBench. \n  - This is a minor point, but I don't favor using statistics of **Huggingface downloads** as some sort of evidence in the introduction (lines 37-38). Regardless of whether you are trying to use the number to support MVBench or question its reliability, it's better to appreciate **scientific merits** instead of **popularity metrics** in academic writing."}},"nonreaders":[],"tmdate":1731429373058,"tcdate":1730588609719,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission11521/Reviewer_A5py"],"signatures":["ICLR.cc/2025/Conference/Submission11521/Reviewer_A5py"],"forum":"DrNN5qx66Z","number":2,"license":"CC BY 4.0","cdate":1730588609719,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission11521/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731429373058,"domain":"ICLR.cc/2025/Conference","replyto":"DrNN5qx66Z","id":"OqDeFoJTob","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"TLDR":{"value":"We reveal the limitations of existing video-language benchmarks and present TVBench, a benchmark designed to truly assess temporal video-language understanding."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video-Language evaluation","Video-Language benchmark"]},"primary_area":{"value":"datasets and benchmarks"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Large language models have demonstrated impressive performance when integrated with vision models even enabling video understanding. However, evaluating these video models presents its own unique challenges, for which several benchmarks have been proposed. In this paper, we show that the currently most used video-language benchmarks can be solved without requiring much temporal reasoning. We identified three main issues in existing datasets: (i) static information from single frames is often sufficient to solve the tasks (ii) the text of the questions and candidate answers is overly informative, allowing models to answer correctly without relying on any visual input (iii) world knowledge alone can answer many of the questions, making the benchmarks a test of knowledge replication rather than visual reasoning. In addition, we found that open-ended question-answering benchmarks for video understanding suffer from similar issues while the automatic evaluation process with LLMs is unreliable, making it an unsuitable alternative. As a solution, we propose TVBench, a novel open-source video multiple-choice question-answering benchmark, and demonstrate through extensive evaluations that it requires a high level of temporal understanding. Surprisingly, we find that most recent state-of-the-art video-language models perform similarly to random performance on TVBench, with only a few models such as Qwen2-VL, and Tarsier clearly surpassing this baseline."},"_bibtex":{"value":"@misc{\ncores2025tvbench,\ntitle={{TVB}ench: Redesigning Video-Language Evaluation},\nauthor={Daniel Cores and Michael Dorkenwald and Manuel Mucientes and Cees G. M. Snoek and Yuki M Asano},\nyear={2025},\nurl={https://openreview.net/forum?id=DrNN5qx66Z}\n}"},"title":{"value":"TVBench: Redesigning Video-Language Evaluation"},"pdf":{"value":"/pdf/d10e747940b697aa51aeb790b853e347a10f151d.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"cores|tvbench_redesigning_videolanguage_evaluation"},"authorids":{"value":["~Daniel_Cores1","~Michael_Dorkenwald1","~Manuel_Mucientes1","~Cees_G._M._Snoek1","~Yuki_M_Asano1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Daniel Cores","Michael Dorkenwald","Manuel Mucientes","Cees G. M. Snoek","Yuki M Asano"]}},"version":2},{"content":{"summary":{"value":"This paper proposes Flow-Anchored Consistency Models (FACM), which incorporates a flow matching objective into consistency model training to stabilize training dynamics. The authors demonstrate that the key to effective joint optimization of flow matching and consistency model training is decoupling of flow velocity prediction and mean velocity prediction.\n\nSpecifically, in contrast to MeanFlow and its variants which are trained to predict flow velocity at time $t$ when given time condition $(t,t)$ and mean velocity from time $t$ to $r$ when given time condition $(t,r)$, FACMs are trained to predict flow velocity at $t$ given $2-t$ and mean velocity from time $t$ to $1$ given $t$. In contrast to MeanFlow, whose time conditions are coupled, i.e., $r = t$ for flow velocity, FACM uses decoupled time conditions, i.e., $r \\neq t$ for flow velocity. Additional techniques such as interpolation of the shortcut target with the current EMA network output and scalable chain-JVP implementation are provided to further accelerate training.\n\nThe authors verify the scalability of FACM on CIFAR-10, ImageNet 256x256, and a Text-to-Image dataset. FACM is shown to consistently out-perform several fast baselines such as IMM, MeanFlow, and sCM."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"**[Q1] Can the authors provide ablations mentioned in [W2] starting from a MeanFlow baseline?** In particular, I am curious how generative performance changes if (a) one fixes $r = 1$. I expect we would obtain a stronger few-step model, as it is solving an easier task compared to MeanFlow, which learns flow maps for all pairs $(t,r)$.\n\n**[Q2] Can the authors provide theoretical or experimental evidence of the hypothesis that instability in CM is caused by the lack of instantaneously velocity field supervision?** The authors could, for instance, plot the variance of the derivative term $dF_{\\theta^{-}}/dt$ without and with flow matching loss.\n\n**[Q3] What is the purpose the cosine similarity term in Eq. (9)?** I believe this loss is redundant, as a cosine similarity is already implicitly contained in the first flow matching objective. If this term plays a non-trivial role, then it should also be included in the ablations."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"- **[S1] This paper is original in the aspect that it provides a new perspective into training instability of CMs.** Specifically, the authors hypothesize that CM training instability arises from missing flow velocity supervision, and that one should decouple flow and mean velocity time conditions to mitigate conflict between the two velocities.\n\n- **[S2] This paper is significant in the aspect that it provides a number of techniques for scaling CMs.** The authors provide a number of practical techniques, such as interpolation of shortcut target and current prediction and chain-JVP, which are shown to work on difficult tasks such as ImageNet 256x256 and text-to-image generation."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- **[W1] The paper lacks theoretical novelty, in the sense that it is a special case of MeanFlow.** MeanFlow learns a flow map between all time pairs $(t,r)$ for $0 \\leq t < r \\leq 1$, along with a flow matching loss at $t = r$. FACM is a special instance of MeanFlow where a flow map is learned only for time pairs $(t,1)$ for $0 \\leq t \\leq 1$, also with flow matching loss at $t = r$, where $r = 2 - c\\_{FM}$ if one uses expanded time interval proposed in Section 3.3.2. Under this perspective, FACM may be viewed just as a collection of techniques for improved training of MeanFlow.\n\n- **[W2] The paper is missing key ablations of the proposed techniques.** Under the MeanFlow perspective written in [W1], key techniques proposed in this paper can be categorized into (1) fixing $r = 1$, i.e., training only with time pairs $(t,1)$, (2) expanded time interval, and (3) other extraneous techniques such as shortcut target interpolation, residual clamping, and tuning $\\alpha(t)$,$\\beta(t)$ in Eq. (11) and (13). However, authors only provide an ablation of (2), so with all techniques intertwined, it is difficult to judge what is the largest contributor to the final generative performance.\n\n- **[W3] The key hypothesis for instability of CMs, claimed by the authors, is unsupported.** At lines 171-175, the authors write \"However, without a stable anchor in the underlying flow, the model's output $F_\\theta$ quickly begins to drift. ... the derivative term in the identity grows to dominate the ground-truth velocity $v$, effectively diluting its supervisory signal.\" However, there is no supporting theory or experiment to verify this claim."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762916362273,"tcdate":1762254668799,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission2756/Reviewer_AryK"],"signatures":["ICLR.cc/2026/Conference/Submission2756/Reviewer_AryK"],"forum":"k9BpW1c4in","number":4,"license":"CC BY 4.0","cdate":1762254668799,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission2756/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762916362273,"domain":"ICLR.cc/2026/Conference","replyto":"k9BpW1c4in","id":"aSiO7sGzdS","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Image Generation","Consistency Model"]},"primary_area":{"value":"generative models"},"abstract":{"value":"Continuous-time Consistency Models (CMs) promise efficient few-step generation but face significant challenges with training instability. We argue this instability stems from a fundamental conflict: Training the network exclusively on a shortcut objective leads to the catastrophic forgetting of the instantaneous velocity field that defines the flow. Our solution is to explicitly anchor the model in the underlying flow, ensuring high trajectory fidelity during training. We introduce the Flow-Anchored Consistency Model (FACM), where a Flow Matching (FM) task serves as a dynamic anchor for the primary CM shortcut objective. Key to this Flow-Anchoring approach is a novel expanded time interval strategy that unifies optimization for a single model while decoupling the two tasks to ensure stable, architecturally-agnostic training. By distilling a pre-trained LightningDiT model, our method achieves a state-of-the-art FID of 1.32 with two steps (NFE=2) and 1.70 with just one step (NFE=1) on ImageNet 256$\\times$256. To address the challenge of scalability, we develop a memory-efficient Chain-JVP that resolves key incompatibilities with FSDP. This method allows us to scale FACM training on a 14B parameter model (Wan 2.2), accelerating its Text-to-Image inference from 2$\\times$40 to 2-8 steps. Our code and pretrained models:\nhttps://github.com/ali-vilab/FACM."},"_bibtex":{"value":"@inproceedings{\npeng2026facm,\ntitle={{FACM}: Flow-Anchored Consistency Models},\nauthor={Yansong Peng and Kai Zhu and Yu Liu and Pingyu Wu and Hebei Li and Xiaoyan Sun and Feng Wu},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=k9BpW1c4in}\n}"},"title":{"value":"FACM: Flow-Anchored Consistency Models"},"pdf":{"value":"/pdf/ca9d0d6460567a55222c5e6cdebf7f6e020c6ba3.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"peng|facm_flowanchored_consistency_models"},"authorids":{"value":["~Yansong_Peng1","~Kai_Zhu4","~Yu_Liu23","~Pingyu_Wu1","~Hebei_Li1","~Xiaoyan_Sun1","~Feng_Wu1"]},"authors":{"value":["Yansong Peng","Kai Zhu","Yu Liu","Pingyu Wu","Hebei Li","Xiaoyan Sun","Feng Wu"]}},"version":2},{"content":{"Primary_Area":{"value":"computational cognitive science / cognitive modeling"},"venue":{"value":"CCN 2026 Extended Abstracts Poster"},"abstract":{"value":"Looking at scenes engages multiple scene-selective regions in the brain, including the parahippocampal place area (PPA), retrosplenial cortex (RSC), and occipital place area (OPA). However, what distinguishes their computations remains less well understood. Prior work has largely relied on cognitive dissociations, which identify differences in when regions may respond but not what visual features drive functional differences. Here, we employ a novel approach: we use computational models to generate images predicted to differentially modulate activity between scene-selective regions. We built encoding models based on CLIP embeddings for the PPA, OPA, and RSC and found that they predicted responses across subjects, while also showing substantial region specificity.  We then used the directions in CLIP embedding space, together with a diffusion-based generative method, to synthesize visually similar image pairs (i.e., minimal pairs) predicted to enhance responses in one region while suppressing it in another. We found that the predicted differential response modulation generalizes across subjects, particularly for comparisons involving RSC, while dissociations between OPA and PPA were weaker. The resulting images suggest that the relevant visual feature differences depend on the pair of regions being compared. Our framework demonstrates how models can be taken beyond their predictivity to directly construct new images or hypotheses that may help reveal what different brain regions compute."},"_bibtex":{"value":"@inproceedings{\nwang2026modelderived,\ntitle={Model-derived Minimal Image Pairs Predict Differential Responses in Human Scene-Selective Cortex},\nauthor={Junxia Wang and Mainak Deb and Haider Al-Tahan and Diego Garcia Cerdas and Iris Groen and Apurva Ratan Murty},\nbooktitle={9th Annual Conference on Cognitive Computational Neuroscience},\nyear={2026},\nurl={https://openreview.net/forum?id=jOMIpRGhWh}\n}"},"title":{"value":"Model-derived Minimal Image Pairs Predict Differential Responses in Human Scene-Selective Cortex"},"Additional_Areas":{"value":["artificial intelligence / machine learning","psychological / behavioral research"]},"pdf":{"value":"/pdf/9cb25290396f47ff58ce63047299cb15d611cc20.pdf"},"venueid":{"value":"ccneuro.org/CCN/2026/Extended_Abstracts"},"paperhash":{"value":"wang|modelderived_minimal_image_pairs_predict_differential_responses_in_human_sceneselective_cortex"},"authorids":{"value":["~Junxia_Wang2","~Mainak_Deb1","~Haider_Al-Tahan2","~Diego_Garcia_Cerdas1","~Iris_Groen1","~Apurva_Ratan_Murty1"]},"authors":{"value":["Junxia Wang","Mainak Deb","Haider Al-Tahan","Diego Garcia Cerdas","Iris Groen","Apurva Ratan Murty"]}},"tmdate":1785520325204,"pdate":1778780220459,"tcdate":1775187744630,"writers":["ccneuro.org/CCN/2026/Extended_Abstracts","ccneuro.org/CCN/2026/Extended_Abstracts/Submission421/Authors"],"signatures":["ccneuro.org/CCN/2026/Extended_Abstracts/Submission421/Authors"],"forum":"jOMIpRGhWh","license":"CC BY 4.0","number":421,"cdate":1775187744630,"readers":["everyone"],"invitations":["ccneuro.org/CCN/2026/Extended_Abstracts/-/Submission","ccneuro.org/CCN/2026/Extended_Abstracts/-/Post_Submission","ccneuro.org/CCN/2026/Extended_Abstracts/-/Edit","ccneuro.org/CCN/2026/Extended_Abstracts/Submission421/-/Camera-Ready_Submission"],"mdate":1785520325204,"odate":1785520325169,"domain":"ccneuro.org/CCN/2026/Extended_Abstracts","id":"jOMIpRGhWh","version":2},{"content":{"summary":{"value":"This paper propose TAR-TVG, a novel timestamp anchor-constrained reasoning framework for temporal video grounding, which includes a efficient reinforcement learning strategy for extracting high-quality reasoning traces. The experiments reveal the improved performance for temporal video grounding with verifiable reasoning chains for progressively refined temporal estimations."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"1. See weakness.\n\n2. Why adopt a three-stage training process in the order of RL-SFT-RL instead of processes(such as SFT-RL only) in other orders? This may need to be better clarified by more ablation experiments with both evaluation performance and training cost.\n\n3. The paper may require a small adjustment of compilation format, such as citation font color in the main text and the underline in the reference, which is different from papers from previous years and other reviewed papers, and may be caused by the compilation."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. The proposed method adopts a three-stage training process with reinforcement learning, improving the interpretability and accuracy of temporal video grounding.\n\n2. The experiment result is solid with high performance on several evaluation benchmarks for temporal video grounding.\n\n3. Convincing visualization examples are provided to prove the effectiveness of the proposed method."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Only testing on temporal video grounding tasks would be limited for the proposed method adopted with LVLMs. Many temporal-aware LVLMs also demonstrate effective generalization ability for related video understanding tasks, not limited to temporal video grounding only. I encourage the authors to evaluate the proposed method on more temporally related video understanding benchmarks.\n\n2. The challenges claimed by this paper, that ‘ the prompts of the used method only implicitly guide the model to output timestamp tags, often leading to missing, incorrect-formatted, or irrelevant tags’, are relatively weak. Since quite a few LVLM-free methods, such as FlashVTG, achieve good temporal video grounding performance and do not have such challenges. The authors should reorganize the statement of addressed challenges."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762921477593,"tcdate":1761294690899,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission10093/Reviewer_EVwv"],"signatures":["ICLR.cc/2026/Conference/Submission10093/Reviewer_EVwv"],"forum":"XOuHTBLrPP","number":1,"license":"CC BY 4.0","cdate":1761294690899,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission10093/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762921477593,"domain":"ICLR.cc/2026/Conference","replyto":"XOuHTBLrPP","id":"1HVyq8frUz","forumContent":{"TLDR":{"value":"TAR-TVG is a reinforcement learning framework that improves Temporal Video Grounding by introducing timestamp anchors to guide reasoning, boosting both accuracy and interpretability."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Reinforcement Learning","Temporal Video Grounding"]},"primary_area":{"value":"foundation or frontier models, including LLMs"},"abstract":{"value":"Temporal video grounding aims to localize relevant video segments based on a given query. Large Vision-Language Models (LVLMs) can address this by taking a video and query as input and outputting the time duration. Recently, some methods fine-tune LVLMs with reinforcement learning (RL), encouraging them to generate reasoning traces for better interpretability. They also prompt the model to include `<timestamp></timestamp>` tags into the reasoning process to strengthen the connection between the reasoning and the final output. However, these prompts only implicitly guide the model to output timestamp tags, often leading to missing, incorrect-formatted, or irrelevant tags. To address this issue, we propose Timestamp Anchor-constrained Reasoning for Temporal Video Grounding (TAR-TVG). By designing reinforcement learning reward functions, we explicitly enforce the inclusion of timestamp tags as anchors within the reasoning traces, providing explicit format control and accuracy validation based on soft IoU. Furthermore, when multiple timestamp anchors appear, the reward function is designed to ensure that the accuracy of these anchors progressively improves, thereby mimicking the human-like thought process of refining from coarse to fine. These additional constraints on timestamp anchors encourage the model to better understand the task of temporal video grounding, thereby improving its grounding performance. Additionally, we first run an RL stage purely for data collection. The collected samples are then used to SFT a fresh base model, and we finally apply RL fine-tuning to the SFT-initialized model. Experiments show that our model achieves state-of-the-art performance while producing verifiable reasoning chains with progressively refined temporal estimations."},"_bibtex":{"value":"@misc{\nguo2026tartvg,\ntitle={{TAR}-{TVG}: Enhancing {LVLM}s with Timestamp Anchor-Constrained Reasoning for Temporal Video Grounding},\nauthor={chaohong guo and Xun Mo and Yongwei Nie and Xuemiao Xu and Chao Xu and Fei Yu and Chengjiang Long},\nyear={2026},\nurl={https://openreview.net/forum?id=XOuHTBLrPP}\n}"},"title":{"value":"TAR-TVG: Enhancing LVLMs with Timestamp Anchor-Constrained Reasoning for Temporal Video Grounding"},"pdf":{"value":"/pdf/2c8eb70ba64afd9b12c4a57bd96656552f10520b.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"guo|tartvg_enhancing_lvlms_with_timestamp_anchorconstrained_reasoning_for_temporal_video_grounding"},"authorids":{"value":["~chaohong_guo1","~Xun_Mo1","~Yongwei_Nie1","~Xuemiao_Xu1","~Chao_Xu18","~Fei_Yu13","~Chengjiang_Long1"]},"authors":{"value":["chaohong guo","Xun Mo","Yongwei Nie","Xuemiao Xu","Chao Xu","Fei Yu","Chengjiang Long"]}},"version":2},{"content":{"summary":{"value":"This paper presents a method for constructing an audio-video generative model with low computational cost by leveraging pre-trained single-modal diffusion models for audio and video. The authors propose a lightweight joint guidance module that aligns audio and video outputs by adjusting scores to approximate the joint distribution, computed through the gradient of an optimal discriminator that distinguishes real and fake audio-video pairs."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. How is the low computational cost demonstrated in this work? Are there specific experiments or data that quantify this claim?\n2. While the quality of the generated 2-second samples is good, how does the model perform on longer video samples?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"1. This paper tackles an important problem in multi-modal generative modeling by introducing a cost-effective approach to synchronize audio and video generation, offering practical value for applications constrained by limited computational resources.\n2. The custom loss function helps to steady the gradient of the discriminator. Extensive evaluations on benchmark datasets highlight enhanced fidelity for individual modalities and better alignment between audio and video, all achieved with minimal additional parameters, underscoring the efficiency of the proposed method."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Although I understand that the experiments are based on MM-Diffusion, the AIST and Landscape datasets are relatively small. The authors should discuss the generalizability of their method on larger datasets.\n2. The paper claims that the method is model-agnostic; however, only one baseline is tested, which makes it difficult to confirm its general applicability.\n3. Besides quantitative metrics, a subjective evaluation of the generated samples, would be beneficial to better assess the quality from a perceptual standpoint."}},"nonreaders":[],"tmdate":1732534052993,"tcdate":1730852873041,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission5876/Reviewer_pfZF"],"signatures":["ICLR.cc/2025/Conference/Submission5876/Reviewer_pfZF"],"forum":"agbiPPuSeQ","number":4,"license":"CC BY 4.0","cdate":1730852873041,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission5876/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732534052993,"domain":"ICLR.cc/2025/Conference","replyto":"agbiPPuSeQ","id":"sCc6dbtRyA","forumContent":{"TLDR":{"value":"Building audio-video joint generative model with minimal computational cost via multimodal discriminator on the top of pre-trained single-modal generative models"},"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Diffusion models","Multi-modal data","Audio-visual generative models"]},"supplementary_material":{"value":"/attachment/002a0a08a630cdbec4ef8021c406d0e8b4216e96.zip"},"primary_area":{"value":"generative models"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"This study aims to construct an audio-video generative model with minimal computational cost by leveraging pre-trained single-modal generative models for audio and video.\nTo achieve this, we propose a novel method that guides single-modal models to cooperatively generate well-aligned samples across modalities. \nSpecifically, given two pre-trained base diffusion models, we train a lightweight joint guidance module to adjust scores separately estimated by the base models to match the score of joint distribution over audio and video. \nWe show that this guidance can be computed using the gradient of the optimal discriminator, which distinguishes real audio-video pairs from fake ones independently generated by the base models. \nBased on this analysis, we construct a joint guidance module by training this discriminator.\nAdditionally, we adopt a loss function to stabilize the discriminator's gradient and make it work as a noise estimator, as in standard diffusion models. \nEmpirical evaluations on several benchmark datasets demonstrate that our method improves both single-modal fidelity and multimodal alignment with relatively few parameters.\nThe code is available at: [https://github.com/SonyResearch/MMDisCo](https://github.com/SonyResearch/MMDisCo)."},"_bibtex":{"value":"@inproceedings{\nhayakawa2025mmdisco,\ntitle={{MMD}isCo: Multi-Modal Discriminator-Guided Cooperative Diffusion for Joint Audio and Video Generation},\nauthor={Akio Hayakawa and Masato Ishii and Takashi Shibuya and Yuki Mitsufuji},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=agbiPPuSeQ}\n}"},"title":{"value":"MMDisCo: Multi-Modal Discriminator-Guided Cooperative Diffusion for Joint Audio and Video Generation"},"pdf":{"value":"/pdf/d2c6e85ea8b816d429d4916770364984d9ec58b1.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"hayakawa|mmdisco_multimodal_discriminatorguided_cooperative_diffusion_for_joint_audio_and_video_generation"},"authorids":{"value":["~Akio_Hayakawa1","~Masato_Ishii1","~Takashi_Shibuya1","~Yuki_Mitsufuji1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Akio Hayakawa","Masato Ishii","Takashi Shibuya","Yuki Mitsufuji"]}},"version":2},{"content":{"summary":{"value":"This paper proposes TS-Mamba, a state-of-the-art method for online video super-resolution (VSR), leveraging trajectory-aware shifted state space models (SSMs) for efficient spatio-temporal information aggregation. The model addresses challenges in long-term temporal modeling for real-time applications while keeping computational complexity low. The approach combines token-level spatio-temporal aggregation with a novel trajectory-aware shifted Mamba aggregation (TSMA) module. TS-Mamba constructs trajectories within video frames to select similar tokens from previous frames and aggregates them using shifted SSM blocks. This enables improved video frame restoration with reduced computational overhead. The proposed model is evaluated on several benchmark datasets (REDS, Vimeo-90K-T, Vid4) and outperforms five state-of-the-art online VSR methods in terms of PSNR/SSIM while reducing complexity by over 22.7%."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"See the weakness."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. Performance: The model achieves superior performance in terms of PSNR/SSIM and visual quality across multiple benchmark datasets (REDS, Vid4, Vimeo-90K-T) and degradation types (BI and BD), demonstrating its robustness in real-world video restoration scenarios.\n\n2. Computational Efficiency: TS-Mamba successfully reduces complexity by 22.7% in terms of MACs compared to existing methods, making it a strong candidate for real-time online VSR applications. The model is also one of the fastest among the tested methods.\n\n3. Comprehensive Ablation Study: The authors provide an extensive ablation study validating the importance of trajectory-aware components and the shifted SSM blocks, as well as the impact of different design choices (e.g., token number, shift operations). This strengthens the validity of their claims and helps illustrate the contributions of each part of the model."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Lack of Comparison. The proposed scheme fails to discuss or compare with several recent restoration schemes that leverage state-space models (SSMs, e.g. Mamba) for superior performance. For instance, MambaIR and MambaIRv2 introduced a residual Mamba-based backbone (with convolution and channel attention) to capture global dependencies in image super-resolution and denoising, outperforming a SwinIR Transformer baseline. The more recent TAMambaIR improves efficiency by modulating the state-space transition for complex textures and using multi-directional scanning, achieving state-of-the-art results across image restoration tasks (e.g. super-resolution, deraining, low-light enhancement). In the video domain, VSRM proposes dual Spatial-to-Temporal and Temporal-to-Spatial Mamba blocks for long-range spatio-temporal feature extraction and a deformable cross-Mamba alignment module for flexible frame alignment, yielding new state-of-the-art performance on VSR benchmarks. Likewise, The authors should incorporate and evaluate these methods to ensure a comprehensive comparison with the current state-of-the-art in video super-resolution. （I believe that the VSR method, when transformed into the online VSR setting, can be compared.）\n\n2. Insufficient Theoretical Justification: While the paper presents an innovative approach, the theoretical justification for the trajectory-aware shifted state space model is somewhat lacking. Specifically, a formal analysis of the long-term spatio-temporal aggregation process and a comparison with existing models using similar principles (e.g., in flow-guided deformable attention models) would provide more clarity. Additionally, the authors mention the use of shifted operations to enhance spatial continuity but could further explore the theoretical implications of this operation on model behavior.\n\n3. Insufficient Validation. A notable weakness is the absence of experiments on real-world VSR datasets. The method’s results are only reported on standard synthetic benchmarks (REDS4, Vid4, Vimeo-90K) with bicubic or simulated blur degradation, but no evaluations on real degraded videos were provided . As real applications involve unknown and complex degradations (compression artifacts, sensor noise, motion blur, etc.), the paper should have validated the approach on established real-world video SR benchmarks. For example, testing on datasets like RealVSR (ICCV 2021), which is built with a dual-camera system on the iPhone 11 Pro Max, would demonstrate the model’s robustness to in-the-wild conditions\n\n4. Potential on Certain Datasets: The model performs exceptionally well on the benchmarks but lacks a more nuanced discussion of failure cases or scenarios where the method might struggle. A brief mention of potential limitations (e.g., when frames are highly dynamic or occlusions are prevalent) would enhance the robustness of the paper's claims."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762919222114,"tcdate":1761894628265,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission7012/Reviewer_2xrU"],"signatures":["ICLR.cc/2026/Conference/Submission7012/Reviewer_2xrU"],"forum":"RygnSGcV49","number":4,"license":"CC BY 4.0","cdate":1761894628265,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission7012/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762919222114,"domain":"ICLR.cc/2026/Conference","replyto":"RygnSGcV49","id":"YoywqqaWkw","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Video Super-resolution","Online","Mamba","Trajectory"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Online video super-resolution (VSR) is an important technique for many real-world video processing applications, which aims to restore the current high-resolution video frame based on temporally previous frames. Most of the existing online VSR methods solely employ one neighboring previous frame to achieve temporal alignment, which limits long-range temporal modeling of videos. Recently, state space models (SSMs) have been proposed with linear computational complexity and a global receptive field, which significantly improve computational efficiency and performance. In this context, this paper presents a novel online VSR method based on Trajectory-aware Shifted SSMs (TS-Mamba), leveraging both long-term trajectory modeling and low-complexity Mamba to achieve efficient spatio-temporal information aggregation. Specifically, TS-Mamba first constructs the trajectories within a video to select the most similar tokens from the previous frames. Then, a Trajectory-aware Shifted Mamba Aggregation (TSMA) module consisting of proposed shifted SSMs blocks is employed to aggregate the selected tokens. The shifted SSMs blocks are designed based on Hilbert scannings and corresponding shift operations to compensate for scanning losses and strengthen the spatial continuity of Mamba. Additionally, we propose a trajectory-aware loss function to supervise the trajectory generation, ensuring the accuracy of token selection when training our model. Extensive experiments on three widely used VSR test datasets demonstrate that compared with six online VSR benchmark models, our TS-Mamba achieves state-of-the-art performance in most cases and over 22.7% complexity reduction (in MACs). The source code for TS-Mamba is available at https://github.com/QZ1-boy/TS-Mamba."},"_bibtex":{"value":"@inproceedings{\nzhu2026trajectoryaware,\ntitle={Trajectory-aware Shifted State Space Models for Online Video Super-Resolution},\nauthor={Qiang Zhu and Xiandong MENG and Yuxuan Jiang and Fan Zhang and David Bull and Shuyuan Zhu and Bing Zeng and Ronggang Wang},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=RygnSGcV49}\n}"},"title":{"value":"Trajectory-aware Shifted State Space Models for Online Video Super-Resolution"},"pdf":{"value":"/pdf/63207ba1dfa93f1f2594d2cd3d80db2064dc0722.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"zhu|trajectoryaware_shifted_state_space_models_for_online_video_superresolution"},"authorids":{"value":["~Qiang_Zhu6","~Xiandong_MENG1","~Yuxuan_Jiang8","~Fan_Zhang6","~David_Bull1","~Shuyuan_Zhu1","~Bing_Zeng1","~Ronggang_Wang1"]},"authors":{"value":["Qiang Zhu","Xiandong MENG","Yuxuan Jiang","Fan Zhang","David Bull","Shuyuan Zhu","Bing Zeng","Ronggang Wang"]}},"version":2},{"content":{"summary":{"value":"The paper introduces EndoAssistant, a large-scale vision-language dataset designed to enhance understanding of endoscopic surgery scenes. It addresses the limitations of existing datasets, which are small in scale and diversity, by providing a significantly larger collection of 590 videos, 65,844 unique images, 30,002 captions, and 157,589 image-caption/question-answer pairs. The dataset focuses on improving tasks like cross-modal retrieval, visual question answering (VQA), and image classification within the surgical context. The data curation process involves keyframe extraction, ASR transcription, hierarchical image classification, and rigorous text cleaning with clinical validation. EndoAssistant's vision-language data pipeline includes EndoCaption (image-caption pairs) and EndoQA (image-question-answer pairs), both of which are shown to improve baseline model performance across multiple benchmarks."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"It is recommended to include in the limitations a discussion on the challenging conditions often faced in endoscopic surgery, such as inconsistent lighting, obstructed views, interference from bodily fluids, as well as data biases arising from differences in hospitals, types of surgery, anatomical regions, or patient demographics."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":3},"strengths":{"value":"(1)EndoAssistant is the first large-scale, open-source vision-language dataset explicitly tailored for endoscopic surgery, surpassing previous datasets like Cholec80 in scale and semantic diversity. By integrating multiple existing models (CLIP, Whisper, GPT-4) into a surgeon-in-the-loop framework, this dataset provides a novel approach to generating diverse, medically relevant Q&A data from endoscopic videos.\n\n(2)The paper clearly outlines each stage of the data pipeline, from video collection to model evaluation. The inclusion of figures detailing the dataset creation process and examples of the Q&A pairs and image captions adds clarity. Each stage is accompanied by performance metrics that demonstrate the impact of EndoAssistant on downstream tasks."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"(1)Although EndoAssistant is curated for endoscopic tasks, some baseline models used (e.g., CLIP) are pre-trained on general vision-language datasets, which might limit their performance in highly specialized domains like medical imagery. Fine-tuning on similar medical datasets could make the evaluation more aligned with the dataset's intended use.\n\n[1] Hecvl: Hierarchical video-language pretraining for zero-shot surgical phase recognition\n[2] Procedure-Aware Surgical Video-language Pretraining with Hierarchical Knowledge Augmentation\n\n(2) How does EndoAssistant perform on surgical tasks beyond classification and VQA, such as surgical phase recognition or anomaly detection?\n\n(3) While the dataset draws from multiple open sources, there is a limited analysis of potential biases within the data. Different hospitals, surgical types, anatomical regions, or patient demographics could introduce significant variability, impacting the generalizability of the model.\n\n(4) The dataset relies on relatively straightforward image-text pairing and may not fully capture deeper semantic alignment between the visual and language modalities (e.g., multi-level semantic alignment or co-occurrence patterns). Surgical procedures often involve subtle contextual changes, and certain tools or anatomical structures may carry different meanings across procedural stages."}},"nonreaders":[],"tmdate":1732895222178,"tcdate":1730346420830,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission7850/Reviewer_A9cJ"],"signatures":["ICLR.cc/2025/Conference/Submission7850/Reviewer_A9cJ"],"forum":"voYshhbWeJ","number":1,"license":"CC BY 4.0","cdate":1730346420830,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission7850/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732895222178,"domain":"ICLR.cc/2025/Conference","replyto":"voYshhbWeJ","id":"WKRYEx1Vot","forumContent":{"TLDR":{"value":"We present a large-scale, meticulously curated dataset from surgical endoscopic videos, designed to use image-text pairs to facilitate medical scene understanding."},"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Medical image","endoscopy","vision-language model"]},"primary_area":{"value":"datasets and benchmarks"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Endoscopic interventions offer a minimally invasive approach, minimizing patient discomfort and facilitating expedited recovery. Proficient training of junior surgeons necessitates the ability to analyze and interpret endoscopic scenes through questioning and answering. Consequently, the development of a robust foundation model for endoscopic visual language understanding holds immense value for medical training and surgical education. However, existing endoscopy vision-language datasets are limited in scale and diversity, consisting of only 50 videos sourced from a few clinical sites, thus posing a significant hurdle to the advancement of generalized and robust artificial intelligence models for endoscopic surgical applications. To address this challenge, we present a large-scale, meticulously curated image-text dataset of surgical endoscopic scenes from expert surgeons, designed to propel a vision-language assistant in medical scene understanding. Encompassing 590 open-source videos spanning more than 91 hours, our curated dataset includes 65,844 unique images, 30,002 unique captions, and 157,589 image-caption/question-answering pairs. This dataset aims to assist the development of automated systems to support medical professionals by mitigating repetitive tasks. We present a comprehensive endoscopic surgery assisting pipeline, (1) a first-ever image-caption dataset specifically for endoscopic scenes; (2) an image-question-answer dataset that offers greater size and diversity compared to existing collections; (3) rigorous evaluation demonstrating its efficacy in downstream surgical endoscopic scene comprehension tasks like classification, retrieval and visual question answering."},"_bibtex":{"value":"@misc{\ngong2025endoassistant,\ntitle={EndoAssistant: A Large-scale Vision-Language Dataset for Endoscopic Surgery Understanding from Open-Source Videos},\nauthor={Xuan Gong and Balu Harshavardan Koduru and Yuanhao Zhai and Shun Liu and Nan Xi and Xi Tang and Yuan Zhang and Tenzin Lhakpa and Yunjie Tian and Yuxuan Sun and Tianyu Luan and Ziqing Xue and Junsong Yuan and David Doermann},\nyear={2025},\nurl={https://openreview.net/forum?id=voYshhbWeJ}\n}"},"title":{"value":"EndoAssistant: A Large-scale Vision-Language Dataset for Endoscopic Surgery Understanding from Open-Source Videos"},"pdf":{"value":"/pdf/b187719ea9a245e64eadc2a235fa89f67d271bd3.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"gong|endoassistant_a_largescale_visionlanguage_dataset_for_endoscopic_surgery_understanding_from_opensource_videos"},"authorids":{"value":["~Xuan_Gong1","~Balu_Harshavardan_Koduru1","~Yuanhao_Zhai1","~Shun_Liu1","~Nan_Xi1","~Xi_Tang2","~Yuan_Zhang31","~Tenzin_Lhakpa1","~Yunjie_Tian1","~Yuxuan_Sun3","~Tianyu_Luan1","~Ziqing_Xue1","~Junsong_Yuan2","~David_Doermann2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Xuan Gong","Balu Harshavardan Koduru","Yuanhao Zhai","Shun Liu","Nan Xi","Xi Tang","Yuan Zhang","Tenzin Lhakpa","Yunjie Tian","Yuxuan Sun","Tianyu Luan","Ziqing Xue","Junsong Yuan","David Doermann"]}},"version":2},{"content":{"summary":{"value":"This paper introduces InfiniBench, a video understanding benchmark dataset featuring the longest video duration (average 52.59 minutes per video) and the largest number of question-answer pairs (108.2K) to evaluate 9 different video understanding tasks.\n\nThe authors conducted comprehensive evaluations of existing large multimodal models (including commercial models like GPT-4V, Gemini 1.5 Flash, and open-source models). Experiments show that even leading AI models still face challenges in long video understanding, with the best models GPT-4V and Gemini 1.5 Flash achieving average accuracy rates of only 49.16% and 42.72% respectively."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"1. Add references and discussions of related work.\n2. It would be better to evaluate more long-video models (e.g., Qwen2VL) and different input frame rates (1, 8, 32, 128, and more).\n3. Since most question-answer pairs are generated by GPT-4o, could this lead to inflated evaluation results for GPT-4o? Analysis is needed regarding dataset quality, hallucination rates, and potential information leakage issues."},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The questions are comprehensive and well-structured, covering multiple dimensions and employing diverse construction strategies for different types of questions.\n2. The evaluation methods are reasonable, adopting different assessment metrics for multiple-choice and open-ended questions."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The paper lacks discussion of related work. For example, benchmarks proposed in Video-MME, LVBench, and Long VideoBench published in June 2024 are very similar to InfiniBench.\n\n2. Most of the question-answer pairs are generated by GPT-4o. Although multiple information sources were used as input, it's difficult to guarantee the quality of the dataset.\n\n3. Part of the data comes from IMDB content, which likely appeared multiple times in the training corpus of LLMs used by video models, potentially leading to dataset leakage issues."}},"nonreaders":[],"tmdate":1733157158709,"tcdate":1730628877621,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission2250/Reviewer_6qyh"],"signatures":["ICLR.cc/2025/Conference/Submission2250/Reviewer_6qyh"],"forum":"2D0uXQbntW","number":5,"license":"CC BY 4.0","cdate":1730628877621,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission2250/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1733157158709,"domain":"ICLR.cc/2025/Conference","replyto":"2D0uXQbntW","id":"cltowg8Aq8","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["video understanding","benchmark","long video benchmark","long video understanding"]},"supplementary_material":{"value":"/attachment/e4286990af860ccb3f16a8fce15b1da8635ef0b5.zip"},"primary_area":{"value":"datasets and benchmarks"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Understanding long videos, ranging from tens of minutes to several hours, presents unique challenges in video comprehension. Despite the increasing importance of long-form video content, existing benchmarks primarily focus on shorter clips. To address this gap, we introduce InfiniBench a comprehensive benchmark for very long video understanding which presents 1)very long video duration, averaging 52.59 minutes per video 2)The largest number of question-answer pairs, 108.2K 3) Diversity in questions that examine nine different skills and include both multiple-choice questions and open-ended questions 4) Memory questions, such as Global Appearance that require remembering and tracking the visual aspects through the video. Using InfiniBench, we comprehensively evaluate existing Large Multi-Modality Models (LMMs) on each skill, including the commercial models such as GPT-4o and Gemini 1.5 Flash and the recent open-source models. \nThe evaluation shows significant challenges in our benchmark.\nOur findings reveal that even leading AI models like GPT-4o and Gemini 1.5 Flash face challenges in achieving high performance in long video understanding, with average accuracies of just 56.01 % and 43.32 %, and average scores of 3.25 and 2.79 out of 5, respectively.\nQwen2-VL matches Gemini's performance in the MCQ skills but lags significantly in open-ended question tasks.\nWe hope this benchmark will stimulate the LMMs community towards long video and human-level understanding."},"_bibtex":{"value":"@misc{\nataallah2025infinibench,\ntitle={InfiniBench: A Comprehensive Benchmark for Large Multimodal Models in Very Long Video Understanding},\nauthor={Kirolos Ataallah and Chenhui Gou and Eslam Mohamed BAKR and Khushbu Pahwa and Jian Ding and Mohamed Elhoseiny},\nyear={2025},\nurl={https://openreview.net/forum?id=2D0uXQbntW}\n}"},"title":{"value":"InfiniBench: A Comprehensive Benchmark for Large Multimodal Models in Very Long Video Understanding"},"pdf":{"value":"/pdf/032fb6c70b723929924e625fabf7cebb80b48405.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"ataallah|infinibench_a_comprehensive_benchmark_for_large_multimodal_models_in_very_long_video_understanding"},"authorids":{"value":["~Kirolos_Ataallah1","~Chenhui_Gou1","~Eslam_Mohamed_BAKR1","~Khushbu_Pahwa1","~Jian_Ding3","~Mohamed_Elhoseiny1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Kirolos Ataallah","Chenhui Gou","Eslam Mohamed BAKR","Khushbu Pahwa","Jian Ding","Mohamed Elhoseiny"]}},"version":2},{"content":{"venue":{"value":"CoRR 2026"},"pdf":{"value":"https://arxiv.org/pdf/2606.12087v1"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"deng|fortsearcher_synthesizing_shortcutresistant_search_tasks_for_training_deep_search_agents"},"html":{"value":"https://doi.org/10.48550/arXiv.2606.12087"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2606-12087,\n  publtype={informal},\n  author={Jia Deng and Yimeng Chen and Xiaoqing Xiang and Ziyang Zeng and Shuo Tang and Wayne Xin Zhao and Feng Chang and Chuan Hao and Yuan Wei and Ran Tao and Bryan Dai and Ji-Rong Wen},\n  title={FORT-Searcher: Synthesizing Shortcut-Resistant Search Tasks for Training Deep Search Agents},\n  year={2026},\n  month={June},\n  cdate={1780272000000},\n  journal={CoRR},\n  volume={abs/2606.12087},\n  url={https://doi.org/10.48550/arXiv.2606.12087}\n}\n"},"abstract":{"value":"Training deep search agents requires verifiable questions whose answers remain unavailable until sufficient evidence has been acquired through search. Existing synthesis methods often increase apparent difficulty by enriching graph structures, but structural complexity alone does not guarantee realized search difficulty: the intended search process can collapse through a cheaper identifying route. We formalize this gap with a shortcut-aware difficulty framework and identify four actionable shortcut risks: evidence co-coverage, single-clue selectivity, exposed constants, and prior-knowledge binding. To diagnose their realized effects, we use trajectory signatures including solving cost, answer hit time, and prior-shortcut rate. Guided by this framework, we introduce FORT, a Framework of Shortcut-Resistant Training-Data Synthesis. FORT constructs shortcut-resistant training data by controlling shortcut risks across entity selection, evidence graph construction, question formulation, and adversarial refinement. Experiments show that FORT induces longer pre-answer search and fewer shortcut patterns than existing open-source deep search datasets. Using the resulting trajectories, we train FORT-Searcher with supervised fine-tuning (SFT) only, and it achieves the best overall performance among comparable-size open-source search agents on challenging deep search benchmarks. Relevant resources will be made available at https://github.com/RUCAIBox/FORT-Searcher."},"title":{"value":"FORT-Searcher: Synthesizing Shortcut-Resistant Search Tasks for Training Deep Search Agents"},"authors":{"value":[{"fullname":"Jia Deng","username":""},{"fullname":"Yimeng Chen","username":""},{"fullname":"Xiaoqing Xiang","username":""},{"fullname":"Ziyang Zeng","username":"~Ziyang_Zeng2"},{"fullname":"Shuo Tang","username":""},{"fullname":"Wayne Xin Zhao","username":""},{"fullname":"Feng Chang","username":""},{"fullname":"Chuan Hao","username":""},{"fullname":"Yuan Wei","username":""},{"fullname":"Ran Tao","username":""},{"fullname":"Bryan Dai","username":""},{"fullname":"Ji-Rong Wen","username":""}]}},"tmdate":1785330671574,"pdate":1798675200000,"externalIds":["dblp:journals/corr/abs-2606-12087"],"tcdate":1785330668344,"writers":["~","OpenReview.net/Public_Article/DBLP.org","OpenReview.net/Support"],"signatures":["~Ziyang_Zeng2"],"forum":"fDmQ8mvUog","license":"CC BY-SA 4.0","number":111630,"cdate":1780272000000,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/DBLP.org/-/Record","OpenReview.net/Public_Article/-/Edit"],"mdate":1785330671574,"domain":"OpenReview.net/Public_Article","id":"fDmQ8mvUog","version":2},{"content":{"venue":{"value":"IEEE Transactions on Circuits and Systems for Video Technology"},"pdf":{"value":"https://ieeexplore.ieee.org/iel8/76/10857816/10666730.pdf"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"hou|bidirectional_erroraware_fusion_network_for_video_inpainting"},"html":{"value":"https://doi.org/10.1109/TCSVT.2024.3454641"},"abstract":{"value":"Existing video inpainting approaches tend to adopt vision transformers with rare customized designs, which poses two limitations. Firstly, the conventional self-attention mechanism treats tokens from invalid and valid regions equally and mingles them, which may incur blurriness. Secondly, these approaches merely employ forward frames as references, while ignoring the past inpainted frames, which are also valuable in enhancing temporal consistency and offering more available information. In this paper, we propose a new video inpainting network, called Bidirectional Error-Aware Fusion Network (BEAF-Net). Concretely, on one hand, we propose a tailored Error-Aware Transformer (EAT) that discerns different tokens by assigning dynamic weights to bridle the use of erroneous tokens. Meanwhile, each EAT is equipped with a Spatial Feature Enhancement (SFE) layer to synthesize features with multi-scales. On the other hand, we apply a pair of EATs to utilize forward reference frames and past inpainted frames simultaneously, and a proposed Bidirectional Fusion (BiF) layer is exerted to blend the aggregation results adaptively. By coupling these novel designs, our proposed BEAF-Net completely leverages the location priors, multi-scale perception, and past predictions to produce more faithful and consistent inpainting results. We corroborate our BEAF-Net on two commonly-used video inpainting datasets: DAVIS and Youtube-VOS, where the experimental results demonstrate BEAF-Net compares favorably with state-of-the-art solutions. Video examples can be found at https://github.com/JCATCV/BEAF-Net."},"title":{"value":"Bidirectional Error-Aware Fusion Network for Video Inpainting"},"authors":{"value":[{"fullname":"Jiacheng Hou","username":"~Jiacheng_Hou1"},{"fullname":"Zhong Ji"},{"fullname":"Jinyu Yang"},{"fullname":"Feng Zheng"}]}},"tmdate":1789090352909,"pdate":1735689600000,"externalIds":["doi:10.1109/tcsvt.2024.3454641"],"tcdate":1769339557327,"writers":["~","OpenReview.net/Public_Article/ORCID.org","OpenReview.net/Support"],"signatures":["~Jiacheng_Hou1"],"forum":"oMlKmABmRd","license":"CC BY-SA 4.0","number":38054,"cdate":1727694438757,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/ORCID.org/-/Record","OpenReview.net/Public_Article/-/Edit","OpenReview.net/-/Edit"],"mdate":1789090352909,"domain":"OpenReview.net/Public_Article","id":"oMlKmABmRd","version":2},{"content":{"comment":{"value":"We sincerely thank the ethics reviewer for highlighting important concerns regarding bias and fairness in LVLM-based VAD.\n\n**Clarification and Acknowledgment:**  \nWe fully agree that reliance on statistical shortcuts (e.g., visual-textual co-occurrence patterns) can lead to biased predictions in surveillance scenarios, potentially resulting in disproportionate impacts on certain groups or contexts. In fact, our work was explicitly motivated by this very concern. The hallucination phenomenon we study in Section 3 (and Appendix A.1–A.2) reveals that LVLMs often misclassify semantically normal scenes as anomalous due to high-frequency visual triggers such as \"fire\" or \"police car\"—a shortcut behavior that poses clear fairness risks if left unaddressed.  \n\nWe acknowledge that our main paper did not sufficiently emphasize how such hallucinations could disproportionately affect certain groups or environments. We now explicitly recognize that biased model behavior—if unaddressed—could lead to unfair deployment outcomes, particularly in sensitive applications like surveillance.\n\n**Response and Future Plans:**  \nWhile VAD-DPO effectively mitigates shortcut reliance through preference-based optimization, we agree that it does not yet address *fairness from a group-sensitive perspective*. In future work, we plan to explore fairness-aware extensions to VAD-DPO, including:\n\n- **Incorporating demographic-sensitive fairness objectives**, when annotations include person-level attributes (e.g., age, ethnicity, clothing) in synthetic or real datasets;\n- **Analyzing disparate false positive rates** across demographic groups or geographic regions using real-world surveillance data;\n- **Integrating VAD-DPO with fairness-aware training techniques**, such as adversarial debiasing or group reweighting, to jointly enhance robustness and equity.\n\n**Planned Revision:**  \n- Explicitly acknowledge the fairness risks of shortcut-driven hallucinations in the *Limitations* section;  \n- Add an in-depth discussion of ethical implications in the *Broader Societal Impacts* section, including the risk that shortcut-driven false positives in sensitive applications (e.g., surveillance) may disproportionately harm certain communities by mislabeling their normal activities as suspicious;  \n- Directly link this societal risk to our technical findings on co-occurrence-based hallucinations, and discuss potential bias mitigation strategies beyond VAD-DPO.\n\nWe appreciate the reviewer’s insightful feedback and believe that addressing these points will strengthen the ethical grounding and broader impact of our work."},"title":{"value":"Response to  Ethics Reviewer LzhQ: Fairness and Bias in Shortcut-Driven Predictions"}},"parentInvitations":"NeurIPS.cc/2025/Conference/-/Official_Comment","tmdate":1761723078817,"tcdate":1754018774433,"writers":["NeurIPS.cc/2025/Conference","NeurIPS.cc/2025/Conference/Submission933/Authors"],"signatures":["NeurIPS.cc/2025/Conference/Submission933/Authors"],"forum":"crPlJvwHhS","number":1,"license":"CC BY 4.0","cdate":1754018774433,"readers":["everyone"],"invitations":["NeurIPS.cc/2025/Conference/Submission933/-/Official_Comment","NeurIPS.cc/2025/Conference/-/Edit"],"mdate":1761723078817,"domain":"NeurIPS.cc/2025/Conference","replyto":"wtn2YbQQiS","id":"jUINUIcaXH","forumContent":{"venue":{"value":"NeurIPS 2025 poster"},"TLDR":{"value":"This paper reveals that LVLMs in video anomaly detection rely on pre-trained statistical shortcuts instead of scene-aware reasoning."},"keywords":{"value":["Video Anomaly Detection","Video Anomaly Reasoning"]},"primary_area":{"value":"applications"},"flagged_for_ethics_review":{"value":true},"abstract":{"value":"Large Vision-Language Models (LVLMs) pretrained on large-scale multimodal data have shown promising capabilities in Video Anomaly Detection (VAD). However, their ability to reason about abnormal events based on scene semantics remains underexplored. In this paper, we investigate LVLMs’ behavior in VAD from a visual-textual co-occurrence perspective, focusing on whether their decisions are driven by statistical shortcuts between visual instances and textual phrases. By analyzing visual-textual co-occurrence in pretraining data and conducting experiments under different data settings, we reveal a hallucination phenomenon: LVLMs tend to rely on co-occurrence patterns between visual instances and textual phrases associated with either normality or abnormality, leading to incorrect predictions when these high-frequency objects appear in semantically mismatched contexts. To address this issue, we propose VAD-DPO, a direct preference optimization method supervised with counter-example pairs. By constructing visually similar but semantically contrasting video clips, VAD-DPO encourages the model to align its predictions with the semantics of scene rather than relying on co-occurrence patterns. Extensive experiments on six benchmark datasets demonstrate the effectiveness of VAD-DPO in enhancing both anomaly detection and reasoning performance, particularly in scene-dependent scenarios."},"_bibtex":{"value":"@inproceedings{\nzhang2025do,\ntitle={Do {LVLM}s Truly Understand Video Anomalies? Revealing Hallucination via Co-Occurrence Patterns},\nauthor={Menghao Zhang and Huazheng Wang and Pengfei Ren and Kangheng Lin and Qi Qi and Haifeng Sun and Zirui Zhuang and Lei Zhang and Jianxin Liao and Jingyu Wang},\nbooktitle={The Thirty-ninth Annual Conference on Neural Information Processing Systems},\nyear={2025},\nurl={https://openreview.net/forum?id=crPlJvwHhS}\n}"},"title":{"value":"Do LVLMs Truly Understand Video Anomalies? Revealing Hallucination via Co-Occurrence Patterns"},"pdf":{"value":"/pdf/aa61db404a3612db834accb6d2176f8e78a7e6f8.pdf"},"venueid":{"value":"NeurIPS.cc/2025/Conference"},"paperhash":{"value":"zhang|do_lvlms_truly_understand_video_anomalies_revealing_hallucination_via_cooccurrence_patterns"},"authorids":{"value":["~Menghao_Zhang2","~Huazheng_Wang2","~Pengfei_Ren2","~Kangheng_Lin1","~Qi_Qi1","~Haifeng_Sun2","~Zirui_Zhuang1","~Lei_Zhang67","~Jianxin_Liao2","~Jingyu_Wang1"]},"authors":{"value":["Menghao Zhang","Huazheng Wang","Pengfei Ren","Kangheng Lin","Qi Qi","Haifeng Sun","Zirui Zhuang","Lei Zhang","Jianxin Liao","Jingyu Wang"]}},"version":2},{"content":{"venue":{"value":"CoRR 2025"},"pdf":{"value":"http://arxiv.org/pdf/2504.14096v2"},"venueid":{"value":"dblp.org/journals/CORR/2025"},"paperhash":{"value":"kulkarni|videopasta_7k_preference_pairs_that_matter_for_videollm_alignment"},"authorids":{"value":["","https://dblp.org/search/pid/api?q=author:Pooyan_Fazli:"]},"html":{"value":"https://doi.org/10.48550/arXiv.2504.14096"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2504-14096,\n  publtype={informal},\n  author={Yogesh Kulkarni and Pooyan Fazli},\n  title={VideoPASTA: 7K Preference Pairs That Matter for Video-LLM Alignment},\n  year={2025},\n  month={April},\n  cdate={1743465600000},\n  journal={CoRR},\n  volume={abs/2504.14096},\n  url={https://doi.org/10.48550/arXiv.2504.14096}\n}\n"},"abstract":{"value":"Video-language models (Video-LLMs) excel at understanding video content but struggle with spatial relationships, temporal ordering, and cross-frame continuity. To address these limitations, we introduce VideoPASTA (Preference Alignment with Spatio-Temporal-Cross Frame Adversaries), a framework that enhances Video-LLMs through targeted preference optimization. VideoPASTA trains models to distinguish accurate video representations from carefully crafted adversarial examples that deliberately violate spatial, temporal, or cross-frame relationships. With only 7,020 preference pairs and Direct Preference Optimization, VideoPASTA enables models to learn robust representations that capture fine-grained spatial details and long-range temporal dynamics. Experiments demonstrate that VideoPASTA is model agnostic and significantly improves performance, for example, achieving gains of up to 3.8% on LongVideoBench, 4.1% on VideoMME, and 4.0% on MVBench, when applied to various state-of-the-art Video-LLMs. These results demonstrate that targeted alignment, rather than massive pretraining or architectural modifications, effectively addresses core video-language challenges. Notably, VideoPASTA achieves these improvements without any human annotation or captioning, relying solely on 32-frame sampling. This efficiency makes our approach a scalable plug-and-play solution that seamlessly integrates with existing models while preserving their original capabilities."},"title":{"value":"VideoPASTA: 7K Preference Pairs That Matter for Video-LLM Alignment"},"authors":{"value":["Yogesh Kulkarni","Pooyan Fazli"]}},"tmdate":1770677535345,"pdate":1735689600000,"tcdate":1751481307494,"writers":["~"],"signatures":["~Yogesh_Kulkarni1"],"forum":"1Fgb7sDrA9","license":"CC BY-SA 4.0","number":565706,"cdate":1743465600000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1770677535345,"domain":"DBLP.org","id":"1Fgb7sDrA9","version":2},{"content":{"venue":{"value":"EMNLP 2025"},"pdf":{"value":"https://aclanthology.org/2025.emnlp-main.1647.pdf"},"venueid":{"value":"dblp.org/conf/EMNLP/2025"},"paperhash":{"value":"kulkarni|videopasta_7k_preference_pairs_that_matter_for_videollm_alignment"},"authorids":{"value":["~Yogesh_Kulkarni1","https://dblp.org/search/pid/api?q=author:Pooyan_Fazli:"]},"html":{"value":"https://doi.org/10.18653/v1/2025.emnlp-main.1647"},"_bibtex":{"value":"@inproceedings{DBLP:conf/emnlp/KulkarniF25,\n  author={Yogesh Kulkarni and Pooyan Fazli},\n  title={VideoPASTA: 7K Preference Pairs That Matter for Video-LLM Alignment},\n  year={2025},\n  cdate={1735689600000},\n  pages={32354-32379},\n  url={https://doi.org/10.18653/v1/2025.emnlp-main.1647},\n  booktitle={EMNLP},\n  crossref={conf/emnlp/2025}\n}\n"},"abstract":{"value":"Video-language models (Video-LLMs) excel at understanding video content but struggle with spatial relationships, temporal ordering, and cross-frame continuity. To address these limitations, we introduce VideoPASTA (Preference Alignment with Spatio-Temporal-Cross Frame Adversaries), a framework that enhances Video-LLMs through targeted preference optimization. VideoPASTA trains models to distinguish accurate video representations from carefully crafted adversarial examples that deliberately violate spatial, temporal, or cross-frame relationships. With only 7,020 preference pairs and Direct Preference Optimization, VideoPASTA enables models to learn robust representations that capture fine-grained spatial details and long-range temporal dynamics. Experiments demonstrate that VideoPASTA is model agnostic and significantly improves performance, for example, achieving gains of up to + 3.8 percentage points on LongVideoBench, +4.1 on VideoMME, and +4.0 on MVBench, when applied to various state-of-the-art Video-LLMs. These results demonstrate that targeted alignment, rather than massive pretraining or architectural modifications, effectively addresses core video-language challenges. Notably, VideoPASTA achieves these improvements without any human annotation or captioning, relying solely on 32-frame sampling. This efficiency makes our approach a scalable plug-and-play solution that seamlessly integrates with existing models while preserving their original capabilities."},"title":{"value":"VideoPASTA: 7K Preference Pairs That Matter for Video-LLM Alignment"},"authors":{"value":["Yogesh Kulkarni","Pooyan Fazli"]}},"tmdate":1770677520016,"pdate":1767139200000,"externalIds":["dblp:conf/emnlp/KulkarniF25"],"tcdate":1770677510445,"writers":["~"],"signatures":["~Yogesh_Kulkarni1"],"forum":"B3i0Y4RFs0","license":"CC BY-SA 4.0","number":821350,"cdate":1735689600000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1770677520016,"domain":"DBLP.org","id":"B3i0Y4RFs0","version":2},{"content":{"summary":{"value":"This paper introduces SALSA-V, a model for video-to-audio (V2A) generation, designed to synthesize high-fidelity, long-form audio that is precisely synchronized with video content. The authors propose three main contributions: (1) a shortcut-augmented training objective to enable high-quality audio generation in very few sampling steps; (2) a masked flow matching approach that allows the model to perform audio-conditioned generation and outpainting, thereby enabling the creation of audio for long-form videos through iterative extension; and (3) a new contrastive audio-visual synchronization model built upon a strong pre-trained vision backbone to yield high-resolution alignment features."},"soundness":{"value":4},"confidence":{"value":5},"questions":{"value":"1.  **On Novelty**: The paper's core components (architecture, masked training, batch size insights) appear to be adapted from prior work. Could the authors clarify the primary technical novelty beyond the successful integration of these known techniques?\n\n2.  **On Experimental Baselines**: The experimental comparison is limited. Could the authors justify the exclusion of several recent and relevant baselines like V-AURA, AudioX, and Frieren, and perhaps provide a comparative analysis on key metrics?\n\n3.  **On Few-Step Generation Performance**: The \"shortcut loss\" is a key claimed contribution, but SOTA results are shown at 32 steps. Could the authors provide key metrics (e.g., DeSync, FAD) for the 8-step generation case against baselines to demonstrate the practical effectiveness of this feature?\n\n4.  **On Reproducibility**: The evaluation benchmark is vaguely described. For reproducibility, could the authors provide a precise composition and list of identifiers for their custom test set?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":1},"strengths":{"value":"*   **Clear Presentation and Structured Evaluation**: The paper is clearly written, with a well-defined problem motivation and a structured presentation of its methods. The evaluation is methodical, employing a range of established objective metrics alongside a human listening study.\n\n*   **Focus on Practical V2A Challenges**: The work addresses relevant practical limitations in the video-to-audio (V2A) domain, particularly the efficiency of the sampling process and the synthesis of audio for longer-form videos. This focus is pertinent to improving the applicability of such generative models.\n\n*   **Achieves Strong Temporal Synchronization**: A notable strength of the proposed model is its ability to generate tightly synchronized audio. The paper reports state-of-the-art results on the DeSync metric, and this quantitative improvement in temporal alignment is also corroborated by the human evaluation study."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The paper, while presenting a well-engineered system, suffers from several weaknesses that question the significance and novelty of its contribution.\n\n1.  **Limited Novelty and Incremental Contribution**: The primary weakness of this work is its reliance on existing techniques across its entire pipeline, from the architectural framework and training methods to the experimental conclusions. This makes the overall contribution feel incremental rather than innovative.\n    *   The model's core design, which combines semantic and high-resolution synchronization features for conditioning, directly follows the framework established by **MMAudio[1]**.\n    *   The use of a masked training objective for audio conditioning is a known technique, with similar approaches having been explored in works like **AudioX[2]** and **MultiFoley[3]**.\n    *   The key insight regarding the choice of batch size for training the synchronization model is also acknowledged to be a direct replication of findings from **Synchformer [4]**.\n    While the integration of these parts is functional, the paper does not introduce a new fundamental concept, algorithm, or a significant insight to the field.\n\n2.  **Insufficient Experimental Baseline and Missing Citations**: The paper's experimental comparison is narrow, failing to benchmark against a sufficient number of relevant contemporary models. This makes it difficult to accurately assess its performance in the broader context of the field.\n    *   The main quantitative comparison (Table 1) only includes two other models, **FoleyCrafter** and **MMAudio**. Other highly relevant V2A methods, such as the autoregressive model **V-AURA [5]** and other diffusion-based models like **AudioX** and **Frieren[6]**, are not included in the benchmark.\n    *   A more comprehensive comparison against a wider array of recent models is necessary to robustly support the claims of state-of-the-art performance.\n\n3.  **Contradictory Results and Lack of Significant Improvement**: Several claims in the paper are not fully supported by its own results, and the overall improvement is not compelling.\n    *   **Lower Subjective Audio Quality**: For a generative model, perceptual quality is paramount. The human evaluation in Table 1 shows that SALSA-V's **Audio Quality** score (2.96) is lower than the baseline MMAudio (3.16). This key result undermines the paper's claim of outperforming existing methods. The overall subjective improvement is not significant.\n    *   **Potentially Flawed Long-Form Evaluation**: The paper compares its iterative, chunk-based generation with MMAudio's one-shot, full-sequence approach. As the models operate with different context windows and inference strategies, this direct comparison of metrics could be misleading.\n    *   **Disconnect Between Main Contribution and SOTA Results**: The \"shortcut loss\" for few-step sampling is highlighted as a major contribution. However, the main results in Table 1, which establish the model's SOTA synchronization, are based on 32 sampling steps. The paper does not demonstrate that this claimed SOTA performance is retained in the few-step regime, thus disconnecting the novel claim from the primary comparative results.\n\n4.  **Lack of Clarity in Experimental Details and Reproducibility Issues**: The description of the experimental setup lacks the necessary detail for reproducibility.\n    *   The test set used for evaluation is vaguely described as a composition of \"a holdout set of in-the-wild videos, the VGGSound test set, and UnAV-100\". Without precise details on the composition, data splits, and preprocessing of this custom benchmark, the results are not verifiable or reproducible by the community.\n\n[1] Cheng H K, Ishii M, Hayakawa A, et al. MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis[C]//Proceedings of the Computer Vision and Pattern Recognition Conference. 2025: 28901-28911.\n\n[2] Tian Z, Jin Y, Liu Z, et al. Audiox: Diffusion transformer for anything-to-audio generation[J]. arXiv preprint arXiv:2503.10522, 2025.\n\n[3] Chen Z, Seetharaman P, Russell B, et al. Video-guided foley sound generation with multimodal controls[C]//Proceedings of the Computer Vision and Pattern Recognition Conference. 2025: 18770-18781.\n\n[4] Iashin V, Xie W, Rahtu E, et al. Synchformer: Efficient synchronization from sparse cues[C]//ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2024: 5325-5329.\n\n[5] Viertola I, Iashin V, Rahtu E. Temporally aligned audio for video with autoregression[C]//ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2025: 1-5.\n\n[6] Wang Y, Guo W, Huang R, et al. Frieren: Efficient video-to-audio generation network with rectified flow matching[J]. Advances in Neural Information Processing Systems, 2024, 37: 128118-128138."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762922765497,"tcdate":1761783174508,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission11727/Reviewer_jLCz"],"signatures":["ICLR.cc/2026/Conference/Submission11727/Reviewer_jLCz"],"forum":"3FcjKYmNY5","number":3,"license":"CC BY 4.0","cdate":1761783174508,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission11727/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762922765497,"domain":"ICLR.cc/2026/Conference","replyto":"3FcjKYmNY5","id":"YYtyogEjnV","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["video-to-audio","audio generation","diffusion"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"We propose SALSA-V, a multimodal video-to-audio generation model capable of synthesizing highly synchronized, high-fidelity long-form audio from silent video content. Our approach introduces a masked diffusion objective, enabling audio-conditioned generation and the seamless synthesis of audio sequences of unconstrained length. Additionally, by integrating a shortcut loss into our training process, we achieve rapid generation of high-quality audio samples in as few as eight sampling steps, paving the way for near-real-time applications without requiring dedicated fine-tuning or retraining. We demonstrate that SALSA-V significantly outperforms existing state-of-the-art methods in both audiovisual alignment and synchronization with video content in quantiative evaluation and a human listening study. Furthermore, our use of random masking during training enables our model to match spectral characteristics of reference audio samples, broadening its applicability to professional audio synthesis tasks such as Foley generation and sound design."},"_bibtex":{"value":"@misc{\ndellali2026salsav,\ntitle={{SALSA}-V: Shortcut-Augmented Long-form Synchronized Audio from Videos},\nauthor={Amir Dellali and Luca A Lanzend{\\\"o}rfer and Florian Gr{\\\"o}tschla and Roger Wattenhofer},\nyear={2026},\nurl={https://openreview.net/forum?id=3FcjKYmNY5}\n}"},"title":{"value":"SALSA-V: Shortcut-Augmented Long-form Synchronized Audio from Videos"},"pdf":{"value":"/pdf/780501365dd6cdc6dc1c38131015c6ed7be54b52.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"dellali|salsav_shortcutaugmented_longform_synchronized_audio_from_videos"},"authorids":{"value":["~Amir_Dellali1","~Luca_A_Lanzendörfer1","~Florian_Grötschla1","~Roger_Wattenhofer1"]},"authors":{"value":["Amir Dellali","Luca A Lanzendörfer","Florian Grötschla","Roger Wattenhofer"]}},"version":2},{"content":{"summary":{"value":"The authors present an autoregressive video diffusion model that augmented with a masked inverse dynamics model.  The authors show that by grounding the video prediction of the video diffusion model with action-relevant masks, they obtain fast and accurate closed-loop control.  The authors claim the learned masks prioritizes the video generation quality of the action-relevant regions. The masked inverse dynamics model is from an earlier paper called Vidar.  The authors show that the model when pre-trained on one million cross-embodiment episodes outperforms baselines in success rates and latency, even on unseen robotic platforms."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"- Figure 1, the left hand side should be drawn in a classical feedback control diagram that is popular in robotics rather than having an execution and environment feedback at the top.  A diagram that shows observation o_t leading to o_t^' producing a_t which then results in o_{t+1} that is feedback is better than the current diagram - which is confusing as the bottom o_1, o_2, o_3 feedback is not clear.  (This is my opinion.)  The figure on the right is disconnected from the figure on the left - this can be revised and the caption can be modified to better connect the two subfigures.\n\n- Figure 2 is missing the action a_1.  Typically, o_{t+1} is produced based on a_t that is computed as a function of o_t.  Here, o_1 -> o_2^' -> a_2, which misses a_1.\n\n- Sec. 2.1: v_\\theta : V x R X C -> V, here the spaces V and C are not defined.  Is V the space of the video x and C the space of the condition c?\n\n- The diffusion loss in Eq. (1) looks at the difference of the vector field v_\\theta and the mean vector field (x_0 - x_1).  v_theta is the vector field at x_t at t, where as (x_0 - x_1) is just the mean vector field.  How does getting these two to match up be sufficient for diffusion?  Can you double check this loss.\n\n- In Eq. (5), why not have a_t = I(o_t) instead of I(\\hat o_t) ?  (I understand that ideally o_t and \\hat o_t should be the same, but they typically are not due to training loss not being zero.)"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"- The video generation model incorporates an embodiment-aware loss to make the model focus generation on action-relevant regions.\n\n- Paper produces a method for low-latency video generation for closed-loop control.\n\n- Results work on unseen robotic platforms."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The key strength of this paper was to enable the video generation to focus generation on action-relevant regions.  However, this is possible due to the prior work on masked inverse dynamics model (from Vidar?).  This makes the novelty of this work weaker."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762916473093,"tcdate":1762098180913,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission2976/Reviewer_fS78"],"signatures":["ICLR.cc/2026/Conference/Submission2976/Reviewer_fS78"],"forum":"gsvjCTIYPb","number":3,"license":"CC BY 4.0","cdate":1762098180913,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission2976/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762916473093,"domain":"ICLR.cc/2026/Conference","replyto":"gsvjCTIYPb","id":"Fb2eUTuQhP","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Robotics","Video Diffusion Model","Computer Vision"]},"supplementary_material":{"value":"/attachment/8ce7790fc4e592a31bf85c4f6cc642b29e9534e9.zip"},"primary_area":{"value":"applications to robotics, autonomy, planning"},"abstract":{"value":"Robotic arm manipulation in data-scarce settings is a highly challenging task due to the complex embodiment dynamics and diverse contexts. Recent video-based approaches have shown great promise in capturing and transferring the temporal and physical interactions by pre-training on Internet-scale video data. However, such methods are often not optimized for the embodiment-specific closed-loop control, typically suffering from high latency and insufficient grounding. In this paper, we present Vidarc (Video Diffusion for Action Reasoning and Closed-loop Control), a novel autoregressive embodied video diffusion approach augmented by a masked inverse dynamics model. By grounding video predictions with action-relevant masks and incorporating real-time feedback through cached autoregressive generation, Vidarc achieves fast, accurate closed-loop control. Pre-trained on one million cross-embodiment episodes, Vidarc surpasses state-of-the-art baselines, achieving at least a 15% higher success rate in real-world deployment and a 91% reduction in latency. We also highlight its robust generalization and error correction capabilities across previously unseen robotic platforms."},"_bibtex":{"value":"@misc{\nfeng2026vidarc,\ntitle={Vidarc: Low Latency Embodied Video Diffusion Model with Closed-loop Control},\nauthor={Yao Feng and Chendong Xiang and Xinyi Mao and Hengkai Tan and Zuyue Zhang and Shuhe Huang and Kaiwen Zheng and Haitian Liu and Hang Su and Jun Zhu},\nyear={2026},\nurl={https://openreview.net/forum?id=gsvjCTIYPb}\n}"},"title":{"value":"Vidarc: Low Latency Embodied Video Diffusion Model with Closed-loop Control"},"pdf":{"value":"/pdf/b6ebd3124a61d8dfba2342153bc18da637196e92.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"feng|vidarc_low_latency_embodied_video_diffusion_model_with_closedloop_control"},"authorids":{"value":["~Yao_Feng2","~Chendong_Xiang1","~Xinyi_Mao1","~Hengkai_Tan1","~Zuyue_Zhang1","~Shuhe_Huang1","~Kaiwen_Zheng2","~Haitian_Liu2","~Hang_Su3","~Jun_Zhu2"]},"authors":{"value":["Yao Feng","Chendong Xiang","Xinyi Mao","Hengkai Tan","Zuyue Zhang","Shuhe Huang","Kaiwen Zheng","Haitian Liu","Hang Su","Jun Zhu"]}},"version":2},{"content":{"summary":{"value":"This paper introduces a framework, VSNLS, for Video Scene Graph Generation (VidSGG) that leverages natural language supervision from video captions to reduce the high cost of manual annotation. Unlike existing methods, VSNLS uses two modules tailored for video data: a Temporality-aware Caption Segmentation (TCS) module to capture time markers in captions, and an Action Duration Variability-aware Caption-Frame Alignment (ADV) module to align captions with frames based on action duration. This approach enables the model to learn from weak supervision, allowing it to predict dynamic relationships and even generalize to unseen actions. The proposed method is shown to be effective on the Action Genome dataset, demonstrating improved performance over traditional image-based weak supervision techniques."},"soundness":{"value":2},"confidence":{"value":5},"questions":{"value":"The results are reported on the Action Genome dataset, but it is unclear how well the framework would perform on other VidSGG datasets with different types of actions or domain-specific captions. How about VidOR dataset?\n\nHow does VSNLS handle noisy or ambiguous captions? Since the method depends on the quality of video captions, it would be helpful to understand how robust the framework is to variations in caption detail and accuracy."},"rating":{"value":6},"details_of_ethics_concerns":{"value":"n/a"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"VSNLS addresses the high annotation cost of fully supervised Video Scene Graph Generation (VidSGG) by using weak supervision from video captions. This method eliminates the need for extensive manual annotation of all frames, which is both time-consuming and costly."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"-Dependency on Co-occurrence Priors: VSNLS, like many video-based scene graph models, often relies on co-occurrence priors to estimate relationships. This dependency can reduce the generalizability of the model to unseen object combinations, potentially limiting its performance on diverse or novel datasets where the objects and relations deviate from the training distribution.\n\n-Handling Long-term Relations: Clip-based methods in VSNLS might struggle with capturing relations that unfold over extended sequences, which require a larger receptive field. If clip length is limited (e.g., fixed at 30 frames), the model may miss detecting relations that span longer temporal windows, thereby impacting its ability to accurately model long-term dependencies.\n\n-Occlusion and Tracking Limitations: Although clip-based approaches are often advocated for mitigating long-term tracking issues like occlusion, the use of short-term tubelets might still face challenges in fragmented tracking. Occlusions or visual artifacts can still lead to lost tracklets or misalignments, reducing the reliability of relationship detection over long video durations.\n\n-High Computational Cost for Short Clips: Analyzing shorter clip sequences may require repeated processing of overlapping frames to maintain temporal context, leading to higher computational and memory demands. This could make the model less efficient and limit scalability for very long video sequences where full contextual understanding is crucial.\n\n-Important References Missing:\n[1] Video visual relation detection via iterative inference. ACM MM 2021.\n[2] Winner: Weakly-supervised hierarchical decomposition and alignment for spatio-temporal video grounding. CVPR 2023.\n[3] In Defense of Clip-based Video Relation Detection. TIP 2024."}},"nonreaders":[],"tmdate":1732520619912,"tcdate":1730689332694,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission3326/Reviewer_aNnm"],"signatures":["ICLR.cc/2025/Conference/Submission3326/Reviewer_aNnm"],"forum":"GQgPj1H4pO","number":3,"license":"CC BY 4.0","cdate":1730689332694,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission3326/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732520619912,"domain":"ICLR.cc/2025/Conference","replyto":"GQgPj1H4pO","id":"etRPaSeqXc","forumContent":{"TLDR":{"value":"We propose  a weakly-supervised video scene graph generation framework that aims to relieve the annotation costs by training a model using natural language supervision."},"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video Scene Understanding","Weakly Supervised Learning","Large Language Model"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Existing Video Scene Graph Generation (VidSGG) studies are trained in a fully supervised manner, which requires all frames in a video to be annotated, thereby incurring high annotation cost compared to Image Scene Graph Generation (ImgSGG). Although the annotation cost of VidSGG can be alleviated by adopting a weakly supervised approach commonly used for ImgSGG (WS-ImgSGG) that uses image captions, there are two key reasons that hinder such a naive adoption: 1) Temporality within video captions, i.e., unlike image captions, video captions include temporal markers (e.g., before, while, then, after) that indicate time-related details, and 2) Variability in action duration, i.e., unlike human actions in image captions, human actions in video captions unfold over varying duration. To address these issues, we propose a Natural Language-based Video Scene Graph Generation (NL-VSGG) framework that only utilizes the readily available video captions for training a VidSGG model. NL-VSGG consists of two key modules: Temporality-aware Caption Segmentation (TCS) module and Action Duration Variability-aware caption-frame alignment (ADV) module. Specifically, TCS segments the video captions into multiple sentences in a temporal order based on a Large Language Model (LLM), and ADV aligns each segmented sentence with appropriate frames considering the variability in action duration. Our approach leads to a significant enhancement in performance compared to simply applying the WS-ImgSGG pipeline to VidSGG on the Action Genome dataset. As a further benefit of utilizing the video captions as weak supervision, we show that the VidSGG model trained by NL-VSGG is able to predict a broader range of action classes that are not included in the training data, which makes our framework practical in reality."},"_bibtex":{"value":"@inproceedings{\nkim2025weakly,\ntitle={Weakly Supervised Video Scene Graph Generation via Natural Language Supervision},\nauthor={Kibum Kim and Kanghoon Yoon and Yeonjun In and Jaehyeong Jeon and Jinyoung Moon and Donghyun Kim and Chanyoung Park},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=GQgPj1H4pO}\n}"},"title":{"value":"Weakly Supervised Video Scene Graph Generation via Natural Language Supervision"},"pdf":{"value":"/pdf/6082acd7ebb0364ab36782cd36c1bb838a5bbf63.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"kim|weakly_supervised_video_scene_graph_generation_via_natural_language_supervision"},"authorids":{"value":["~Kibum_Kim1","~Kanghoon_Yoon2","~Yeonjun_In1","~Jaehyeong_Jeon1","~Jinyoung_Moon1","~Donghyun_Kim2","~Chanyoung_Park1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Kibum Kim","Kanghoon Yoon","Yeonjun In","Jaehyeong Jeon","Jinyoung Moon","Donghyun Kim","Chanyoung Park"]}},"version":2},{"content":{"summary":{"value":"This paper identifies and defines Shortcut Alignment as a key failure mode in Large Reasoning Models (LRMs). The authors posit that models learn to bypass their internal Chain-of-Thought (CoT) processes, instead issuing templated refusals based solely on surface cues from the input. They argue this shortcut is the root cause of widespread over-refusal on benign queries, which degrades the model's general helpfulness. To address this, the paper proposes a new training method called DIFT, centered on a CMI-Loss."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"See above."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"The paper's  contribution is the identification and formalization of \"Shortcut Alignment.\" This is a precise and insightful diagnosis of why modern safety-aligned LRMs suffer from over-refusal. Instead of relying on the generic \"safety vs. helpfulness trade-off,\" the authors provide a mechanistic explanation: the model's refusal decision becomes decoupled from its internal reasoning."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The paper's \"Shortcut Alignment\" premise is undermined by its own evidence in Appendix G.3. The failure case (Case E1) does not show a \\mathbit{c}\\ \\rightarrow\\ \\mathbit{y} decoupling; rather, the refusal (\\mathbit{y}) is perfectly faithful to a flawed CoT (\\mathbit{c}) that misinterprets the query. This suggests the root cause of over-refusal may be poor CoT quality, not a dependency on the CoT. The proposed CMI-Loss, which enforces \\mathbit{c}\\ \\rightarrow\\ \\mathbit{y} dependency, may therefore be misdiagnosing the problem.\n\n2. The CMI-Loss mechanism enforces the answer's (\\mathbit{y}) dependency on the CoT (\\mathbit{c}), which implicitly assumes the CoT is robust. However, the CoT itself is a known attack vector (e.g., H-cot [1]、mousetrap [2].). By training the model to trust its CoT more, the method risks amplifying the impact of CoT poisoning or hijacking attacks, forcing the model to faithfully execute a compromised reasoning chain. The paper lacks a robustness analysis against this critical, emergent threat.\n\n- [1] Martin Kuo, Jianyi Zhang, Aolin Ding, Qinsi Wang, Louis DiValentin, Yujia Bao, Wei Wei, Hai Li, and Yiran Chen. H-cot: Hijacking the chain-of-thought safety reasoning mechanism to jailbreak large reasoning models, including openai o1/o3, deepseek-r1, and gemini 2.0 flash thinking. arXiv preprint arXiv:2502.12893, 2025.\n\n- [2] Yang Yao, Xuan Tong, Ruofan Wang, Yixu Wang, Lujundong Li, Liang Liu, Yan Teng, and Yingchun Wang. A mousetrap: Fooling large reasoning models for jailbreak with chain of iterative chaos. arXiv preprint arXiv:2502.15806, 2025.\n\n3. The central claim of \"preserving safety\" is not fully supported by the quantitative data in Table 1. In several instances, the 'ours' method scores lower on key safety benchmarks (e.g., WildJB, WildChat) than the baseline. For instance, DeepSeek-32B's WildJB score drops from 90.8 to 88.8, and Qwen3-14B's drops from 94.4 to 92.8. This consistent (though small) degradation suggests the method introduces a \"safety tax\" and is, in effect, trading safety for reduced over-refusal (NOR), which should be explicitly acknowledged."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762920832316,"tcdate":1761913817104,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission9146/Reviewer_GvHB"],"signatures":["ICLR.cc/2026/Conference/Submission9146/Reviewer_GvHB"],"forum":"3qHILWiEob","number":2,"license":"CC BY 4.0","cdate":1761913817104,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission9146/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762920832316,"domain":"ICLR.cc/2026/Conference","replyto":"3qHILWiEob","id":"IIidKatSGn","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["safety alignment","large reasoning model","over refusal"]},"supplementary_material":{"value":"/attachment/28cdf7b1b5c255402cc6319fab3039bdc650b07a.zip"},"primary_area":{"value":"alignment, fairness, safety, privacy, and societal considerations"},"abstract":{"value":"Large Reasoning models (LRMs) with reasoning capabilities have demonstrated remarkable performance on complex tasks, yet achieving robust safety alignment remains a significant challenge. Supervised fine-tuning (SFT) with safety data is a widely-used approach to improve the models' safety, however, we identify that current safety alignment methods with SFT often induce a phenomenon we term $\\textbf{Shortcut Alignment}$. \nIn this case, the model learns to recognize the patterns in harmful inputs and emit templated refusals (e.g., \"I'm sorry...\") while decoupling the final response from its internal chain-of-thought (CoT) reasoning. This superficiality leads to two critical problems: (i) refusals without reasoning carry no informative value, and (ii) models become overly cautious, leading to excessive false refusals on benign queries and thereby degrading their general helpfulness.\nTo understand this behavior, we formalize it through the lens of conditional mutual information (CMI), hypothesizing that when the information gain from CoT is low, such shortcuts become low-resistance solutions that reduce training loss with little cost. We empirically verify this hypothesis via probe experiments that estimate the gap between predictions with and without CoT on harmful versus benign data. \nMotivated by these insights, we propose Deep Instruct Fine-tuning  (DIFT), which uses $\\textbf{CMI-Loss}$, explicitly penalizing shortcut predictions while preserving original instruct-tuning on benign examples. Through theoretical analysis and empirical evidence, we show that our method offers a better solution. It alleviates erroneous refusals while preserving safety. Our work bridges theory and practice, offering the first fine-grained alignment method that explicitly targets shortcut alignment in LRMs."},"_bibtex":{"value":"@misc{\nliu2026beyond,\ntitle={Beyond Refusals: Fine-grained Safety Alignment for Reasoning {LLM}s},\nauthor={Zhendong Liu and Baihui Zheng and Hongqiong Zhong and Boren Zheng and Yingshui Tan and Xiaoyong Zhu and Bo Zheng},\nyear={2026},\nurl={https://openreview.net/forum?id=3qHILWiEob}\n}"},"title":{"value":"Beyond Refusals: Fine-grained Safety Alignment for Reasoning LLMs"},"pdf":{"value":"/pdf/10085969ef4991399b9fa16e5b35be31c94e5e8b.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"liu|beyond_refusals_finegrained_safety_alignment_for_reasoning_llms"},"authorids":{"value":["~Zhendong_Liu6","~Baihui_Zheng2","~Hongqiong_Zhong3","~Boren_Zheng1","~Yingshui_Tan1","~Xiaoyong_Zhu1","~Bo_Zheng5"]},"authors":{"value":["Zhendong Liu","Baihui Zheng","Hongqiong Zhong","Boren Zheng","Yingshui Tan","Xiaoyong Zhu","Bo Zheng"]}},"version":2},{"content":{"summary":{"value":"## Summary\n\nThis paper introduces MLLM-PRUNER, an efficient post-training pruning framework that adapts an **activation-aware** $\\ell_2$-norm metric to multimodal large language models (MLLMs). By leveraging a small calibration dataset to estimate activation importance, the method aims to reduce model size and theoretical FLOPs with minimal performance loss on various VQA benchmarks."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"## Questions\n\n1.  **Q1: Multimodal Justification for the Core Metric.**\n    The activation-aware metric in MLLM-Pruner is directly ported from **unimodal LLMs**. Visual and text tokens have a **vast difference** in activation distribution and feature dimension. How can the authors prove, from a **theoretical or information-theoretic** standpoint, that a simple combination of the $\\ell_2$-norms of these two modalities accurately captures the critical information that influences the **fragility of multimodal fusion** within the **cross-attention layer**? Without this proof, is the application of this metric in MLLMs merely **heuristic** rather than principled?\n\n2.  **Q2: Efficiency and Overhead of the \"Calibration Dataset.\"**\n    The paper claims to be an efficient **post-training pruning** method. However, the overhead of using a **calibration dataset** to perform a full network pass for $\\ell_2$-norm estimation is **not quantified**. How does the time and resource cost (GPU hours) required for this \"calibration\" process compare to the overhead of performing **lightweight fine-tuning** (e.g., via LoRA) with a simple magnitude-based pruning method? If the calibration overhead is higher, how does MLLM-Pruner justify its superiority in **practical efficiency**?\n\n3.  **Q3: Completeness of Performance Evaluation and Missing Multiturn Dialogue.**\n    The robustness of pruned MLLMs on **complex tasks** is crucial. The paper's evaluation is primarily limited to a subset of VQA tasks. Can the authors provide performance data on **Multiturn Conversation** datasets (e.g., VisDial or related benchmarks)? Given that contextual dependency and accumulated visual redundancy in these tasks are highly sensitive to pruning, doesn't the lack of this validation imply that MLLM-Pruner **cannot guarantee** its robustness in practical dialogue systems?\n\n4.  **Q4: Disconnect between Efficiency Claims and Real Latency Data.**\n    The main evidence for efficiency is the **theoretical reduction in FLOPs**. However, for this **Unstructured Pruning**, theoretical FLOPs reduction often does not translate directly to **actual latency speedup**. Can the authors provide quantified data on the **real End-to-End Latency Speedup** during inference on standard hardware (e.g., V100 or A100)? If real latency data is missing, doesn't the **\"efficient\"** claim lack essential hardware-level support?\n\n\n## Suggestions\n\n1.  **G1: Strengthen the Theoretical Foundation of the Multimodal Metric.**\n    Given that MLLM-Pruner's core metric is ported directly from unimodal LLMs, the authors should attempt to **validate** or **refine** the $\\ell_2$-norm combination using **causal or information-theoretic methods** (e.g., inspiration from Transfer Entropy or PID). This can be achieved via an **ablation study** to prove the metric's superior \\textbf{sensitivity to cross-modal interactions} compared to simple weight magnitude or pure $\\ell_2$ activation, thereby grounding its **principled application** in MLLMs.\n\n2.  **G2: Quantify and Mitigate the Real Overhead of the Calibration Process.**\n    To fully substantiate MLLM-Pruner's efficiency claim, the authors **must** quantify the time and resource cost (e.g., GPU hours on an A100) required for the **calibration dataset pass**. Furthermore, the authors should investigate the robustness of using **smaller, domain-agnostic** data subsets for calibration, or explore ways to **batch or approximate** the $\\ell_2$-norm estimation to minimize the impact of calibration on deployment efficiency.\n\n3.  **G3: Adopt Fixed Token Retention Rates for Intuitive Comparison and Update Baselines.**\n    To provide a more intuitive and persuasive comparison of pruning efficacy, it is recommended that the experiments include performance data at **fixed token retention rates** (e.g., 11.1% or 22.2% visual tokens, or 64/128 tokens for LLaVANext), and compare against **advanced token pruning methods** like VisionZip or VSCAN. This presentation style, commonly used in prior token pruning works, would more directly demonstrate MLLM-Pruner's **optimal performance advantage** at specific, relevant pruning rates, and address the issue of using older baselines."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"## Strengths\n\n1.  **S1: Adapting Activation-Aware Pruning for Multimodal LLMs.**\n    The paper successfully adapts an effective activation-aware pruning metric—combining weight magnitude and input activation $\\ell_2$-norm—from the unimodal LLM domain to the complex MLLM structure. This approach represents a principled effort to move beyond simple magnitude pruning by incorporating runtime activation information.\n\n2.  **S2: Post-Training Efficiency and Broad Model Coverage.**\n    MLLM-PRUNER is a post-training method, which avoids costly fine-tuning. The authors demonstrate its effectiveness across a relatively broad spectrum of modern MLLM architectures (like LLaVA and mPLUG-Owl), showcasing its potential as a plug-and-play tool for efficient MLLM deployment."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"## Weaknesses\n\n1.  **W1: Lack of Theoretical Basis and Multimodal Inapplicability of the Core Metric.**\n    The paper's core approach is simply porting the $\\ell_2$-norm activation-aware metric from unimodal LLMs. This strategy *fundamentally ignores two core challenges* of multimodal pruning: a) the vast difference in activation distributions between vision and text tokens; and b) the specific fragility of cross-modal attention post-pruning. The lack of any justification that this simple $\\ell_2$ combination captures the *critical information alignment* specific to multimodal fusion *fundamentally challenges* the metric's validity and effectiveness.\n\n2.  **W2: Calibration Overhead Contradicts Post-Training Efficiency Claim.**\n    MLLM-Pruner relies on a calibration dataset to obtain activation $\\ell_2$-norms. For large MLLMs, even a single calibration pass can incur *significant computational and time cost*.\n    * If the overhead of this \"calibration\" process substantially exceeds the cost of minimal fine-tuning (e.g., via LoRA), the method *loses its appeal* as an \"efficient\" post-training pruning technique.\n    * The paper fails to provide the time and resource cost required for calibration, leaving its efficiency claim *unsubstantiated*.\n\n3.  **W3: Insufficient Performance Evaluation and Missing Multiturn Dialogue.**\n    The paper relies on a limited subset of VQA tasks for evaluation, which *fails to fully reflect* real-world MLLM performance.\n    * *Missing* is the robustness analysis on challenging benchmarks like Multiturn Conversation and complex visual reasoning.\n    * Since these tasks are far more sensitive to pruning-induced errors, the lack of this crucial validation limits confidence in MLLM-Pruner's *generalizability and robustness* in practical applications."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762917291685,"tcdate":1761980667009,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission4314/Reviewer_Ayii"],"signatures":["ICLR.cc/2026/Conference/Submission4314/Reviewer_Ayii"],"forum":"jmQKr47S77","number":3,"license":"CC BY 4.0","cdate":1761980667009,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission4314/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762917291685,"domain":"ICLR.cc/2026/Conference","replyto":"jmQKr47S77","id":"rbjqUDzvPn","forumContent":{"TLDR":{"value":"Efficient Activation-aware  MLLM Pruning."},"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Post-training Pruning","Activation-aware Pruning","MLLM"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Multimodal large language models (MLLMs) have demonstrated impressive performance across a wide range of vision-language tasks. However, the increasing scale of these models leads to significant challenges in deployment costs. Post-training pruning emerges as an effective compression technique to address these challenges. \nRecent pruning studies on large language models (LLMs) has shown that activation-aware pruning strategies that combine weight magnitude with the $\\ell_2$-norm of input activations can achieve superior performance. Nevertheless, directly applying these approaches to MLLMs often leads to substantial performance degradation. This is because the $\\ell_2$-norm assumes all activations contribute equally, while in MLLMs, visual and textual tokens exhibit divergent activation patterns. Moreover, textual-only calibration datasets used in LLM pruning are inadequate for capturing modality-specific dependencies, which further limits their ability to evaluate the importance of weight. In this paper, we propose MLLM-Pruner, a novel activation-aware pruning framework specifically tailored for MLLMs.    To address these issues, MLLM-Pruner introduces two key innovations: (1) we construct a representative multimodal calibration dataset comprising general-domain text, Instruction Tuning, and Visual Instruction Tuning data to comprehensively preserve language generation, instruction-following, and visual reasoning abilities for MLLMs. (2), we design a modality-sensitive importance estimation metric that leverages the Singular Value Decomposition (SVD) of attention distributions to reweight the input activations, effectively captures the activation contribution across modalities, and reduces the pruning error.\n Our MLLM-Pruner does not rely on an expensive iterative reconstruction and re-training process. Extensive experiments on LLaVA-based MLLMs across various benchmarks demonstrate that MLLM-Pruner consistently outperforms state-of-the-art pruning methods while maintaining efficient compression. Our code, model weights, and multimodal calibration dataset will be made publicly available upon publication."},"_bibtex":{"value":"@misc{\nding2025mllmpruner,\ntitle={{MLLM}-Pruner: Efficient Activation-aware Pruning for Multimodal {LLM}s},\nauthor={Yunan Ding and Yan Tai and Siqi Luo and Xiaohong Liu and Guodong Guo and Bo Yang},\nyear={2025},\nurl={https://openreview.net/forum?id=jmQKr47S77}\n}"},"title":{"value":"MLLM-Pruner: Efficient Activation-aware Pruning for Multimodal LLMs"},"pdf":{"value":"/pdf/9c7a49262423e3bf84c0120423ec09c3bda7b92d.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"ding|mllmpruner_efficient_activationaware_pruning_for_multimodal_llms"},"authorids":{"value":["~Yunan_Ding1","~Yan_Tai1","~Siqi_Luo2","~Xiaohong_Liu2","~Guodong_Guo1","~Bo_Yang7"]},"authors":{"value":["Yunan Ding","Yan Tai","Siqi Luo","Xiaohong Liu","Guodong Guo","Bo Yang"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a plug-and-play module for control of camera pose. This module can be integrated to various video diffusion models. This module leverages Plücker embeddings to encode camera movements, offering detailed spatial representation compared to traditional numerical values (e.g., extrinsic matrix)."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"* Why are static scene video datasets mainly used without general video datasets? (For example, one could use general video datasets for training along with estimated camera trajectories from various models like particleSFM.)\n* This didn't affect my rating, but I'm curious about the authors' opinions on whether this method can be generally applied to DiT-based methods that do not separate temporal and spatial aggregation."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":3},"strengths":{"value":"* This is one of the early works that enables adding camera control to existing video models. I believe it will have a positive impact on various downstream tasks.\n* While the way of achieving camera control itself isn't technically very novel, the paper presents an intuitive design and provides extensive experimental comparisons on various conceivable design choices (e.g., where to inject the control and different camera representations).\n* It is interesting that camera control can be achieved with a lightweight additional adapter without further tuning the video model."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"* Since it is fine-tuned on static scene data, it tends to show somewhat reduced ability in generating dynamic motions compared to the original video diffusion baseline (as shown in ODD). But, it shows better performance compared to other methods.\n* In object-centric video generation, it appears that the movement of the object is minimal compared to the background. Ideally, models fine-tuned on Objaverse or MVImgNet should be able to generate camera trajectories covering 360 degrees of the object. I'm curious if this model can effectively generate videos when provided with such camera trajectories as input (e.g., generating a video from a frame showing the front of a person to one showing the back)."}},"nonreaders":[],"tmdate":1731427419035,"tcdate":1730514747810,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission1390/Reviewer_7FvC"],"signatures":["ICLR.cc/2025/Conference/Submission1390/Reviewer_7FvC"],"forum":"Z4evOUYrk7","number":3,"license":"CC BY 4.0","cdate":1730514747810,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission1390/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427419035,"domain":"ICLR.cc/2025/Conference","replyto":"Z4evOUYrk7","id":"lnJGbNCNlz","forumContent":{"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["camera viewpoints control in Video Generation"]},"supplementary_material":{"value":"/attachment/945a433f90edacf1ceefee55868fe3ba2f373ac9.zip"},"primary_area":{"value":"generative models"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Controllability plays a crucial role in video generation, as it allows users to create and edit content more precisely. Existing models, however, lack control of camera pose that serves as a cinematic language to express deeper narrative nuances. To alleviate this issue, we introduce \\method, enabling accurate camera pose control for video diffusion models. Our approach explores effective camera trajectory parameterization along with a plug-and-play camera pose control module that is trained on top of a video diffusion model, leaving other modules of the base model untouched. Moreover, a comprehensive study on the effect of various training datasets is conducted, suggesting that videos with diverse camera distributions and similar appearance to the base model indeed enhance controllability and generalization. Experimental results demonstrate the effectiveness of \\method in achieving precise camera control with different video generation models, marking a step forward in the pursuit of dynamic and customized video storytelling from textual and camera pose inputs."},"_bibtex":{"value":"@inproceedings{\nhe2025cameractrl,\ntitle={CameraCtrl: Enabling Camera Control for Video Diffusion Models},\nauthor={Hao He and Yinghao Xu and Yuwei Guo and Gordon Wetzstein and Bo Dai and Hongsheng Li and Ceyuan Yang},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=Z4evOUYrk7}\n}"},"title":{"value":"CameraCtrl: Enabling Camera Control for Video Diffusion Models"},"pdf":{"value":"/pdf/681b7a0b97770cc7f4f6d567ad538ce599d0a8a6.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"he|cameractrl_enabling_camera_control_for_video_diffusion_models"},"authorids":{"value":["~Hao_He7","~Yinghao_Xu1","~Yuwei_Guo1","~Gordon_Wetzstein3","~Bo_Dai2","~Hongsheng_Li3","~Ceyuan_Yang2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Hao He","Yinghao Xu","Yuwei Guo","Gordon Wetzstein","Bo Dai","Hongsheng Li","Ceyuan Yang"]}},"version":2},{"content":{"summary":{"value":"This paper addresses first-frame-guided video editing with diffusion models, proposing a training-time, mask-aware low-rank adaptation of off-the-shelf image-to-video (I2V) models. By inserting learnable LoRA modules into attention layers and conditioning the denoising network on a user-defined spatio-temporal mask, the model is taught (i) to preserve pixels where the mask is 1 and (ii) to synthesise new content where the mask is 0, either by copying motion from the source clip or by adopting appearance from an extra reference image. At inference, only the edited first frame and the same mask are needed to propagate the edit through the whole sequence. Extensive experiments against recent baselines (I2VEdit, AnyV2V, VACE, Kling1.6) show superior or comparable DEQA, CLIP-score and user-study rankings."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"- Provide a direct ablation that keeps everything else identical except mask conditioning is missing.\n- Provide frame-count vs metric curves on Benchmark to verify scalability.\n- How does performance degrade under automatic segmentation errors\n- Provide some failure case"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"- Two-stage mask scheduling (motion learning → appearance learning) cleanly disentangles dynamics from look\n- Solid empirical protocol: two tasks, two backbones (Wan2.1 & HunyuanVideo), three metrics plus user study, ablations for mask usage and extra reference frames.\n- Equations are minimal and directly show the change of conditioning, easing reproducibility.\n- Delivers a lightweight plug-in that practitioners can apply to any I2V model; potential to become a default extension for personalised video creation tools."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- I2VEdit already finetunes LoRA per video and uses attention masks to reduce background leakage. The paper must quantify the marginal gain of feeding the mask into the network input instead of using it only at feature level. \n- 20 self-collected clips plus an undisclosed I2VEdit test set is small. No long-form videos, fast motion, or multi-person scenes are reported. \n- Perfect masks are assumed. Evaluate with noisy masks (erosion / dilation, IoU = 0.80–0.90) or segmentation-model predictions."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762918436267,"tcdate":1762173836474,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission6060/Reviewer_4aXZ"],"signatures":["ICLR.cc/2026/Conference/Submission6060/Reviewer_4aXZ"],"forum":"xkRMJ1Y7Um","number":3,"license":"CC BY 4.0","cdate":1762173836474,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission6060/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762918436267,"domain":"ICLR.cc/2026/Conference","replyto":"xkRMJ1Y7Um","id":"PrZMl36Dii","forumContent":{"TLDR":{"value":"The paper introduces a mask-based LoRA tuning method for highly flexible video editing using the pre-trained Image-to-Video model."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Video Editing"]},"supplementary_material":{"value":"/attachment/034635334583436f67f6ff70e86f40fbc1976fcf.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Video editing using diffusion models has achieved remarkable results in generating high-quality edits for videos. However, current methods often rely on large-scale pretraining, limiting flexibility for specific edits. First-frame-guided editing provides control over the first frame, but lacks fine-grained control over the edit's subsequent temporal evolution. To address this, we propose a mask-based LoRA (Low-Rank Adaptation) tuning method that adapts pretrained Image-to-Video models for flexible video editing.\nOur key innovation is using a spatiotemporal mask to strategically guide the LoRA fine-tuning process. This teaches the model two distinct skills: first, to interpret the mask as a command to either preserve content from the source video or generate new content in designated regions. Second, for these generated regions, LoRA learns to synthesize either temporally consistent motion inherited from the video or novel appearances guided by user-provided reference frames.\nThis dual-capability LoRA grants users control over the edit's entire temporal evolution, allowing complex transformations like an object rotating or a flower blooming. Experimental results show our method achieves superior video editing performance compared to baseline methods. The code and video results are available at our project website: https://cjeen.github.io/LoRAEdit."},"_bibtex":{"value":"@inproceedings{\ngao2026controllable,\ntitle={Controllable First-Frame-Guided Video Editing via Mask-Aware Lo{RA} Fine-Tuning},\nauthor={Chenjian Gao and Lihe Ding and Xin Cai and Zhanpeng Huang and Zibin Wang and Tianfan Xue},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=xkRMJ1Y7Um}\n}"},"title":{"value":"Controllable First-Frame-Guided Video Editing via Mask-Aware LoRA Fine-Tuning"},"pdf":{"value":"/pdf/a7fc2c198d3de6a267b714b5b442ae140e592d4c.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"gao|controllable_firstframeguided_video_editing_via_maskaware_lora_finetuning"},"authorids":{"value":["~Chenjian_Gao1","~Lihe_Ding1","~Xin_Cai2","~Zhanpeng_Huang1","~Zibin_Wang1","~Tianfan_Xue2"]},"authors":{"value":["Chenjian Gao","Lihe Ding","Xin Cai","Zhanpeng Huang","Zibin Wang","Tianfan Xue"]}},"version":2},{"content":{"summary":{"value":"This paper introduces P²-DPO, a self-correcting framework for reducing hallucination in Large Vision-Language Models. It targets perceptual processing bottlenecks rather than perception failures, generating on-policy, vision-aware preference pairs to improve attention and robustness. With a Calibration Loss and Dynamic Deficit-Weighting, P²-DPO enhances visual grounding and robustness, outperforming human- and AI-feedback DPO baselines without external supervision."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Please refer to the weakness part."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- The paper provides a novel perspective on hallucination in LVLMs by distinguishing between perception failure and perceptual processing failure. It highlights the latter as an overlooked yet solvable issue that can be addressed through self-correction within the model, offering new insight into the root causes of hallucination.\n\n- The proposed on-policy strategy, which generates vision-aware contrastive preference pairs from the model itself, effectively removes the dependence on human or AI feedback used in traditional DPO frameworks."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Limited experimental scope: The experiments are primarily conducted on LLaVA-1.5-7B and Qwen2.5-VL-3B, but it is unclear why the authors did not include the more commonly used Qwen2.5-VL-7B model. This omission limits the completeness of the evaluation and raises questions about the method’s scalability and consistency across model sizes.\n\n- Potential self-reinforcement bias in on-policy data: Although the on-policy preference generation strategy avoids the off-policy issue, self-generated preference pairs in the early training stage may inherit the model’s existing biases or hallucinations.\n\n- Lack of qualitative interpretation of improvements: While quantitative results are comprehensive, the paper lacks intuitive visual evidence explaining how the Calibration Loss and Dynamic Deficit-Weighting improve the model’s attention distribution or semantic grounding. Visualizations such as attention map changes would make the effectiveness of these mechanisms clearer and more convincing."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762920614788,"tcdate":1762363300566,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission8851/Reviewer_m9BD"],"signatures":["ICLR.cc/2026/Conference/Submission8851/Reviewer_m9BD"],"forum":"ekOwxTn65Y","number":3,"license":"CC BY 4.0","cdate":1762363300566,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission8851/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762920614788,"domain":"ICLR.cc/2026/Conference","replyto":"ekOwxTn65Y","id":"uWwlrmllGA","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["LLMs ; MLLMs ; Hallucination ; DPO"]},"primary_area":{"value":"foundation or frontier models, including LLMs"},"abstract":{"value":"Hallucination has recently garnered significant research attention in Large Vision-Language Models (LVLMs). Direct Preference Optimization (DPO) aims to learn directly from the corrected preferences provided by humans, thereby addressing the hallucination issue. Despite its success, this paradigm has yet to specifically target the perceptual bottleneck in attended regions or address insufficient Visual Robustness against image degradation. Furthermore, existing preference pairs are often vision-agnostic and their inherently off-policy nature limits their effectiveness in guiding model learning. To address these challenges, we propose Perceptual Processing Direct Preference Optimization (P$^2$-DPO), a novel training paradigm in which the model generates and learns from its own preference pairs, thereby directly addressing the identified visual bottlenecks while inherently avoiding the issues of vision-agnostic and off-policy data. It introduces: (1) an on-policy preference pairs construction method targeting Focus-and-Enhance perception and Visual Robustness, and (2) a well-designed Calibration Loss to precisely align visual signals with the causal generation of text. Experimental results demonstrate that with a comparable amount of training data and cost, P$^2$-DPO outperforms strong baselines that rely on costly human feedback on benchmarks. Furthermore, evaluations on Attention Region Fidelity (ARF) and image degradation scenarios validate the effectiveness of P$^2$-DPO in addressing perceptual bottleneck in attended regions and improving Visual Robustness against degraded inputs."},"_bibtex":{"value":"@inproceedings{\nzhang2026pdpo,\ntitle={P\\${\\textasciicircum}2\\$-{DPO}: Grounding Hallucination in Perceptual Processing via Calibration Direct Preference Optimization},\nauthor={ruipeng zhang and Zhihao Li and Haozhang Yuan and C.L.Philip Chen and Tong Zhang},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=ekOwxTn65Y}\n}"},"title":{"value":"P$^2$-DPO: Grounding Hallucination in Perceptual Processing via Calibration Direct Preference Optimization"},"pdf":{"value":"/pdf/14ba92ba4d75afc9ea61f6dca935bd5078b584ef.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"zhang|p^2dpo_grounding_hallucination_in_perceptual_processing_via_calibration_direct_preference_optimization"},"authorids":{"value":["~Ruipeng_Zhang4","~Zhihao_Li21","~Haozhang_Yuan1","~C.L.Philip_Chen1","~Tong_Zhang14"]},"authors":{"value":["Ruipeng Zhang","Zhihao Li","Haozhang Yuan","C.L.Philip Chen","Tong Zhang"]}},"version":2},{"content":{"venue":{"value":"IEEE Trans. Circuits Syst. Video Technol. 2024"},"pdf":{"value":"https://ieeexplore.ieee.org/iel7/76/10422872/10168142.pdf"},"venueid":{"value":"dblp.org/journals/TCSV/2024"},"paperhash":{"value":"menon|jndaware_twopass_pertitle_encoding_scheme_for_adaptive_live_streaming"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Vignesh_V._Menon:","https://dblp.org/search/pid/api?q=author:Prajit_T._Rajendran:","https://dblp.org/search/pid/api?q=author:Christian_Feldmann:","https://dblp.org/search/pid/api?q=author:Klaus_Schoeffmann:","https://dblp.org/search/pid/api?q=author:Mohammed_Ghanbari_0001:","~Christian_Timmerer1"]},"html":{"value":"https://doi.org/10.1109/TCSVT.2023.3290725"},"_bibtex":{"value":"@article{DBLP:journals/tcsv/MenonRFSGT24,\n  author={Vignesh V. Menon and Prajit T. Rajendran and Christian Feldmann and Klaus Schoeffmann and Mohammed Ghanbari and Christian Timmerer},\n  title={JND-Aware Two-Pass Per-Title Encoding Scheme for Adaptive Live Streaming},\n  year={2024},\n  month={February},\n  cdate={1706745600000},\n  journal={IEEE Trans. Circuits Syst. Video Technol.},\n  volume={34},\n  number={2},\n  pages={1281-1294},\n  url={https://doi.org/10.1109/TCSVT.2023.3290725}\n}\n"},"abstract":{"value":"Adaptive live video streaming applications utilize a predefined collection of bitrate-resolution pairs, known as a bitrate ladder, for simplicity and efficiency, eliminating the need for additional run-time to determine the optimal pairs during the live streaming session. These applications do not incorporate two-pass encoding methods due to increased latency. However, an optimized bitrate ladder could result in lower storage and delivery costs and improved Quality of Experience (QoE). This paper presents a Just Noticeable Difference (JND)-aware constrained Variable Bitrate (cVBR) Two-pass Per-title encoding Scheme (JTPS) designed specifically for live video streaming. JTPS predicts a content- and JND-aware bitrate ladder using low-complexity features based on Discrete Cosine Transform (DCT) energy and optimizes the constant rate factor (CRF) for each representation using random forest-based models. The effectiveness of JTPS is demonstrated using the open source video encoder x265, with an average bitrate reduction of 18.80% and 32.59% for the same PSNR and VMAF, respectively, compared to the standard HTTP Live Streaming (HLS) bitrate ladder using Constant Bitrate (CBR) encoding. The implementation of JTPS also resulted in a 68.96% reduction in storage space and an 18.58% reduction in encoding time for a JND of six VMAF points."},"title":{"value":"JND-Aware Two-Pass Per-Title Encoding Scheme for Adaptive Live Streaming"},"authors":{"value":["Vignesh V. Menon","Prajit T. Rajendran","Christian Feldmann","Klaus Schoeffmann","Mohammed Ghanbari","Christian Timmerer"]}},"tmdate":1741250331198,"pdate":1704067200000,"tcdate":1741250322155,"writers":["~"],"signatures":["~Christian_Timmerer1"],"forum":"DlX4Mw5X6z","license":"CC BY-SA 4.0","number":359755,"cdate":1706745600000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1741250331198,"domain":"DBLP.org","id":"DlX4Mw5X6z","version":2},{"content":{"venue":{"value":"EMNLP 2023 Main"},"TLDR":{"value":"Proposal of practical and challenging multimodal video summarization task."},"keywords":{"value":["Video Summarization","Multimodality"]},"abstract":{"value":"This paper proposes a practical multimodal video summarization task setting and a dataset to train and evaluate the task. The target task involves summarizing a given video into a predefined number of keyframe-caption pairs and displaying them in a listable format to grasp the video content quickly. This task aims to extract crucial scenes from the video in the form of images (keyframes) and generate corresponding captions explaining each keyframe's situation. This task is useful as a practical application and presents a highly challenging problem worthy of study. Specifically, achieving simultaneous optimization of the keyframe selection performance and caption quality necessitates careful consideration of the mutual dependence on both preceding and subsequent keyframes and captions. To facilitate subsequent research in this field, we also construct a dataset by expanding upon existing datasets and propose an evaluation framework.\nFurthermore, we develop two baseline systems and report their respective performance."},"_bibtex":{"value":"@inproceedings{\nkudo2023a,\ntitle={A Challenging Multimodal Video Summary: Simultaneously Extracting and Generating Keyframe-Caption Pairs from Video},\nauthor={Keito Kudo and Haruki Nagasawa and Jun Suzuki and Nobuyuki Shimizu},\nbooktitle={The 2023 Conference on Empirical Methods in Natural Language Processing},\nyear={2023},\nurl={https://openreview.net/forum?id=YvzA0hFCF3}\n}"},"title":{"value":"A Challenging Multimodal Video Summary: Simultaneously Extracting and Generating Keyframe-Caption Pairs from Video"},"Submission_Type":{"value":"Regular Long Paper"},"pdf":{"value":"/attachment/be51c4fea332e2ffffbbfe9fd261ac00d4b30d77.pdf"},"Submission_Track":{"value":"Speech and Multimodality"},"venueid":{"value":"EMNLP/2023/Conference"},"paperhash":{"value":"kudo|a_challenging_multimodal_video_summary_simultaneously_extracting_and_generating_keyframecaption_pairs_from_video"},"authorids":{"value":["~Keito_Kudo1","~Haruki_Nagasawa1","~Jun_Suzuki1","~Nobuyuki_Shimizu1"]},"authors":{"value":["Keito Kudo","Haruki Nagasawa","Jun Suzuki","Nobuyuki Shimizu"]}},"tmdate":1701459740970,"pdate":1696709262381,"tcdate":1686887473533,"writers":["EMNLP/2023/Conference","EMNLP/2023/Conference/Submission2443/Authors"],"signatures":["EMNLP/2023/Conference/Submission2443/Authors"],"forum":"YvzA0hFCF3","number":2443,"cdate":1686887473533,"mdate":1701459740970,"readers":["everyone"],"invitations":["EMNLP/2023/Conference/-/Submission","EMNLP/2023/Conference/-/Post_Submission","EMNLP/2023/Conference/Submission2443/-/Revision","EMNLP/2023/Conference/-/Edit","EMNLP/2023/Conference/Submission2443/-/Camera_Ready_Revision"],"odate":1701459740956,"domain":"EMNLP/2023/Conference","id":"YvzA0hFCF3","version":2},{"content":{"summary":{"value":"The paper introduces BLADE, a framework that integrates Adaptive Block-Sparse Attention (ASA) with step distillation for efficient video generation. It proposes a data-free joint training approach, leveraging ASA to generate dynamic, content-aware sparsity masks and sparsity-aware Trajectory Distribution Matching (TDM) to enhance quality. Experiments on CogVideoX-5B and Wan2.1-1.3B demonstrate significant speedups (up to 14.10×) and quality improvements, validated by VBench-2.0 and human evaluations."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Please refer to the **Weaknesses** above."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"- Integration of adaptive block-sparse attention with step distillation, enabling data-free joint training for efficient video generation.\n- ASA mechanism dynamically generates content-aware sparsity masks that enable high sparsity levels, achieving hardware-friendly acceleration without quality loss when combined with distillation training.\n- Demonstrates substantial speedups (up to 14.10×) on diverse models like CogVideoX-5B and Wan2.1-1.3B, with consistent quality improvements on VBench."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"This paper lacks details on experimental settings and comparative results, for example:\n- Lack of reporting on specific GPU hours, training batch size, and memory usage for the 100-200 distillation iterations.\n- Lack of inference results demonstrating video quality across low-to-high sparsity levels to illustrate the impact."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762921564146,"tcdate":1761797917905,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission10202/Reviewer_asWM"],"signatures":["ICLR.cc/2026/Conference/Submission10202/Reviewer_asWM"],"forum":"O9J20MsmRl","number":4,"license":"CC BY 4.0","cdate":1761797917905,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission10202/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762921564146,"domain":"ICLR.cc/2026/Conference","replyto":"O9J20MsmRl","id":"XWDrRKmJda","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["sparse attention; video generation; step distillation"]},"supplementary_material":{"value":"/attachment/ac5e9cebf5419d07332d99362d65e157ee2320bc.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Diffusion transformers currently lead the field in high-quality video generation, but their slow iterative denoising process and prohibitive quadratic attention costs for long sequences create significant inference bottlenecks. While both step distillation and sparse attention mechanisms have shown promise as independent acceleration strategies, effectively combining these approaches presents critical challenges---training-free integration yields suboptimal results, while separately training sparse attention after step distillation requires prohibitively expensive high-quality video data. To overcome these limitations, we propose $\\textit{BLADE}$, an innovative data-free joint training framework that introduces: (1) an Adaptive Block-Sparse Attention (ASA) mechanism for dynamically generating content-aware sparsity masks to focus computation on salient spatiotemporal features, and (2) a sparsity-aware step distillation paradigm, built upon Trajectory Distribution Matching (TDM), directly incorporates sparsity into the distillation process rather than treating it as a separate compression step and features fast convergence. We validate BLADE on text-to-video models like CogVideoX-5B and Wan2.1-1.3B, and our framework demonstrates remarkable efficiency gains across different scales. On Wan2.1-1.3B, BLADE achieves a 14.10$\\times$ end-to-end inference acceleration over a 50-step baseline. Moreover, on models such as CogVideoX-5B with short video sequence lengths, our framework delivers a robust 8.89$\\times$ speedup. Crucially, the acceleration is accompanied by a consistent quality improvement. On the VBench-2.0 benchmark, BLADE boosts the score of CogVideoX-5B to 0.569 (from 0.534) and Wan2.1-1.3B to 0.570 (from 0.563), results that are further corroborated by superior ratings in human evaluations."},"_bibtex":{"value":"@inproceedings{\ngu2026blade,\ntitle={{BLADE}: Block-Sparse Attention Meets Step Distillation for Efficient Video Generation},\nauthor={Youping Gu and XIAOLONG LI and Yuhao Hu and Chen Minqi and Bohan Zhuang},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=O9J20MsmRl}\n}"},"title":{"value":"BLADE: Block-Sparse Attention Meets Step Distillation for Efficient Video Generation"},"pdf":{"value":"/pdf/b689331c6a64100dc1434dd52af9f3b12384a9b9.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"gu|blade_blocksparse_attention_meets_step_distillation_for_efficient_video_generation"},"authorids":{"value":["~Youping_Gu2","~XIAOLONG_LI15","~Yuhao_Hu1","~Chen_Minqi1","~Bohan_Zhuang1"]},"authors":{"value":["Youping Gu","XIAOLONG LI","Yuhao Hu","Chen Minqi","Bohan Zhuang"]}},"version":2},{"content":{"summary":{"value":"This paper proposed AVHBench, a benchmark to assess cross-modal hallucinations in audio-visual LLMs. The benchmark is built on existing datasets VALOR and AudioCaps, and contains about 6k QnA pairs and 1k audio-visual captions across four cross-modal tasks, including audio-driven video hallucination, video-driven audio hallucination, audio-visual matching and audio-visual captioning. The authors design a semi-automatic pipeline for data anntation. Several open-source audio-visual LLMs are evaluated on AVHBench, and the results show that most existing audio-visual LLMs suffer from cross-modal hallucinations. To alleviate this problem, the authors further enhances Video-LLaMA through audio feature alignment and LoRA fine-tuning, proving that audio-visual hallucinations might come from insufficient training on paired audio-visual data."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"- After LoRA fine-tuning, the results of the model on AVHBench become significantly better. Does this suggest that the model may just not be able to do the judgement questions, or that it hasn't seen the mismatched audio/video and therefore performs poorly on that test set, rather than the model having a large number of hallucinations?\n\n- Do the authors consider providing Gemini's results on the benchmark?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"- This paper proposed a audio-visual hallucination benchmark mainly focusing on cross-modal hallucination evaluation, which is an aspect that has received little attention.\n\n- Some tasks in the benchmark provide new perspectives on the study of multi-modal hallucinations, like *Audio-driven Video Hallucination* and *Video-driven Audio Hallucination*. The inability of the model to distinguish between information from audio or video may be a vital reason that causes multi-modal hallucination.\n\n- The paper is well-written, clear and easy to understand."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The benchmark uses existing datasets VALOR and AudioCaps, which may introduce biases inherent to those datasets, potentially affecting the generalizability and validity of the evaluation.\n\n- In this benchmark, for human speech, the authors seem to have only considered the event of \"one person is speaking,\" instead of taking into account the content of what is being said. The content of speech contains a wealth of information and is very likely to contribute to the hallucinations of audio-visual LLMs. However, the paper seems to overlook this scenario.\n    \n    For example, consider a scenario where a person in a video is saying to himself \"Yesterday I heard a dog barking\", but neither the dog nor the barking sound appears in the video or audio. Models are prone to hallucinations in this scenario. \n\n- It is recommended to evaluate more recent audio-visual LLMs, like Video-LLaMA 2, or video-SALMONN.\n\n- Since the video is so rich in content, long text is required to descirbe the video completely. However, for the \"audio-visual captioning\" task, only a short caption is provided as the groundtruth. This suggests that the groundtruth caption is likely to contain only the important information in the video and omit the secondary information. However, it is still possible for the model to describe something that is present in the video but is not hallucination. That's why I'm concerned about the correctness of the \"audio-visual captioning\" task of AVHBench."}},"nonreaders":[],"tmdate":1731428261592,"tcdate":1729496047321,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission8517/Reviewer_huna"],"signatures":["ICLR.cc/2025/Conference/Submission8517/Reviewer_huna"],"forum":"jTEKTdI3K9","number":1,"license":"CC BY 4.0","cdate":1729496047321,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission8517/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428261592,"domain":"ICLR.cc/2025/Conference","replyto":"jTEKTdI3K9","id":"WL1nhmB7uy","forumContent":{"TLDR":{"value":"We introduce AVHBench, comprehensive audio-visual hallucination benchmark specifically designed to evaluate the perception and comprehension capabilities of audio-visual LLMs."},"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Multi-modal Large Language Models","Hallucination in Large Language Models","Audio-visual Learning"]},"supplementary_material":{"value":"/attachment/06ba16fb92bf8ab7afc677de74177f8ed80372f0.zip"},"primary_area":{"value":"datasets and benchmarks"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Following the success of Large Language Models (LLMs), expanding their boundaries to new modalities represents a significant paradigm shift in multimodal understanding. Human perception is inherently multimodal, relying not only on text but also on auditory and visual cues for a complete understanding of the world. In recognition of this fact, audio-visual LLMs have recently emerged. Despite promising developments, the lack of dedicated benchmarks poses challenges for understanding and evaluating models. In this work, we show that audio-visual LLMs struggle to discern subtle relationships between audio and visual signals, leading to hallucinations and highlighting the need for reliable benchmarks. To address this, we introduce AVHBench, the first comprehensive benchmark specifically designed to evaluate the perception and comprehension capabilities of audio-visual LLMs. Our benchmark includes tests for assessing hallucinations, \nas well as the cross-modal matching and reasoning abilities of these models. Our results reveal that most existing audio-visual LLMs struggle with hallucinations caused by cross-interactions between modalities, due to their limited capacity to perceive complex multimodal signals and their relationships. Additionally, we demonstrate that simple training with our AVHBench improves robustness of audio-visual LLMs against hallucinations. Dataset: https://github.com/kaist-ami/AVHBench"},"_bibtex":{"value":"@inproceedings{\nsung-bin2025avhbench,\ntitle={{AVHB}ench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models},\nauthor={Kim Sung-Bin and Oh Hyun-Bin and JungMok Lee and Arda Senocak and Joon Son Chung and Tae-Hyun Oh},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=jTEKTdI3K9}\n}"},"title":{"value":"AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models"},"pdf":{"value":"/pdf/6ce4fd6cce092a4d54c6cb38398c934e7eee5952.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"sungbin|avhbench_a_crossmodal_hallucination_benchmark_for_audiovisual_large_language_models"},"authorids":{"value":["~Kim_Sung-Bin1","~Oh_Hyun-Bin1","~JungMok_Lee1","~Arda_Senocak1","~Joon_Son_Chung1","~Tae-Hyun_Oh3"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Kim Sung-Bin","Oh Hyun-Bin","JungMok Lee","Arda Senocak","Joon Son Chung","Tae-Hyun Oh"]}},"version":2},{"content":{"venue":{"value":"CoRR 2025"},"pdf":{"value":"https://arxiv.org/pdf/2510.24954v2"},"venueid":{"value":"dblp.org/journals/CORR/2025"},"paperhash":{"value":"bernstein|reviving_thorups_shortcut_conjecture"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Aaron_Bernstein:","https://dblp.org/search/pid/api?q=author:Henry_L._Fleischmann:","~Maximilian_Probst_Gutenberg1","https://dblp.org/search/pid/api?q=author:Bernhard_Haeupler:","https://dblp.org/search/pid/api?q=author:Gary_Hoppenworth:","https://dblp.org/search/pid/api?q=author:Yonggang_Jiang:","https://dblp.org/search/pid/api?q=author:George_Z._Li:","https://dblp.org/search/pid/api?q=author:Seth_Pettie:","https://dblp.org/search/pid/api?q=author:Thatchaphol_Saranurak:","https://dblp.org/search/pid/api?q=author:Leon_Schiller:"]},"html":{"value":"https://doi.org/10.48550/arXiv.2510.24954"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2510-24954,\n  publtype={informal},\n  author={Aaron Bernstein and Henry L. Fleischmann and Maximilian Probst Gutenberg and Bernhard Haeupler and Gary Hoppenworth and Yonggang Jiang and George Z. Li and Seth Pettie and Thatchaphol Saranurak and Leon Schiller},\n  title={Reviving Thorup's Shortcut Conjecture},\n  year={2025},\n  month={October},\n  cdate={1759276800000},\n  journal={CoRR},\n  volume={abs/2510.24954},\n  url={https://doi.org/10.48550/arXiv.2510.24954}\n}\n"},"abstract":{"value":"We aim to revive Thorup's conjecture [Thorup, WG'92] on the existence of reachability shortcuts with ideal size-diameter tradeoffs. Thorup originally asked whether, given any graph $G=(V,E)$ with $m$ edges, we can add $m^{1+o(1)}$ ``shortcut'' edges $E_+$ from the transitive closure $E^*$ of $G$ so that $\\text{dist}_{G_+}(u,v) \\leq m^{o(1)}$ for all $(u,v)\\in E^*$, where $G_+=(V,E\\cup E_+)$. The conjecture was refuted by Hesse [Hesse, SODA'03], followed by significant efforts in the last few years to optimize the lower bounds. In this paper we observe that although Hesse refuted the letter of Thorup's conjecture, his work~[Hesse, SODA'03] -- and all followup work -- does not refute the spirit of the conjecture, which should allow $G_+$ to contain both new (shortcut) edges and new Steiner vertices. Our results are as follows. (1) On the positive side, we present explicit attacks that break all known shortcut lower bounds when Steiner vertices are allowed. (2) On the negative side, we rule out ideal $m^{1+o(1)}$-size, $m^{o(1)}$-diameter shortcuts whose ``thickness'' is $t=o(\\log n/\\log \\log n)$, meaning no path can contain $t$ consecutive Steiner vertices. (3) We propose a candidate hard instance as the next step toward resolving the revised version of Thorup's conjecture. Finally, we show promising implications. Almost-optimal parallel algorithms for computing a generalization of the shortcut that approximately preserves distances or flows imply almost-optimal parallel algorithms with $m^{o(1)}$ depth for exact shortcut paths and exact maximum flow. The state-of-the-art algorithms have much worse depth of $n^{1/2+o(1)}$ [Rozhoň, Haeupler, Martinsson, STOC'23] and $m^{1+o(1)}$ [Chen, Kyng, Liu, FOCS'22], respectively."},"title":{"value":"Reviving Thorup's Shortcut Conjecture"},"authors":{"value":["Aaron Bernstein","Henry L. Fleischmann","Maximilian Probst Gutenberg","Bernhard Haeupler","Gary Hoppenworth","Yonggang Jiang","George Z. Li","Seth Pettie","Thatchaphol Saranurak","Leon Schiller"]}},"tmdate":1769507064084,"pdate":1767139200000,"externalIds":["dblp:journals/corr/abs-2510-24954"],"tcdate":1769507029498,"writers":["~"],"signatures":["~Maximilian_Probst_Gutenberg1"],"forum":"24zhBIKTQZ","license":"CC BY-SA 4.0","number":807198,"cdate":1759276800000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1769507064084,"domain":"DBLP.org","id":"24zhBIKTQZ","version":2},{"content":{"summary":{"value":"The paper introduces Temporal Preference Optimization (TPO), a post-training framework that enhances temporal reasoning and grounding capabilities in large video-language models (video-LMMs) without requiring manual annotations. TPO generates contrastive preference pairs by comparing model responses to original versus temporally corrupted video clips, then refines models through Direct Preference Optimization (DPO). Experiments on LongVideoBench, MLVU, and Video-MME demonstrate consistent performance gains across multiple backbones."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Typo in FIgure 3-(a); MLUV → MLVU"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. **Clarity of methods.** The contrastive setup between relevant and manipulated frames, combined with LLM-based post-filtering, is simple yet effective.\n2. **Strong results.** TPO demonstrates consistent improvements across diverse benchmarks and models (LongVA-TPO and LLaVA-Video-TPO), outperforming baselines."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. **Limited technical contribution.** The idea is simply generating positive/negative captions pairs and optimizing with DPO, which seems to be limited in their novelty or technical contribution.\n2. **Comparison with recent RL-based approaches.** Direct comparisons with recent reinforcement or segmentation-based optimization methods (e.g., Time-R1, Grounded-VideoLLM) are missing."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762921892083,"tcdate":1761894780611,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission10638/Reviewer_aGyH"],"signatures":["ICLR.cc/2026/Conference/Submission10638/Reviewer_aGyH"],"forum":"vZLZyNxeOa","number":4,"license":"CC BY 4.0","cdate":1761894780611,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission10638/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762921892083,"domain":"ICLR.cc/2026/Conference","replyto":"vZLZyNxeOa","id":"lShOWLaSLA","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Video Understanding","Vision-Language Model","Preference Learning","Post-Training"]},"supplementary_material":{"value":"/attachment/2b50b9c583c3ba1e6ae199b123b50293246f93ca.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Despite recent advancements in video large multimodal models (video-LMMs), accurate temporal grounding remains a key challenge. In this work, we introduce Temporal Preference Optimization (TPO)—a post-training framework that unlocks superior temporal reasoning in video-LMMs without requiring human annotations. TPO enables preference modeling by manipulating video inputs to generate contrastive responses, ensuring that preferred responses are more temporally grounded than dis-preferred ones. Through preference learning, TPO enhances the model’s capability for more comprehensive video understanding with better temporal reasoning. Extensive experiments on LongVideoBench, MLVU, and Video-MME demonstrate that TPO significantly improves temporal grounding across multiple video-LMMs.   Notably, LLaVA-Video-TPO achieves state-of-the-art performance among 7B models on Video-MME, establishing TPO as a scalable and effective solution for advancing temporal understanding in video analysis."},"_bibtex":{"value":"@misc{\nli2025temporal,\ntitle={Temporal Preference Optimization of Large Multimodal Models},\nauthor={Rui Li and Xiaohan Wang and Yuhui Zhang and Orr Zohar and Zeyu Wang and Serena Yeung-Levy},\nyear={2025},\nurl={https://openreview.net/forum?id=vZLZyNxeOa}\n}"},"title":{"value":"Temporal Preference Optimization of Large Multimodal Models"},"pdf":{"value":"/pdf/b782ce9c318b44c952f15ad9541830af11a97f24.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"li|temporal_preference_optimization_of_large_multimodal_models"},"authorids":{"value":["~Rui_Li26","~Xiaohan_Wang2","~Yuhui_Zhang3","~Orr_Zohar1","~Zeyu_Wang1","~Serena_Yeung-Levy1"]},"authors":{"value":["Rui Li","Xiaohan Wang","Yuhui Zhang","Orr Zohar","Zeyu Wang","Serena Yeung-Levy"]}},"version":2},{"content":{"summary":{"value":"The current study investigates the effect of hue variance on video recognition and proposes a data augmentation method called Motion Coherent Augmentation (MCA) that introduces appearance variation to encourage the model to prioritize motion patterns over static appearances. This approach is based on the observation that static appearances are less important in videos with motion information. The proposed method uses an operation called SwapMix to modify video samples' appearances efficiently and Variation Alignment (VA) to resolve distribution shifts caused by SwapMix, enforcing the model to learn appearance-invariant representations. Comprehensive experiments on different architectures and datasets demonstrate the effectiveness and generalization ability of MCA, achieving an average performance gain of 1.95% on the Something-Something V1 dataset compared to the competing method Uniformer."},"presentation":{"value":"3 good"},"contribution":{"value":"3 good"},"soundness":{"value":"3 good"},"strengths":{"value":"The proposed Motion Coherent Augmentation (MCA) method addresses the issue of overfitting in video recognition by introducing an effective data augmentation strategy. The key advantages of MCA are:\n\n- Hue Jittering: MCA leverages Hue Jittering, which is usually overlooked in object recognition, to generate new appearances for videos. This operation helps the model prioritize the motion pattern over static appearance.\n\n- SwapMix: An efficient operation called SwapMix is introduced to modify the appearance of video samples in the RGB space. This operation helps simulate the effect of appearance variation while preserving important color attributes like saturation and lightness.\n\n- Variation Alignment (VA): VA is used to construct training pairs with different appearances and encourage their predictions to align with each other. This helps the model learn appearance-invariant representations.\n\n- Compatibility: MCA can be seamlessly integrated into existing video recognition approaches with minimal code modifications and provides consistent improvement. This indicates that the method is compatible with existing data augmentation techniques.\n\nExperimental validation: Comprehensive experiments on various architectures and benchmarks demonstrate the effectiveness and generalization ability of MCA. The method demonstrates excellent performance and can further enhance the results of competing approaches like Uniformer.\n\nIn summary, the Motion Coherent Augmentation method addresses the issue of overfitting in video recognition by providing an effective data augmentation strategy that leverages Hue Jittering, SwapMix, and Variation Alignment. MCA achieves consistent improvement while being compatible with existing techniques, making it a promising approach for video recognition."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"It is a valid point that the paper lacks a discussion on the generalization capabilities of MCA and how it could potentially improve the transferability of models. Introducing such a discussion could indeed help increase the impact of the work. A possible approach could be to analyze the performance of MCA on tasks that are different from the benchmark datasets used during training, and compare it to other data augmentation methods. Additionally, analyzing the learned representations using techniques like clustering or visualization could provide insights into their generalization potential. Overall, a thorough evaluation of the generalization capabilities of MCA could strengthen the paper and make it more impactful.\n\nThe relationship between MCA and existing video self-supervised methods is not further explored in the given text. It would be valuable to investigate how MCA can be integrated with other video self-supervised methods and examine the potential synergies between them. Understanding how MCA complements or enhances existing techniques could provide valuable insights into the effectiveness and generalization abilities of the combined approach. Further research is needed to explore the relationship between MCA and video self-supervised methods and determine how they can be effectively integrated to improve video recognition tasks."},"confidence":{"value":"3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked."},"questions":{"value":"Whether the MCA data augmentation method provides gains in video domain adaptation or video domain generalization settings, where videos from different domains are involved, is not explicitly discussed in the given context. To assess whether MCA is beneficial in such video migration scenarios, an evaluation of its performance in tasks that involve domain adaptation or generalization would be necessary. Comprehensive experiments comparing MCA to other data augmentation techniques and assessing its effect on domain adaptation or generalization metrics would provide insights into its potential usefulness in video migration settings.\n\n\nThe potential of integrating MCA with existing video self-supervised methods to further enhance the performance is not explicitly discussed in the given text. However, it is worth exploring the compatibility of MCA with other video self-supervised methods and investigating whether their combination can lead to improved results. By combining the strengths of MCA in introducing appearance variation and prioritizing motion patterns with other self-supervised techniques, it is possible to enhance the overall effectiveness of video representation learning. Further research and experimentation are needed to explore the potential synergies and benefits of integrating MCA with existing video self-supervised methods."},"rating":{"value":"6: marginally above the acceptance threshold"},"details_of_ethics_concerns":{"value":"N/A"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636409347,"tcdate":1698804316383,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission4368/Reviewer_SoaL"],"signatures":["ICLR.cc/2024/Conference/Submission4368/Reviewer_SoaL"],"forum":"RIcYTbpO38","number":3,"license":"CC BY 4.0","cdate":1698804316383,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission4368/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636409347,"domain":"ICLR.cc/2024/Conference","replyto":"RIcYTbpO38","id":"3VY1EItKjT","forumContent":{"TLDR":{"value":"We present a motion coherent augmentation for video understanding which could enforce the model to learn appearance invariant representations."},"venue":{"value":"ICLR 2024 poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video Understanding","Data Augmentation"]},"primary_area":{"value":"representation learning for computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Current training pipelines in object recognition neglect Hue Jittering when doing data augmentation as it not only brings appearance changes that are detrimental to classification, but also the implementation is inefficient in practice. In this study, we investigate the effect of hue variance in the context of video understanding and find this variance to be beneficial since static appearances are less important in videos that contain motion information. Based on this observation, we propose a data augmentation method for video understanding, named Motion Coherent Augmentation (MCA), that introduces appearance variation in videos and implicitly encourages the model to prioritize motion patterns, rather than static appearances. Concretely, we propose an operation SwapMix to efficiently modify the appearance of video samples, and introduce Variation Alignment (VA) to resolve the distribution shift caused by SwapMix, enforcing the model to learn appearance invariant representations. Comprehensive empirical evaluation across various architectures and different datasets solidly validates the effectiveness and generalization ability of MCA, and the application of VA in other augmentation methods. Code is available at https://github.com/BeSpontaneous/MCA-pytorch."},"_bibtex":{"value":"@inproceedings{\nzhang2024dont,\ntitle={Don't Judge by the Look: Towards Motion Coherent Video Representation},\nauthor={Yitian Zhang and Yue Bai and Huan Wang and Yizhou Wang and Yun Fu},\nbooktitle={The Twelfth International Conference on Learning Representations},\nyear={2024},\nurl={https://openreview.net/forum?id=RIcYTbpO38}\n}"},"title":{"value":"Don't Judge by the Look: Towards Motion Coherent Video Representation"},"pdf":{"value":"/pdf/df94612d9bb46799cd06d6a004a8040d124c48ae.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference"},"paperhash":{"value":"zhang|dont_judge_by_the_look_towards_motion_coherent_video_representation"},"authorids":{"value":["~Yitian_Zhang1","~Yue_Bai1","~Huan_Wang3","~Yizhou_Wang3","~Yun_Fu1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Yitian Zhang","Yue Bai","Huan Wang","Yizhou Wang","Yun Fu"]}},"version":2},{"content":{"summary":{"value":"This paper addresses the limitations of existing Multimodal Large Language Models (MLLMs) in video spatial reasoning (inferring 3D spatial structures from video frames), which arise from the lack of high-quality task-specific datasets and effective training strategies. Inspired by Reinforcement Learning with Verifiable Reward (RLVR), the authors propose the SpaceR framework, consisting of two core innovations: SpaceR-151k Dataset: A large-scale dataset with 91k spatial reasoning QA pairs (SR-91k, derived from the 3D ScanNet dataset, covering 6 tasks like relative direction and object size) and 60k general multimodal samples (from Video-R1-260k) to preserve general video understanding.\nSpatially-Guided RLVR (SG-RLVR): An RL approach extending Group Relative Policy Optimization (GRPO) with a map imagination mechanism and task-specific verifiable rewards (e.g., format, multi-choice, numerical, map rewards).\nExtensive experiments show SpaceR achieves state-of-the-art performance on spatial reasoning benchmarks (VSI-Bench, STI-Bench, SPAR-Bench)—surpassing GPT-4o by 11.6% on VSI-Bench—and competitive results on general video understanding benchmarks (Video-MME, TempCompass, LongVideoBench). It also matches the proprietary Gemini-2.0-Flash, validating the framework’s effectiveness."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Please refer to the weaknesses part."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"Targeted Dataset Solution: SpaceR-151k fills a critical gap in video spatial reasoning resources by providing 91k high-quality, verifiable QA pairs across diverse spatial tasks (e.g., relative distance, room size), addressing the field’s data scarcity.\n\nInnovative RL Enhancement: The map imagination mechanism explicitly guides models to infer spatial layouts, a novel design that strengthens deep spatial reasoning (ablation studies confirm it boosts performance on spatial benchmarks).\n\nStrong Generalization: SG-RLVR outperforms supervised fine-tuning (SFT) across both spatial reasoning and general video tasks, avoiding SFT’s \"memorization\" limitation and demonstrating broader applicability.\n\nComprehensive Validation: Experiments cover 3 spatial and 3 general video benchmarks, with comparisons to state-of-the-art models (GPT-4o, Gemini series, open-source MLLMs like Qwen2.5-VL), ensuring robust evaluation of effectiveness.\n\nScalability: SG-RLVR works across model sizes (3B, 7B parameters) and architectures (dense VLMs like Qwen2.5-VL, MoE models like Kimi-VL-Thinking), showing flexibility for different MLLM frameworks."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Limited Failure Analysis: While the paper identifies two failure modes (visual perception errors, location misidentification), it lacks deeper investigation into root causes (e.g., why object recognition fails for occluded items) or actionable solutions to address them.\n\nReasoning Efficiency Tradeoffs: The \"think mode\" (structured reasoning) improves spatial performance but introduces noise in general video tasks (slight accuracy drops). The paper does not propose adaptive mechanisms to decide when reasoning is necessary, limiting practicality.\n\nDataset Dependency on ScanNet: SR-91k relies heavily on the ScanNet 3D dataset (indoor scenes), which may restrict the framework’s generalization to non-indoor scenarios (e.g., outdoor, dynamic environments) not covered by ScanNet.\n\nAll performance metrics are automatic (accuracy, ROUGE scores), with no human assessment of reasoning interpretability (e.g., whether cognitive maps align with human spatial intuition), which is critical for trust in real-world use."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762923436662,"tcdate":1760517297539,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission12588/Reviewer_sGvR"],"signatures":["ICLR.cc/2026/Conference/Submission12588/Reviewer_sGvR"],"forum":"zLcTDEqDmS","number":1,"license":"CC BY 4.0","cdate":1760517297539,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission12588/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762923436662,"domain":"ICLR.cc/2026/Conference","replyto":"zLcTDEqDmS","id":"svks7W6ScF","forumContent":{"TLDR":{"value":"We introduce a high-quality dataset SpaceR-151k and propose SpaceR trained via our SG-RLVR framework towards video spatial reasoning."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Reinforcement Learning","Video Spatial Reasoning","Multimodal Large Language Models"]},"supplementary_material":{"value":"/attachment/47ddfa9a04603df8065b94f201ad1e026f8eb9fa.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Video spatial reasoning, which involves inferring the underlying spatial structure from observed video frames, poses a significant challenge for existing Multimodal Large Language Models (MLLMs). This limitation stems primarily from 1) the absence of high-quality datasets for this task, and 2) the lack of effective training strategies to develop spatial reasoning capabilities. Motivated by the success of Reinforcement Learning with Verifiable Reward (RLVR) in unlocking LLM reasoning abilities, this work aims to improve MLLMs in video spatial reasoning through the RLVR paradigm. To this end, we introduce the **SpaceR** framework. First, we present **SpaceR-151k**, a dataset with 91k questions spanning diverse spatial reasoning scenarios with verifiable answers, and 60k samples for maintaining general multimodal understanding. Second, we propose **Spatially-Guided RLVR (SG-RLVR)**, a novel reinforcement learning approach that extends Group Relative Policy Optimization (GRPO) with a novel map imagination mechanism, which encourages the model to infer spatial layouts in the thinking process, thereby facilitating more effective spatial reasoning.\nExtensive experiments demonstrate that SpaceR achieves state-of-the-art performance on spatial reasoning benchmarks (e.g., VSI-Bench, STI-Bench, and SPAR-Bench), while showing competitive results on video understanding benchmarks (e.g., Video-MME, TempCompass, and LongVideoBench). \nRemarkably, SpaceR surpasses the advanced GPT-4o by 11.6\\% accuracy on VSI-Bench and is on par with the leading proprietary model Gemini-2.0-Flash, highlighting the effectiveness of our SpaceR-151k dataset and SG-RLVR in reinforcing spatial reasoning ability of MLLMs."},"_bibtex":{"value":"@misc{\nouyang2026spacer,\ntitle={SpaceR: Reinforcing {MLLM}s in Video Spatial Reasoning},\nauthor={Kun Ouyang and Yuanxin Liu and Haoning Wu and Yi Liu and Hao Zhou and Fandong Meng and Jie Zhou and Xu Sun},\nyear={2026},\nurl={https://openreview.net/forum?id=zLcTDEqDmS}\n}"},"title":{"value":"SpaceR: Reinforcing MLLMs in Video Spatial Reasoning"},"pdf":{"value":"/pdf/e8b1487a340821956e9c877f43953a182ea60f3c.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"ouyang|spacer_reinforcing_mllms_in_video_spatial_reasoning"},"authorids":{"value":["~Kun_Ouyang2","~Yuanxin_Liu1","~Haoning_Wu1","~Yi_Liu18","~Hao_Zhou8","~Fandong_Meng3","~Jie_Zhou8","~Xu_Sun1"]},"authors":{"value":["Kun Ouyang","Yuanxin Liu","Haoning Wu","Yi Liu","Hao Zhou","Fandong Meng","Jie Zhou","Xu Sun"]}},"version":2},{"content":{"venue":{"value":"IEEE Trans. Circuits Syst. Video Technol. 2009"},"pdf":{"value":"https://ieeexplore.ieee.org/iel5/76/5344167/05159442.pdf"},"venueid":{"value":"dblp.org/journals/TCSV/2009"},"paperhash":{"value":"liu|multigraphbased_queryindependent_learning_for_video_search"},"authorids":{"value":["","","","~Xian-Sheng_Hua1"]},"html":{"value":"https://doi.org/10.1109/TCSVT.2009.2026951"},"_bibtex":{"value":"@article{DBLP:journals/tcsv/LiuMWH09,\n  author={Yuan Liu and Tao Mei and Xiuqing Wu and Xian-Sheng Hua},\n  title={Multigraph-Based Query-Independent Learning for Video Search},\n  year={2009},\n  cdate={1230768000000},\n  journal={IEEE Trans. Circuits Syst. Video Technol.},\n  volume={19},\n  number={12},\n  pages={1841-1850},\n  url={https://doi.org/10.1109/TCSVT.2009.2026951}\n}\n"},"abstract":{"value":"Most of the existing learning-based methods for video search take query examples as ¿positive¿ and build a model for each query. These methods, referred to as query-dependent, only achieve limited success as users are mostly reluctant to provide enough query examples. To address this problem, we propose a novel query-independent learning approach based on multigraph to video search, which learns the relevance information existing in the query-shot pairs. The proposed approach, named MG-QIL, is more general and suitable for a real-world video search system as the learned relevance is independent of any queries. Specifically, MG-QIL constructs multiple graphs, including a main-graph covering all the pairs and a set of subgraphs covering the pairs within the same query. The pairs in the main-graph are connected in terms of relational similarity, while the pairs in the subgraphs for the same query are connected in terms of attributional similarity. The relevance labels are then propagated in the multiple graphs until convergence. We conducted extensive experiments on automatic search tasks over the TRECVID 2005-2007 benchmark and the results show a superior performance to state-of-the-art approaches to video search. Furthermore, when applied to video search reranking, MG-QIL can also achieve significant and consistent improvement over a text search baseline."},"title":{"value":"Multigraph-Based Query-Independent Learning for Video Search"},"authors":{"value":["Yuan Liu","Tao Mei","Xiuqing Wu","Xian-Sheng Hua"]}},"tmdate":1772784978393,"pdate":1262217600000,"externalIds":["dblp:journals/tcsv/LiuMWH09"],"tcdate":1772784926072,"writers":["~"],"signatures":["~Xian-Sheng_Hua1"],"forum":"GPPExmEUEZ","license":"CC BY-SA 4.0","number":849149,"cdate":1230768000000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1772784978393,"domain":"DBLP.org","id":"GPPExmEUEZ","version":2},{"content":{"summary":{"value":"This paper presents DraftAttention, a training-free acceleration framework for video diffusion transformers that introduces dynamic sparse attention guided by a low-resolution draft attention map. The approach computes attention at a downsampled resolution to identify spatially and temporally redundant regions, and then applies a block sparse mask to full-resolution attention computations. The method achieves up to 2× speedup in video synthesis with minimal loss in generation quality, outperforming existing sparse attention baselines according to the reported results."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"- How does DraftAttention differ fundamentally from existing pooling-based block-sparse acceleration techniques (e.g., MInference)?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- The presentation is clear and well organized, with thorough motivation, method description, and experimental validation.\n\n- The computational bottlenecks in attention of video diffusion transformers are a pressing issue, and the proposed approach targets a key challenge for scaling video generation systems."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The core idea—using pooling-based  approximations to guide sparse block attention—is conceptually similar to prior work (e.g., MInference for LLMs and SpargeAttention for video diffusion). This overlap weakens the novelty claim.\n\n- The experimental evaluation omits direct comparisons with several relevant sparse or spatially adaptive attention methods such as SpargeAttention, SlidingTileAttention, RadialAttention, and XAttention, which are necessary to situate the method’s performance within the broader literature."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762923123992,"tcdate":1761943731141,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission12174/Reviewer_CeWm"],"signatures":["ICLR.cc/2026/Conference/Submission12174/Reviewer_CeWm"],"forum":"jUNmW3s45i","number":4,"license":"CC BY 4.0","cdate":1761943731141,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission12174/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762923123992,"domain":"ICLR.cc/2026/Conference","replyto":"jUNmW3s45i","id":"Yv8sj9nyKq","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Video Generation","Efficient Video Generation","Sparse Attention"]},"supplementary_material":{"value":"/attachment/84c3b89e822ef44f572ffd8586c4576ba6ccf454.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Video generation models based on diffusion transformers have recently attracted widespread attention for their excellent generation quality.\nDespite recent progress, their computational expense remains the principal bottleneck. In particular, attention alone accounts for more than 80\\% of the overall latency, and the synthesis of only 8 seconds 720p video takes tens of minutes, which severely restricts practical applicability and scalability.\nTo address this, we propose **DraftAttention**, a training-free framework for the acceleration of video diffusion transformers with dynamic sparse attention on GPUs.\nThe key idea is to compute the low-resolution draft attention based on the downsampled low-resolution query and key with minor computational overhead. The draft attention exposes redundancy both spatially within each feature map and temporally across frames, thus identifying the most important areas in the attention map.\nThe resulting low-resolution sparse mask then guides full-resolution sparse attention computations.\nTo align region-level sparsity with token-level computations, we further propose a deterministic reordering of tokens such that entries in each region become contiguous in memory, ensuring hardware-friendly execution of sparse attention.\nOur theoretical analysis demonstrates that the low-resolution draft attention closely approximates the full attention, providing reliable guidance for constructing accurate sparse attention.\nExperimental results show that our method outperforms existing sparse attention approaches in video generation quality and achieves up to 2x end-to-end speedup on GPUs."},"_bibtex":{"value":"@misc{\nshen2025draftattention,\ntitle={DraftAttention: Fast Video Diffusion via Low-Resolution Attention Guidance},\nauthor={Xuan Shen and Chenxia Han and Yufa Zhou and Yanyue Xie and Yifan Gong and Quanyi Wang and Yiwei Wang and Yanzhi Wang and Pu Zhao and Jiuxiang Gu},\nyear={2025},\nurl={https://openreview.net/forum?id=jUNmW3s45i}\n}"},"title":{"value":"DraftAttention: Fast Video Diffusion via Low-Resolution Attention Guidance"},"pdf":{"value":"/pdf/89651ed6f0637a5752cf7b2022469825e9a54ebf.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"shen|draftattention_fast_video_diffusion_via_lowresolution_attention_guidance"},"authorids":{"value":["~Xuan_Shen1","~Chenxia_Han1","~Yufa_Zhou1","~Yanyue_Xie1","~Yifan_Gong2","~Quanyi_Wang1","~Yiwei_Wang2","~Yanzhi_Wang3","~Pu_Zhao1","~Jiuxiang_Gu2"]},"authors":{"value":["Xuan Shen","Chenxia Han","Yufa Zhou","Yanyue Xie","Yifan Gong","Quanyi Wang","Yiwei Wang","Yanzhi Wang","Pu Zhao","Jiuxiang Gu"]}},"version":2},{"content":{"summary":{"value":"The paper explores the use of Denoising Diffusion Models (DDM) for event-based video reconstruction task. The most important contribution, in my opinion, is the improvement in video quality. The researchers adapted an existing model (ModelScope Text-to-Video) for the  event-based video reconstruction task. They introduced ESA, that modiffies a conditional inpainting technique (presented on High-Resolution Image Synthesis with Latent Diffusion Models). Additionally, they used text conditioning to further enhance video quality.\n\nThe innovative contributions of the paper are two. The fist one is the Event-aware Noise Initialization; This technique enables the use of frame I_(t-1) to reconstruct frame I_(t) during the inference stage (autoregressive).\nThe second one is the Event-aware Mask Loss; This new loss function is designed to improve the temporal consistensy.\n\nIn summary, the contributions are:\n\n1. An existing text-to-video model (diffusion-based) was adapted for the event-based recosntruction task.\n2. The model was adapted to event data using a conditional inpainting technique. Also proposed Event-aware Noise Initialization and Mask Loss to improve video quality.\n3. The result is a new state-of-the-art in event-to-video reconstruction.\n\nHowever, the proposed model has several drawbacks. The DDM \"hallucinates\" content, especially in areas with low or no event data, which can be problematic for applications like object detection in the context of self-driving cars."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1) Although traditional metrics (LPIPS, MSE, and SSIM) show an increase in video reconstruction quality, the qualitative results, especially in Figure 3, reveal that the reconstructed frames do not reflect reality in specific areas. These artifacts represent a smaller percentage of the image (in terms of pixels) and are not captured by traditional metrics, how does this affect the overall video quality?\n\n2) In the inference stage, no mention is made of the length (number of frames) that the model can reconstruct. If there is a limit on the number of frames (how many frames and) , how is temporal consistency between reconstructed sections resolved?\n\n3) Despite discussing the model's temporal consistency, no metrics verify this.\n\n4) On line 134, the variable V is introduced, representing the conversion of event data to voxels. However, it is unclear whether V represents a segment with N temporal bins or N segments (the assumption is that V represents N segments). Clarification of this detail is necessary.\n\n5) On line 139, the input latent representation is introduced through the variable \\hat{\\epsilon} . In DDM, \\epsilon variable represents noise, and the noisy input latent representation normaly is denoted with z_t in this case as \\hat{z}^{i}_{t}, since it is the latent image after applying the noise and the representation of the events z^{i}_{e}. I'm not sure if it's correct. \n\n6) The Event-guided Spatio-temporal Attention (ESA) module is very similar to the technique used in SD (High-Resolution Image Synthesis with Latent Diffusion Models) for inpainting. A reference to this would be beneficial.\n\n7) In section 4.2, on line 162, the title \"Event Spatial Attention\" mentions a cross-attention technique. However, the title does not reflect this, causing confusion. It could be changed to \"Event Spatial Cross-Attention\". Similarly, on line 176, the title \"Event Temporal Attention\" maybe should be changed to \"Event Temporal Cross-Attention.\"\n\n7) In equation 8, line 199, the value of \\(\\lambda\\) is not mentioned.\n\n8) No mention is made of the computational cost in the inference stage, nor is the inference time mentioned. Could these data be mentioned?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"As mentioned before, the main contribution of this paper is the improvement in video quality. Also, this paper proves that it is possible to use diffusion-based models for the event-based video reconstruction task."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Diffusion-based models tend to \"hallucinate,\" producing some parts on a frames that are far from reality (in the event-based video reconstruction task), which can be problematic for applications like object detection in the context of safety (self-driving cars). \n\nAnother problem is that the proposed model uses text prompts for video generation. Although these text prompts allow some control over the content, they also introduce a lot of ambiguity. This leads to a trial-and-error process (prompt engineering) until the most realistic reconstruction is achieved according to the user.\n\nAdditionally, the proposed model requires very high computational resources. Inference times are not specified in the paper, but it is known that the model does not run in real-time."},"limitations":{"value":"Due to the DDM's tendency to hallucinate, the model cannot be used in self-driving cars or other computer vision applications where safety is involved.\n\nIt is believed that the model cannot be run in real time, much less on an embedded device."}},"nonreaders":[],"tmdate":1730879331534,"tcdate":1720617650888,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission9695/Reviewer_H8fU"],"signatures":["NeurIPS.cc/2024/Conference/Submission9695/Reviewer_H8fU"],"forum":"3ilqQHBWTf","number":2,"license":"CC BY 4.0","cdate":1720617650888,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission9695/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730879331534,"domain":"NeurIPS.cc/2024/Conference","replyto":"3ilqQHBWTf","id":"DyHt72GW9l","forumContent":{"venue":{"value":"NeurIPS 2024 poster"},"keywords":{"value":["event camera","video reconstruction","diffusion model"]},"supplementary_material":{"value":"/attachment/18d7d7e7074c2d330d75515eeba54e12b1b57dda.zip"},"primary_area":{"value":"machine_vision"},"abstract":{"value":"Event cameras harness advantages such as low latency, high temporal resolution, and high dynamic range (HDR), compared to standard cameras. Due to the distinct imaging paradigm shift, a dominant line of research focuses on event-to-video (E2V) reconstruction to bridge event-based and standard computer vision. However, this task remains challenging due to its inherently ill-posed nature: event cameras only detect the edge and motion information locally. Consequently, the reconstructed videos are often plagued by artifacts and regional blur, primarily caused by the ambiguous semantics of event data. In this paper, we find language naturally conveys abundant semantic information, rendering it stunningly superior in ensuring semantic consistency for E2V reconstruction. Accordingly, we propose a novel framework, called LaSe-E2V,  that can achieve semantic-aware high-quality E2V reconstruction from a language-guided perspective, buttressed by the text-conditional diffusion models. However, due to diffusion models' inherent diversity and randomness, it is hardly possible to directly apply them to achieve spatial and temporal consistency for E2V reconstruction. Thus, we first propose an Event-guided Spatiotemporal Attention (ESA) module to condition the event data to the denoising pipeline effectively. We then introduce an event-aware mask loss to ensure temporal coherence and a noise initialization strategy to enhance spatial consistency. Given the absence of event-text-video paired data, we aggregate existing E2V datasets and generate textual descriptions using the tagging models for training and evaluation. Extensive experiments on three datasets covering diverse challenging scenarios (e.g., fast motion, low light) demonstrate the superiority of our method. Demo videos for the results are attached to the project page."},"_bibtex":{"value":"@inproceedings{\nchen2024laseev,\ntitle={LaSe-E2V: Towards Language-guided Semantic-aware Event-to-Video Reconstruction},\nauthor={Kanghao Chen and Hangyu Li and Jiazhou Zhou and Zeyu Wang and Lin Wang},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=3ilqQHBWTf}\n}"},"title":{"value":"LaSe-E2V: Towards Language-guided Semantic-aware Event-to-Video Reconstruction"},"pdf":{"value":"/pdf/7793a11fd5cecb06c47363429b62f4053e561f68.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"chen|lasee2v_towards_languageguided_semanticaware_eventtovideo_reconstruction"},"authorids":{"value":["~Kanghao_Chen1","~Hangyu_Li5","~Jiazhou_Zhou1","~Zeyu_Wang15","~Lin_Wang2"]},"authors":{"value":["Kanghao Chen","Hangyu Li","Jiazhou Zhou","Zeyu Wang","Lin Wang"]}},"version":2},{"content":{"summary":{"value":"This paper presents Virtual-Eyes, a deterministic, lung-aware 16-bit CT quality-control (QC) pipeline for low-dose CT (LDCT) lung cancer screening, and provides a systematic quantitative validation of its impact across both generalist foundation models and specialist architectures. Using 765 NLST patients, the authors evaluate how Virtual-Eyes affects RAD-DINO, Merlin, Sybil, and a ResNet-18 baseline under a leakage-free evaluation protocol. The key finding is that anatomically targeted QC significantly improves discrimination, calibration, and stability for a generalist radiology foundation model (RAD-DINO), while degrading performance for specialist models that appear to rely on contextual shortcuts in raw clinical data. The work is valuable in demonstrating that preprocessing is not a neutral design choice, and that lung-aware QC can function both as an enabler for foundation models and as a diagnostic stress test for shortcut dependence."},"justification_of_final_rating":{"value":"Justification:\n\nThe authors have thoroughly addressed the reviewer’s concerns and provided clear responses that enhance the manuscript’s rigor and applicability. Their work continues to make a valuable contribution to the medical imaging field by highlighting the significant impact of preprocessing choices on foundation-model pipelines.\n\nHyperparameter Sensitivity: The authors have provided a clear response regarding the sensitivity of the Virtual-Eyes hyperparameters. They explain that the thresholds were selected through systematic sweeps, and preliminary analysis indicates that modest variations do not significantly affect RAD-DINO performance. The planned inclusion of a sensitivity analysis with sweep ranges and AUC results will further strengthen this aspect and provide additional clarity for practitioners.\n\nSybil Degradation: The authors correctly acknowledge that the observed performance degradation in Sybil is likely due to the model’s implicit adaptation to raw NLST volumes and scanner-dependent features. By enforcing strict lung geometry, Virtual-Eyes exposes this domain dependence. The authors expect that retraining or fine-tuning on Virtual-Eyes–processed data would mitigate this degradation, and they have appropriately discussed this in the revised manuscript. This insight is valuable for understanding how specialist models can adapt to preprocessing pipelines.\n\nStatistical Uncertainty: The authors have agreed to add bootstrap confidence intervals for the Brier and KS statistics, which will enhance the statistical robustness of the reported results and further address the reviewer’s request for more detailed uncertainty quantification.\n\nRejected Series Clarification: The authors have clarified the reason for the rejection of series, stating that it is driven primarily by scan length and field-of-view rather than cancer prevalence or acquisition parameters. This additional explanation strengthens the interpretation of the results and helps contextualize the findings more accurately.\n\nOverall, the authors have effectively addressed all of the reviewer’s concerns, adding valuable statistical details, clarifying the impact of hyperparameters and preprocessing on specialist models, and enhancing the robustness of the paper. The manuscript now provides a comprehensive and transparent analysis of how preprocessing choices can affect the performance of foundation models in medical imaging, which is a critical contribution to the field. Therefore, the paper is well-suited for acceptance."},"confidence":{"value":4},"final_rating":{"value":5},"justification_of_the_preliminary_rating":{"value":"This paper makes a clear and timely contribution by rigorously demonstrating that preprocessing is a first-order design choice in foundation-model pipelines rather than a neutral implementation detail. The experimental evidence is convincing, the methodology is transparent, and the insights regarding differential effects on generalist versus specialist models are highly relevant to the MIDL community."},"confidentiality_llm_acknowledgment":{"value":"Yes"},"strengths":{"value":"The primary strength of this paper lies in its clear isolation and quantification of preprocessing effects, an aspect often under-specified or implicitly assumed in medical imaging pipelines. The Virtual-Eyes pipeline is fully deterministic, transparent, and computationally lightweight, avoiding the opacity and failure modes of learned segmentation-based approaches while remaining grounded in established HU-based lung-processing practices.\n\nThe experimental design is rigorous: patient-level splits prevent leakage, evaluation is performed on a fixed held-out test set, and multiple complementary metrics (AUC, DeLong tests, KS statistics, Brier scores, and agreement analyses) are used to characterize both discrimination and calibration. The comparative analysis across generalist and specialist models is particularly insightful, as it reveals that anatomically focused QC can meaningfully stabilize and improve foundation models, while simultaneously exposing shortcut or context-dependent learning in specialist architectures.\n\nFinally, the paper is well written, carefully structured, and places its contributions in context with prior foundation-model and LDCT screening literature. The planned public release of code and reproducibility details further enhance its potential value to the community."},"weaknesses":{"value":"All foundation models are evaluated in a frozen-encoder regime. While this is appropriate for isolating preprocessing effects, it leaves open how Virtual-Eyes interacts with fine-tuning, which is common in practical deployment and may mitigate some of the observed degradation in specialist models.\n\nFinally, although the paper interprets performance drops in Sybil and ResNet-18 as evidence of shortcut reliance, this conclusion remains indirect. Complementary analyses (e.g., attribution maps or controlled context re-injection experiments) could further substantiate the causal link between removed contextual cues and degraded performance."},"detailed_comments":{"value":"Reporting uncertainty or confidence intervals for Brier scores and KS statistics would further strengthen the statistical presentation.\n\nIt may be helpful to explicitly clarify whether rejected series differ systematically in cancer prevalence or acquisition parameters."},"questions_to_address_in_the_rebuttal":{"value":"How sensitive are the reported gains to the specific Virtual-Eyes hyperparameters (e.g., lung-area ratio, minimum block length), and do modest changes materially affect RAD-DINO performance?\n\nDo the authors expect the observed degradation in Sybil to persist after retraining or fine-tuning on Virtual-Eyes–processed data?"},"preliminary_rating":{"value":5}},"parentInvitations":"MIDL.io/2026/Validation_Papers/-/Official_Review","nonreaders":[],"tmdate":1771087811960,"tcdate":1767432121209,"writers":["MIDL.io/2026/Validation_Papers","MIDL.io/2026/Validation_Papers/Submission38/Reviewer_pk2e"],"signatures":["MIDL.io/2026/Validation_Papers/Submission38/Reviewer_pk2e"],"forum":"8c4103IiyO","number":1,"license":"CC BY 4.0","cdate":1767432121209,"readers":["everyone"],"invitations":["MIDL.io/2026/Validation_Papers/Submission38/-/Official_Review","MIDL.io/2026/Validation_Papers/-/Edit"],"mdate":1771087811960,"domain":"MIDL.io/2026/Validation_Papers","replyto":"8c4103IiyO","id":"mVkCjLYrHT","forumContent":{"venue":{"value":"MIDL 2026 - Validation Papers Poster"},"TLDR":{"value":"A lung-aware 16-bit CT preprocessing pipeline that significantly stabilizes and improves foundation-model performance for LDCT cancer risk prediction while revealing shortcut dependence in specialist models."},"keywords":{"value":["Lung Cancer Screening","Foundation Models","Quality Control","CT Preprocessing","Validation"]},"reproducibility":{"value":"Code is publicly available at: https://github.com/Enamul-Hoq/virtual-eyes-ldct-qc-validation.  NLST data are accessible through The Cancer Imaging Archive (TCIA). The repository provides instructions to reproduce preprocessing, dataset splits, and evaluation.  TCIA integration of Virtual-Eyes is under development and will be released soon."},"read_cfp_and_author_instructions":{"value":"Yes"},"originality_policy":{"value":"Yes"},"abstract":{"value":"Robust preprocessing is rarely quantified in deep-learning pipelines for low-dose CT (LDCT) lung cancer screening. We develop and validate \\emph{Virtual-Eyes}, a clinically motivated, 16-bit CT quality-control pipeline for NLST, and measure its differential impact on generalist foundation models (FMs) versus specialist models. Virtual-Eyes enforces strict $512\\times512$ in-plane resolution, rejects short or non-diagnostic series, and extracts a contiguous lung block using Hounsfield-unit filtering and bilateral lung-coverage scoring, while preserving the original 16-bit DICOM grid. Using 765 NLST patients (182 cancer, 583 non-cancer), we compute slice-level embeddings from RAD-DINO and Merlin with frozen encoders and train leakage-free patient-level MLP heads. We also apply Virtual-Eyes to Sybil and a 2D ResNet-18 baseline without retraining their backbones. For RAD-DINO, preprocessing improves slice-level AUC from 0.576 to 0.610 and patient-level AUC from 0.646 to 0.683 (mean pooling) and 0.619 to 0.735 (max pooling). These gains are accompanied by reduced distributional drift between raw and preprocessed outputs (KS $D=0.041$, $p<10^{-80}$) and better calibration (Brier score $0.188 \\to 0.112$). In contrast, Sybil and ResNet-18 degrade under Virtual-Eyes (Sybil AUC $0.886 \\to 0.837$, ResNet-18 $0.571 \\to 0.596$) and show evidence of shortcut or context-dependent learning, with Sybil becoming overconfident (Brier $0.092 \\to 0.145$). Merlin exhibits limited transferability to thoracic risk prediction (AUC $\\approx 0.507$--$0.567$) regardless of preprocessing. To our knowledge, this is the first quantitative validation of lung-aware preprocessing for LDCT foundation-model workflows. Our results highlight that anatomically targeted QC can meaningfully stabilize and improve generalist FMs, but may disrupt specialist models that have adapted to raw clinical context."},"_bibtex":{"value":"@inproceedings{\nhoq2026virtualeyes,\ntitle={Virtual-Eyes: Quantitative Validation of a Lung {CT} Quality-Control Pipeline for Foundation-Model Cancer Risk Prediction},\nauthor={Md. Enamul Hoq and Linda Larson-Prior and Fred Prior},\nbooktitle={Medical Imaging with Deep Learning - Validation Papers},\nyear={2026},\nurl={https://openreview.net/forum?id=8c4103IiyO}\n}"},"title":{"value":"Virtual-Eyes: Quantitative Validation of a Lung CT Quality-Control Pipeline for Foundation-Model Cancer Risk Prediction"},"latex_code":{"value":"/attachment/3098639b02930849ab4d7c8ab4a377acdecbeef4.zip"},"secondary_subject_area":{"value":"Transfer Learning and Domain Adaptation"},"pdf":{"value":"/pdf/713ff63fad763d21350359cfd18a3fc2cb0934dd.pdf"},"copyright_form":{"value":"/attachment/4d5bd090a602ad53055039e2c832a6c2aac448ce.pdf"},"visa":{"value":"No"},"single_blind_notice":{"value":"Yes"},"venueid":{"value":"MIDL.io/2026/Validation_Papers"},"paperhash":{"value":"hoq|virtualeyes_quantitative_validation_of_a_lung_ct_qualitycontrol_pipeline_for_foundationmodel_cancer_risk_prediction"},"primary_subject_area":{"value":"Foundation Models"},"authorids":{"value":["~Md._Enamul_Hoq1","~Linda_Larson-Prior1","fwprior@uams.edu"]},"registration":{"value":"Yes"},"authors":{"value":["Md. Enamul Hoq","Linda Larson-Prior","Fred Prior"]},"llm_policy_acknowledgment":{"value":"Yes"}},"version":2},{"content":{"venue":{"value":"IEEE Trans. Circuits Syst. Video Technol. 2022"},"pdf":{"value":"https://ieeexplore.ieee.org/iel7/76/9973795/09829839.pdf"},"venueid":{"value":"dblp.org/journals/TCSV/2022"},"paperhash":{"value":"peng|lves2d_lowlight_video_enhancement_from_static_to_dynamic"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Bo_Peng_0007:","~Xuanyu_Zhang2","https://dblp.org/search/pid/api?q=author:Jianjun_Lei:","https://dblp.org/search/pid/api?q=author:Zhe_Zhang:","https://dblp.org/search/pid/api?q=author:Nam_Ling:","~Qingming_Huang1"]},"html":{"value":"https://doi.org/10.1109/TCSVT.2022.3190916"},"_bibtex":{"value":"@article{DBLP:journals/tcsv/PengZLZLH22,\n  author={Bo Peng and Xuanyu Zhang and Jianjun Lei and Zhe Zhang and Nam Ling and Qingming Huang},\n  title={LVE-S2D: Low-Light Video Enhancement From Static to Dynamic},\n  year={2022},\n  cdate={1640995200000},\n  journal={IEEE Trans. Circuits Syst. Video Technol.},\n  volume={32},\n  number={12},\n  pages={8342-8352},\n  url={https://doi.org/10.1109/TCSVT.2022.3190916}\n}\n"},"abstract":{"value":"Recently, deep-learning-based low-light video enhancement methods have drawn wide attention and achieved remarkable performance. However, limited by the difficulty in collecting dynamic low-light and well-lighted video pairs in real scenes, how to construct video sequences for supervised learning and design a low-light enhancement network for real dynamic video remains a challenge. In this paper, we propose a simple yet effective low-light video enhancement method (LVE-S2D), which generates dynamic video training pairs from static videos, and enhances the low-light video by mining dynamic temporal information. To obtain low-light and well-lighted video pairs, a sliding window-based dynamic video generation mechanism is designed to produce pseudo videos with rich dynamic temporal information. Then, a siamese dynamic low-light video enhancement network is presented, which effectively utilizes temporal correlation between adjacent frames to enhance the video frames. Extensive experimental results demonstrate that the proposed method not only achieves superior performance on static low-light videos, but also outperforms the state-of-the-art methods on real dynamic low-light videos."},"title":{"value":"LVE-S2D: Low-Light Video Enhancement From Static to Dynamic"},"authors":{"value":["Bo Peng","Xuanyu Zhang","Jianjun Lei","Zhe Zhang","Nam Ling","Qingming Huang"]}},"tmdate":1731503883643,"pdate":1640995200000,"tcdate":1727593757997,"writers":["~"],"signatures":["~Xuanyu_Zhang2"],"forum":"BgARFRyU4C","license":"CC BY-SA 4.0","number":105125,"cdate":1640995200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1731503883643,"domain":"DBLP.org","id":"BgARFRyU4C","version":2},{"content":{"venue":{"value":"Signal Image Video Process. 2022"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/s11760-022-02174-7.pdf"},"venueid":{"value":"dblp.org/journals/SIVP/2022"},"paperhash":{"value":"xu|motionaware_future_frame_prediction_for_video_anomaly_detection_based_on_saliency_perception"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Haitao_Xu:","https://dblp.org/search/pid/api?q=author:Weibin_Liu:","https://dblp.org/search/pid/api?q=author:Weiwei_Xing:","~Xiang_Wei1"]},"html":{"value":"https://doi.org/10.1007/s11760-022-02174-7"},"_bibtex":{"value":"@article{DBLP:journals/sivp/XuLXW22,\n  author={Haitao Xu and Weibin Liu and Weiwei Xing and Xiang Wei},\n  title={Motion-aware future frame prediction for video anomaly detection based on saliency perception},\n  year={2022},\n  cdate={1640995200000},\n  journal={Signal Image Video Process.},\n  volume={16},\n  number={8},\n  pages={2121-2129},\n  url={https://doi.org/10.1007/s11760-022-02174-7}\n}\n"},"abstract":{"value":"The anomaly in videos can be considered as a deviation from regular video sequences. Most existing approaches neglect the imbalanced information distribution between the foreground and the background during the process of reconstruction or prediction. To address this problem, we propose a motion-aware future frame prediction network consisting of a frame prediction branch and a saliency perception branch. In particular, the saliency perception branch is designed to predict the most salient targets in the video frame, and the frame prediction branch is used to predict the RGB future frame with the guidance of saliency perception. Besides, a motion-aware attention module is bridged in the frame prediction branch to improve the representation ability of moving targets. Furthermore, a saliency prediction loss and a saliency-guided appearance loss are designed to optimize saliency prediction frames and constrain the weight of foreground. Experiments on three challenging benchmarks demonstrate our competitive performance with the state-of-the-art approaches."},"title":{"value":"Motion-aware future frame prediction for video anomaly detection based on saliency perception"},"authors":{"value":["Haitao Xu","Weibin Liu","Weiwei Xing","Xiang Wei"]}},"tmdate":1731480898235,"pdate":1640995200000,"tcdate":1731476668200,"writers":["~"],"signatures":["~Xiang_Wei1"],"forum":"H9WAWWdhjU","license":"CC BY-SA 4.0","number":202744,"cdate":1640995200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1731480898235,"domain":"DBLP.org","id":"H9WAWWdhjU","version":2},{"content":{"venue":{"value":"Signal Image Video Process. 2022"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/s11760-022-02174-7.pdf"},"venueid":{"value":"dblp.org/journals/SIVP/2022"},"paperhash":{"value":"xu|motionaware_future_frame_prediction_for_video_anomaly_detection_based_on_saliency_perception"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Haitao_Xu:","https://dblp.org/search/pid/api?q=author:Weibin_Liu:","https://dblp.org/search/pid/api?q=author:Weiwei_Xing:","~Xiang_Wei1"]},"html":{"value":"https://doi.org/10.1007/s11760-022-02174-7"},"_bibtex":{"value":"@article{DBLP:journals/sivp/XuLXW22,\n  author={Haitao Xu and Weibin Liu and Weiwei Xing and Xiang Wei},\n  title={Motion-aware future frame prediction for video anomaly detection based on saliency perception},\n  year={2022},\n  cdate={1640995200000},\n  journal={Signal Image Video Process.},\n  volume={16},\n  number={8},\n  pages={2121-2129},\n  url={https://doi.org/10.1007/s11760-022-02174-7}\n}\n"},"abstract":{"value":"The anomaly in videos can be considered as a deviation from regular video sequences. Most existing approaches neglect the imbalanced information distribution between the foreground and the background during the process of reconstruction or prediction. To address this problem, we propose a motion-aware future frame prediction network consisting of a frame prediction branch and a saliency perception branch. In particular, the saliency perception branch is designed to predict the most salient targets in the video frame, and the frame prediction branch is used to predict the RGB future frame with the guidance of saliency perception. Besides, a motion-aware attention module is bridged in the frame prediction branch to improve the representation ability of moving targets. Furthermore, a saliency prediction loss and a saliency-guided appearance loss are designed to optimize saliency prediction frames and constrain the weight of foreground. Experiments on three challenging benchmarks demonstrate our competitive performance with the state-of-the-art approaches."},"title":{"value":"Motion-aware future frame prediction for video anomaly detection based on saliency perception"},"authors":{"value":["Haitao Xu","Weibin Liu","Weiwei Xing","Xiang Wei"]}},"tmdate":1731478348298,"pdate":1640995200000,"tcdate":1731476645581,"writers":["~"],"signatures":["~Xiang_Wei1"],"forum":"8Z06XuL6GQ","license":"CC BY-SA 4.0","number":202705,"cdate":1640995200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1731478348298,"domain":"DBLP.org","id":"8Z06XuL6GQ","version":2},{"content":{"summary":{"value":"This paper focuses the core challenges of long video understanding with VLMs, i.e., the excessive context length, prohibitive memory consumption, and loss of critical information in long context. The authors propose VideoDetective, a framework with recurrent question-aware memory compression and critical clue summarization. For evaluation, the authors introduce GLVC, a dataset with concrete critical clues and timestamps scattered across long videos. The experimental results show improvements on computation efficiency."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"Why choose to initialize memory tokens with bos token embedding?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. The explored problem is meaningful and the motivation of using recurrent memory compression is clear.\n2. The recurrent sub-segment processing limits peak memory usage and friendly to real-time video interaction device deployment.\n3. The GLVC benchmark is a contribution to the community for comprehensive evaluation of grounded video understanding."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The authors claim dynamic compression ratio for different sub-segments to avoid over- or under-compression, but the technical details are not presented. The experiments only show results with different fixed compression ratio on different benchmarks.\n2. The recurrent memory compression with history memory continuously added to the context is quite similar to [1], with the only difference in question-aware or not. Due to the lack of dynamic compression ratio, the advantage of question-aware compression is not shown in this architecture.\n3. The scalability to extremely long videos is doubted. The growing historical memory puts limitation on ultra-long videos in terms of both computation and information forgetting.\n4. The presentation is not clear in some crucial technical details. For example, in line 269, the authors claim \"memory tokens do not participate in loss calculation\", which is quite confusing since the memory tokens are explicitly in the sequence context in loss computation, and how do you optimize the memory representations.\n\n[1] Qian, R., Dong, X., Zhang, P., Zang, Y., Ding, S., Lin, D., & Wang, J. (2024). Streaming long video understanding with large language models. Advances in Neural Information Processing Systems, 37, 119336-119360."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762918252735,"tcdate":1761969000474,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission5773/Reviewer_rML2"],"signatures":["ICLR.cc/2026/Conference/Submission5773/Reviewer_rML2"],"forum":"9glgMTTZb9","number":3,"license":"CC BY 4.0","cdate":1761969000474,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission5773/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762918252735,"domain":"ICLR.cc/2026/Conference","replyto":"9glgMTTZb9","id":"dsu2c71Dlu","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["MLLM; LongVideo"]},"supplementary_material":{"value":"/attachment/2e42a1c09f6d6f7b74ce5b21f718da8e25d638dd.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Long Video Question-Answering (LVQA) presents a significant challenge for Multi-modal Large Language Models (MLLMs) due to immense context and overloaded information, which could also lead to prohibitive memory consumption. \nWhile existing methods attempt to address these issues by reducing visual tokens or extending model's context length, they may miss useful information or take considerable computation.\nIn fact, when answering given questions, only a small amount of crucial information is required.\nTherefore, we propose an efficient question-aware memory mechanism, enabling MLLMs to recurrently seek these critical clues. Our approach, named VideoDetective, simplifies this task by iteratively processing video sub-segments. For each sub-segment, a question-aware compression strategy is employed by introducing a few special memory tokens to achieve purposefully compression. This allows models to effectively seek critical clues while reducing visual tokens.\nThen, due to history context could have a significant impact, we recurrently aggregate and store these memory tokens to update history context, which would be reused for subsequent sub-segments. \nFurthermore, to more effectively measure model's long video understanding ability, we introduce GLVC (Grounding Long Video Clues), a long video question-answering dataset, which features grounding critical and concrete clues scattered throughout entire videos.\nExperimental results demonstrate our method enables MLLMs with limited context length of 32K to efficiently process 100K tokens (3600 frames, an hour-long video sampled at 1fps), requiring only 2 minutes and 37GB GPU memory usage. Evaluation results across multiple long video benchmarks illustrate our method can more effectively seek critical clues from massive information."},"_bibtex":{"value":"@misc{\ndu2026video,\ntitle={Video Detective: Seek Critical Clues Recurrently to Answer Question from Long Videos},\nauthor={Henghui Du and Chang Zhou and Chunjie Zhang and Xi Chen and Di Hu},\nyear={2026},\nurl={https://openreview.net/forum?id=9glgMTTZb9}\n}"},"title":{"value":"Video Detective: Seek Critical Clues Recurrently to Answer Question from Long Videos"},"pdf":{"value":"/pdf/bf8baa9af35aa0e72e7a5b3d40c3c4f35d3b6d6c.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"du|video_detective_seek_critical_clues_recurrently_to_answer_question_from_long_videos"},"authorids":{"value":["~Henghui_Du1","~Chang_Zhou3","~Chunjie_Zhang6","~Xi_Chen21","~Di_Hu1"]},"authors":{"value":["Henghui Du","Chang Zhou","Chunjie Zhang","Xi Chen","Di Hu"]}},"version":2},{"content":{"summary":{"value":"The paper proposes a method for automatically generating text-physics simulation video pairs that can be used for training text-to-simulation models. \nAt the core of the method is the design of set of attributed stochastic grammar that describes the 3D scene. A 3-step sample stragegy is designed so that different scenes can be sampled while satisfying all the constraints. The scenes are then simulated and rendered. The corresponding text descriptions can also be sampled using rules, and subsequently rewritten by ChatGPT. The 3D objects used for constructing the scenes are obtained from both existing datasets and Shap-E, a text-to-3D model. \nAlthough no quantitative evaluation is provided, the paper compared the generated data with the output of existing text-to-video models, hypothesizing that the data could improve the physics understanding of these models."},"presentation":{"value":"2 fair"},"contribution":{"value":"2 fair"},"soundness":{"value":"1 poor"},"strengths":{"value":"* The proposed 3D scene & text generation pipeline is quite general and it can potentially benefit other data generation tasks that require the generation of scene-text pairs.\n* Introducing physics into text-to-3d or text-to-video model is an understudied problem. And it is of great value to study whether such a problem can be solved solely through synthesized data."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"* The entire paper is based on the assumption that the data generated through the proposed pipeline can improve the physics understanding of visual-language multimodal models. However, there is no experiment that verifies this assumption.\n* Section 5 is flawed and incomplete. Instead of doing a proper AB testing by training a model using the generated data and evaluate the physical understanding of the models with and without the proposed data, the paper chose an extremely unfair setting of comparing the synthesized data directly with the results from an existing text-to-video model.\n* Without any quantitative evaluate, it is not obvious whether the generated data is actually useful for its designated purpose. In fact, it can even be counterproductive due to the lack of photorealism and potentially non-realistic scene setup.\n* Generating 3D scenes using stochastic grammar is not a new technique. It has been extensively studied, for example, in \"Configurable 3D Scene Synthesis and 2D Image Rendering with Per-Pixel Ground Truth using Stochastic Grammars, Jiang et al., IJCV 2018\""},"confidence":{"value":"3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked."},"questions":{"value":"I wonder if there could be other ways to better evaluate the new dataset, even without training a generative model?"},"rating":{"value":"3: reject, not good enough"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699637128573,"tcdate":1699238250690,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission8967/Reviewer_ZNsp"],"signatures":["ICLR.cc/2024/Conference/Submission8967/Reviewer_ZNsp"],"forum":"G7M3f3Ditm","number":3,"license":"CC BY 4.0","cdate":1699238250690,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission8967/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699637128573,"domain":"ICLR.cc/2024/Conference","replyto":"G7M3f3Ditm","id":"tAI3m7moIm","forumContent":{"venue":{"value":"ICLR 2024 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Text to Physics-based Animation","Multimodal Generation"]},"supplementary_material":{"value":"/attachment/5346a2e5938f55d805f64bde5f8c5ff36ac79275.pdf"},"primary_area":{"value":"datasets and benchmarks"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Powered by an enormous amount of paired data from the vision and language domains, Vision-Language (V&L) Multi-Modality (MM) research has achieved remarkable results in both text-driven generation and understanding. However, constrained by the data, the learned Multi-Modality (MM) knowledge space predominantly represents the alignments between text and appearances or shapes, lacking further understanding of the underlying dynamics. In this paper, we aim to expand the Multi-Modality (MM) knowledge space by bridging the gap between text, vision, and real-world physical dynamics from a data-centric perspective, enabling Multi-Modality (MM) models to better estimate these dynamics. We propose an automatic pipeline to generate Text-to-Video/Simulation (T2V/S) data. Each generated scenario comprises a high-resolution 3D physical simulation and a textual description of the physical phenomena. To simulate a diverse set of real-world dynamic phenomena---such as elastic deformations, material fractures, collisions, and turbulence---as faithfully as possible, we take advantage of state-of-the-art physical simulation methods: (i) Incremental Potential Contact (IPC) and (ii) Material Point Method (MPM . Additionally, high-quality, multi-view rendering is integrated into the pipeline. We envision our work as the first step towards fully automatic Text-to-Simulation (T2S), potentially shifting the paradigm towards understanding world dynamics."},"_bibtex":{"value":"@misc{\nqiu2024tpagen,\ntitle={{TPA}-Gen: A Multi-modal Data Generative Method for Text and Physics-based Animation},\nauthor={Yuxing Qiu and Feng Gao and Minchen Li and Yin Yang and Chenfanfu Jiang},\nyear={2024},\nurl={https://openreview.net/forum?id=G7M3f3Ditm}\n}"},"title":{"value":"TPA-Gen: A Multi-modal Data Generative Method for Text and Physics-based Animation"},"pdf":{"value":"/pdf/e189289905d873614396c015d489e6653230f46e.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Withdrawn_Submission"},"paperhash":{"value":"qiu|tpagen_a_multimodal_data_generative_method_for_text_and_physicsbased_animation"},"authorids":{"value":["~Yuxing_Qiu1","~Feng_Gao2","~Minchen_Li1","~Yin_Yang4","~Chenfanfu_Jiang3"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Yuxing Qiu","Feng Gao","Minchen Li","Yin Yang","Chenfanfu Jiang"]}},"version":2},{"content":{"summary":{"value":"This paper targets efficient subject-to-video (S2V) learning without relying on costly video–subject-pair datasets. The authors observe that naively training video models with image-paired data induces catastrophic degradation of temporal coherence due to gradient conflicts. They hypothesize that S2V can be decomposed into two orthogonal objectives: identity learning from images and temporal dynamics from videos. Building on this, they propose a stochastic task-switching strategy that predominantly samples from image datasets while maintaining minimal video replay to preserve temporal ability. Empirically, they show that the gradient inner product between the two tasks converges rapidly toward zero, indicating emergent orthogonalization without explicit projection. Experiments demonstrate superior performance with compute comparable to per-subject tuned methods for single-subject customization, while also providing zero-shot capability and outperforming both per-subject tuned approaches and several existing zero-shot baselines."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"If my concerns are satisfactorily addressed, I would be happy to raise my score."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The manuscript is clearly written and easy to follow.\n2. The motivation is well articulated: addressing the heavy reliance on expensive video data for zero-shot video customization.\n3. The topic is valuable, aligning with personalized user needs and real-world applications.\n4. The paper presents extensive experiments and some insightful empirical analyses."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The decomposition of video customization into ID injection and temporal dynamics appears similar to prior work (e.g., DreamVideo, Still-Moving). Please clarify the conceptual and methodological differences from these approaches, and specify what is genuinely new.\n2. The connection to continual learning (CL) is unclear. In standard CL, models learn sequentially over tasks/data while mitigating catastrophic forgetting and acquiring new capabilities. Here, the missing temporal dynamics seem to stem from fine-tuning on static image data rather than from sequential learning.\n3. The methodological novelty seems limited: the ID injection component does not introduce a new design, and the proposed “video replay” looks like joint training on images and videos rather than a CL-style memory replay mechanism, as the replay buffer is not updated during training.\n4. The tuning-free baselines used for comparison appear dated. Stronger and more recent video customization baselines would make the evaluation more convincing.\n5. Does the proposed method support multi-subject customization?"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762920541596,"tcdate":1761980296396,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission8758/Reviewer_iWod"],"signatures":["ICLR.cc/2026/Conference/Submission8758/Reviewer_iWod"],"forum":"xAaW436epC","number":3,"license":"CC BY 4.0","cdate":1761980296396,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission8758/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762920541596,"domain":"ICLR.cc/2026/Conference","replyto":"xAaW436epC","id":"mWYmaYxJmg","forumContent":{"TLDR":{"value":"We employ continual learning with video replay, and adjust replay ratio dynamically, achieving on-par subject fidelity and motion with lower compute compared to state-of-the-art models."},"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["video generation","customization","personalization","diffusion models","continual learning"]},"supplementary_material":{"value":"/attachment/a075028e26247ef5388ca11118150c5daa1eb0b3.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"We aim to enable efficient subject-to-video (S2V) learning, which otherwise requires expensive video-subject-pair datasets that require tens of thousands of GPU hours for training. While utilizing image-paired datasets to train video models could address this challenge, naively training with image pairs results in catastrophic loss of temporal ability due to gradient conflicts. We hypothesize that S2V generation decomposes into two orthogonal objectives of identity learning from images and temporal dynamics from videos. Based on this orthogonality assumption, we design a stochastic task-switching strategy that predominantly samples from image datasets while maintaining minimal video replay for temporal coherence. Our experiments validate this hypothesis by demonstrating that the gradient inner product between tasks converges exponentially to near-zero, confirming emergent orthogonalization without requiring explicit orthogonal projection. This validated orthogonality enables efficient image-dominant training while preventing catastrophic forgetting through proxy experience replay. We employ regularization techniques including random frame selection and token dropping during video replay to ensure efficient temporal learning. Extensive experiments demonstrate our approach achieves superior performance with  comparable compute to per-subject tuned methods for single subjects, while providing zero-shot capability and outperforming both per-subject tuned methods and some existing zero-shot approaches."},"_bibtex":{"value":"@misc{\nkim2025subjectdriven,\ntitle={Subject-driven Video Generation Emerges from Experience Replays},\nauthor={Daneul Kim and Jingxu Zhang and Wonjoon Jin and Sunghyun Cho and Qi Dai and Jaesik Park and Chong Luo},\nyear={2025},\nurl={https://openreview.net/forum?id=xAaW436epC}\n}"},"title":{"value":"Subject-driven Video Generation Emerges from Experience Replays"},"pdf":{"value":"/pdf/8049c6a47793f0babda63c028ea8ef3e70ae36cf.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"kim|subjectdriven_video_generation_emerges_from_experience_replays"},"authorids":{"value":["~Daneul_Kim1","~Jingxu_Zhang1","~Wonjoon_Jin1","~Sunghyun_Cho3","~Qi_Dai4","~Jaesik_Park3","~Chong_Luo1"]},"authors":{"value":["Daneul Kim","Jingxu Zhang","Wonjoon Jin","Sunghyun Cho","Qi Dai","Jaesik Park","Chong Luo"]}},"version":2},{"content":{"venue":{"value":"IEEE Trans. Circuits Syst. Video Technol. 2017"},"pdf":{"value":"https://ieeexplore.ieee.org/iel7/76/7807363/07508389.pdf"},"venueid":{"value":"dblp.org/journals/TCSV/2017"},"paperhash":{"value":"wei|qosaware_resource_allocation_for_video_transcoding_in_clouds"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Lei_Wei_0008:","https://dblp.org/search/pid/api?q=author:Jianfei_Cai_0001:","https://dblp.org/search/pid/api?q=author:Chuan_Heng_Foh:","~Bingsheng_He1"]},"html":{"value":"https://doi.org/10.1109/TCSVT.2016.2589621"},"_bibtex":{"value":"@article{DBLP:journals/tcsv/WeiCFH17,\n  author={Lei Wei and Jianfei Cai and Chuan Heng Foh and Bingsheng He},\n  title={QoS-Aware Resource Allocation for Video Transcoding in Clouds},\n  year={2017},\n  cdate={1483228800000},\n  journal={IEEE Trans. Circuits Syst. Video Technol.},\n  volume={27},\n  number={1},\n  pages={49-61},\n  url={https://doi.org/10.1109/TCSVT.2016.2589621}\n}\n"},"abstract":{"value":"As the biggest big data, video data streaming in the network contributes the largest portion of global traffic nowadays and in the future. Due to heterogeneous mobile devices, networks, and user preferences, the demands of transcoding source videos into different versions have increased significantly. However, video transcoding is a time-consuming task, and how to guarantee quality-of-service (QoS) for large video data is very challenging, particularly for those real-time applications that hold strict delay requirement such as live TV. In this paper, we propose a cloud-based online video transcoding (COVT) system aiming to offer economical and QoS guaranteed solution for online large-volume video transcoding. COVT utilizes the performance profiling technique to obtain the different performances of transcoding tasks in different infrastructures. Based on the profiles, we model the cloud-based transcoding system as a queue and derive the QoS values of the system based on the queuing theory. With the analytically derived relationship between QoS values and the number of CPU cores required for transcoding workloads, COVT is able to solve the optimization problem and obtain the minimum resource reservation for specific QoS constraints. A task scheduling algorithm is further developed to dynamically adjust the resource reservation and schedule the tasks so as to guarantee the QoS in runtime. We implement a prototype system of COVT and experimentally study the performance on real-world workloads. Experimental results show that the COVT effectively provisions a minimum number of resources for predefined QoS. To validate the effectiveness of our proposed method under large-scale video data, we further perform simulation evaluation, which again shows that the COVT is capable of achieving economical and QoS-aware video transcoding in cloud environment."},"title":{"value":"QoS-Aware Resource Allocation for Video Transcoding in Clouds"},"authors":{"value":["Lei Wei","Jianfei Cai","Chuan Heng Foh","Bingsheng He"]}},"tmdate":1769039351239,"pdate":1514678400000,"externalIds":["dblp:journals/tcsv/WeiCFH17"],"tcdate":1769039322038,"writers":["~"],"signatures":["~Bingsheng_He1"],"forum":"PtYdWEmbLC","license":"CC BY-SA 4.0","number":786035,"cdate":1483228800000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1769039351239,"domain":"DBLP.org","id":"PtYdWEmbLC","version":2},{"content":{"summary":{"value":"This paper introduces CHIMERA, a benchmark for diagnosing shortcut learning in VLMs on diagram understanding tasks. The dataset provides 7,500 Wikipedia-sourced diagrams, automatically annotated with semantic triples and four levels of multiple-choice questions. The authors evaluate 15 open-source VLMs and claim their performance is largely driven by three types of shortcuts: visual memorization, knowledge recall, and Clever-Hans. While the work addresses a relevant problem, significant methodological and conceptual limitations undermine its contributions."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"- How does CHIMERA differ from IconQA, ChartQA, and ERBench? These benchmarks also decompose visual reasoning into levels and test chart/diagram understanding. What unique contribution  of CHIMERA compared to those other works?\n\n- If entity recognition provided accurate diagram content to the model, wouldn't reasoning tasks become trivial? An ablation study can be done in this matter by showing that fixing ER errors would improve downstream task performance."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- The paper states a good question of investigating why VLMs succeed, focusing on shortcut behaviors rather than just performance metrics.\n\n- The four-level task hierarchy (ER → RU → KG → VR) provides a good framework for analyzing different aspects of diagram comprehension.\n\n- Evaluating across three modalities (visual diagrams, semantic graphs, and textual descriptions) is a well-structured approach for isolating modality-specific biases.\n\n- The dataset includes 7,500 diagrams with 20% human validation, demonstrating reasonable quality control efforts."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The finding that VLMs rely on language priors and exhibit superficial pattern matching is well-established in many works, such as VQA, IconQA, VLMs are biased, etc. The paper does not sufficiently differentiate its contributions from these existing works. \n\n- The near 2% gap between visual and semantic modalities is within noise margins and insufficient to claim memorization effects, and it is not a powerful claim. It can be changed in another experiment.\n\n- ER performance itself consists of two steps: OCR text extraction and visual element extraction. It should be discussed and investigated which one affects the low performance. Moreover, successful reasoning fundamentally depends on accurate entity recognition. If ER fails in one of those tasks, then subsequent reasoning operates on corrupted inputs, making high performance on KG/VR tasks despite poor ER performance inherently suspicious and indicative of shortcut behavior.\n\n- The superior performance of the models in Wikipedia data can be caused by the models' training on the same data and memorizing it.\n\n- In Clever-Hans Shortcut, while the blank-image experiment is interesting, the analysis is incomplete. There is a need for an analysis of question-answer correlation biases; there is also a need for a comparison with random baseline or majority class baseline beyond the 25% chance level; and lastly, the evaluation also misses an analysis of whether performance correlates with question linguistic features.\n\n- The paper is purely diagnostic with no proposed methods to mitigate identified shortcuts. As an example, a work published in ICLR 2025, \"Chain-of-Region,\" used the OpenCV library to extract the visual data and feed it as text to the models."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762924090950,"tcdate":1761882302277,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission13473/Reviewer_HN5s"],"signatures":["ICLR.cc/2026/Conference/Submission13473/Reviewer_HN5s"],"forum":"q3eB3PhtqD","number":3,"license":"CC BY 4.0","cdate":1761882302277,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission13473/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762924090950,"domain":"ICLR.cc/2026/Conference","replyto":"q3eB3PhtqD","id":"TvvbHiJRdF","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["vision-language model","diagram understanding","multimodal reasoning","visual question answering","shortcut learning","knowledge grounding","benchmark dataset","multimodal evaluation"]},"primary_area":{"value":"interpretability and explainable AI"},"abstract":{"value":"Diagrams convey symbolic information in a visual format rather than a linear stream of words, making them especially challenging for AI models to process. While recent evaluations suggest that vision-language models (VLMs) perform well on diagram-related benchmarks, their reliance on knowledge, reasoning, or modality shortcuts raises concerns about whether they genuinely understand and reason over diagrams.\nTo address this gap, we introduce Chimera, a comprehensive test suite comprising 7,500 high-quality diagrams sourced from Wikipedia; each diagram is annotated with its symbolic content represented by semantic triples along with multi-level questions designed to assess four fundamental aspects of diagram comprehension: entity recognition, relation understanding, knowledge grounding, and visual reasoning.\nWe use Chimera to measure the presence of three types of shortcuts in visual question answering: \n(1) the visual-memorization shortcut, where VLMs rely on memorized visual patterns;\n(2) the knowledge-recall shortcut, where models leverage memorized factual knowledge instead of interpreting the diagram; and\n(3) the Clever-Hans shortcut, where models exploit superficial language patterns or priors without true comprehension. We evaluate 15 open-source VLMs from 7 model families on Chimera and find that their seemingly strong performance largely stems from shortcut behaviors: visual-memorization shortcuts have slight impact, knowledge-recall shortcuts play a moderate role, and Clever-Hans shortcuts contribute significantly.\nThese findings expose critical limitations in current VLMs and underscore the need for more robust evaluation protocols that benchmark genuine comprehension of complex visual inputs (e.g., diagrams) rather than question-answering shortcuts."},"_bibtex":{"value":"@misc{\nchi2026chimera,\ntitle={Chimera: Diagnosing Shortcut Learning in Visual-Language Understanding},\nauthor={Ziheng Chi and Yifan Hou and Chenxi Pang and Shaobo Cui and Mubashara Akhtar and Mrinmaya Sachan},\nyear={2026},\nurl={https://openreview.net/forum?id=q3eB3PhtqD}\n}"},"title":{"value":"Chimera: Diagnosing Shortcut Learning in Visual-Language Understanding"},"pdf":{"value":"/pdf/f92e46808334228be1f33c8c27812a84e924f889.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"chi|chimera_diagnosing_shortcut_learning_in_visuallanguage_understanding"},"authorids":{"value":["~Ziheng_Chi1","~Yifan_Hou1","~Chenxi_Pang1","~Shaobo_Cui1","~Mubashara_Akhtar1","~Mrinmaya_Sachan3"]},"authors":{"value":["Ziheng Chi","Yifan Hou","Chenxi Pang","Shaobo Cui","Mubashara Akhtar","Mrinmaya Sachan"]}},"version":2},{"content":{"summary":{"value":"This paper studies streaming video understanding. The authors propose a new benchmark named StreamBench to evaluate streaming video understanding across diverse media types and interactive scenarios, including multi-turn interactions and complex reasoning tasks. They also propose StreamChat, a training-free framework for streaming video reasoning and conversational interaction, with a complex memory mechanism. Their method enhances processing speed and reduces latency, ensuring robust performance in real-world applications. Extensive evaluations on StreamBench and some public benchmarks demonstrate that StreamChat outperforms the selected\nbaselines."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"Please respond to Weaknesses."},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. The proposed StreamBench may be the first benchmark for streaming video understanding.\n2. The proposed StreamChat outperforms the selected baselines on StreamBench.\n3. The processing speed of StreamChat significantly outperforms those of baselines."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The proposed dataset is too small, only 306 videos and 1.8K  question-answer pairs are collected. The current video benchmark typically has at least thousands of videos and tens of thousands of QA pairs. This version of the benchmark is not ready for release.\n2. State-of-the-art video LLMs are not included in the benchmark, such as MiniCPM-V 2.6 [1], InternLM-XComposer2.5 [2], VILA [3], and InternVL2 [4]. The effectiveness of the proposed method is unclear.\n3. The Selective Frame Stacking seems may ignore small objects in the video and only focus on the global frame feature.\n4. The proposed memory mechanism is very complicated and discards a large amount of information in the video.\n5. The authors test the processing speed of models on two NVIDIA Tesla A800 80GB GPUs, which is not a typical scenario of model development in real-world applications.\n\n\n[1] Yao, Y., Yu, T., Zhang, A., Wang, C., Cui, J., Zhu, H., ... & Sun, M. (2024). Minicpm-v: A gpt-4v level mllm on your phone. arXiv preprint arXiv:2408.01800.\n[2] Zhang, P., Dong, X., Zang, Y., Cao, Y., Qian, R., Chen, L., ... & Wang, J. (2024). Internlm-xcomposer-2.5: A versatile large vision language model supporting long-contextual input and output. arXiv preprint arXiv:2407.03320.\n[3] Lin, J., Yin, H., Ping, W., Molchanov, P., Shoeybi, M., & Han, S. (2024). Vila: On pre-training for visual language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 26689-26699).\n[4] Chen, Z., Wang, W., Tian, H., Ye, S., Gao, Z., Cui, E., ... & Qiao, Y. (2024). How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites. arXiv preprint arXiv:2404.16821."}},"nonreaders":[],"tmdate":1731427310356,"tcdate":1730475431848,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission743/Reviewer_yVyo"],"signatures":["ICLR.cc/2025/Conference/Submission743/Reviewer_yVyo"],"forum":"JbPb6RieNC","number":1,"license":"CC BY 4.0","cdate":1730475431848,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission743/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427310356,"domain":"ICLR.cc/2025/Conference","replyto":"JbPb6RieNC","id":"tKDwBz0OMj","forumContent":{"venue":{"value":"ICLR 2025 Poster"},"TLDR":{"value":"A Vertasile Approach and Benchmark for Streaming Video Understanding and Multi-round Interaction"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Streaming Video Understanding; Video MLLM; Hierarchical Memory System"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Recent advances in Large Language Models (LLMs) have enabled the development of Video-LLMs, advancing multimodal learning by bridging video data with language tasks. However, current video understanding models struggle with processing long video sequences, supporting multi-turn dialogues, and adapting to real-world dynamic scenarios. To address these issues, we propose StreamChat, a training-free framework for streaming video reasoning and conversational interaction. StreamChat leverages a novel hierarchical memory system to efficiently process and compress video features over extended sequences, enabling real-time, multi-turn dialogue. Our framework incorporates a parallel system scheduling strategy that enhances processing speed and reduces latency, ensuring robust performance in real-world applications. Furthermore, we introduce StreamBench, a versatile benchmark that evaluates streaming video understanding across diverse media types and interactive scenarios, including multi-turn interactions and complex reasoning tasks.  Extensive evaluations on StreamBench and other public benchmarks demonstrate that  StreamChat significantly outperforms existing\nstate-of-the-art models in terms of accuracy and response times, confirming its effectiveness for streaming video understanding. Code is available at StreamChat."},"_bibtex":{"value":"@inproceedings{\nxiong2025streaming,\ntitle={Streaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge},\nauthor={Haomiao Xiong and Zongxin Yang and Jiazuo Yu and Yunzhi Zhuge and Lu Zhang and Jiawen Zhu and Huchuan Lu},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=JbPb6RieNC}\n}"},"title":{"value":"Streaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge"},"pdf":{"value":"/pdf/6b4412176320c83956b2b47bd354a372b72935a8.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"xiong|streaming_video_understanding_and_multiround_interaction_with_memoryenhanced_knowledge"},"authorids":{"value":["~Haomiao_Xiong1","~Zongxin_Yang1","~Jiazuo_Yu1","~Yunzhi_Zhuge1","~Lu_Zhang7","~Jiawen_Zhu1","~Huchuan_Lu1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Haomiao Xiong","Zongxin Yang","Jiazuo Yu","Yunzhi Zhuge","Lu Zhang","Jiawen Zhu","Huchuan Lu"]}},"version":2},{"content":{"summary":{"value":"The paper introduces LongViTU for video understanding, which comprises approximately 121k question-answer pairs across 900 hours of video content, focusing on long-context videos that require rich knowledge and reasoning. The authors propose a hierarchical pipeline for generating high-quality QA pairs with explicit timestamp labels, catering to diverse real-world scenarios. LongViTU is curated to support fine-grained and open-ended QA. The paper also presents experiments demonstrating the performance gap between open-source and commercial models on this benchmark and the effectiveness of SFT on LongViTU."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"See weaknesses."},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"- The paper is easy to follow, and the experiments are clearly described.\n- The dataset is of high quality, featuring a large number of QA pairs and encompassing a variety of diverse scenarios."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Figure 1: The icons, while visually appealing, come across as unprofessional and occupy space that could be better utilized to present more information.\n- Ablation Studies: The paper lacks ablation studies for different-level captions. For instance, it would be beneficial to know if event-level captions can be skipped without significant detriment.\n- Results: Additional results are necessary to clarify the performance of different Multi-modal Large Language Models (MLLMs) on LongViTU videos with varying durations.\n- Comparison with ShareGPT4Video[1]: The authors of ShareGPT4Video present a progressive framework that generates detailed captions for diverse videos. In contrast, LongViTU focuses solely on ego-centric videos due to its dependence on human annotation, which potentially limits its application and robustness for general QA, as evidenced in Table 3.\n\n---\nReference:\n\n[1] Chen, Lin et al. “ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.” ArXiv abs/2406.04325 (2024): n. pag."}},"nonreaders":[],"tmdate":1731428925742,"tcdate":1730730837734,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission7604/Reviewer_KRZ6"],"signatures":["ICLR.cc/2025/Conference/Submission7604/Reviewer_KRZ6"],"forum":"4j9plQoOH1","number":4,"license":"CC BY 4.0","cdate":1730730837734,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission7604/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428925742,"domain":"ICLR.cc/2025/Conference","replyto":"4j9plQoOH1","id":"zkCyubB3OY","forumContent":{"TLDR":{"value":"We propose a large-scale instruction-tuning dataset for long-form video understanding."},"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["vision language models","instruction-tuning","long-form video understanding"]},"primary_area":{"value":"datasets and benchmarks"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"This paper presents LongViTU, a large-scale (~121k QA pairs, ~900h videos), automatically generated dataset for long-form video understanding. Our key idea is inspired by the success of Large Language Models (LLMs) and Multimodal Language Models (MLMs) that are fueled by machine-generated instruction-following data (*e.g.*, InstructGPT, LLaVA). We developed a *systematic* approach to produce massive question-answeringing pairs tailored to virtually unbounded long videos by organizing them into a ***hierarchical tree***, incorporating ***self-revision*** mechanisms to guarantee high quality. We curate LongViTU for each QA pair: 1) involves a long context (average *certificate length* of 4.6 minutes); 2) requires rich knowledge and condensed reasoning (commonsense, causality, planning, *etc.*); 3) explicit labels the timestamps of relevant events throughout the entire video. Furthermore, LongViTU provides a benchmark to facilitate future research in instruction-following for long-form videos. Our experiments first reveal the performance gap between open-source video MLMs and their commercial counterparts (*e.g.*, Gemini-1.5-Pro) on this benchmark. Supervised Fine-Tuning (SFT) on open-source models led to Video-LLaVA achieving the best performance, with a GPT-4 score of $50.7$, closely following $52.3$ by the leading closed-source model Gemini-1.5-Pro, underscoring the substantial challenge posed by our benchmark. Further SFT on LongViTU with Video-LLaVA resulted in improvements of $30.7$% on the In-Distribution (ID) benchmark EgoSchema; $12.9$% and $0.6$% on the Out-of-Distribution (OOD) benchmarks WorldQA and VideoMME, respectively. These outcomes demonstrate the effectiveness and robust OOD generalizability of our proposed instruction-tuning scheme for long-form video understanding. The dataset, SFT models, and code are publicly available on the anonymous page [LongViTU](https://longvitu.github.io)."},"_bibtex":{"value":"@misc{\nwu2024longvitu,\ntitle={LongVi{TU}: Instruction Tuning for Long-Form Video Understanding},\nauthor={Rujie Wu and Xiaojian Ma and Hai Ci and Yue Fan and Yuxuan Wang and Haozhe Zhao and Qing Li and Yizhou Wang},\nyear={2024},\nurl={https://openreview.net/forum?id=4j9plQoOH1}\n}"},"title":{"value":"LongViTU: Instruction Tuning for Long-Form Video Understanding"},"pdf":{"value":"/pdf/e663a2eb9e041444826a666f95acc8764c6e736b.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"wu|longvitu_instruction_tuning_for_longform_video_understanding"},"authorids":{"value":["~Rujie_Wu2","~Xiaojian_Ma1","~Hai_Ci1","~Yue_Fan2","~Yuxuan_Wang4","~Haozhe_Zhao1","~Qing_Li1","~Yizhou_Wang1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Rujie Wu","Xiaojian Ma","Hai Ci","Yue Fan","Yuxuan Wang","Haozhe Zhao","Qing Li","Yizhou Wang"]}},"version":2},{"content":{"summary":{"value":"This paper enhances video anomaly detection by using single timestamp labels that indicate the start of an anomaly, providing minimal yet effective supervision. Leveraging this prior information, the authors propose Gaussian-prior Normal Pattern Modeling (GNPM) to capture normal patterns in the frames preceding the timestamp within anomalous videos. Additionally, they introduce Similarity-based Abnormal Pattern Modeling (SAPM) to model abnormal patterns effectively based on the single frame annotation. Together, these methods improve the model’s ability to distinguish between normal and abnormal sequences, achieving robust anomaly detection with minimal annotation effort."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"See the weaknesses."},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The experimental results show that the proposed method outperforms existing weakly-supervised approaches, demonstrating its effectiveness in video anomaly detection even with minimal supervision.\n\n2. The method is intuitive and easy to understand, making the concepts and implementation accessible. This simplicity, coupled with effective results, underscores its practical applicability.\n\n3. The use of a single timestamp to mark the start of an anomaly is a novel approach that significantly reduces annotation workload, enabling robust anomaly detection with minimal supervision. This innovation makes the method both efficient and practical for real-world applications."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The ablation study shows that AUC scores are sensitive to the parameter settings in both GNPM and SAPM across different datasets, which could indicate a limitation in the generalization ability of the proposed method. This sensitivity suggests that the model may require careful parameter tuning for optimal performance on new datasets.\n\n2. It would strengthen the paper to include comparisons with fully-supervised methods on fully annotated datasets. Since the proposed method relies only on the start frame as an annotation, such a comparison would better illustrate its effectiveness and efficiency. Limiting comparisons to only weakly-supervised and semi-supervised methods leaves an incomplete assessment of the model’s overall performance.\n\n3. There are minor typos in the paper that could be addressed for clarity. For example, in Figure 6, the caption misses “SAPM,” and on line 537, the word “temporal” is repeated."}},"nonreaders":[],"tmdate":1731428623903,"tcdate":1730786001689,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission6299/Reviewer_hnoM"],"signatures":["ICLR.cc/2025/Conference/Submission6299/Reviewer_hnoM"],"forum":"A18zU6cgQ0","number":5,"license":"CC BY 4.0","cdate":1730786001689,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission6299/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428623903,"domain":"ICLR.cc/2025/Conference","replyto":"A18zU6cgQ0","id":"yus6aa64js","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video Anomaly Detection","Inexact Supervision","Single Frame Supervision"]},"supplementary_material":{"value":"/attachment/0dac630e28277622baabc210d20bb04adefe3880.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Video Anomaly Detection (VAD) aims to identify anomalous frames in given videos. Existing fully-supervised VAD encounters substantial annotation cost and weakly-supervised VAD suffers from the deficiency of weak labels. In this paper, we propose a more effective Single Frame supervised VAD (SF-VAD), which leverages single abnormal frame as label. We argue that single abnormal frame provides precise dual references to abnormal and normal frames, which facilitates dependable anomaly and normality modeling, and it can be obtained with negligible extra cost. Under this setting, we propose similarity-based abnormal pattern modeling, to learn inclusive abnormal patterns reliably from mined abnormal frames, guided by similarity-based abnormal probability. And we introduce Gaussian-prior normal pattern modeling to decouple normal patterns in abnormal videos, by learning normal patterns in preceding frames, guided by Gaussian-prior normal probability. In inference, we additionally design temporal decoupling and boundary refining modules to reveal discriminative abnormal characters of temporal features. Extensive experiments show our SF-VAD method outperforms state-of-the-art VAD methods and achieves an optimal performance-cost trade-off. We construct and release three SF-VAD datasets to support future research."},"_bibtex":{"value":"@misc{\nchen2025video,\ntitle={Video Anomaly Detection via Single Frame Supervision},\nauthor={Junxi Chen and Liang Li and Li Su and Yunbin Tu and Zhe Xue and Qingming Huang},\nyear={2025},\nurl={https://openreview.net/forum?id=A18zU6cgQ0}\n}"},"title":{"value":"Video Anomaly Detection via Single Frame Supervision"},"pdf":{"value":"/pdf/d54f36fc796e4f0daeee0e89769a6dbdf57f2bde.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"chen|video_anomaly_detection_via_single_frame_supervision"},"authorids":{"value":["~Junxi_Chen3","~Liang_Li3","~Li_Su4","~Yunbin_Tu1","~Zhe_Xue2","~Qingming_Huang2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Junxi Chen","Liang Li","Li Su","Yunbin Tu","Zhe Xue","Qingming Huang"]}},"version":2},{"content":{"venue":{"value":"ICLR 2027 Conference Submission"},"keywords":{"value":["Encrypted Traffic Classification; Shortcut Learning; Invariant Representation Learning; Mutual Information Minimization; Pre-training Framework"]},"primary_area":{"value":"alignment, fairness, safety, privacy, and societal considerations"},"abstract":{"value":"Pre-trained models on raw traffic bytes report encrypted traffic classification (NTC) accuracies above 98\\% on mainstream benchmarks, yet recent evidence attributes much of this progress to shortcut learning: models exploit dataset-specific artifacts that correlate with labels in collected data but carry no application semantics, and degrade under distribution shift, field obfuscation, and adaptive evasion. Existing mitigations either remove a pre-defined list of suspect fields---missing dataset-specific shortcuts---or diagnose shortcut reliance post hoc without feeding it back into training. We present InvarFlow, a pre-training framework that encodes shortcut invariance directly into the training objective, combining per-corpus shortcut profiling via adjusted mutual information, a differentiable suppression term built on a mutual-information upper bound, a cross-environment invariant-risk penalty, and a shortcut-aware masking curriculum. We formalize shortcut invariance as a constrained optimization problem and prove that, under an explicit semantic/shortcut decomposition, any representation achieving exact invariance attains the environment-robust optimum---and that approximate invariance bounds the cross-environment risk gap linearly in the residual dependence. A prerequisite gate experiment on synthetic ground truth shows that the deployed mutual-information upper bound is loose but directionally valid (monotone, above-truth), which we turn into an explicit design rule: the bound drives training while calibrated probe statistics drive evaluation. In a controlled toy study with an injected, abolished-at-test shortcut, the invariance objective reduces the cross-environment accuracy gap by 71\\% relative to ERM at a 3-point in-distribution cost, matching the trend predicted by our bounds. On a large-scale benchmark under our pre-registered protocol, InvarFlow cuts the cross-environment gap by 72--74\\% relative to ERM while staying within 0.4--0.9 F1 of the best published in-distribution result per dataset, and its accuracy survives field-level decontamination with only 1.2--2.0 points of degradation---evidence that the gains come from genuine semantic reliance rather than artifact fitting."},"_bibtex":{"value":"@inproceedings{\nanonymous2026invarflow,\ntitle={InvarFlow: Shortcut-Invariant Pre-training for Generalizable Encrypted Traffic Classification},\nauthor={Anonymous},\nbooktitle={Submitted to The Fifteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=JQSzQOFcTz},\nnote={under review}\n}"},"title":{"value":"InvarFlow: Shortcut-Invariant Pre-training for Generalizable Encrypted Traffic Classification"},"pdf":{"value":"/pdf/969a3d879bba6fd52c98f5940f18f0fa4fbb9bac.pdf"},"venueid":{"value":"ICLR.cc/2027/Conference/Submission"}},"tmdate":1791235052743,"tcdate":1789799166477,"writers":["ICLR.cc/2027/Conference","ICLR.cc/2027/Conference/Submission55637/Authors"],"signatures":["ICLR.cc/2027/Conference/Submission55637/Authors"],"forum":"JQSzQOFcTz","license":"CC BY 4.0","number":55637,"cdate":1789799166477,"readers":["everyone"],"invitations":["ICLR.cc/2027/Conference/-/Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Bidding","ICLR.cc/2027/Conference/Submission55637/-/Full_Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Reviewing","ICLR.cc/2027/Conference/-/Edit"],"mdate":1791235052743,"odate":1791031328286,"domain":"ICLR.cc/2027/Conference","id":"JQSzQOFcTz","version":2},{"content":{"summary":{"value":"The paper introduces Instruct-Video-Continuation (InstructVC), a two-stage inference-time framework for multi-action long video generation that adds temporal causality to pretrained video diffusion models. Stage 1 (Temporal Action Binding) uses an MLLM with in-context examples to decompose and expand a user prompt into a scene description plus a sequence of action-duration pairs. Stage 2 (Causal Video Continuation) converts an Image-to-Video (I2V) diffusion model into a history-aware autoregressive continuer via Video Path Integral (VPI), which integrates multiple I2V “video paths” from sampled historical frames to bias future trajectories toward history-consistent directions. The practical instance SteinsGate adds three optimizations—Guidance Interval, History-aligned Redistribution, and Path Convergence Guidance—to reduce compute and improve alignment. Experiments on a constructed InstructVC benchmark and ablations report improved temporal control, motion smoothness, and multi-action continuity vs. several inference-time baselines and show competitive performance with heavier training-based methods."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. Can the authors precisely define the notation and dimensionality used in Eq. (5) for Video Path Integral, including what each index (i, j, K, H) represents and how histories are sampled in practice?\n\n2. How exactly is the mapping performed from image-conditioned I2V vector fields to the history-conditioned vector fields ve(Zt | xj); give the algorithmic steps used at each denoising iteration.\n\n3. What is the precise schedule for Guidance Interval: is it a fixed fraction of the continuous time variable, and how was 0.3 chosen?\n\n4. When you say the Video Path Integral approximates p(Zi | Zh) by product of frame-wise conditionals (Eq. 6), what independence assumptions are implied and when do they break down?\n\n5. How many prompts and total videos are in the InstructVC benchmark (train/val/test splits), and what is the distribution of number-of-actions, scene types, and durations?\n\n6. What are the exact definitions and implementations for CSCV+, Motion Smoothness, and Text-Image Alignment metrics, and how were thresholds or evaluation pipelines calibrated?\n\n7. Were human raters used for causal correctness or action completeness? If so, how many raters, what instructions, and inter-rater agreement (e.g., Cohen’s kappa) were measured?\n\n8. For quantitative comparisons, were compute budgets and sampling steps matched across methods (same sampler, same number of denoising steps)? Please provide latency and GPU-memory numbers.\n\n9. Please provide a sensitivity analysis for the number of history frames K, the selection strategy for those frames, and the history-length ratio used for different action durations.\n\n10. How robust is SteinsGate to noisy or incorrect history (e.g., partially corrupted frames, temporal jitter, or mismatched last-frame poses)?\n\n11. How does performance degrade over very long sequences beyond 30s, and what empirical evidence supports the claim that VPI reduces error accumulation compared to I2V-AR?\n\n12. What failure modes are most common (motion reversal, skipped actions, object disappearance), and can you quantify their frequencies per baseline?\n\n13. What MLLM model/version and temperature/prompting strategy was used for Temporal Action Binding, and what in-context examples were provided (exact prompts/examples)?\n\n13. Please provide quantitative measures of MLLM output quality: rate of hallucinated or physically implausible action sequences, distribution shift vs. the TI2V training captions, and any post-filtering applied.\n\n14. Will code, exact prompts, model checkpoints (Wan2.1 config and weights), and InstructVC benchmark splits be released?"},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The paper isolates two well-motivated, complementary challenges for long video generation—temporal causality and temporal control—and proposes a coherent two-stage solution.\n\n2. Video Path Integral provides a principled, interpretable way to propagate temporal information from multiple history frames and to bias generation toward causally consistent continuations without retraining the foundation model.\n\n3. The three SteinsGate optimizations (Guidance Interval, History-aligned Redistribution, Path Convergence Guidance) address computational cost and estimation noise, making the method more viable in practice.\n\n4. Temporal Action Binding leverages MLLM capabilities for temporally grounding and duration prediction, enabling fine-grained temporal control over multi-action sequences.\n\n5. The paper includes qualitative continuation examples, a system-level comparison for multi-action long generation, and ablations that quantify contribution of each component.\n\n6. The method can be applied at inference-time to pretrained TI2V/DiT models, which increases practical applicability and lowers barrier for adoption."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The VPI argument is mostly conceptual and analogy-driven (path integral intuition); formal bounds or rigorous analysis of when/why VPI converges to the true conditional distribution are missing.\n\n2. Performance and long-term consistency appear sensitive to how much and which historical frames are used, but the paper provides only heuristic choices and limited analysis of failure modes for varying history lengths.\n\n3. Although in-context learning is used to reduce implausible decompositions, reliance on an MLLM can still produce action sequences or durations that are out-of-distribution for the pretrained video model; mitigation and quantitative assessment are limited.\n\n4. Benchmark construction and evaluation transparency: the InstructVC benchmark is constructed by expanding dense-caption datasets, but details on dataset size, prompt diversity, split protocols, evaluation metrics definitions, and human evaluation procedures are sparse.\n\n5. Reported baselines are appropriate, but the paper lacks comparisons against recent large-scale training-based long-video systems (or more thorough hyperparameter-matched baselines) that could better contextualize headroom and limitations.\n\n6. The metrics focus on smoothness and CLIP alignment; deeper semantic correctness (action completeness, causal correctness judged by humans) and statistical significance of improvements are not fully reported."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762928264042,"tcdate":1761683392729,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission18554/Reviewer_inkc"],"signatures":["ICLR.cc/2026/Conference/Submission18554/Reviewer_inkc"],"forum":"8WS5nDWIWE","number":1,"license":"CC BY 4.0","cdate":1761683392729,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission18554/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762928264042,"domain":"ICLR.cc/2026/Conference","replyto":"8WS5nDWIWE","id":"N0A8u6g8vx","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Generative Models","Video Generation","Diffusion Guidance"]},"supplementary_material":{"value":"/attachment/325404a03779eedb98dfecccd83ad4d515610887.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Video generation has advanced rapidly, but current models remain limited to short clips, far from the length and complexity of real-world narratives. Long video generation is thus both important and challenging. Existing approaches either attempt to extend the modeling length of video diffusion models directly or merge short clips via shared frames. However, due to the lack of temporal causality modeling for video data, they achieve only limited extensions, suffer from discontinuous or even contradictory actions, and fail to support flexible and fine-grained temporal control. Thus, we propose Instruct-Video-Continuation (InstructVC), combining Temporal Action Binding for fine-grained temporal control and Causal Video Continuation for natural long-term simulation. Temporal Action Binding decomposes complex long videos by temporal causality into scene descriptions and action sequences with predicted durations, while Causal Video Continuation autoregressively generates coherent video narratives from the text story. We further introduce SteinsGate, an inference-time instance of InstructVC that uses an MLLM for Temporal Action Binding and Video Path Integral to enforce causality between actions, converting a pre-trained TI2V diffusion model into an autoregressive video continuation model. Benchmark results demonstrate the advantages of SteinsGate and InstructVC in achieving accurate temporal control and generating natural, smooth multi-action long videos."},"_bibtex":{"value":"@inproceedings{\nhuang2026steinsgate,\ntitle={SteinsGate: Adding Causality to Diffusions for Long Video Generation via Path Integral},\nauthor={Yufei Huang and Liangyu Yuan and Changxi Chi and Yunfan Liu and Cheng Tan and Siyuan Li and Jingbo Zhou and Haitao Lin and Chang Yu and Stan Z. Li},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=8WS5nDWIWE}\n}"},"title":{"value":"SteinsGate: Adding Causality to Diffusions for Long Video Generation via Path Integral"},"pdf":{"value":"/pdf/72d2476c464517c6b0b3d5b1e46565ebd799bdf9.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"huang|steinsgate_adding_causality_to_diffusions_for_long_video_generation_via_path_integral"},"authorids":{"value":["~Yufei_Huang4","~Liangyu_Yuan1","~Changxi_Chi1","~Yunfan_Liu2","~Cheng_Tan1","~Siyuan_Li6","~Jingbo_Zhou2","~Haitao_Lin2","~Chang_Yu1","~Stan_Z._Li2"]},"authors":{"value":["Yufei Huang","Liangyu Yuan","Changxi Chi","Yunfan Liu","Cheng Tan","Siyuan Li","Jingbo Zhou","Haitao Lin","Chang Yu","Stan Z. Li"]}},"version":2},{"content":{"venue":{"value":"IEEE Transactions on Circuits and Systems for Video Technology"},"pdf":{"value":"https://ieeexplore.ieee.org/iel8/76/10949577/10755967.pdf"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"zhao|finegrained_modality_relationaware_network_for_video_moment_retrieval"},"html":{"value":"https://doi.org/10.1109/TCSVT.2024.3494744"},"abstract":{"value":"Video moment retrieval (VMR) involves localizing video segments semantically aligned with given queries within videos. Despite the development of numerous methods for VMR in recent years, there remains a need to better incorporate fine-grained modality relation-aware information both in intra-modality and cross-modality. To address these challenges, we propose a Fine-grained Modality Relation-Aware Network (FMRN) tailored for the video moment retrieval task. FMRN effectively explores fine-grained modality relation-aware information within text queries, videos, and proposals. Our approach begins with a semantic graph encoder to capture deep semantic relations in intra-modality. Besides, we introduce a novel fine-grained cross-modality interaction module comprising a cross-similarity weighting module, an intra-modality weighting module, and an adaptive fusion module. These components comprehensively exploit fine-grained relation information within intra-modality and cross-modality contexts. Specifically, the cross-similarity weighting module leverages similarities between text queries and video snippets, as well as between videos and query words. The intra-modality weighting module determines the importance of words and snippets, while the adaptive fusion module combines cross-similarity weighting and intra-modality weighting. Additionally, we design a proposal relation module to enhance retrieval by capturing fine-grained proposals-relation information in videos. Extensive experiments demonstrate that the proposed method can outperform all state-of-the-art methods on the TACoS dataset and obtain comparable results on the Charades-STA and ActivityNet-Captions datasets. Compared with MCMN (TCSVT2024) and DPHANet (TMM2024), FMRN can achieve average improvements of 3.61 % and 5.44 % on the TACoS dataset, respectively."},"title":{"value":"Fine-Grained Modality Relation-Aware Network for Video Moment Retrieval"},"authors":{"value":[{"fullname":"Yibo Zhao","username":"~Yibo_Zhao9"},{"fullname":"Zan Gao","username":"~Zan_Gao1"},{"fullname":"Chunjie Ma"},{"fullname":"Weili Guan"},{"fullname":"Riwei Wang"},{"fullname":"Shengyong Chen"}]}},"tmdate":1789090346887,"pdate":1743465600000,"externalIds":["doi:10.1109/tcsvt.2024.3494744"],"tcdate":1768547110753,"writers":["~","OpenReview.net/Public_Article/ORCID.org","OpenReview.net/Support"],"signatures":["~Zan_Gao1"],"forum":"oSR1IQKnU9","license":"CC BY-SA 4.0","number":30729,"cdate":1731961659555,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/ORCID.org/-/Record","OpenReview.net/Public_Article/-/Edit","OpenReview.net/Public_Article/-/Authorship_Claim","OpenReview.net/-/Edit"],"mdate":1789090346887,"domain":"OpenReview.net/Public_Article","id":"oSR1IQKnU9","version":2},{"content":{"summary":{"value":"The paper proposes a model called Motion-Catcher. This model is a diffusion model designed to enhance motion and content consistency in multi-sequence video generation. The model addresses common issues in long video generation, such as motion inconsistency and content degradation, by introducing two main components: a motion capture module that leverages optical flow information for enhanced motion continuity, and a dynamic content prior module to mitigate content degradation over time. Experimental results demonstrate that Motion-Catcher significantly improves video quality, stability, and coherence compared to other models, with applicability as a plug-and-play enhancement for other video diffusion models."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Please address my concerns above. Thank you!"},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. The proposed method introduces solutions to longstanding issues in video generation, namely motion inconsistency and content degradation in long video sequences.\n2. The motion capture and dynamic content prior modules are well-designed, providing complementary functions to achieve both motion and content stability across video sequences.\n3. The paper includes qualitative and quantitative evaluations, as well as ablation studies, which validate the model's effectiveness over existing methods.\n4. The model is versatile, with components designed to be integrated into other video diffusion models, potentially broadening its applicability.\n5. Motion-Catcher consistently outperforms baseline models on standard metrics, such as MSE, SSIM, and temporal consistency, indicating its robustness."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The paper writing needs to be improved. It would be better to pay more attention to the meaning of this paper, including task definition, motivation of experiment design, notations in methodology, etc.\n2. This model needs some well-designed user studies since the motion consistency should be evaluated by humans.\n3. The model is designed to modify the generated video clip into a more consistent one. Please describe why not design modules to improve the motion consistency for the generated video clip the first time. I believe it is a hard task to find the anti-fact details and fix them.\n4. The dataset, AIGC (Fan et al., 2024), used in this paper is not well-known. It would be better to use some widely used datasets (e.g., MSCOCO, LAION-2B, UCF-101, Cityscapes).\n5. It would be better if the discussion in related works included methods for video generation with optical flow (e.g., [1r, 2r, 3r]).\n6. Minor:\n(1) Line 432: \"Table 1: compares Motion-Catcher against ...\" -> \"Comparison of Motion-Catcher against ...\"\n(2) Line 441, Line 485: \"Figure. 6\" -> \"Figure 6\". The latex code should be \"Figure~\\\\\\ref{xxx}.\"\n(3) Line 449: \"we\" -> \"We\"\n\n\n[1r] Liang, Feng, et al. \"Flowvid: Taming imperfect optical flows for consistent video-to-video synthesis.\" Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2024.\n\n[2r] Liang, Jingyun, et al. \"MoVideo: Motion-Aware Video Generation with Diffusion Model.\" European Conference on Computer Vision. Springer, Cham, 2024.\n\n[3r] Ni, Haomiao, et al. \"Conditional image-to-video generation with latent flow diffusion models.\" Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2023."}},"nonreaders":[],"tmdate":1731428549324,"tcdate":1731042381578,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission5993/Reviewer_m11L"],"signatures":["ICLR.cc/2025/Conference/Submission5993/Reviewer_m11L"],"forum":"6hsnpDXgHC","number":3,"license":"CC BY 4.0","cdate":1731042381578,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission5993/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428549324,"domain":"ICLR.cc/2025/Conference","replyto":"6hsnpDXgHC","id":"j5PgipgSuB","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Diffusion models","video generation"]},"supplementary_material":{"value":"/attachment/65fab445ce41feb8dc5bed46bae547780264f948.zip"},"primary_area":{"value":"generative models"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Recent developments in diffusion models have significantly advanced the field of video generation. However, technical challenges still exist in terms of spatiotemporal continuity and content consistency in long video generation. In this paper, we propose Motion-Catcher, a diffusion model-based method for multi-sequence video generation that aims to address the issues of motion inconsistency and content degradation. By incorporating a motion capture module, the model leverages optical flow information from video sequences to capture both local and global movements, enhancing the motion consistency of the videos. Furthermore, a dynamic content prior module is proposed to monitor regions prone to degradation, which helps maintain content consistency throughout the generated videos. Extensive experiments have validated that the proposed Motion-Catcher can generate videos with higher quality in terms of motion continuity and consistency. The source code and additional experimental results are available at https://github.com/YuukiGong/Motion-Catcher."},"_bibtex":{"value":"@misc{\ngong2024motioncatcher,\ntitle={Motion-Catcher: Upholding Motion and Content Consistency in Multi-Sequence Video Generation},\nauthor={Zhicheng Gong and Fangzhou Yi and Qi Zhou and HuiZeng},\nyear={2024},\nurl={https://openreview.net/forum?id=6hsnpDXgHC}\n}"},"title":{"value":"Motion-Catcher: Upholding Motion and Content Consistency in Multi-Sequence Video Generation"},"pdf":{"value":"/pdf/0d556ea1658400349abbd29af962d324cb3aefa7.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"gong|motioncatcher_upholding_motion_and_content_consistency_in_multisequence_video_generation"},"authorids":{"value":["~Zhicheng_Gong1","~Fangzhou_Yi1","~Qi_Zhou12","~HuiZeng1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Zhicheng Gong","Fangzhou Yi","Qi Zhou","HuiZeng"]}},"version":2},{"content":{"summary":{"value":"The paper describes an automatically generated dataset of 10k annotations applied to videos from existing video datasets. It shows that fine-tuning Qwen2.5-VL-7B-Instruct on this data by using a two-stage training strategy consisting of SFT followed by GRPO yields strong performance on various video understanding benchmarks. The automatically generated dataset is composed of two types of data: i) data with temporally assigned captions, and ii) data with global instead of temporal questions and answer pairs."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"Is the performance on Video-Holmes (but the question could apply similarly to the other benchmark results) based on the same test-set as the results on the official Leaderboard? Does the model described in this paper currently not appear there to retain anonymity of the submission and will appear it there after anonymity is lifted? \n\nI do not quite understand the statement in Line 320 onward: “For the in-domain evaluation, since the TutorialVQA (…) training set contains only 76 samples, we do not construct a corresponding test set. Instead, we derive held-out test sets from the five training datasets…” First, I do not understand how and why the limitation of TutorialVQA affects the choice of test-set selection for the other datasets. Can you elaborate? Second, I wonder whether the performance figures reported in the paper (for example, Table 1) are based on the identical train-test splits across all models or not. Can you please clarify?\n\nDo you expect the choice of source datasets, annotation scheme and training approach detailed in the paper to potentially degrade rather than improve performance on certain video-related tasks? Or do you expect these to be \"universally relevant\" to most if not all video-understanding benchmark tasks one can imagine? It would be nice to better understand the potential trade-offs and limitations besides the performance-improvements on existing benchmarks.\n\nHow were good values for the hyperparameters (such as beta, weight decay, data mix, etc.) determined?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"The paper introduces a carefully created dataset of annotations yielding strong performance results on several video understanding benchmarks. The code is (or will be) made publicly available."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The paper describes a data engineering approach to improving performance on a variety of video understanding benchmarks. While the performance appears to be strong overall, I do not find the paper particularly scientifically insightful or revealing. Specifically, I am not surprised that for the given set of video benchmark tasks (Video-Holmes, CG-Bench-Reasoning and VRBench), a careful selection of DeepSeek-R1-assisted and Gemini-assisted annotations on a careful selection of existing video datasets can improve the performance over the Qwen2.5-VL-7B-Instruct baseline and starting point. Importantly, I am a bit confused about some statements made in the paper (see questions below)."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762927239878,"tcdate":1761854403816,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission17303/Reviewer_Hmd7"],"signatures":["ICLR.cc/2026/Conference/Submission17303/Reviewer_Hmd7"],"forum":"ofNbGPV6Ve","number":2,"license":"CC BY 4.0","cdate":1761854403816,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission17303/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762927239878,"domain":"ICLR.cc/2026/Conference","replyto":"ofNbGPV6Ve","id":"IGBbc0yIhw","forumContent":{"TLDR":{"value":"We introduce Video-Thinker, a novel approach that empowers MLLMs to think with videos by autonomously leveraging their intrinsic grounding and captioning capabilities to generate reasoning clues throughout the inference process."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Video reasoning","Multimodal large language model","Thinking with videos"]},"primary_area":{"value":"foundation or frontier models, including LLMs"},"abstract":{"value":"Recent advances in image reasoning methods, particularly \"Thinking with Images\", have demonstrated remarkable success in Multimodal Large Language Models (MLLMs); however, this dynamic reasoning paradigm has not yet been extended to video reasoning tasks. In this paper, we propose Video-Thinker, which empowers MLLMs to think with videos by autonomously leveraging their intrinsic \"grounding\" and \"captioning\" capabilities to generate reasoning clues throughout the inference process. To spark this capability, we construct Video-Thinker-10K, a curated dataset featuring autonomous tool usage within chain-of-thought reasoning sequences. Our training strategy begins with Supervised Fine-Tuning (SFT) to learn the reasoning format, followed by Group Relative Policy Optimization (GRPO) to strengthen this reasoning capability. Through this approach, Video-Thinker enables MLLMs to autonomously navigate grounding and captioning tasks for video reasoning, eliminating the need for constructing and calling external tools. Extensive experiments demonstrate that Video-Thinker achieves significant performance gains on both in-domain tasks and challenging out-of-domain video reasoning benchmarks, including Video-Holmes, CG-Bench-Reasoning, and VRBench. Our Video-Thinker-7B substantially outperforms existing baselines such as Video-R1 and establishes state-of-the-art performance among 7B-sized MLLMs."},"_bibtex":{"value":"@misc{\nwang2026videothinker,\ntitle={Video-Thinker: Sparking ''Thinking with Videos'' via Reinforcement Learning},\nauthor={Shijian Wang and Jiarui Jin and Xingjian Wang and Linxin Song and Runhao Fu and Hecheng Wang and Zongyuan Ge and Yuan Lu and Xuelian Cheng},\nyear={2026},\nurl={https://openreview.net/forum?id=ofNbGPV6Ve}\n}"},"title":{"value":"Video-Thinker: Sparking \"Thinking with Videos\" via Reinforcement Learning"},"pdf":{"value":"/pdf/2434c327ac439ca9afeb57c45297d6f58cbecc77.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"wang|videothinker_sparking_thinking_with_videos_via_reinforcement_learning"},"authorids":{"value":["~Shijian_Wang1","~Jiarui_Jin1","~Xingjian_Wang2","~Linxin_Song1","~Runhao_Fu1","~Hecheng_Wang1","~Zongyuan_Ge1","~Yuan_Lu7","~Xuelian_Cheng2"]},"authors":{"value":["Shijian Wang","Jiarui Jin","Xingjian Wang","Linxin Song","Runhao Fu","Hecheng Wang","Zongyuan Ge","Yuan Lu","Xuelian Cheng"]}},"version":2},{"content":{"comment":{"value":"Dear Reviewer,\n\nThank you for your valuable and very nuanced feedback with pointers to recent relevant papers!\n\nWe will address each comment, as well as our respective changes to the paper and experiments. To make it easy to spot changes in the updated paper PDF, we highlight changes between red brackets “[[...]]”. \n\n### **Define & illustrate what “shortcut” means:**\nWe added a brief subsection “2.1 Shortcut solutions in video understanding” to the paper where we define shortcuts and common ways to detect them in our paper’s context, as well as provide a short illustrative figure:\n\n*What is classified as a shortcut depends on the skills a task is meant to test, e.g., whether a feature is \\textit{causal} or merely \\textit{spurious} \\citep{geirhos2020shortcut}; detecting and mitigating shortcut learning therefore remains an open problem \\citep{geirhos2020shortcut}.\nIn our setting, …*\n\n### **Additional task formats (captioning, localization, temporal ordering) and analyzing failures:**\n\nOther task formats, such as localization, have distinct advantages that may allow pinpointing the exact temporal or spatial failures (Cheng et.al.,V-STaR: Benchmarking Video-LLMs on Video Spatio-Temporal Reasoning, arXiv, 2025), however they do not allow the same ease of adoption and flexibility: applying a VideoLLM to such task (e.g. V-STaR) requires careful attention to the CoT prompts for different what/when/where sub-tasks, and involves a mix of more complex metrics such as “mean visual Intersection over Union.” Designing MVP as a multiple-choice QA task ensures a simple unified setup that is easy to adopt by the community. At the same time, multiple-choice QA is highly flexible since most tasks can be reformulated as QA; note that many videos we source examples from were not originally designed for QA, such as Language Table.\n\nOur work is complementary to V-STAR. We provide a framework for gathering minimal pairs and a concrete large-scale benchmark to address shortcuts. On the other hand,V-STaR does not leverage minimal change pairs, but instead decomposes questions into sub-steps e.g. what/when/where to build insight into models’ reasoning process. While V-STaR’s fine-grained grounding in spatial or temporal position helps locate where a model fails, it would not strictly control for spurious cues which is the focus of our work.  Finally, TUMTraffic-VideoQA is a valuable resource for advanced video understanding in complex/chaotic real-world scenarios. Orthogonal to this work, we primarily rely on sources that test more atomic “simpler” skills, that might perhaps not require chaining together as many tasks and reasoning steps but instead measure for a basic  understanding of intuitive physics and simple physical interactions.\n\nTo make all this more clear to the reader, we now describe the advantages and disadvantages of task formats to better motivate our decision in the paper at the beginning of Section 3.1 (“Task Definition”).\n\n### **Quantity and proportion of noisy or ambiguous dataset cases:**\n\nTo quantify the level of noise in our data, we had collected human accuracy on MVP reported in the main results Table 4. Additionally we now provide the explicit human accuracy for each of the 9 data sources in the appendix, and further analyze certain subsets in detail: We conduct a small manual annotation on the examples humans struggled with to quantify non-mutually exclusive answer options (a rare issues stemming from our automated pipeline). This is added at the end of Section 3. To summarize, we find that in some rare cases (less than 10% of all examples) two answer candidates can both be true at the same time, or otherwise that the question is ambiguous.\n\n### **More detailed analysis of where model fails:**\nSpecifically for the intuitive physics subsets, we provide a short discussion of recent literature in the updated paper (p. 9). We also expand our existing “Fine-grained failure analysis” section in Section 4.1 where we discuss various low subsets that stand out, and show some examples from one of the intuitive physics subsets (GRASP).\nFinally, while not explicitly requested by this reviewer, we also note that we added additional shortcut baselines to the main results Table 4 on MVP, in line with our shortcut analysis of existing benchmarks: the video-only and Socratic LLM baselines, as well as a third single-frame baseline.\n(For now these were run on MVP-mini only to speed up experiments but we will run on the full dataset for the final paper.)\n\n### **Illustrate diversity and phenomena in MVP benchmark with examples in appendix:**\n\nWe now added 18 examples in the appendix, two samples for each of the nine sources we use. For STAR examples we covered faces for privacy reasons.\n\nWe have addressed your main concerns and suggestions, and have updated the paper to reflect these changes. We are happy to discuss further and address any remaining questions in the following days."},"title":{"value":"Addressing main suggestions such as more detailed failure analysis"}},"parentInvitations":"TMLR/-/Official_Comment","tmdate":1758579615203,"tcdate":1758579615203,"writers":["TMLR","TMLR/Paper5382/Authors"],"signatures":["TMLR/Paper5382/Authors"],"forum":"gvFgNJcSw1","number":10,"license":"CC BY 4.0","cdate":1758579615203,"readers":["everyone"],"invitations":["TMLR/Paper5382/-/Official_Comment"],"mdate":1758579615203,"domain":"TMLR","replyto":"N4iJxQjoYj","id":"vpoqlnOrOq","forumContent":{"submission_length":{"value":"Regular submission (no more than 12 pages of main content)"},"venue":{"value":"Accepted by TMLR"},"abstract":{"value":"Existing benchmarks for assessing the spatio-temporal understanding and reasoning abilities of video language models are susceptible to score inflation due to the presence of shortcut solutions based on superficial visual or textual cues. This paper mitigates the challenges in accurately assessing model performance by introducing the Minimal Video Pairs (MVP) benchmark, a simple shortcut-aware video QA benchmark for assessing the physical understanding of video language models. The benchmark is comprised of 55K high-quality multiple-choice video QA examples focusing on physical world understanding. Examples are curated from nine video data sources, spanning first-person egocentric and exocentric videos, robotic interaction data, and cognitive science intuitive physics benchmarks. To mitigate shortcut solutions that rely on superficial visual or textual cues and biases, each sample in MVP has a minimal-change pair — a visually similar video accompanied by an identical question but an opposing answer. To answer a question correctly, a model must provide correct answers for both examples in the minimal-change pair; as such, models that solely rely on visual or textual biases would achieve below random performance. Human performance on MVP is 92.9%, while the best open-source state-of-the- art video-language model achieves 40.2% compared to random performance at 25%."},"_bibtex":{"value":"@article{\nkrojer2025a,\ntitle={A Shortcut-aware Video-{QA} Benchmark for Physical Understanding via Minimal Video Pairs},\nauthor={Benno Krojer and Mojtaba Komeili and Candace Ross and Quentin Garrido and Koustuv Sinha and Nicolas Ballas and Mido Assran},\njournal={Transactions on Machine Learning Research},\nissn={2835-8856},\nyear={2025},\nurl={https://openreview.net/forum?id=gvFgNJcSw1},\nnote={}\n}"},"title":{"value":"A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs"},"pdf":{"value":"/pdf/f01dd06430bcd7b2aba068243d47011dcc96be6e.pdf"},"venueid":{"value":"TMLR"},"paperhash":{"value":"krojer|a_shortcutaware_videoqa_benchmark_for_physical_understanding_via_minimal_video_pairs"},"authorids":{"value":["~Benno_Krojer1","~Mojtaba_Komeili1","~Candace_Ross1","~Quentin_Garrido1","~Koustuv_Sinha1","~Nicolas_Ballas1","~Mido_Assran1"]},"assigned_action_editor":{"value":"~Tae-Hyun_Oh3"},"authors":{"value":["Benno Krojer","Mojtaba Komeili","Candace Ross","Quentin Garrido","Koustuv Sinha","Nicolas Ballas","Mido Assran"]}},"version":2},{"content":{"summary":{"value":"The authors propose to study whether current large multimodal models entail adequate reasoning capabilities for even short duration videos (in comparison to growing focus on long form videos) with counterfactual caption pairs . Accordingly, authors introduce a benchmark named Vinoground (similar in spirit to the Winoground benchmark) which comprises 500 counterfactual text pairs with corresponding videos of max 10 seconds. GPT is used to generate counterfactual captions of different categories, and matching videos are then obtained from VATEX dataset and similarity matching. Human annotators (the authors themselves) perform quality control to check if caption is a good match for the video. For evaluation, a 'text score' and 'video score' to test temporal reasoning capabilities across different input forms, and a cumulative 'group score' metric."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"Please see weaknesses."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. The paper studies an interesting problem on whether modern LMMs are capable at temporal reasoning capabilities within even short videos (where models can't suffer from long context capabilities)\n2. The methodology regarding dataset construction and evaluation is mostly adequately described.\n3. Results show that models perform fairly poorly on the benchmark and lag behind humans on the composite score. A wide number of models are considered and human performances are also noted for comparison."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. There have been multiple previous works that have explored creation of temporal counterfactual question-video pairs to assess basic reasoning capabilities of models. These include \"Test of Time: Instilling Video-Language Models with a Sense of Time\" [1] and CLAVI (Complements in Language and Video) [2] which have largely done similar work in creating complementary/counterfactual pairs of text and video, and demonstrating how video-language models show inherent biases and weak performances on simple temporal reasoning scenarios. These works are not mentioned and should be differentiated from. Currently, the novelty of this work seems to be limited and it seems an extension of these works evaluated for more recent LMM models. \n\n2. A more detailed analysis into limitations of current models on constructed benchmark can be useful. This could help differentiate the work further from previous similar benchmarks, authors could consider doing a more detailed error analysis and study into how different components of open source models (such as Qwen or InternVL) impact the model's reasoning ability. E.g. some questions to explore could be: \n- do limitations or biases arise from the backbone LLM of these models or is it the visual backbone model? \n- does the way attention is performed b/w image and video inputs (e.g. cross attention vs concatenated attention) influence results?\n- does chain-of-thought or other prompting strategies impact performances, and is produced chain of thought sensical/reasonable?\n\n(To make space for point 2, less important details on experimental setup and dataset construction can be moved to appendix)\n\n\n[1] Bagad, Piyush, Makarand Tapaswi, and Cees GM Snoek. \"Test of time: Instilling video-language models with a sense of time.\" Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2023.\n\n[2] Rawal, Ishaan Singh, et al. \"Dissecting Multimodality in VideoQA Transformer Models by Impairing Modality Fusion.\" International Conference on Machine Learning. PMLR, 2024."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762931663379,"tcdate":1762075607635,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission19810/Reviewer_jeu5"],"signatures":["ICLR.cc/2026/Conference/Submission19810/Reviewer_jeu5"],"forum":"WAk6tf8VkQ","number":4,"license":"CC BY 4.0","cdate":1762075607635,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission19810/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762931663379,"domain":"ICLR.cc/2026/Conference","replyto":"WAk6tf8VkQ","id":"FKyEZGzZYr","forumContent":{"TLDR":{"value":"Modern SoTA LMMs still demonstrates subpar performance at temporal reasoning with our temporal counterfactual benchmark composed of natural videos."},"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["short videos","temporal understandings","large multimodal models"]},"supplementary_material":{"value":"/attachment/3d0bca0c56260ea829d0358e4f2b058110d67562.zip"},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"There has been growing sentiment recently that modern large multimodal models (LMMs) have addressed most of the key challenges related to short video comprehension. As a result, both academia and industry are gradually shifting their attention towards the more complex challenges posed by understanding long-form videos. However, is this really the case? Our studies indicate that LMMs still lack many fundamental reasoning capabilities even when dealing with short videos. We introduce Vinoground, a temporal counterfactual LMM evaluation benchmark encompassing 1000 short and natural video-caption pairs. We demonstrate that existing LMMs severely struggle to distinguish temporal differences between different actions and object transformations. For example, the best model o3 only obtains ~50% on our group score metric, showing a large gap compared to the human baseline of ~90%. All open-source multimodal models and CLIP-based models perform much worse, producing mostly random chance performance. Through this work, we shed light onto the fact that temporal reasoning in short videos is a problem yet to be fully solved. We will publicly share our benchmark."},"_bibtex":{"value":"@misc{\nzhang2025vinoground,\ntitle={Vinoground: Today{\\textquoteright}s {LMM}s Don{\\textquoteright}t Understand Short Counterfactual Videos},\nauthor={Jianrui Zhang and Mu Cai and Yong Jae Lee},\nyear={2025},\nurl={https://openreview.net/forum?id=WAk6tf8VkQ}\n}"},"title":{"value":"Vinoground: Today’s LMMs Don’t Understand Short Counterfactual Videos"},"pdf":{"value":"/pdf/7f9aa6ef8c93c10b4bcd6f3da82713e3f5391479.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"zhang|vinoground_todays_lmms_dont_understand_short_counterfactual_videos"},"authorids":{"value":["~Jianrui_Zhang1","~Mu_Cai1","~Yong_Jae_Lee2"]},"authors":{"value":["Jianrui Zhang","Mu Cai","Yong Jae Lee"]}},"version":2},{"content":{"pdf":{"value":"/pdf/e27fcc4d77a797a8ad5545cbf59a7eccd3da45db.pdf"},"venueid":{"value":"OpenReview.net/Archive"},"paperhash":{"value":"liu|investigating_and_defending_shortcut_learning_in_personalized_diffusion_models"},"authorids":{"value":["~Yixin_Liu4","~Ruoxi_Chen1","~Lichao_Sun1"]},"html":{"value":"https://arxiv.org/pdf/2406.18944"},"abstract":{"value":"Personalized diffusion models have gained popularity for adapting pre-trained\ntext-to-image models to generate images of specific topics with minimal training\ndata. However, these models are vulnerable to minor adversarial perturbations,\nleading to degraded performance on corrupted datasets. Such vulnerabilities are\nfurther exploited to craft protective perturbations on sensitive images like portraits\nthat prevent unauthorized generation. In response, diffusion-based purification\nmethods have been proposed to remove these perturbations and retain generation\nperformance. However, existing works lack detailed analysis of the fundamental\nshortcut learning vulnerability of personalized diffusion models and also turn to\nover-purifying the images, which causes information loss. In this paper, we take a\ncloser look at the fine-tuning process of personalized diffusion models through the\nlens of shortcut learning. And we propose a hypothesis explaining the manipulation\nmechanisms of existing perturbation methods, demonstrating that perturbed images\nsignificantly deviate from their original prompts in the CLIP-based latent space.\nThis misalignment during fine-tuning causes models to associate noisy patterns\nwith identifiers, resulting in performance degradation. Based on these insights,\nwe introduce a systematic approach to maintain training performance through\npurification. Our method first purifies the images to realign them with their original\nsemantic meanings in latent space. Then, we introduce contrastive learning with\nnegative tokens to decouple the learning of clean identities from noisy patterns,\nwhich shows a strong potential capacity against adaptive perturbation. Our study\nuncovers shortcut learning vulnerabilities in personalized diffusion models and\nprovides a firm evaluation framework for future protective perturbation research.\nCode is available at https://github.com/liuyixin-louis/DiffShortcut."},"title":{"value":"Investigating and Defending Shortcut Learning in Personalized Diffusion Models"},"authors":{"value":["Yixin Liu","Ruoxi Chen","Lichao Sun"]}},"tmdate":1724390717161,"pdate":1723094705473,"tcdate":1724390717161,"writers":["~Yixin_Liu4","~Ruoxi_Chen1","~Lichao_Sun1"],"signatures":["~Yixin_Liu4"],"forum":"ddZFJQvYBC","license":"CC BY 4.0","number":27847,"cdate":1724390717161,"readers":["everyone"],"invitations":["OpenReview.net/Archive/-/Direct_Upload"],"mdate":1724390717161,"domain":"OpenReview.net/Archive","id":"ddZFJQvYBC","version":2},{"content":{"summary":{"value":"This paper proposes ApoAvatar, a DiT-based framework for audio-driven video generation, aiming to address the common issue in existing methods where body motions are poorly synchronized with speech rhythm, resulting in stiff and unnatural animations. ApoAvatar introduces two key innovations: 1) an Audio-Pose Prior Refocusing mechanism that dynamically modulates the strength of pose guidance based on frame-level audio intensity—strong accents amplify gesture magnitude, while quiet segments suppress unnecessary motion, thus aligning gesture dynamics with speaking style. 2) a Frame-Wise Audio–Video Interaction module that employs a bidirectional cross-attention within an Audio DiT adapter, enabling audio features to be refined using the current visual context and the refocused pose prior, producing \"pose-aware\" audio embeddings. The framework supports unified inference with or without pose input. Experiments on the EMTD and HDTF datasets demonstrate that ApoAvatar outperforms baselines in lip-audio synchronization, gesture expressiveness, motion naturalness, and identity preservation."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"* Regarding the training procedure: Could the authors clarify the overall training pipeline and detailed training configurations? Specifically, how are the text labels in the training data obtained, and how is the influence of text conditioning on pose generation balanced?\n* Regarding comparisons: Are the text inputs kept consistent across all compared methods? Additionally, could the authors provide more generation results via an anonymous link to better support the claims in the paper?\n* Role of pose conditioning: When pose guidance is available, how does the proposed method compare in performance against recent pose-driven approaches such as RealisDance-DiT[1] and X-UniMotion[2]?\n* Evolution of the pose branch: Omni-Human1.5[3] deliberately removed pose conditioning in its upgrade from Omni-Human1[4], as challenges in body motion generation—such as body turning, finger articulation, and dance movements—remain difficult. In contrast, this work weakens the role of the text branch and reintroduces the pose branch. What deeper insights or design considerations motivate this architectural choice?\n\n[1]Zhou, Jingkai, et al. \"RealisDance-DiT: Simple yet Strong Baseline towards Controllable Character Animation in the Wild.\" arXiv preprint arXiv:2504.14977 (2025).\n\n[2]Song, Guoxian, et al. \"X-UniMotion: Animating Human Images with Expressive, Unified and Identity-Agnostic Motion Latents.\" arXiv preprint arXiv:2508.09383 (2025).\n\n[3]Jiang, Jianwen, et al. \"Omnihuman-1.5: Instilling an active mind in avatars via cognitive simulation.\" arXiv preprint arXiv:2508.19209 (2025).\n\n[4]Lin, Gaojie, et al. \"Omnihuman-1: Rethinking the scaling-up of one-stage conditioned human animation models.\" arXiv preprint arXiv:2502.01061 (2025)."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"* This paper identifies a critical yet often overlooked limitation in most existing works: audio and pose features are typically modeled as independently contributing modalities with insufficient interaction. To address this issue, the authors propose well-designed solutions, whose effectiveness is thoroughly validated through ablation studies. This work offers the community a novel and valuable perspective for optimizing audio-driven video generation.\n* This work introduces an audio-aware pose prior refocusing mechanism that acts as a pose feature refiner. By dynamically modulating the retention level of pose embeddings based on prosodic cues, the method naturally scales gesture amplitude according to speech intonation. It amplifies movements during accented syllables and suppresses redundant motions during silent segments, thereby generating more expressive and rhythmically coherent full-body animations.\n* Similarly, the paper designs a frame-wise audio-video interaction architecture that serves as an audio feature refiner. At each denoising step, the model updates audio representations by conditioning on both the current video state and the refocused pose prior, transforming audio features from static inputs into dynamically evolving, \"video-aware\" signals. This significantly enhances short-term synchronization and motion smoothness."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"* The experimental setup is insufficiently described: Section 4.1 lacks essential details regarding the experimental configuration, including—but not limited to—the number of training stages, and whether full-parameter fine-tuning, LoRA, or SFT was used for parameter updates.\n* Insufficient discussion on the text modality: The backbone model used in this work is an image-to-video (I2V) model, in which the text input inherently provides strong conditioning. However, the paper does not adequately discuss the role of the text branch, such as how the text labels in the training data are obtained or what specific text inputs are used during comparison with other methods. Notably, methods like MultiTalk[1] and OmniAvatar[2] do not use pose conditioning; hence, when text descriptions are insufficient, their generated body motions may naturally be less expressive. This oversight may unintentionally give the proposed method an unfair advantage in comparative evaluation.\n* Qualitative comparison is inadequate: The paper only provides nine qualitative cases in the appendix, which do not even include the example shown in Figure 4. Moreover, while OmniAvatar[2] is included in quantitative comparisons, it is missing from the qualitative results. These issues raise concerns that the presented results may be cherry-picked, failing to fully demonstrate the effectiveness and superiority of the proposed approach.\n* Concerns regarding reproducibility: The work builds upon Goku[3], a foundational model that is neither open-sourced nor commercially available. Despite initial promises by the Goku[3] authors to release it, the model remains inaccessible to the public as of now. This severely undermines the reproducibility of the work, as readers cannot verify whether the generated results stem from the base model's capabilities or from the contributions introduced in this paper. The experimental results would be significantly more convincing if the proposed method were implemented and evaluated on an open-source foundation model such as Wan[4], which is already used by several compared methods (e.g., FantasyTalking[5], MultiTalk[1], InfiniteTalk[6], OmniAvatar[2]).\n\n[1]Kong, Zhe, et al. \"Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation.\" arXiv preprint arXiv:2505.22647 (2025).\n\n[2]Gan, Qijun, et al. \"OmniAvatar: Efficient Audio-Driven Avatar Video Generation with Adaptive Body Animation.\" arXiv preprint arXiv:2506.18866 (2025).\n\n[3]Chen, Shoufa, et al. \"Goku: Flow based video generative foundation models.\" Proceedings of the Computer Vision and Pattern Recognition Conference. 2025.\n\n[4]Wan, Team, et al. \"Wan: Open and advanced large-scale video generative models.\" arXiv preprint arXiv:2503.20314 (2025).\n\n[5]Wang, Mengchao, et al. \"Fantasytalking: Realistic talking portrait generation via coherent motion synthesis.\" arXiv preprint arXiv:2504.04842 (2025).\n\n[6]Yang, Shaoshu, et al. \"InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing.\" arXiv preprint arXiv:2508.14033 (2025)."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1764360165382,"tcdate":1760948203300,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission5696/Reviewer_DSoK"],"signatures":["ICLR.cc/2026/Conference/Submission5696/Reviewer_DSoK"],"forum":"9fw3g2jFbc","number":1,"license":"CC BY 4.0","cdate":1760948203300,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission5696/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1764360165382,"domain":"ICLR.cc/2026/Conference","replyto":"9fw3g2jFbc","id":"tvR2aA9DB1","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Video Generation","Audio Driven Avatar Animation"]},"supplementary_material":{"value":"/attachment/7a249a2bb77cd0a8632652a68082b1b883bfec30.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Audio-driven human video generation has greatly improved lip synchronization. However, most methods still use audio mainly to control the mouth, while the relationship between speech rhythm and body motion remains weak. This often makes generated characters look unnatural. We present \\textbf{ApoAvatar}, a diffusion-based framework that ties speaking style to motion dynamics. We introduce an Audio–Pose Prior Refocusing mechanism, which adjusts pose guidance based on audio intensity. Strong accents increase gesture magnitude, while quiet parts suppress unnecessary motion. We also design a frame-wise audio–video interaction module. It updates audio features using the current visual context and the refocused pose prior through a designed bidirectional cross-attention. This yields better short-term synchronization and motion coherence. The framework supports both pose-controlled and pose-free inference within one model. Extensive experiments on EMTD and HDTF show clear gains over strong baselines in lip–audio synchronization, gesture expressiveness, and overall motion naturalness."},"_bibtex":{"value":"@misc{\nlin2026apoavatar,\ntitle={ApoAvatar: Expressive Audio-Driven Avatar Generation via Refocused Audio-Pose Priors},\nauthor={Jingyu Lin and Chao Zhang and Wei Feng and Donghao Zhou and Shilei Wen and Lan Du and Cunjian Chen},\nyear={2026},\nurl={https://openreview.net/forum?id=9fw3g2jFbc}\n}"},"title":{"value":"ApoAvatar: Expressive Audio-Driven Avatar Generation via Refocused Audio-Pose Priors"},"pdf":{"value":"/pdf/62e5ea3fe5038f933137ef04d5192bb4d49728cb.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"lin|apoavatar_expressive_audiodriven_avatar_generation_via_refocused_audiopose_priors"},"authorids":{"value":["~Jingyu_Lin5","~Chao_Zhang40","~Wei_Feng18","~Donghao_Zhou1","~Shilei_Wen1","~Lan_Du1","~Cunjian_Chen2"]},"authors":{"value":["Jingyu Lin","Chao Zhang","Wei Feng","Donghao Zhou","Shilei Wen","Lan Du","Cunjian Chen"]}},"version":2},{"content":{"summary":{"value":"This paper introduces CamPilot, a video diffusion framework designed to improve camera controllability through a reward feedback learning strategy. The authors propose a camera-aware 3D decoder that decodes video latents into 3D Gaussian representations to evaluate camera-video alignment efficiently. The model achieves better camera control and 3D consistency compared to existing methods on the RealEstate10K and WorldScore benchmarks."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"See WeakNess"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":1},"contribution":{"value":3},"strengths":{"value":"1. Using 3DGS to enhance camera-guided video generation is a good starting point"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The writing quality of this paper needs improvement. Many expressions are verbose and not concise enough. For example, Sections 2.2 and 2.3 contain multiple repetitions of earlier content, resulting in unnecessary length. Moreover, the core comparison experiment showing how much computational cost is reduced compared to VAE decoding is placed only in Appendix A.4. From the main text alone, the experimental details are unclear.\n\n2. The baseline methods used for comparison are somewhat outdated. The paper lacks comparisons with several recent and highly relevant approaches, such as CamI2V [1], OmniCam [2], RealCam-I2V [3], and ReCamMaster [4].\n\n3. The proposed method employs a multi-stage reward scoring strategy for evaluating camera trajectories. Although the authors claim that this method is efficient, I am curious whether it could lead to error accumulation across stages. Compared with approaches that estimate camera trajectories from decoded videos (e.g., via VAE decoding) and directly compare them to ground-truth trajectories, how large is the accuracy gap between these two strategies?\n\n4. Minor revision suggestions:  Please correctly use the `\\citep` command for references, e.g., Line 208 for *UniFL*.\n\n[1] CamI2V: Camera-Controlled Image-to-Video Diffusion Model\n\n[2] OmniCam: Unified Multimodal Video Generation via Camera Control\n\n[3] RealCam-I2V: Real-World Image-to-Video Generation with Interactive Complex Camera Control\n\n[4] ReCamMaster: Camera-Controlled Generative Rendering from A Single Video"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762916483037,"tcdate":1760694738139,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission2989/Reviewer_n4a2"],"signatures":["ICLR.cc/2026/Conference/Submission2989/Reviewer_n4a2"],"forum":"qqij8fCGDl","number":2,"license":"CC BY 4.0","cdate":1760694738139,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission2989/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762916483037,"domain":"ICLR.cc/2026/Conference","replyto":"qqij8fCGDl","id":"j4EY5TFsco","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Video generation; 3D scene exploration; Reward feedback learning"]},"supplementary_material":{"value":"/attachment/190335a20c10687ec00345e0bee18e333335cb61.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Recent advancements in camera-controlled video diffusion models have significantly improved video-camera alignment and enabled more accurate 3D scene generation, driven by potential downstream applications such as virtual reality.\nHowever, we reveal that existing approaches often struggle to precisely adhere to the given camera conditions, leading to inconsistencies in the 3D geometry.\nInspired by Reward Feedback Learning in diffusion models, which has demonstrated strong potential in aligning model outputs with task-specific objectives, we build upon this paradigm and aim to further improve camera controllability.\nDirectly borrowing existing ReFL approaches faces several challenges. First, current reward models lack the capacity to assess video-camera alignment. Second, decoding latent into RGB videos for reward computation introduces substantial computational overhead. Third, 3D geometric information is typically neglected during video decoding.\nTo address these limitations, we introduce a camera-aware 3D decoder that efficiently decodes video latent into 3D representations for reward computation. Specifically, we project the video latent and camera pose into 3D Gaussians, which supports efficient rendering from arbitrary views. \nIn this process, the camera pose not only acts as an input variable but also serves as a projection parameter for determining the mean of each Gaussian.\nIf the generated video does not match the camera conditions, the 3D structure becomes geometrically inconsistent, leading to blurry rendered images.\nBased on this property, we explicitly optimizing pixel-level consistency between rendered novel views and ground-truth ones as reward feedback.\nTo accommodate the stochastic nature, we further introduce a visibility term that selectively supervises only deterministic regions derived via geometric warping.\nExtensive experiments conducted on the RealEstate10K and WorldScore benchmarks demonstrate the effectiveness of our proposed method in enhancing both camera controllability and generation quality."},"_bibtex":{"value":"@misc{\nge2025campilot,\ntitle={CamPilot: Improving Camera Control in Video Diffusion Model with Efficient Camera Reward Feedback},\nauthor={Wenhang Ge and Guibao Shen and Jiawei Feng and Luozhou Wang and Hao LU and Xingye Tian and Xin Tao and Pengfei Wan and Ying-Cong Chen},\nyear={2025},\nurl={https://openreview.net/forum?id=qqij8fCGDl}\n}"},"title":{"value":"CamPilot: Improving Camera Control in Video Diffusion Model with Efficient Camera Reward Feedback"},"pdf":{"value":"/pdf/3919cc68511ac2752505bc0f7b4e7f570a2cd725.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"ge|campilot_improving_camera_control_in_video_diffusion_model_with_efficient_camera_reward_feedback"},"authorids":{"value":["~Wenhang_Ge1","~Guibao_Shen1","~Jiawei_Feng1","~Luozhou_Wang2","~Hao_LU8","~Xingye_Tian2","~Xin_Tao3","~Pengfei_Wan1","~Ying-Cong_Chen1"]},"authors":{"value":["Wenhang Ge","Guibao Shen","Jiawei Feng","Luozhou Wang","Hao LU","Xingye Tian","Xin Tao","Pengfei Wan","Ying-Cong Chen"]}},"version":2},{"content":{"summary":{"value":"This paper proposes the Visionary-R1 framework, which tackles the shortcut learning problem in vision-language models through reinforcement learning. The core innovation lies in introducing a “Describe–Reason–Answer” output format, forcing the model to first generate a detailed visual description before performing reasoning. Trained on only 273K visual question-answer pairs without reasoning-chain annotations, the model outperforms large commercial systems such as GPT-4o and Claude 3.5 on benchmarks like MathVista."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"see the weaknesses"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"1. By enforcing the “Describe–Reason–Answer” output format, the model must first produce a detailed image description before reasoning. This clever design ensures the model deeply understands image content rather than relying on superficial patterns.\n\n2. The dataset is broad, the evaluation benchmarks are comprehensive, and the paper provides extensive ablation studies and hyperparameter analyses, which strengthen the credibility of the conclusions."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The proposed approach still relies on a VLM to generate reasoning chains during training. How is this fundamentally different from methods that extract reasoning chains from proprietary models such as GPT-4o?\n\n2. The discussion of related work is incomplete — the paper does not adequately cover other recent approaches to avoiding shortcut learning (e.g., causal intervention or disentangled representation learning) and lacks a thorough comparison with similar “intermediate supervision” methods (e.g., generating visual descriptions before reasoning).\n\n3. Demonstrating the shortcut problem in GRPO only through the qualitative analysis in Figure 1 is insufficient. The authors should provide quantitative evidence, such as blind experiments, to confirm the existence of the issue."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762925261617,"tcdate":1761812074508,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission14918/Reviewer_tjFT"],"signatures":["ICLR.cc/2026/Conference/Submission14918/Reviewer_tjFT"],"forum":"bya3KOdLeS","number":3,"license":"CC BY 4.0","cdate":1761812074508,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission14918/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762925261617,"domain":"ICLR.cc/2026/Conference","replyto":"bya3KOdLeS","id":"cn3fAfHErf","forumContent":{"TLDR":{"value":"We addressed a shortcut problem in applying reinforcement learning to VLMs with Visionary-R1, trained on 273K CoT-free visual question-answer pairs  using only reinforcement learning."},"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Large Vision Language Model","Visual Reasoning","Reinforcement Learning"]},"primary_area":{"value":"foundation or frontier models, including LLMs"},"abstract":{"value":"Learning general-purpose reasoning capabilities has long been a challenging problem in AI. Recent research in large language models (LLMs), such as DeepSeek-R1, has shown that reinforcement learning techniques like GRPO can enable pre-trained LLMs to develop reasoning capabilities using simple question-answer pairs. In this paper, we aim to train visual language models (VLMs) to perform reasoning on image data through reinforcement learning and visual question-answer pairs, without any explicit chain-of-thought (CoT) supervision. Our findings indicate that simply applying reinforcement learning to a VLM---by prompting the model to produce a reasoning chain before providing an answer---can lead the model to develop shortcuts from easy questions, thereby reducing its ability to generalize across unseen data distributions. We argue that the key to mitigating shortcut learning is to encourage the model to interpret images prior to reasoning. Therefore, we train the model to adhere to a caption-reason-answer output format: initially generating a detailed caption for an image, followed by constructing an extensive reasoning chain. When trained on 273K CoT-free visual question-answer pairs and using only reinforcement learning, our model, named Visionary-R1, outperforms strong multimodal models, such as GPT-4o, Claude3.5-Sonnet, and Gemini-1.5-Pro, on multiple visual reasoning benchmarks. Code and models will be publicly released."},"_bibtex":{"value":"@misc{\nxia2025visionaryr,\ntitle={Visionary-R1: Mitigating Shortcuts in Visual Reasoning with Reinforcement Learning},\nauthor={Jiaer Xia and Yuhang Zang and Peng Gao and Yixuan Li and Kaiyang Zhou},\nyear={2025},\nurl={https://openreview.net/forum?id=bya3KOdLeS}\n}"},"title":{"value":"Visionary-R1: Mitigating Shortcuts in Visual Reasoning with Reinforcement Learning"},"pdf":{"value":"/pdf/3dd28e833f99b895e121c749459d21c527822646.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"xia|visionaryr1_mitigating_shortcuts_in_visual_reasoning_with_reinforcement_learning"},"authorids":{"value":["~Jiaer_Xia1","~Yuhang_Zang1","~Peng_Gao3","~Yixuan_Li1","~Kaiyang_Zhou1"]},"authors":{"value":["Jiaer Xia","Yuhang Zang","Peng Gao","Yixuan Li","Kaiyang Zhou"]}},"version":2},{"content":{"venue":{"value":"CoRR 2024"},"pdf":{"value":"http://arxiv.org/pdf/2403.09488v3"},"venueid":{"value":"dblp.org/journals/CORR/2024"},"paperhash":{"value":"jang|rectifying_demonstration_shortcut_in_incontext_learning"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Joonwon_Jang:","~Sanghwan_Jang1","~Wonbin_Kweon2","~Minjin_Jeon1","~Hwanjo_Yu1"]},"html":{"value":"https://doi.org/10.48550/arXiv.2403.09488"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2403-09488,\n  publtype={informal},\n  author={Joonwon Jang and Sanghwan Jang and Wonbin Kweon and Minjin Jeon and Hwanjo Yu},\n  title={Rectifying Demonstration Shortcut in In-Context Learning},\n  year={2024},\n  cdate={1704067200000},\n  journal={CoRR},\n  volume={abs/2403.09488},\n  url={https://doi.org/10.48550/arXiv.2403.09488}\n}\n"},"abstract":{"value":"Large language models (LLMs) are able to solve various tasks with only a few demonstrations utilizing their in-context learning (ICL) abilities. However, LLMs often rely on their pre-trained semantic priors of demonstrations rather than on the input-label relationships to proceed with ICL prediction. In this work, we term this phenomenon as the 'Demonstration Shortcut'. While previous works have primarily focused on improving ICL prediction results for predefined tasks, we aim to rectify the Demonstration Shortcut, thereby enabling the LLM to effectively learn new input-label relationships from demonstrations. To achieve this, we introduce In-Context Calibration, a demonstration-aware calibration method. We evaluate the effectiveness of the proposed method in two settings: (1) the Original ICL Task using the standard label space and (2) the Task Learning setting, where the label space is replaced with semantically unrelated tokens. In both settings, In-Context Calibration demonstrates substantial improvements, with results generalized across three LLM families (OPT, GPT, and Llama2) under various configurations."},"title":{"value":"Rectifying Demonstration Shortcut in In-Context Learning"},"authors":{"value":["Joonwon Jang","Sanghwan Jang","Wonbin Kweon","Minjin Jeon","Hwanjo Yu"]}},"tmdate":1768447326744,"pdate":1704067200000,"tcdate":1718675610467,"writers":["~"],"signatures":["~Hwanjo_Yu1"],"forum":"i9lkREhuhP","license":"CC BY-SA 4.0","number":31911,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1768447326744,"domain":"DBLP.org","id":"i9lkREhuhP","version":2},{"content":{"venue":{"value":"NAACL-HLT 2024"},"pdf":{"value":"https://aclanthology.org/2024.naacl-long.242.pdf"},"venueid":{"value":"dblp.org/conf/NAACL/2024"},"paperhash":{"value":"jang|rectifying_demonstration_shortcut_in_incontext_learning"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Joonwon_Jang:","~Sanghwan_Jang1","~Wonbin_Kweon2","~Minjin_Jeon1","~Hwanjo_Yu1"]},"html":{"value":"https://doi.org/10.18653/v1/2024.naacl-long.242"},"_bibtex":{"value":"@inproceedings{DBLP:conf/naacl/JangJKJY24,\n  author={Joonwon Jang and Sanghwan Jang and Wonbin Kweon and Minjin Jeon and Hwanjo Yu},\n  title={Rectifying Demonstration Shortcut in In-Context Learning},\n  year={2024},\n  cdate={1704067200000},\n  pages={4294-4321},\n  url={https://doi.org/10.18653/v1/2024.naacl-long.242},\n  booktitle={NAACL-HLT},\n  crossref={conf/naacl/2024}\n}\n"},"abstract":{"value":"Large language models (LLMs) are able to solve various tasks with only a few demonstrations utilizing their in-context learning (ICL) abilities.However, LLMs often rely on their pre-trained semantic priors of demonstrations rather than on the input-label relationships to proceed with ICL prediction. In this work, we term this phenomenon as the ‘Demonstration Shortcut’.While previous works have primarily focused on improving ICL prediction results for predefined tasks, we aim to rectify the Demonstration Shortcut, thereby enabling the LLM to effectively learn new input-label relationships from demonstrations.To achieve this, we introduce In-Context Calibration, a demonstration-aware calibration method.We evaluate the effectiveness of the proposed method in two settings: (1) the Original ICL Task using the standard label space and (2) the Task Learning setting, where the label space is replaced with semantically unrelated tokens.In both settings, In-Context Calibration demonstrates substantial improvements, with results generalized across three LLM families (OPT, GPT, and Llama2) under various configurations."},"title":{"value":"Rectifying Demonstration Shortcut in In-Context Learning"},"authors":{"value":["Joonwon Jang","Sanghwan Jang","Wonbin Kweon","Minjin Jeon","Hwanjo Yu"]}},"tmdate":1768447326710,"pdate":1704067200000,"tcdate":1747308011920,"writers":["~"],"signatures":["~Hwanjo_Yu1"],"forum":"O7GW4h9jru","license":"CC BY-SA 4.0","number":483120,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1768447326710,"domain":"DBLP.org","id":"O7GW4h9jru","version":2},{"content":{"summary":{"value":"The paper studies two-hop reasoning with distractors. Empirically, LLMs (e.g., OLMo-7B) guess uniformly over candidate ends when multiple two-hop chains co-occur; brief fine-tuning yields near-perfect accuracy and length generalization. \nTo explain this, the authors train and reverse-engineer a 3-layer attention-only Transformer on a symbolic task, observing a phase transition from random guessing to a sequential query mechanism: layer-1 copies parent→child, layer-2 retrieves the bridge tied to the source, layer-3 aligns the query to the correct end. They also sketch a minimal 3-parameter analytical model capturing the loss/attention dynamics."},"soundness":{"value":1},"confidence":{"value":4},"questions":{"value":"none"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":1},"strengths":{"value":"- The sequential-query circuit is intuitively appealing and matches observed attention maps, which is not surprising. \n\n- Figures and explanations are pedagogically strong, easy to follow, even for OOD readers."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The dataset itself introduces a shortcut bias that undermines the very reasoning claims. In other words, a synthetic symbolic dataset is too clean and shortcut-prone. Such design encourages shortcut learning (memorizing surface token positions or n-gram spans) rather than genuine compositional reasoning.\n\n- Strong mechanistic write-up and neat minimal model, but conceptual novelty is limited relative to existing multi-hop analyses and depth theory; external validity beyond symbolics is unclear. Strengthening comparisons, adding tuned-lens checks, and demonstrating transfer to natural prompts would move this toward acceptance.\n\n- Toy 3-layer model offers limited insight into real decoder-only LLMs. The analyzed model is trained from scratch on synthetic data and lacks positional embedding and other complex architecture designs, as well as no noisy pretraining data.\nIt is therefore unclear whether the observed “sequential query” reflects processes in real decoder-only models like Llama or Gema, which are trained on much more noise and complex data with many more layers. \n\n- Empirical depth claim not grounded in formal theory. Also, the statement that “three layers are minimal for two-hop reasoning” is only observed empirically, or more like accommodates explaining the findings."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762941730319,"tcdate":1760731049622,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission21375/Reviewer_Hs31"],"signatures":["ICLR.cc/2026/Conference/Submission21375/Reviewer_Hs31"],"forum":"DPdev1Gg1o","number":2,"license":"CC BY 4.0","cdate":1760731049622,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission21375/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762941730319,"domain":"ICLR.cc/2026/Conference","replyto":"DPdev1Gg1o","id":"LlWsqYUx5m","forumContent":{"TLDR":{"value":"We study how Transformers make two-hop reasoning based on the context."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["two-hop reasoning","in-context learning"]},"supplementary_material":{"value":"/attachment/eb6d5785386a07eea393631007f831f1fd546159.pdf"},"primary_area":{"value":"foundation or frontier models, including LLMs"},"abstract":{"value":"``Socrates is human. All humans are mortal. Therefore, Socrates is mortal.'' \nThis form of argument illustrates a typical pattern of two-hop reasoning.\nFormally, two-hop reasoning refers to the process of inferring a conclusion by making two logical steps, each connecting adjacent concepts, such that the final conclusion depends on the integration of both steps.\nIt is one of the most fundamental components of human reasoning and plays a crucial role in both formal logic and everyday decision-making.\nDespite recent progress in large language models (LLMs), we surprisingly find that they are vulnerable when distractors are present. \nWe observe on a synthetic dataset that pre-trained LLMs often resort to random guessing among all plausible conclusions. \nHowever, after few steps of fine-tuning, models achieve near-perfect accuracy and exhibit strong length generalization. \nTo understand the underlying mechanisms, we train a 3-layer Transformer from scratch on a synthetic two-hop reasoning task and reverse-engineer its internal information flow.\nWe observe a clear progression in the attention logits throughout training.\nThis pictures a sharp phase transition from an initial stage of random guessing to the emergence of a structured sequential query mechanism, where the model first retrieves the preceding and the bridge concepts in the early layers and then uses them to infer the final answer. \nFinally, we show that these dynamics can be captured by a minimal three-parameter attention-only network."},"_bibtex":{"value":"@misc{\nguo2026how,\ntitle={How Do Transformers Perform Two-Hop Reasoning in Context?},\nauthor={Tianyu Guo and Hanlin Zhu and Ruiqi Zhang and Jiantao Jiao and Song Mei and Michael I. Jordan and Stuart Russell},\nyear={2026},\nurl={https://openreview.net/forum?id=DPdev1Gg1o}\n}"},"title":{"value":"How Do Transformers Perform Two-Hop Reasoning in Context?"},"pdf":{"value":"/pdf/a52a748231c4a815cec1bc8bcf6b499d5ea68a7a.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"guo|how_do_transformers_perform_twohop_reasoning_in_context"},"authorids":{"value":["~Tianyu_Guo4","~Hanlin_Zhu2","~Ruiqi_Zhang2","~Jiantao_Jiao1","~Song_Mei1","~Michael_I._Jordan2","~Stuart_Russell1"]},"authors":{"value":["Tianyu Guo","Hanlin Zhu","Ruiqi Zhang","Jiantao Jiao","Song Mei","Michael I. Jordan","Stuart Russell"]}},"version":2},{"content":{"summary":{"value":"This paper introduces VideoSafetyEval — a large-scale, real-world benchmark for video LLM safety (11.4k video–query pairs, 19 categories, 10 languages). Building on VideoSafetyEval, the authors propose a post-training defense framework, VideoSafety-R1, \nwhich integrates: (1) VideoSafetyThinking (VST) for chain-of-thought safety reasoning (46k samples), (2) Alarm Token-Guided Safety Fine-Tuning (AT-SFT), and (3) Safety-Guided GRPO with rule-based rewards."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1. How are the hyperparameters, such as $\\lambda_1$ and $\\lambda_2$, determined? Do their values significantly affect the performance of the alarm token and the GRPO strategy?\n\n2. Does fine-tuning video LLMs on safety dimensions impact their performance in other aspects, such as hallucination or overall video understanding ability (e.g., on Video-MME)?\n\n3. Could the authors further analyze the diversity of reasoning chains in VST (e.g., average number of logical steps, template proportion) and the model’s performance across different languages?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. To address the safety vulnerabilities of video LLMs, the authors introduce the VideoSafetyThinking dataset, which includes 46K video–query–reasoning triples and provides a valuable resource for the research community.\n2. The defense strategy is well-designed: it follows a coherent end-to-end pipeline — from perception (alarm token) to reasoning (CoT) to reinforcement learning (GRPO with dynamic rewards) — offering a clear approach that convincingly shifts from passive refusal to safety-driven reasoning."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Overstatement: The authors claim that VideoSafetyEval is the first real-world safety evaluation benchmark for video LLMs. However, prior works such as VideoSafetyBench, SafeVid, and Trust-VideoLLMs have already explored safety evaluation in this domain.\n2. Data collection, query generation, initial harmful labeling, and evaluation heavily rely on Qwen-Max, Qwen-Long, and GPT-4o, with limited enough human verification. Since these models have inherent reliability issues in video understanding, the resulting dataset may share similar distributions and biases with them, undermining the credibility of the evaluation."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762925540801,"tcdate":1761488364360,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission15243/Reviewer_JMVT"],"signatures":["ICLR.cc/2026/Conference/Submission15243/Reviewer_JMVT"],"forum":"QuW5RUDwMo","number":1,"license":"CC BY 4.0","cdate":1761488364360,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission15243/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762925540801,"domain":"ICLR.cc/2026/Conference","replyto":"QuW5RUDwMo","id":"DHsAkOdLdk","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Video Large Language Model","Safety of Multimodal Large Language Model","Safety Alignment","RLHF"]},"primary_area":{"value":"alignment, fairness, safety, privacy, and societal considerations"},"abstract":{"value":"While the safety risks of image-based large language models (Image LLMs) have been extensively studied, their video-based counterparts (Video LLMs) remain critically under-examined. To systematically study this problem, we introduce \\textbf{VideoSafetyEval} -- a large-scale, real-world benchmark for Video LLM safety, which comprises 11.4k video-query pairs and spans 19 principal risk categories. Based on this, \\textit{we reveal that integrating video modality degrades safety performance by an average of 34.2\\%, thereby exposing systemic risks in multimodal attack exploitation.}\nTo address this vulnerability, we propose \\textbf{VideoSafety-R1}, a dual-stage framework achieving unprecedented safety gains through three innovations: (1) VideoSafetyThinking dataset contains 46k video-query–thinking response triplets.  (2) Alarm Token-Guided Safety Fine-Tuning (AT-SFT) injects learnable alarm tokens into visual and textual sequences, enabling explicit harm perception across modalities via multitask objectives. (3) Safety-guided GRPO enhances defensive reasoning through dynamic policy optimization with rule-based rewards derived from dual-modality verification. These components synergize to shift safety alignment from harm perception to active reasoning. The framework achieves a 71.1\\% improvement on VSE-HH, and improves by 59.1\\%, 44.3\\%, and 15.0\\% on the image safety datasets MMBench, VLGuard, and FigStep, respectively. Our code and dataset are  available at \\url{https://github.com/Emiya-syw/VideoSafety-R1.git}.\n\\textcolor{red}{Note: This paper contains harmful language and image examples, and reader discretion is recommended.}"},"_bibtex":{"value":"@inproceedings{\nsun2026from,\ntitle={From Evaluation to Defense: Advancing Safety in Video Large Language Models},\nauthor={Yiwei Sun and Peiqi Jiang and Chuanbin Liu and Luohao Lin and Zhiying Lu and Hongtao Xie},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=QuW5RUDwMo}\n}"},"title":{"value":"From Evaluation to Defense: Advancing Safety in Video Large Language Models"},"pdf":{"value":"/pdf/02077ad527117ebc8263104650597329d84c5fe2.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"sun|from_evaluation_to_defense_advancing_safety_in_video_large_language_models"},"authorids":{"value":["~Yiwei_Sun3","~Peiqi_Jiang3","~Chuanbin_Liu2","~Luohao_Lin1","~Zhiying_Lu1","~Hongtao_Xie2"]},"authors":{"value":["Yiwei Sun","Peiqi Jiang","Chuanbin Liu","Luohao Lin","Zhiying Lu","Hongtao Xie"]}},"version":2},{"content":{"venue":{"value":"IEEE Trans. Circuits Syst. Video Technol. 2024"},"pdf":{"value":"https://ieeexplore.ieee.org/iel7/76/10550083/10347233.pdf"},"venueid":{"value":"dblp.org/journals/TCSV/2024"},"paperhash":{"value":"chen|hasi_hierarchical_attentionaware_spatiotemporal_interaction_for_videobased_person_reidentification"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Si_Chen_0002:","~Hui_Da2","~Da-Han_Wang1","https://dblp.org/search/pid/api?q=author:Xu-Yao_Zhang:","https://dblp.org/search/pid/api?q=author:Yan_Yan_0001:","https://dblp.org/search/pid/api?q=author:Shunzhi_Zhu:"]},"html":{"value":"https://doi.org/10.1109/TCSVT.2023.3340428"},"_bibtex":{"value":"@article{DBLP:journals/tcsv/ChenDWZYZ24,\n  author={Si Chen and Hui Da and Da-Han Wang and Xu-Yao Zhang and Yan Yan and Shunzhi Zhu},\n  title={HASI: Hierarchical Attention-Aware Spatio-Temporal Interaction for Video-Based Person Re-Identification},\n  year={2024},\n  month={June},\n  cdate={1717200000000},\n  journal={IEEE Trans. Circuits Syst. Video Technol.},\n  volume={34},\n  number={6},\n  pages={4973-4988},\n  url={https://doi.org/10.1109/TCSVT.2023.3340428}\n}\n"},"abstract":{"value":"Video-based person re-identification (re-ID) aims to match the same pedestrian of video sequences across non-overlapping cameras. Video re-ID methods generally adopt frame-level feature extraction for different video frames, but they still lack effective spatio-temporal interaction, easily leading to the multi-frame misalignment problem. In this paper, we propose a Hierarchical Attention-aware Spatio-temporal Interaction (HASI) network, including an Attention-aware Temporal Interaction (ATI) module and a Hierarchical Local-spatial Enhancement (HLE) module for video-based person re-ID. In order to avoid the spatial misalignment between video frames, the ATI module employs multiple Frame-to-Frame Temporal Interaction (2FTI) blocks with the Multi-head Inter-frame Alignment Attention (MIAA) to make the current frame iteratively interact with each rest frame of a video in a positive single-cycle manner, rather than only interacting with the adjacent frame or directly building the relationship of all frames at once. This module can not only obtain the long-range non-adjacent temporal information, but also learn the pairwise frame-to-frame relationships. Moreover, the HLE module is designed to enhance the local fine-grained features from multiple Transformer layers, whilst delivering low-level information to further enrich middle-level and high-level semantic knowledge. Thus, our method can learn multi-perspective pedestrian information, including inter-frame long-range interaction information and intra-frame multi-layer global and local information. Extensive experiments demonstrate the superiority of the proposed HASI method compared with the state-of-the-art methods on the three challenging video-based re-ID datasets, i.e., MARS, iLIDS-VID, and PRID-2011."},"title":{"value":"HASI: Hierarchical Attention-Aware Spatio-Temporal Interaction for Video-Based Person Re-Identification"},"authors":{"value":["Si Chen","Hui Da","Da-Han Wang","Xu-Yao Zhang","Yan Yan","Shunzhi Zhu"]}},"tmdate":1731463439868,"pdate":1704067200000,"tcdate":1731288362929,"writers":["~"],"signatures":["~Da-Han_Wang1"],"forum":"4mvrwtBmgU","license":"CC BY-SA 4.0","number":178501,"cdate":1717200000000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1731463439868,"domain":"DBLP.org","id":"4mvrwtBmgU","version":2},{"content":{"TLDR":{"value":"Measuring the extent to which shortcuts are utilized by a deep learning disease classification models through counterfactually removing potential shortcut attributes from the layers."},"venue":{"value":"MIDL 2025 Oral"},"midl_latex_submission_checklist":{"value":["The paper compiles correctly using the pdflatex compiler.","Replace NNN with your OpenReview submission ID.","The LaTeX file includes the correct header commands before \\title.","Hyperref is preloaded; do not modify it.","If included, must contact program chairs immediately.","Format names correctly and avoid Unicode characters.","Title page, citations, repositories, data sources, and acknowledgements must reflect correct authorship.","Did not use \\begin{thebibliography} directly to insert references.","Avoid using \\scalebox; prefer \\resizebox{\\textwidth}{!}{...}.","Zip archive should only contain needed files.","Clean final version without visual artifacts.","Ensure proper LaTeX encoding of characters.","No separate supplementary PDF uploads.","Check page limit and content structure carefully."]},"keywords":{"value":["Shortcuts","counterfactual","Parkinson's disease","bias mitigation"]},"abstract":{"value":"Although deep learning models can surpass human performance in many medical image analysis tasks, they remain vulnerable to algorithmic shortcuts, where spurious correlations in the data are exploited, which may lead to reduced trust in their predictions/classifications. This issue is especially concerning when models rely on protected attributes (e.g., sex, race, or site) as shortcuts. Such shortcut reliance not only impairs their ability to generalize to unseen datasets but also raises fairness concerns, ultimately undermining their purpose for computer-aided diagnosis. Previous techniques for analyzing protected attributes, such as supervised prediction layer information tests, only highlight the presence of protected attributes in the feature space but do not confirm their role in solving the primary task. Determining the impact of protected attributes as shortcuts is particularly challenging, as it requires knowing how a model would perform without those attributes — a counterfactual scenario typically unattainable in real-world data. As a workaround, researchers have addressed the absence of counterfactuals by generating synthetic datasets with and without protected attributes. In this study, we propose a novel approach to evaluate real-world datasets and determine the extent to which each protected attribute is used as a shortcut in a classification task. Therefore, we define and train a causal generative model to produce causally-grounded counterfactuals, removing protected attributes from activations and allowing us to measure their impact on model performance. Employing T1-weighted MRI data from 9 sites (835 subjects: 426 with Parkinson’s disease (PD) and 409 healthy), we demonstrate that counterfactually removing the 'site' attribute from the penultimate layer of a trained classification model reduced the AUROC for PD classification from 0.74 to 0.65, indicating a 9% performance improvement achieved by using 'site' as a shortcut. In contrast, counterfactually removing the 'sex' attribute had minimal impact on performance, with only a slight change of 0.004, indicating that 'sex' was not utilized as a shortcut by the classification model. The proposed method offers a robust framework for assessing shortcut utilization in medical image classification, paving the way for improved bias detection and mitigation in medical imaging tasks. The code for this work is available on https://github.com/vibujithan/shortcut-analysis."},"_bibtex":{"value":"@inproceedings{\nvigneshwaran2025evaluating,\ntitle={Evaluating Shortcut Utilization in Deep Learning Disease Classification through Counterfactual Analysis},\nauthor={Vibujithan Vigneshwaran and Emma A.M. Stanley and Raissa Souza and Erik Ohara and Matthias Wilms and Nils Forkert},\nbooktitle={Medical Imaging with Deep Learning},\nyear={2025},\nurl={https://openreview.net/forum?id=vxSo5TJxlB}\n}"},"title":{"value":"Evaluating Shortcut Utilization in Deep Learning Disease Classification through Counterfactual Analysis"},"latex_code":{"value":"/attachment/02d2eb9a84a26a0d1dcb26d3cf7c4a6f17664a25.zip"},"paper_type":{"value":"Both"},"secondary_subject_area":{"value":"Causality"},"pdf":{"value":"/pdf/f1671e38a3fefda63b04165aee4595c396c667aa.pdf"},"Reproducibility":{"value":"https://github.com/vibujithan/shortcut-analysis"},"copyright_form":{"value":"/attachment/4f5cbf1807bdf0225b66f1ce7eb2e3fa8f86085c.pdf"},"visa":{"value":"Yes"},"venueid":{"value":"MIDL.io/2025/Conference"},"paperhash":{"value":"vigneshwaran|evaluating_shortcut_utilization_in_deep_learning_disease_classification_through_counterfactual_analysis"},"primary_subject_area":{"value":"Fairness and Bias"},"authorids":{"value":["~Vibujithan_Vigneshwaran1","~Emma_A.M._Stanley2","~Raissa_Souza1","erik.ohara@ucalgary.ca","~Matthias_Wilms1","~Nils_Forkert1"]},"registration":{"value":"Yes"},"authors":{"value":["Vibujithan Vigneshwaran","Emma A.M. Stanley","Raissa Souza","Erik Ohara","Matthias Wilms","Nils Forkert"]}},"tmdate":1748872107296,"pdate":1743082528869,"tcdate":1737062617560,"writers":["MIDL.io/2025/Conference","MIDL.io/2025/Conference/Submission127/Authors"],"signatures":["MIDL.io/2025/Conference/Submission127/Authors"],"forum":"vxSo5TJxlB","license":"CC BY 4.0","number":127,"cdate":1737062617560,"readers":["everyone"],"invitations":["MIDL.io/2025/Conference/-/Submission","MIDL.io/2025/Conference/-/Post_Submission","MIDL.io/2025/Conference/Submission127/-/Full_Submission","MIDL.io/2025/Conference/-/Edit","MIDL.io/2025/Conference/Submission127/-/Camera_Ready"],"mdate":1748872107296,"odate":1737206940415,"domain":"MIDL.io/2025/Conference","id":"vxSo5TJxlB","version":2},{"content":{"summary":{"value":"The paper introduces AuroraCap, an LLM-based video captioning model that uses token merging to reduce visual tokens. AuroraCap is evaluated across multiple image and video captioning datasets. To assess models' ability to generate detailed captions, the paper also proposes a new benchmark, Video Detailed Captioning (VDC), along with an LLM-assisted metric, VDCscore. VDC includes over 1,000 annotated structured captions, while VDCscore transforms long captions into multiple short question-answer pairs. Various existing models are evaluated on the VDC benchmark."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Please see Weaknesses."},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"* The presentation of the paper is clear. It's easy to follow.\n* The proposed VDC benchmark is a valuable contribution to the community. The detailed video captioning capability of video LLMs is an area that has been relatively underexplored, and introducing VDCscore adds a meaningful component to the evaluation pipeline. The evaluation process of VDCscore aligns well with human reasoning process.\n* Extensive experiments are conducted."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"* Certain aspects of AuroraCap are not fully explored. For instance, while it performs well on the ANet dataset in video QA tasks, it performs poorly on MSVD, which is typically an easier benchmark. What might be causing this discrepancy?\n* Unfair comparison:\n  * Table 2 refer to the baseline models as the SoTA methods on image captioning benchmarks under zero-shot setting. However, I believe there are other LLM-based models with stronger captioning performance. For example, the BLIP family—BLIP2 [1], which achieves a 121.6 CIDEr score on NoCaps (val set), and InstructBLIP [2], which is even stronger. It would be beneficial to clarify the criteria used to select these baselines.\n  * Similar concerns apply to the video captioning and video QA task comparisons.\n* I notice that the VDCscore relies heavily on the use of GPT-4o:\n  * As also pointed out by other work [3], different versions of GPT can affect the results of GPT assisted evaluation heavily. Although I may have missed it, I did not see which version of GPT-4o was used for the evaluation in the paper. Standardizing the evaluation method is essential for consistency in future studies.\n  * Given the multiple variables that can influence GPT-assisted evaluation, and since VDCscore already uses phrased answers, what are the potential drawbacks of using an automatic metric based on, e.g. n-gram matching, to evaluate the final score? Could a variant of VDCscore be devised that assesses final triplets (i.e., <question, correct answer, predicted answer>) without relying on GPT?\n* minor: Several instances of \"BELU\" should be corrected to \"BLEU\"; small grammar errors like in line 203: \"are only contains\" --> \"only contain\"\n\n\n\n[1] Junnan Li, Dongxu Li, Silvio Savarese, Steven Hoi. BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models. ICML 2023.\n\n[2] Wenliang Dai, Junnan Li, Dongxu Li, Anthony Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, Steven Hoi. InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning. NeurIPS 2023.\n\n[3] Wenhao Wu. Freeva: Offline mllm as training-free video assistant. arXiv 2024."}},"nonreaders":[],"tmdate":1732279225301,"tcdate":1730588817999,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission640/Reviewer_7VHc"],"signatures":["ICLR.cc/2025/Conference/Submission640/Reviewer_7VHc"],"forum":"tTDUrseRRU","number":4,"license":"CC BY 4.0","cdate":1730588817999,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission640/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732279225301,"domain":"ICLR.cc/2025/Conference","replyto":"tTDUrseRRU","id":"PGcDY8X9gr","forumContent":{"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video Captioning","Benchmark","Multimodel Large Language Model"]},"primary_area":{"value":"datasets and benchmarks"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Video detailed captioning is a key task which aims to generate comprehensive and coherent textual descriptions of video content, benefiting both video understanding and generation. In this paper, we propose AuroraCap, a video captioner based on a large multimodal model. We follow the simplest architecture design without additional parameters for temporal modeling. To address the overhead caused by lengthy video sequences, we implement the token merging strategy, reducing the number of input visual tokens. Surprisingly, we found that this strategy results in little performance loss. AuroraCap shows superior performance on various video and image captioning benchmarks, for example, obtaining a CIDEr of 88.9 on Flickr30k, beating GPT-4V (55.3) and Gemini-1.5 Pro (82.2). However, existing video caption benchmarks only include simple descriptions, consisting of a few dozen words, which limits research in this field. Therefore, we develop VDC, a video detailed captioning benchmark with over one thousand carefully annotated structured captions. In addition, we propose a new LLM-assisted metric VDCscore for bettering evaluation, which adopts a divide-and-conquer strategy to transform long caption evaluation into multiple short question-answer pairs. With the help of human Elo ranking, our experiments show that this benchmark better correlates with human judgments of video detailed captioning quality."},"_bibtex":{"value":"@inproceedings{\nchai2025auroracap,\ntitle={AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark},\nauthor={Wenhao Chai and Enxin Song and Yilun Du and Chenlin Meng and Vashisht Madhavan and Omer Bar-Tal and Jenq-Neng Hwang and Saining Xie and Christopher D Manning},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=tTDUrseRRU}\n}"},"title":{"value":"AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark"},"pdf":{"value":"/pdf/54378994740e36257bfa53e767f50f28a4e36483.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"chai|auroracap_efficient_performant_video_detailed_captioning_and_a_new_benchmark"},"authorids":{"value":["~Wenhao_Chai1","~Enxin_Song2","~Yilun_Du1","~Chenlin_Meng1","~Vashisht_Madhavan3","~Omer_Bar-Tal2","~Jenq-Neng_Hwang1","~Saining_Xie2","~Christopher_D_Manning1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Wenhao Chai","Enxin Song","Yilun Du","Chenlin Meng","Vashisht Madhavan","Omer Bar-Tal","Jenq-Neng Hwang","Saining Xie","Christopher D Manning"]}},"version":2},{"content":{"summary":{"value":"The paper effectively highlights the challenge of training models on biased datasets and the potential pitfalls of spurious correlations. The proposed method based on multi-task learning (MTL) is an innovative approach to addressing the problem of multiple biases in training data. This introduces a new perspective on debiased training. The introduction of a new real-image dataset, MultiCelebA, is a valuable contribution. It allows for evaluation under more realistic and challenging scenarios compared to existing synthetic-image datasets."},"presentation":{"value":"3 good"},"contribution":{"value":"3 good"},"soundness":{"value":"3 good"},"strengths":{"value":"1. Application of multitask learning in the context of multiple shortcuts is a novel idea.\n2.  MultiCelebA dataset can be instrumental for future research on evaluating shortcut learning algorithms."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Limitation:\n1. No related works on shortcuts. Relevant literature can be found in\nDiscover and Cure: Concept-aware Mitigation of Spurious Correlation Wu et al. ICML 2023.\n\n2. The paper is not easy to follow. The writing could have been better.\n\n3. No code, so limited reproducibility.\n\n4. The major issue with the approach is knowing so many different subgroups a priori. In a more challenging setting, it is almost impossible to know all possible different subgroups beforehand to design the training strategy. I would like see an experiment where out of 3 subgroups the authors include two in their training, leaving one unidentified and how their method performs.\n\n5. There is a recent notion of difficulty in shortcuts. For example, some shortcuts are easy to learn, and some are difficult to learn. If the authors are using different subgroups to design multitask losses, the losses should be weighted corresponding to the shorcut difficulty. For example, a hard shortcut loss will be penalized less than an easy shortcut loss, as the mode is more prone to latching on the easy shotcut. The paper on shortcut difficulty is as follows:\nBeyond Distribution Shift: Spurious Features Through the Lens of Training Dynamics. Murali et al. TMLR 2023.\n\nThis setting will be more realistic.\n\n6. The dataset MultiCelebA is good for evaluation, but I would like to see a more realistic dataset like NIH-chesttube in the shortcut paper in #5.\n\n7. Also, with multiple groups involved in the multitask loss, the overall performance may drop.\n\n8. This is related to #4. All the possible groups may not be a shortcut. In this regard, can the shortcut discovery be aligned with the notion of slice discovery (ex DOMINO) to detect if really a spurious correlation going on before applying their method? This is a nice-to-have comment. I request the authors to think about this as a future work.\n\n\n**Post rebuttal**\n\nThanks for the response. I agree about point 5 and 6 but request the authors to think about it. I would like to thank the authors for  considering the point 4 and would like to give a score of 7. Unfortunately, I cant assign that so i would keep my score."},"confidence":{"value":"3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked."},"questions":{"value":"See the weakness"},"rating":{"value":"6: marginally above the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1700890509380,"tcdate":1698796770447,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission1800/Reviewer_4jDT"],"signatures":["ICLR.cc/2024/Conference/Submission1800/Reviewer_4jDT"],"forum":"rJKlmCpOQ7","number":2,"license":"CC BY 4.0","cdate":1698796770447,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission1800/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1700890509380,"domain":"ICLR.cc/2024/Conference","replyto":"rJKlmCpOQ7","id":"Be2VlNHKNG","forumContent":{"venue":{"value":"Submitted to ICLR 2024"},"TLDR":{"value":"We present a novel training algorithm mitigating multiple biases of training data based on a theory of multi-task learning, along with a new real-image multi-bias dataset to facilitate future research in this direction."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Debiasing","spurious correlation","multiple biases","shortcut learning"]},"primary_area":{"value":"societal considerations including fairness, safety, privacy"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"We consider the problem of training an unbiased and accurate model using a biased dataset with multiple biases. This problem is challenging since the multiple biases cause multiple undesirable shortcuts during training, and even worse, mitigating one of them may exacerbate another. To address this challenge, we introduce a novel method connecting the problem to multi-task learning (MTL). Our method divides training data into several groups according to their effects on the model bias and defines each task of MTL as solving the target problem for each group. It in turn trains a single model for all the tasks with a weighted sum of task-wise losses as the training objective, while optimizing the weights as well as the model parameters. At the heart of our method lies the weight adjustment algorithm, which is rooted in a theory of multi-objective optimization and guarantees a Pareto-stationary solution. In addition, we also present a new real-image dataset with multiple biases, dubbed MultiCelebA, for evaluating debiased training methods under realistic and challenging scenarios. Our method achieved the state of the art on three datasets with multiple biases including MultiCelebA, and demonstrated superior performance on conventional single-bias datasets."},"_bibtex":{"value":"@misc{\nkim2024removing,\ntitle={Removing Multiple Shortcuts through the Lens of Multi-task Learning},\nauthor={Nayeong Kim and Juwon Kang and Sungsoo Ahn and Jungseul Ok and Suha Kwak},\nyear={2024},\nurl={https://openreview.net/forum?id=rJKlmCpOQ7}\n}"},"title":{"value":"Removing Multiple Shortcuts through the Lens of Multi-task Learning"},"pdf":{"value":"/pdf/5fbdee79645a2e61ef6888b4c70a38b05f94aab9.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Rejected_Submission"},"paperhash":{"value":"kim|removing_multiple_shortcuts_through_the_lens_of_multitask_learning"},"authorids":{"value":["~Nayeong_Kim1","~Juwon_Kang1","~Sungsoo_Ahn1","~Jungseul_Ok2","~Suha_Kwak3"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Nayeong Kim","Juwon Kang","Sungsoo Ahn","Jungseul Ok","Suha Kwak"]}},"version":2},{"content":{"venue":{"value":"IEEE Trans. Circuits Syst. Video Technol. 2022"},"pdf":{"value":"https://ieeexplore.ieee.org/iel7/76/9812753/09632538.pdf"},"venueid":{"value":"dblp.org/journals/TCSV/2022"},"paperhash":{"value":"xu|learning_an_occlusionaware_network_for_video_deblurring"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Qian_Xu_0014:","https://dblp.org/search/pid/api?q=author:Jinshan_Pan:","~Yuntao_Qian1"]},"html":{"value":"https://doi.org/10.1109/TCSVT.2021.3132102"},"_bibtex":{"value":"@article{DBLP:journals/tcsv/XuPQ22,\n  author={Qian Xu and Jinshan Pan and Yuntao Qian},\n  title={Learning an Occlusion-Aware Network for Video Deblurring},\n  year={2022},\n  cdate={1640995200000},\n  journal={IEEE Trans. Circuits Syst. Video Technol.},\n  volume={32},\n  number={7},\n  pages={4312-4323},\n  url={https://doi.org/10.1109/TCSVT.2021.3132102}\n}\n"},"abstract":{"value":"Video deblurring is a challenging task since the blur is caused by camera shake, object motions, etc. The success of the state-of-the-art methods stems mainly from exploiting the temporal information of neighboring frames through alignment. When there exists occlusion among the sequence, these approaches become less effective for inaccurate alignment. In this paper, we propose an effective occlusion-aware network to handle the occlusion for video deblurring. The proposed module first generates a coarse pixel-wise alignment filter to explore the temporal information and then learns an adaptive affine transformation to deal with the occluded areas. In addition, a self-attention mechanism is developed to better model the occluded pixels. To further improve the performance, we progress a multi-scale strategy and train the network in an end-to-end manner. Both quantitative and qualitative experimental results show that the proposed method achieves favorable performance against state-of-the-art methods on the benchmark datasets. The code and trained models are available at: https://github.com/XQLuck/code.git"},"title":{"value":"Learning an Occlusion-Aware Network for Video Deblurring"},"authors":{"value":["Qian Xu","Jinshan Pan","Yuntao Qian"]}},"tmdate":1762322434631,"pdate":1640995200000,"externalIds":["dblp:journals/tcsv/XuPQ22"],"tcdate":1762322387465,"writers":["~"],"signatures":["~Yuntao_Qian1"],"forum":"EuL95YFEho","license":"CC BY-SA 4.0","number":655999,"cdate":1640995200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1762322434631,"domain":"DBLP.org","id":"EuL95YFEho","version":2},{"content":{"summary":{"value":"The paper ChineseVideoBench: Benchmarking Multimodal Large Models for Chinese Video Question Answering presents a new benchmark to evaluate multimodal large language models (MLLMs) on Chinese video understanding tasks. It addresses the lack of culturally and linguistically relevant datasets for non-English scenarios. The dataset contains 1,625 CC0-licensed Chinese videos and 6,507 human-annotated multiple-choice QA pairs, covering 8 main tasks and 12 sub-tasks such as world knowledge, scene understanding, temporal localization, and logical reasoning."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"Why don't the authors test the videoQA without removing the audio, as an ablation?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"The paper introduces the first large-scale benchmark for Chinese VideoQA, covering 1,625 CC0-licensed videos and 6,507 manually annotated QA pairs\nThe benchmark exposes systematic failure patterns in temporal localization and fine-grained spatiotemporal grounding."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The paper does not report inter-annotator consistency, distractor calibration, or difficulty-level validation, leaving the annotation quality claims insufficiently quantified.\n\nThe paper emphasizes long-video evaluation but does not provide convincing evidence that its videos require long-term temporal reasoning; average durations appear modest, weakening this claim"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762922469055,"tcdate":1761807124137,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission11335/Reviewer_VNBP"],"signatures":["ICLR.cc/2026/Conference/Submission11335/Reviewer_VNBP"],"forum":"Aqc5PRH2KF","number":3,"license":"CC BY 4.0","cdate":1761807124137,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission11335/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762922469055,"domain":"ICLR.cc/2026/Conference","replyto":"Aqc5PRH2KF","id":"vGrbulYtUF","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["MLLM","Benchmark"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"This paper introduces ChineseVideoBench, a pioneering benchmark specifically designed for evaluating Multimodal Large Language Models (MLLMs) in Chinese Video Question Answering. The growing demand for sophisticated video analysis capabilities highlights the critical need for comprehensive, culturally-aware evaluation frameworks. ChineseVideoBench addresses this gap by providing a robust dataset and tailored evaluation metrics, enabling rigorous assessment of state-of-the-art MLLMs on complex Chinese video content. Specifically, ChineseVideoBench comprises 8 main classes and 12 sub-classes, encompassing tasks that demand both deep video understanding and nuanced Chinese linguistic and cultural awareness. Our empirical evaluations reveal that ChineseVideoBench presents a significant challenge to current MLLMs. Among the models assessed, Gemini 2.5 Pro achieves the highest performance with an overall score of 77.9%, while Intern-VL-38B emerges as the most competitive open-source model."},"_bibtex":{"value":"@misc{\nnie2026chinesevideobench,\ntitle={ChineseVideoBench: Benchmarking Multi-modal Large Models for Chinese Video Question Answering},\nauthor={Yuxiang Nie and Han Wang and Yongjie Ye and Haiyang Yu and WeitaoJia and Tao Zeng and Hao Feng and Xiang Fei and Yang Li and Xiaohui Lv and Guozhi Tang and Jingqun Tang and Jinghui Lu and Zehui Dai and Jiacong Wang and Dingkang Yang and An-Lan Wang and Can Huang and ChaoFeng and Ran Jiao},\nyear={2026},\nurl={https://openreview.net/forum?id=Aqc5PRH2KF}\n}"},"title":{"value":"ChineseVideoBench: Benchmarking Multi-modal Large Models for Chinese Video Question Answering"},"pdf":{"value":"/pdf/8279d37605d22dd7f7425ca8e164d569b9d600e4.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"nie|chinesevideobench_benchmarking_multimodal_large_models_for_chinese_video_question_answering"},"authorids":{"value":["~Yuxiang_Nie2","~Han_Wang18","~Yongjie_Ye1","~Haiyang_Yu5","~WeitaoJia1","~Tao_Zeng4","~Hao_Feng4","~Xiang_Fei1","~Yang_Li59","~Xiaohui_Lv1","~Guozhi_Tang1","~Jingqun_Tang1","~Jinghui_Lu2","~Zehui_Dai1","~Jiacong_Wang1","~Dingkang_Yang1","~An-Lan_Wang1","~Can_Huang1","~ChaoFeng1","~Ran_Jiao1"]},"authors":{"value":["Yuxiang Nie","Han Wang","Yongjie Ye","Haiyang Yu","WeitaoJia","Tao Zeng","Hao Feng","Xiang Fei","Yang Li","Xiaohui Lv","Guozhi Tang","Jingqun Tang","Jinghui Lu","Zehui Dai","Jiacong Wang","Dingkang Yang","An-Lan Wang","Can Huang","ChaoFeng","Ran Jiao"]}},"version":2},{"content":{"summary":{"value":"The paper proposes a method for continuous video interpolation that generates intermediate frames at any frame rate. Given a start and end frame, the method interpolates frames at arbitrary timestamps. To achieve this, the authors assign timestamps 0 and 1 to the start and end frames, respectively, normalize intermediate timestamps, and introduce a timestamp-aware RoPE. They further propose appearance–motion decoupled conditioning, enabling smooth and continuous segment-level interpolation. On VBench, the method achieves significantly superior video quality compared to prior approaches."},"soundness":{"value":2},"confidence":{"value":2},"questions":{"value":"- Beyond VFI, it would be valuable to verify whether appearance and motion have indeed been properly decoupled. If the decoupling is effective, could one borrow motion from other objects and use it as input for controlled motion transfer?\n- Can the model also predict consecutive frames at non-equidistant timestamps? For example, at timestamps such as [0, 0.1, 0.4, 0.5, 0.9, 1]"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- **Novel problem setting:** This is the first work, to the best of my knowledge, to utilize generative models for continuous interpolation between two frames.\n- **Clear and intuitive writing:** The paper is easy to follow and provides intuitive explanations. In particular, the concise extension from standard RoPE to timestamp-aware RoPE (TaRoPE) is elegant and highly effective. The method demonstrates strong performance, significantly outperforming existing approaches."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- While it is novel to apply generative modeling to VFI, methods that use continuous timestamps already exist in the non-generative video frame interpolation literature [1]. The paper does not compare against these methods. Including such comparisons would more clearly position this work relative to deterministic approaches.\n- The authors combine timestamp-aware RoPE with a training dataset containing a wide timestamp range (1/2–1/240), which likely contributes significantly to performance. However, it is unclear whether the baselines were trained under the same settings. If baselines were trained only at fixed rates, it becomes difficult to attribute performance gains solely to TaRoPE rather than to broader timestamp coverage. It would be beneficial to train at least one baseline using the same diverse-rate data to clarify this.\n\n[1] Super slomo: High quality estimation of multiple intermediate frames for video interpolation, Huaizu, et al., 2018"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762925316443,"tcdate":1761831916741,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission14985/Reviewer_E5JG"],"signatures":["ICLR.cc/2026/Conference/Submission14985/Reviewer_E5JG"],"forum":"eKGkb4cFRe","number":1,"license":"CC BY 4.0","cdate":1761831916741,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission14985/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762925316443,"domain":"ICLR.cc/2026/Conference","replyto":"eKGkb4cFRe","id":"zg6oJ2WLOW","forumContent":{"TLDR":{"value":"A novel generative VFI paradigm that enables interpolation at any timestamp and of any length."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Video Frame Interpolation","ROPE","Video Generation"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Generative Video Frame Interpolation (VFI), which synthesizes intermediate frames from a given pair of start and end frames, plays a pivotal role in video creation. However, existing generative VFI methods are constrained to producing a fixed number of intermediate frames, which significantly limits the flexibility in adjusting the frame rate or duration of videos during the creation process. In this work, we present \\textbf{ArbInterp}, a novel generative VFI framework that enables efficient interpolation at any timestamp and of any length. Specifically, to support interpolation at any timestamp, we propose the Timestamp-aware Rotary Position Embedding (TaRoPE), which modulates positions in temporal RoPE to align generated frames with target normalized timestamps. This design enables fine-grained control over frame timestamps, addressing the inflexibility of fixed-position paradigms in prior work. For any-length interpolation, we decompose long-sequence generation into segment-wise frame synthesis. We further design a novel appearance-motion decoupled conditioning strategy: it leverages prior segment endpoints to enforce appearance consistency and temporal semantics to maintain motion coherence, ensuring seamless spatiotemporal transitions across segments. Experimentally, we build comprehensive benchmarks for multi-scale frame interpolation (2× to 32×) to assess generalizability across arbitrary interpolation factors. Results show that ArbInterp outperforms prior methods across all scenarios with higher fidelity and more seamless spatiotemporal continuity. Video demos are provided on the website: https://mcg-nju.github.io/ArbInterp-Web."},"_bibtex":{"value":"@inproceedings{\nzhang2026arbitrary,\ntitle={Arbitrary Generative Video Interpolation},\nauthor={Guozhen Zhang and Haiguang Wang and Chunyu Wang and Yuan Zhou and Qinglin Lu and Limin Wang},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=eKGkb4cFRe}\n}"},"title":{"value":"Arbitrary Generative Video Interpolation"},"pdf":{"value":"/pdf/aaacd0c7f77411fa844576342f4459817cdde9eb.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"zhang|arbitrary_generative_video_interpolation"},"authorids":{"value":["~Guozhen_Zhang2","~Haiguang_Wang1","~Chunyu_Wang1","~Yuan_Zhou12","~Qinglin_Lu2","~Limin_Wang1"]},"authors":{"value":["Guozhen Zhang","Haiguang Wang","Chunyu Wang","Yuan Zhou","Qinglin Lu","Limin Wang"]}},"version":2},{"content":{"summary":{"value":"This paper proposed a synthetic dataset LLaVA-Video-178K, which consists of 178510 videos with detailed annotations, open-ended questions and multiple-choice questions. To build the dataset, the authors select the most dynamic videos from 10 major video data sources, and use a recurrent caption generation pipeline to generate video captions. The authors define 16 question types and generate question-answer pairs using GPT-4o. Based on LLaVA-Video-178K, the authors fine-tuned LLaVA-OneVision on the combination of LLaVA-Video-178K and other four public datasets to obtain the model called LLaVA-Video. Experiments show that the model trained with LLaVA-Video-178K will have a performance gain on a wide range of video benchmarks."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1. In Table 3, what is the training data settings for +LLaVA-Video-178K, +Three Q&A datasets, and +LLaVA-OV (images)? From line 481 to line 485, it seems that LLaVA-Video-178K, three Q&A datasets and LLaVA-OV(images) are incrementally added for training. If so, what is the performance gain if LLaVA-Video-178K is the last dataset added for training? If datasets are trained separately in three settings, then why the performance of LLaVA-Video-178K is lower than LLaVA-OV (images)?\n2. In Table 4, are there any insights on the out-of-domain performance loss of LLaVA-Video-178K compared to LLaVA-Hound on EgoSchema?\n3. Figure 4 illustrates an interesting video. I am wondering whether the generated questions are only about the “facts” in the video such as “How many steps does “normal people” climb?”. I think people are more curious about whether the model understand the humor in the video. Whether the captions of the video can generate such questions?"},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. The paper proposes a caption generation pipeline, which recurrently generates and refines the captions of video in three temporal levels.\n2. To filter out the static videos that can be summarized by a single video frame, the authors construct the dataset using dynamic videos which are selected by detecting the number of scenes in the videos. \n3. The paper proposes 16 different question types on the video, which are more comprehensive compared to existing benchmarks."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The level-1 and level-2 description are generated on fixed time intervals (10s and 30s), which are not event-specific or scene-specific. \n2. Most of the videos in the dataset are 180 seconds long, which may hinder the dataset's effectiveness in the field of long video understanding. \n3. Several typos and errors exist in the paper. It seems that the paper is not carefully proofread and checked. For example:\nline 184 & 186: condtion->condition\nline 400: should be [M/(4*p^2)] for the fast frames\nline 415: considegreen->considered as"}},"nonreaders":[],"tmdate":1731427614748,"tcdate":1730944434216,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission2539/Reviewer_qq6R"],"signatures":["ICLR.cc/2025/Conference/Submission2539/Reviewer_qq6R"],"forum":"8Livf4oZxz","number":4,"license":"CC BY 4.0","cdate":1730944434216,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission2539/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427614748,"domain":"ICLR.cc/2025/Conference","replyto":"8Livf4oZxz","id":"bFNBOekudo","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"TLDR":{"value":"The largest high-quality synthetic dataset specifically for video instruction-following"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video instruction dataset","video-language model"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"The development of video large multimodal models (LMMs) has been hindered by the difficulty of curating large amounts of high-quality raw data from the web. To address this, we consider an alternative approach, creating a high-quality synthetic dataset specifically for video instruction-following, namely LLaVA-Video-178K. This dataset includes key tasks such as detailed captioning, open-ended question-answering (QA), and multiple-choice QA. By training on this proposed dataset, in combination with existing visual instruction tuning data, we introduce LLaVA-Video, a new video LMM. Our experiments demonstrate that LLaVA-Video achieves strong performance across various video benchmarks, highlighting the effectiveness of our dataset. We plan to release the dataset, its generation pipeline, and the model checkpoints."},"_bibtex":{"value":"@misc{\nzhang2024video,\ntitle={Video Instruction Tuning with Synthetic Data},\nauthor={Yuanhan Zhang and Jinming Wu and Wei Li and Bo Li and Zejun MA and Ziwei Liu and Chunyuan Li},\nyear={2024},\nurl={https://openreview.net/forum?id=8Livf4oZxz}\n}"},"title":{"value":"Video Instruction Tuning with Synthetic Data"},"pdf":{"value":"/pdf/98c983fa704b94de7f370bfa5fcf2b7e8e7db9e8.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"zhang|video_instruction_tuning_with_synthetic_data"},"authorids":{"value":["~Yuanhan_Zhang1","~Jinming_Wu1","~Wei_Li78","~Bo_Li23","~Zejun_MA1","~Ziwei_Liu1","~Chunyuan_Li1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Yuanhan Zhang","Jinming Wu","Wei Li","Bo Li","Zejun MA","Ziwei Liu","Chunyuan Li"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a video super-resolution algorithm called VEnhancer, which enhances video resolution in both spatial and temporal dimensions. The method is based on a pretrained video generation model, with a Spatio-Temporal (ST) Controller trained specifically for the task. The video to be enhanced is provided as a condition input to the ST-Controller. The authors constructed a test dataset, AIGC2023, to validate the effectiveness of the algorithm."},"soundness":{"value":2},"confidence":{"value":5},"questions":{"value":"1. In terms of experimental results, the proposed method shows only minor improvements in Aesthetic Quality Dynamic, Degree of Motion, and Smoothness, while achieving relatively significant improvements in MUSIQ and DOVER. What causes this discrepancy?\n2. As shown in Table 3, the improvement of the proposed method on CogVideoX-5B is relatively minor and not as significant as on VideoCrafter-2. Could the authors explain the reasons for this?\n3. How many videos are there in the AIGC2023 dataset? What are the approximate resolutions and lengths of these videos? Is there any statistical information available? Will the authors release this test set in the future?"},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. From the visual results provided by the authors, the proposed method enhances the video's spatial resolution and FPS. The details of objects are richer compared to the pre-enhancement version, and the smoothness of the video has also improved to some extent.\n2. This method can achieve both temporal and spatial video super-resolution simultaneously."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The proposed method's overall idea is quite similar to ControlNet: both introduce a control branch by copying the encoder of the UNet model to the control branch and initializing the output layer to zero. The difference lies in their applications—this method uses the framework for video super-resolution, while ControlNet is designed for controllable generation. In this paper, the condition is the video to be super-resolved, whereas in ControlNet, conditions include modalities like depth, edge, and mask. The authors should explicitly discuss how their approach differs from or improves upon ControlNet for the specific task of video enhancement. \n2. For the video-aware conditioning (section 4.3), the noise augmentation technique proposed by the authors is a common trick that was first introduced by *Cascaded Diffusion Models* [Jonathan Ho et al., 2021] for multi-stage high-resolution image generation. Essentially, it incrementally upscales from low to high resolution, which is quite similar to this paper's application. Additionally, using the time embedding as a condition is a standard approach in diffusion models. Conditioning on the downscale factor is somewhat analogous to using FPS as a condition in video generation [Make-A-Video, Uriel Singer et al., 2022], allowing the generation process to better incorporate additional conditioning information. This approach isn't particularly novel in itself. The authors should clarify what specific innovations or improvements their approach offers over these existing techniques. \n3. The proposed method is for video super-resolution and, theoretically, should be applicable to both AI-generated videos and real videos. However, the authors only conducted experiments on AI-generated videos (section 5.3), making this comparison less comprehensive. I suggest that the authors include experiments on real-world videos or explain why their method might not be suitable for such videos if that's the case."}},"nonreaders":[],"tmdate":1731428645254,"tcdate":1730033619174,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission10660/Reviewer_Sk5p"],"signatures":["ICLR.cc/2025/Conference/Submission10660/Reviewer_Sk5p"],"forum":"Ysdo3fyD4Q","number":1,"license":"CC BY 4.0","cdate":1730033619174,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission10660/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428645254,"domain":"ICLR.cc/2025/Conference","replyto":"Ysdo3fyD4Q","id":"6Arvl0PwII","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Diffusion models","Video Generation","Generative Video enhancement","video super-resolution","frame interpolation","space-time super-resolution"]},"supplementary_material":{"value":"/attachment/00330b4a79a9fd4b3296c2a4573c03d4de0b4f63.zip"},"primary_area":{"value":"generative models"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"We present \\emph{VEnhancer}, a generative space-time enhancement method that can improve the existing AI-generated videos spatially and temporally through one video diffusion model. Given a generated low-quality video, our approach can increase its spatial and temporal resolution simultaneously with arbitrary up-sampling space and time scales by adding more details in spatial domain and synthesize detailed motion in temporal domain. Furthermore, VEnhancer is able to remove generated spatial artifacts and temporal flickering of generated videos.  \nTo achieve this, basing on a pretrained generative video prior, we train a \\textbf{S}pace-\\textbf{T}ime Controller and inject it to the prior as a condition on low-frame-rate and low-resolution videos. To effectively train this ST-Controller, we design \\textit{space-time data augmentation} to create diversified video training pairs as well as \\textit{video-aware conditioning} for realizing different augmentation parameters in both spatial and temporal dimensions.\nBenefiting from the above designs, VEnhancer can be end-to-end trained to enable multi-function in one single model. \nExtensive experiments show that VEnhancer\nsurpasses existing state-of-the-art video super-resolution and space-time super-resolution methods in enhancing AI-generated videos. Moreover, VEnhancer is able to greatly improve the performance of open-source state-of-the-art text-to-video methods on video generation benchmark, VBench."},"_bibtex":{"value":"@misc{\nhe2025venhancer,\ntitle={{VE}nhancer: Generative Space-Time Enhancement for Video Generation},\nauthor={Jingwen He and Tianfan Xue and Dongyang Liu and Xinqi Lin and Peng Gao and Dahua Lin and Yu Qiao and Wanli Ouyang and Ziwei Liu},\nyear={2025},\nurl={https://openreview.net/forum?id=Ysdo3fyD4Q}\n}"},"title":{"value":"VEnhancer: Generative Space-Time Enhancement for Video Generation"},"pdf":{"value":"/pdf/8e459f2f620978655e892017e981fe3deb5dc4df.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"he|venhancer_generative_spacetime_enhancement_for_video_generation"},"authorids":{"value":["~Jingwen_He1","~Tianfan_Xue2","~Dongyang_Liu1","~Xinqi_Lin1","~Peng_Gao3","~Dahua_Lin1","~Yu_Qiao1","~Wanli_Ouyang1","~Ziwei_Liu1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Jingwen He","Tianfan Xue","Dongyang Liu","Xinqi Lin","Peng Gao","Dahua Lin","Yu Qiao","Wanli Ouyang","Ziwei Liu"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a new algorithm to extract cardinal-minimal sufficient explanations for Neural Additive Models (NAMs).\nIt does so by exploiting key design choices of NAMs, showing how this family of models supports explanations with guarantees.\n\nThis is achieved as follows. First, the paper introduces a method to rank features based on how much they influence the final prediction. Then, after this ranking is obtained, an algorithm is discussed to exploit this order to efficiently explore which features to remove from the current sufficient explanation until a cardinal-minimal explanation is obtained."},"soundness":{"value":4},"confidence":{"value":3},"questions":{"value":"**Q1:** In line 2 of Alg. 2, how is the operation $\\\\in$ implemented? Is this a uniform random sampling?\n\n**Q2:** How is the proposed algorithm applicable to NAMs for data types different from tabular data, like images and graphs ?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":4},"strengths":{"value":"- Making minimal sufficient explanations scaling is a timely and valuable research direction.\n\n- The proposed idea is simple but effective\n\n- The paper flows well\n\n- The paper is formal and precise in its claims"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"### Proof of Proposition 3 is unclear\n\n- **W1:** In line 842, the authors write *Let 1 ≤ l ≤ n represent the last feature added to $S$ in line ?? of Alg. 3*. However, I do not understand what this statement refers to, as in no part of Alg. 3 features are added to $S$.\n\n- **W2:** In line 843, the authors write *Then, for $S^{'}= S \\setminus \\\\{ l \\\\}$, it follows that: $suff(f, x, S, \\epsilon)$ does not hold true, implying that $S^{'}$ is not a sufficient explanation*, but it is unclear why this is true. I feel there could be some writing issues in this paragraph that prevent me from understanding the proof.\n\n\n\n### Definition of Sufficient explanation\n\n- **W3:** Definition 1 expresses the condition $\\\\forall \\\\tilde{x} \\\\in B_p^{\\\\epsilon_p}(x)$, where $B_p^{\\\\epsilon_p}(x) = \\\\{\\\\tilde{x} \\\\in R^{n} | ...\\\\}$. However, this condition seems to be impossible to satisfy, as real numbers are dense and therefore would require to run the sufficiency check for an infinite number of $\\\\tilde{x}$.\n\n\n- **W4:** The set of allowed perturbations is defined over an $\\\\epsilon$-ball around $x$. Nonetheless, defining a fixed $\\\\epsilon$ for every feature could hinder some feature-specific behaviors induced by, for example, different magnitudes of features. For example, let us consider a binary feature representing the gender of a person (0=male, 1=female), and a function $f_i(x_i)$ shaped as: $f_i(x_i)=100$ if $x_i < 0.5$, and  $f_i(x_i)=-100$ if $x_i >= 0.5$. Then a value of $\\\\epsilon=0.1$ will not be able to check for the sufficiency of the explanation when the gender of the person is switched from male to female, as this feature will be assigned zero importance even if its true impact is actually very high.\n\n\n### Figure 1 unclear\n\n- **W5:** Figure 1 is difficult to interpret, and its description is not self-contained. I suggest improving its description. For example, the part on *users might wrongly believe that only feature 1 yields negative outputs, while feature 2 can also flip the classification* is non-trivial to a non-expert reader.\n\n\n**Minors**\n\n- Propositions 2 and 3 are in reverse order in the Appendix.\n\n- The link between lines 46 and 47 is a bit unclear. In fact, while the first sentence in the paragraph is referred to sufficient explanations, the following sentence refers only to cardinal-minimal explanations, and not cardinal-minimal sufficient explanations, making the context of the sentence less clear.\n\n- In line 52, the authors write: *Consequently, existing methods focus on (locally) subsetminimal explanations which are typically suboptimal in size, potentially large, and thus less informative than their globally minimal counterparts*, which is, however, not clear why this is true without further context. I would recommend adding an explanation of why this is the case.\n\n\n- I personally find the wording local vs global sufficient explanation a bit misleading, as this could be confused with standard notions of local (instance-level) and global (model-level) explanations. I believe cardinal-minimal and subset-minimal are already discriminative enough and may not need further quantifications."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762931560659,"tcdate":1760718178623,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission19723/Reviewer_uiyA"],"signatures":["ICLR.cc/2026/Conference/Submission19723/Reviewer_uiyA"],"forum":"040ClRXMf3","number":1,"license":"CC BY 4.0","cdate":1760718178623,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission19723/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762931560659,"domain":"ICLR.cc/2026/Conference","replyto":"040ClRXMf3","id":"MthqrzzFcv","forumContent":{"TLDR":{"value":"Our approach constructs provably sufficient and (globally) cardinal-minimal explanations for neural additive models with improved runtime complexity."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["explainability","XAI","explainable AI","formal verification","sufficient explanations"]},"supplementary_material":{"value":"/attachment/688a5ff66ccb15d28a06f568b0f04b60f4413e61.zip"},"primary_area":{"value":"interpretability and explainable AI"},"abstract":{"value":"Despite significant progress in post-hoc explanation methods for neural\nnetworks, many remain heuristic and lack provable guarantees. A key approach\nfor obtaining explanations with provable guarantees is by identifying a\ncardinally-minimal subset of input features which by itself is provably\nsufficient to determine the prediction. However, for standard neural networks,\nthis task is often computationally infeasible, as it demands a worst-case\nexponential number of verification queries in the number of input features,\neach of which is NP-hard.\n  In this work, we show that for Neural Additive Models (NAMs), a recent and\nmore interpretable neural network family, we can efficiently generate\nexplanations with such guarantees. We present a new model-specific algorithm\nfor NAMs that generates provably cardinally-minimal explanations using only a\nlogarithmic number of verification queries\n  in the number of input features, after a parallelized preprocessing step with\nlogarithmic runtime in the required precision is applied to each small\nunivariate NAM component.\n  Our algorithm not only makes the task of obtaining cardinally-minimal\nexplanations feasible, but even outperforms existing algorithms designed to\nfind the relaxed variant of subset-minimal explanations - which may be larger\nand less informative but easier to compute - despite our algorithm solving a\nmuch more difficult task.\n  Our experiments demonstrate that, compared to previous algorithms, our\napproach provides provably smaller explanations than existing works and\nsubstantially reduces the computation time. Moreover, we show that our\ngenerated provable explanations offer benefits that are unattainable by\nstandard sampling-based techniques typically used to interpret NAMs."},"_bibtex":{"value":"@inproceedings{\nbassan2026provably,\ntitle={Provably Explaining Neural Additive Models},\nauthor={Shahaf Bassan and Yizhak Yisrael Elboher and Tobias Ladner and Volkan {\\c{S}}ahin and Jan Kretinsky and Matthias Althoff and Guy Katz},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=040ClRXMf3}\n}"},"title":{"value":"Provably Explaining Neural Additive Models"},"pdf":{"value":"/pdf/24a6ba7196e865623a6c47c236096913b01e68ba.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"bassan|provably_explaining_neural_additive_models"},"authorids":{"value":["~Shahaf_Bassan1","~Yizhak_Yisrael_Elboher1","~Tobias_Ladner1","~Volkan_Şahin1","~Jan_Kretinsky1","~Matthias_Althoff1","~Guy_Katz1"]},"authors":{"value":["Shahaf Bassan","Yizhak Yisrael Elboher","Tobias Ladner","Volkan Şahin","Jan Kretinsky","Matthias Althoff","Guy Katz"]}},"version":2},{"content":{"summary":{"value":"The paper proposes a long video understanding benchmark for multimodal large language models (MLLMs), focussing on videos over 30 minutes. The benchmark includes ~100 diverse videos with ~1550 QA paiers that test six core capabilities: temporal grounding, summarization, reasoning, entity recognition, event understanding, and key information retrieval. The paper studies several state-of-the-art MLLMs, both that focus on long videos natively and ones that don't. They key focus is on understanding the capabilities (or lack thereof) of current models in understanding lengthy video content."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"Please see the weaknesses listed.\n\nMinor questions:\n- Clarify how the \"Clue duration\" ground-truth was collected.\n- InternVL2-40B results: Why's there a discrepancy between Table 2 (39.8) vs. Table 3 (39.5).\n- Table 3 -- why can't these numbers be provided for all method studied as opposed to just the two best methods?"},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":2},"strengths":{"value":"+ The paper introduces a video understanding benchmark that emphasizes long videos, longer than prior works. \n+ The paper does a commendable job of enumerating different capabilities required for long video understanding\n+ The paper evaluates a total of 15 different MLLMs, which is a reasonably comprehensive list."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Benchmark:\n- While I appreciate the effort to create a benchmark with longer videos, as emphasized in the paper and Table 1, without enough scale, the benchmark can prove to be of limited utility. As noted in Table 1, the benchmark as 1549 QA pairs, making it the second smallest benchmark in terms of number of QA, with only the older ActivityNet-QA (Yu et al., 2019) having fewer QA pairs. It only has 103 videos, which is considered quite small by current standards. In fact, number of videos should have been a column in Table 1 to compare with existing  benchmarks, which would reveal how small the dataset is. While the videos in the proposed benchmark are longer, a large number of videos is still necessary for a credible video understanding benchmark, especially since the QA pairs in this paper are video-level rather than fixed-time-clip-level.\n- Although the proposed benchmark includes 21 video subcategories, it still lacks many common and significant types of videos, such as video games, non-cartoon educational videos, travel videos/vlogs, health and fitness videos, and technology videos, etc. Not necessarily a negative, but the paper should discuss other potential categories it does not cover so folks who use the results on benchmark should understand what it works or doesn't work on.\n- On a similar note, the benchmark is heavily skewed towards object and action recognition -- arguably two of the most studied tasks in recognition (Figure 3, >80% of questions come from Entity and Event recognition categories). \n- Another baseline missing is the following. Using video (or clip) caption generation methods to describe the video and then use LLMs to answer the questions using those descriptions/captions as input. \n\nExposition\n- The paper emphasizes several times that the questions are designed to be compositional, which can flexibly combine multiple skills to construct complex queries. However, from what I gather, all the questions belong to one of the categories listed in Section 3.2, and there's no composition of questions. It might be a language issue but compositionality of questions has a specific meaning and I'd recommend paraphrasing to not imply compositionality.\n- I like the \"Clue duration\" analysis, however, I don't understand how they ground-truth clue duration was collected? Were annotators asked to label this or was this estimated in an automated way?\n\nExperimentation\n- The paper discusses some details in L302-4 regarding how they adapt models not designed long videos. Example, by sampling 32 or 96 frames to provide as input. Now, I understand that if the models are not designed for long videos and/or have limited context length, then it's not easy to adapt MLLMs for long videos. However, providing 32 or 96 frames for avg. ~4000 seconds, without any sampling of frames, is a futile exercise when it comes to understanding long videos. The paper should either propose ways in which to adapt models that natively do not support long videos that is more meaningful or not have so many of them and focus on models that do support longer videos.\n- I like the analysis in Section 4.2.2, but I it is still unclear how much of the answer/distribution comes from actually understanding videos vs. LLMs inherent biases. It would help to use the LLMs for each model to answer the questions and provide that as lower bound for the results. I understand the authors did try filtering out questions using two LLM models, but for this to make sense, it has to be done using the LLM underlying the MLLM.\n- InternVL2-40B results: Why's there a discrepancy between Table 2 (39.8) vs. Table 3 (39.5).\n- Table 3 -- why can't these numbers be provided for all method studied as opposed to just the two best methods? It would help understand the video category axes of performance for all methods, and doesn't require any more experimentation.\n\nFindings:\n- While a pure benchmark and result paper is suffice, it would be good if the paper provided some discussion on potential directions to improve current MLLMs using the proposed benchmark. Another shortcoming is straight foward application of MLLMs not designed for long video understanding. I'd have preferred to see some effort to make these MLLMs work better on long video understanding.\n\nMinor typos:\nSection 3.2, all enumerated items are missing a space before the parenthesis (Grounding(TG) --> Grounding (TG))."}},"nonreaders":[],"tmdate":1731427871036,"tcdate":1730700561007,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission8887/Reviewer_n5MG"],"signatures":["ICLR.cc/2025/Conference/Submission8887/Reviewer_n5MG"],"forum":"uHgVrGF2Wn","number":3,"license":"CC BY 4.0","cdate":1730700561007,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission8887/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427871036,"domain":"ICLR.cc/2025/Conference","replyto":"uHgVrGF2Wn","id":"vK6GCMzb0t","forumContent":{"TLDR":{"value":"We introduce LVBench, a benchmark specifically designed for long video understanding."},"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video understanding","Multimodal learning","Visual question answering","Long-form video","Datasets and benchmarking"]},"primary_area":{"value":"datasets and benchmarks"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Recent progress in multimodal large language models has markedly enhanced the understanding of short videos (typically under one minute), and several evaluation datasets have emerged accordingly. However, these advancements fall short of meeting the demands of real-world applications such as embodied intelligence for long-term decision-making, in-depth movie reviews and discussions, and live sports commentary, all of which require comprehension of long videos spanning several hours. To address this gap, we introduce LVBench, a benchmark specifically designed for long video understanding. Our dataset comprises publicly sourced videos and encompasses a diverse set of tasks aimed at long video comprehension and information extraction. LVBench is designed to challenge multimodal models to demonstrate long-term memory and extended comprehension capabilities. Our extensive evaluations reveal that current multimodal models still underperform on these demanding long video understanding tasks. Through LVBench, we aim to spur the development of more advanced models capable of tackling the complexities of long video comprehension."},"_bibtex":{"value":"@misc{\nwang2024lvbench,\ntitle={{LVB}ench: An Extreme Long Video Understanding Benchmark},\nauthor={Weihan Wang and Zehai He and Wenyi Hong and Yean Cheng and Xiaohan Zhang and Ji Qi and Ming Ding and Xiaotao Gu and Shiyu Huang and Bin Xu and Yuxiao Dong and Jie Tang},\nyear={2024},\nurl={https://openreview.net/forum?id=uHgVrGF2Wn}\n}"},"title":{"value":"LVBench: An Extreme Long Video Understanding Benchmark"},"pdf":{"value":"/pdf/841779fddd82bcb7aa920cc3028ce329537a9552.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"wang|lvbench_an_extreme_long_video_understanding_benchmark"},"authorids":{"value":["~Weihan_Wang2","~Zehai_He1","~Wenyi_Hong1","~Yean_Cheng1","~Xiaohan_Zhang6","~Ji_Qi3","~Ming_Ding1","~Xiaotao_Gu1","~Shiyu_Huang2","~Bin_Xu1","~Yuxiao_Dong1","~Jie_Tang1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Weihan Wang","Zehai He","Wenyi Hong","Yean Cheng","Xiaohan Zhang","Ji Qi","Ming Ding","Xiaotao Gu","Shiyu Huang","Bin Xu","Yuxiao Dong","Jie Tang"]}},"version":2},{"content":{"summary":{"value":"This paper introduces PPLLaVA, a novel video understanding framework that improves both efficiency and performance by compressing visual tokens based on user prompts. The core idea is to reduce video redundancy using a prompt-guided 3D pooling mechanism, which dynamically compresses visual tokens while preserving instruction-relevant information. The model includes three key components: a CLIP-based vision-prompt alignment module to identify relevant visual regions; a prompt-guided pooling mechanism that performs 3D convolution-style pooling using prompt-based attention weights; a CLIP context extension module to support longer text inputs for multi-turn dialogues.  \nPPLLaVA achieves up to 18× token compression, and consistently outperforms strong baselines (e.g., LLaVA-OneVision, InternVL3) across 7 video understanding benchmarks, including long-form video tasks. It is also plug-and-play, transferable across different base models (LLaVA-Next, InternVL3, etc.)."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. Have you considered making the prompt-guided pooling weights learnable (e.g., via lightweight attention or gating mechanisms) instead of directly using CLIP similarity?\n2. How does the model perform when user prompts are vague or generic (e.g., “What is happening?”)? Is there a quantitative analysis of performance degradation in such cases?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. Innovative and Practical Solution: The prompt-guided pooling mechanism effectively tackles visual token redundancy in long video understanding, also enhancing performance in image-only tasks, indicating broad applicability.\n2. Strong Empirical Performance: PPLLaVA achieves state-of-the-art results across multiple benchmarks (e.g., Video-MME, NextQA, EgoSchema), often using fewer tokens and enabling faster inference.\n3. High Generalizability and Efficiency: The method demonstrates strong transferability across various base models (image-only, video-only, and unified VLMs), significantly improving throughput—up to 3× faster—with minimal to no accuracy loss."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The prompt-guided pooling weights, derived directly from CLIP similarity scores without learnable parameters or adaptive gating, may restrict expressiveness and adaptability.\n2. The paper lacks direct comparisons with other prompt-aware compression methods (e.g., VideoAgent, VideoTree) in terms of accuracy and efficiency, despite numerous SOTA model comparisons."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762919510336,"tcdate":1761901214820,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission7387/Reviewer_u53s"],"signatures":["ICLR.cc/2026/Conference/Submission7387/Reviewer_u53s"],"forum":"LOLhTA51tr","number":4,"license":"CC BY 4.0","cdate":1761901214820,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission7387/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762919510336,"domain":"ICLR.cc/2026/Conference","replyto":"LOLhTA51tr","id":"REtpsf0qi6","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Video LLM","Prompt-guided Pooling","PPLLaVA"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"The past year has witnessed the significant advancement of video-based large language models. However, the challenge of developing a unified model for both short and long video understanding remains unresolved. Most existing video LLMs cannot handle hour-long videos, while methods custom for long videos tend to be ineffective for shorter videos and images. In this paper, we identify the key issue as the redundant content in videos. To address this, we propose a novel pooling strategy that simultaneously achieves token compression and instruction-aware visual feature aggregation. Our model is termed Prompt-guided Pooling LLaVA, or PPLLaVA for short. Specifically, PPLLaVA consists of three core components: the CLIP-based visual-prompt alignment that extracts visual information relevant to the user's instructions, the prompt-guided pooling that compresses the visual sequence to arbitrary scales using convolution-style pooling, and the clip context extension designed for lengthy prompt common in visual dialogue. Extensive experiments have validated the performance of our model. With superior throughput, PPLLaVA achieves better results on image benchmarks as a video LLM, while achieving state-of-the-art performance across various video benchmarks, excelling in tasks ranging from caption generation to multiple-choice questions, and handling video lengths from seconds to hours."},"_bibtex":{"value":"@inproceedings{\nsun2026ppllava,\ntitle={{PPLL}a{VA}: Varied Video Sequence Understanding With Prompt Guidance},\nauthor={Shangkun Sun and Ruyang Liu and Haoran Tang and Yixiao Ge and Haibo Lu and Jiankun Yang and Chen Li},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=LOLhTA51tr}\n}"},"title":{"value":"PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance"},"pdf":{"value":"/pdf/38c0e8e608ff660b862b3f27a2ed2041e8f50c9c.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"sun|ppllava_varied_video_sequence_understanding_with_prompt_guidance"},"authorids":{"value":["~Shangkun_Sun3","~Ruyang_Liu1","~Haoran_Tang4","~Yixiao_Ge2","~Haibo_Lu1","~Jiankun_Yang1","~Chen_Li34"]},"authors":{"value":["Shangkun Sun","Ruyang Liu","Haoran Tang","Yixiao Ge","Haibo Lu","Jiankun Yang","Chen Li"]}},"version":2},{"content":{"summary":{"value":"This paper aims to tackle the generalizability and counterfactual robustness of foundation models on theory-of-mind (ToM) reasoning tasks. They first design a principled filtering approach that identifies the shortcut issue in the existing ToM benchmarks, especially for the datasets like Hi-ToM higher-order queries. After identifying the shortcut-free dataset, the authors further make a comprehensive comparison among the performance of zero-shot and different post-training approaches, including SFT and RFT (with or without thinking tokens). The results demonstrate that RFT with thinking outperforms other post-training approaches in most tasks. RFT enjoys not only the best accuracy on different ToM benchmarks, but also better generalizability to higher-order reasoning cases and better causal consistency to counterfactual probing. The qualitative studies on the attention map also reveal that the model after RFT is more capable of capturing the semantics in the contexts."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"See the weakness section for the main questions I have and the comments. I would consider re-adjust my assessment if the majority of them get resolved during the rebuttal phase."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"* The paper is generally well-written and easy to follow. The clarity on the experiment settings is good. \n* The shortcut perspective on the current ToM benchmark is interesting and worth studying deeper for the ToM research community. \n* The empirical results on different post-training approaches, as well as the qualitative studies are clearly presented and can justify the major claim of contribution in this paper."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"* **Incoherent contribution in section 2 and 3:** It's a little bit unclear what is the necessary logical connection between section 2 and 3. Section 2 only identifies those dataset that potentially has confounded shortcuts that should be precluded from post-training. However, these datasets with shortcut are also held away from evaluation set in the final experiments with generalization. \n* **Missing baselines:** There exists quite a few baselines in solving ToM-QA tasks, such as SimToM [1], AutoToM [2]. However, the paper misses these baselines that are currently leading approaches in the MMToM benchmark. \n* **Limited analysis on the higher-order generalizability:** Since the authors only use shortcut-free dataset like OpenToM to conduct post-training and evaluation, the claim on 'generalizability to higher-order queries' does not seem to be sufficiently supported as they only care about the generalization from first-order to second-order ToM. A more comprehensive studies on third-order and fourth-order ToM (like the ones in Hi-ToM) will be interesting. \n* **Missing evaluation on cross-dataset generalizability:** It is hard to justify whether the results the authors present are coming from better overfitting to a single dataset, or indeed a better capability in general reasoning across different ToM datasets. Therefore, it is necessary to add the evaluation on different ToM datasets after the SFT/RFT, rather than a small test split which is less different in the QA domain. \n* **Limited scale of evaluation:** The OpenToM dataset, the authors only select 100 samples from each category as evaluation. They also missed the evaluation on the test split provided by the MMToM leaderboard. These limits the contribution of the proposed method. \n* **Limited scale of qualitative analysis**: The author only demonstrates one pair of qualitative comparisons on the attention map. More qualitative examples can be provided in the appendix to make the claim of causality-coherent reasoning more solid. \n* **Missing analysis on the 'spurious correlation'**: The authors mentioned their motivation comes from the observation that the model 'simply exploiting spurious correlations'. However, the analysis terminates after section 2 right after they preclude the Hi-COM and other datasets in the finetuning and evaluation dataset. It will be more reasonable if they can conduct **quantitative** analysis on how (a) the data filtering (judged by simple rules and lexical association), as well as (b) the SFT/RFT-style post-training, can help mitigate such spurious correlation exploitation, even on those benchmarks with shortcuts. \n\n> [1] Wilf, Alex, et al. \"Think twice: Perspective-taking improves large language models' theory-of-mind capabilities.\" ACL 2024.\n>\n> [2] Zhang, Zhining, et al. \"Autotom: Automated bayesian inverse planning and model discovery for open-ended theory of mind.\" *ICLR 2025 Workshop on Foundation Models in the Wild*. 2025."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1764360073906,"tcdate":1761599349187,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission1862/Reviewer_MqzA"],"signatures":["ICLR.cc/2026/Conference/Submission1862/Reviewer_MqzA"],"forum":"BsEYEXjkVO","number":2,"license":"CC BY 4.0","cdate":1761599349187,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission1862/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1764360073906,"domain":"ICLR.cc/2026/Conference","replyto":"BsEYEXjkVO","id":"FRsUmlmKje","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["theory of mind","reasoning","reinforcement finetuning","large language model"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Theory of Mind (ToM) is a must-acquire skill for modern foundation model systems to operate effectively and safely in the real world. Recent works have explored honing ToM via post-training; however, we show that such progress is confounded by a pervasive “shortcut” issue: tasks can reach up to 99% accuracy by simply exploiting spurious causal correlations, leading to a false sense of ToM. Motivated by this, we first develop a framework to systematically examine ToM datasets for shortcuts and provide guidance for future development. We find that questions reducible to pure state tracking (e.g., “belief”) are especially shortcut-prone compared to mind questions (e.g., “intention”) where reasoning beyond tracking is required. Using four shortcut-free datasets across three ToM contexts, we then comprehensively study whether reinforcement-learning fine-tuning with verifiable rewards and explicit reasoning (Thinking-RFT) elevates ToM beyond supervised fine-tuning (SFT). Our key findings are: 1) Thinking-RFT effectively improves ToM in all scenarios (+6% vs. SFT), particularly in complex higher-order reasoning (+10% vs. SFT) and multimodal cases (+7% vs. SFT), and generalizes notably better to unseen domains and higher-order queries while being more robust to counterfactuals. 2) ToM benefits specifically from the joint effect of reasoning and RL: Thinking-RFT outperforms No-Thinking-RFT by 7% on average. 3) RFT works by learning to ground its reasoning on anchor cues (keywords/state changes) that correspond to causal factors. We believe our study is useful for developing effective and robust ToM post-training datasets and advancing critical ToM capabilities in foundation models."},"_bibtex":{"value":"@misc{\nzhong2026from,\ntitle={From Shortcuts to Reasoning: Robust Post-Training of Theory of Mind with Reinforcement Learning},\nauthor={Jike Zhong and Yuxiang Lai and Ming Li and Yuheng Li and Wuao Liu and Behzad Dariush and Konstantinos Psounis and Shao-Yuan Lo},\nyear={2026},\nurl={https://openreview.net/forum?id=BsEYEXjkVO}\n}"},"title":{"value":"From Shortcuts to Reasoning: Robust Post-Training of Theory of Mind with Reinforcement Learning"},"pdf":{"value":"/pdf/81759a9364a3e76b0e9cc8d4b9ce2a7b855c88a7.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"zhong|from_shortcuts_to_reasoning_robust_posttraining_of_theory_of_mind_with_reinforcement_learning"},"authorids":{"value":["~Jike_Zhong1","~Yuxiang_Lai1","~Ming_Li28","~Yuheng_Li4","~Wuao_Liu1","~Behzad_Dariush2","~Konstantinos_Psounis1","~Shao-Yuan_Lo1"]},"authors":{"value":["Jike Zhong","Yuxiang Lai","Ming Li","Yuheng Li","Wuao Liu","Behzad Dariush","Konstantinos Psounis","Shao-Yuan Lo"]}},"version":2},{"content":{"summary":{"value":"This paper explores the use of video diffusion transformers (Video DiTs) as feature backbones for point tracking, aiming to leverage their global 3D attention and large-scale real-world pretraining to improve robustness under large motion, occlusion, and motion blur. The authors propose a simple two-stage adaptation: (i) an upsampler module to recover spatial detail and fuse multi-layer DiT features; (ii) an iterative refiner adopted from common point trackers for high-precision trajectory estimation. Experiments on TAP-Vid DAVIS and Kinetics datasets are reported."},"soundness":{"value":2},"confidence":{"value":5},"questions":{"value":"See weaknesses."},"rating":{"value":2},"details_of_ethics_concerns":{"value":"None"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":1},"strengths":{"value":"* Diffusion models are indeed a natural candidate for geometric perception tasks, and this work provides an initial empirical exploration of how their temporal reasoning could benefit pixel-level correspondence.\n* Conceptually simple and clear idea. The paper integrates video DiTs into a point tracking framework with minimal modifications and clear motivation.\n* Well-written and easy to follow, with coherent structure and extensive visualizations (cost volumes, qualitative examples)."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"* While the paper claims that Video DiTs inherently capture temporal consistency, results (Table 3) show that they underperform most image-based backbones, even when equipped with the upsampler and refiner. Given that image-based backbones model temporal relations after feature extraction, Video DiTs should have an advantage in these zero-shot settings. Yet, their performance remains lower, contradicting the claimed benefit of 3D attention in Section 3.2.\n* The proposed upsampler leads to inconsistent gains across datasets: improvements on DAVIS but a drop on Kinetics. This suggests that the design may not generalize well and the gains are dataset-specific. The paper should analyze why the upsampler struggles in long or diverse sequences.\n* Evaluation is limited to two datasets (DAVIS, Kinetics). To establish the generality of diffusion-based backbones, additional domains, such as RoboTAP (robotics), should be tested. These would better reflect the claim of robustness to real-world artifacts.\n* The scale gap between Video DiTs (billions of parameters) and ResNet-based CoTracker3 backbones (tens of millions) raises fairness concerns. The current results may conflate improvements from model capacity rather than architectural superiority. The authors should either (a) compare against equally large vision backbones (eg DINOv3 7B, SwinV2 3B), or (b) use smaller DiTs to isolate the temporal modeling effect.\n* Video DiTs are known for high memory and computational cost. Section A mentions 4 A6000 GPUs for 15 K iterations for training, but memory usage and runtime are not quantified for inference.\n\n### Overall\nThe paper explores a promising direction, integrating video diffusion transformers into point tracking, but lacks strong empirical evidence and insight on why and how this integration improves performance. Most observed gains are modest or dataset-specific, and core claims about temporal reasoning and real-world robustness are not convincingly demonstrated. If extended with deeper analysis of DiT adaptation, this line of work could meaningfully contribute to understanding how generative video priors support geometric perception."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762917511562,"tcdate":1761649456603,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission4685/Reviewer_3P9a"],"signatures":["ICLR.cc/2026/Conference/Submission4685/Reviewer_3P9a"],"forum":"bhuFwR5rOS","number":3,"license":"CC BY 4.0","cdate":1761649456603,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission4685/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762917511562,"domain":"ICLR.cc/2026/Conference","replyto":"bhuFwR5rOS","id":"pPGqHmdFfV","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Video Diffusion Models","Point Tracking","Visual Correspondence"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Point tracking aims to estimate pixel trajectories across video frames but remains challenging under large displacements, occlusion, and real-world artifacts. Conventional trackers, built on image-centric backbones and synthetic training, often fail in these settings. We revisit this problem through the lens of video diffusion models based on Diffusion Transformers (DiTs), whose 3D global attention structure and large-scale training naturally provide global temporal context and real-world priors. We first analyze the intrinsic robustness of video DiT features, showing stronger correlation maps than supervised ResNet backbones even under occlusion and motion blur. To fully exploit these properties, we introduce an upsampler that restores spatial detail while fusing multi-layer features, followed by an iterative refiner for high-precision trajectories. Extensive experiments on TAP-Vid benchmarks demonstrate that our framework achieves superior robustness and accuracy compared to existing backbones, establishing video DiTs as powerful foundations for point tracking."},"_bibtex":{"value":"@misc{\nson2025video,\ntitle={Video Diffusion Model for Point Tracking},\nauthor={Soowon Son and Honggyu An and Chaehyun Kim and Jung Yi and Hyunah Ko and Jisu Nam and Jaewon Min and Dahyun Chung and Siyoon Jin and Jiyoung Kim and Junhwa Hur and Seungryong Kim},\nyear={2025},\nurl={https://openreview.net/forum?id=bhuFwR5rOS}\n}"},"title":{"value":"Video Diffusion Model for Point Tracking"},"pdf":{"value":"/pdf/0466899ffea0da630f5f75bed5b771bebc69ee9b.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"son|video_diffusion_model_for_point_tracking"},"authorids":{"value":["~Soowon_Son1","~Honggyu_An1","~Chaehyun_Kim1","~Jung_Yi1","~Hyunah_Ko1","~Jisu_Nam1","~Jaewon_Min1","~Dahyun_Chung1","~Siyoon_Jin1","~Jiyoung_Kim2","~Junhwa_Hur1","~Seungryong_Kim1"]},"authors":{"value":["Soowon Son","Honggyu An","Chaehyun Kim","Jung Yi","Hyunah Ko","Jisu Nam","Jaewon Min","Dahyun Chung","Siyoon Jin","Jiyoung Kim","Junhwa Hur","Seungryong Kim"]}},"version":2},{"content":{"summary":{"value":"The paper proposes UNIC, a simple, adapter-free framework that unifies diverse video-editing tasks, such as ID insert/delete/swap, stylization, first-frame propagation, and re-camera control within one model by concatenating three token types (noisy video latent, reference-video tokens, and multi-modal condition tokens) into a single sequence processed by a DiT with full attention. UNIC features two key designs: Condition Bias (task-type embeddings) and Task-aware RoPE (per-task positional indexing), which mitigate task ambiguity and alignment conflicts. On a six-task benchmark, UNIC reports competitive or superior results to task specialists, shows emergent task composition, and includes ablations and efficiency analyses (e.g., step-cache)."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. In appendix E.3, \"All finetuning experiments start from a pre-trained model with 1B parameters and 28 sequential standard Diffusion Transformer (DiT) blocks.\" What is this base pretrained model?"},"rating":{"value":6},"details_of_ethics_concerns":{"value":"N/A"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. UNIC reformulates many video editing tasks as tokenized conditions concatenated with noisy and reference tokens.\n2. UNIC supports six representative tasks and demonstrates emergent compositions (e.g. stylization + re-camera)."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. I think the paper could benefit from reorganizing its structure for improved presentation. One of the most important parts of training video editing models is in its data construction pipeline, since paired before-and-after editing data is often difficult to obtain. Right now the main text barely has any description of training data, and all of descriptions are in the appendix.\n2. Following weakness 1, there are no statistics about the quantity of the training data in the main text. In Appendix D, there are statistics for some tasks, but not all tasks.\n3. While it is clear that both condition bias and task-aware RoPE benefit model performance, it seems that combining these two designs leads to weaker performance on the style transfer task, and the results of ArtFID and CFSD are even worse than not having the two designs at all. Can the author provide more intuition on this issue?"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762922904531,"tcdate":1761862039549,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission11891/Reviewer_CmCf"],"signatures":["ICLR.cc/2026/Conference/Submission11891/Reviewer_CmCf"],"forum":"Vb4nE3WWf5","number":2,"license":"CC BY 4.0","cdate":1761862039549,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission11891/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762922904531,"domain":"ICLR.cc/2026/Conference","replyto":"Vb4nE3WWf5","id":"ServXaYb4s","forumContent":{"TLDR":{"value":"a parameter-efficient and unified framework for video editing tasks"},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["video editing; video generation; diffusion models"]},"primary_area":{"value":"generative models"},"abstract":{"value":"Recent advances in text-to-video generation have sparked interest in generative video editing tasks. Previous methods often rely on task-specific architectures (e.g., additional adapter modules) or dedicated customizations (e.g., DDIM inversion), which limit the integration of versatile editing conditions and the unification of various editing tasks. In this paper, we introduce UNified In-Context Video Editing (UNIC), a simple yet effective framework that unifies diverse video editing tasks within a single model in an in-context manner. To achieve this unification, we represent the inputs of various video editing tasks as three types of tokens: the source video tokens, the noisy video latent, and the multi-modal conditioning tokens that vary according to the specific editing task. Based on this formulation, our key insight is to integrate these three types into a single consecutive token sequence and jointly model them using the native attention operations of DiT, thereby eliminating the need for task-specific adapter designs. Nevertheless, direct task unification under this framework is challenging, leading to severe token collisions and task confusion due to the varying video lengths and diverse condition modalities across tasks. To address these, we introduce task-aware RoPE to facilitate consistent temporal positional encoding, and condition bias that enables the model to clearly differentiate different editing tasks. This allows our approach to adaptively perform different video editing tasks by referring the source video and varying condition tokens \"in context\", and support flexible task composition. To validate our method, we construct a unified video editing benchmark containing six representative video editing tasks. Results demonstrate that our unified approach achieves comparable performance with task specialists and exhibits emergent task composition abilities."},"_bibtex":{"value":"@inproceedings{\nye2026unified,\ntitle={Unified In-Context Video Editing},\nauthor={Zixuan Ye and Xuanhua He and Quande Liu and Qiulin Wang and Xintao Wang and Pengfei Wan and Di ZHANG and Kun Gai and Qifeng Chen and Wenhan Luo},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=Vb4nE3WWf5}\n}"},"title":{"value":"Unified In-Context Video Editing"},"pdf":{"value":"/pdf/b037cc220716c285a2c27112cbc052e5b1e7ef62.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"ye|unified_incontext_video_editing"},"authorids":{"value":["~Zixuan_Ye1","~Xuanhua_He1","~Quande_Liu1","~Qiulin_Wang1","~Xintao_Wang1","~Pengfei_Wan1","~Di_ZHANG3","~Kun_Gai1","~Qifeng_Chen1","~Wenhan_Luo1"]},"authors":{"value":["Zixuan Ye","Xuanhua He","Quande Liu","Qiulin Wang","Xintao Wang","Pengfei Wan","Di ZHANG","Kun Gai","Qifeng Chen","Wenhan Luo"]}},"version":2},{"content":{"venue":{"value":"ICLR 2027 Conference Submission"},"keywords":{"value":["360° Video Editing;Video Generation;Video Removal"]},"supplementary_material":{"value":"/attachment/a420b18a43ada953c32454db0b439d46a7ab8dd2.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Panoramic video object removal aims to remove targets and associated effects while reconstructing occluded content across the viewing sphere. However, existing perspective-based methods struggle with panoramic videos, where ERP-induced seam discontinuities and polar distortions challenge both spherical reconstruction and target-effect association, while paired panoramic removal data remain scarce. To address these challenges, we construct PANOVOR, a paired panoramic video removal dataset containing 1.53M panoramic frames spanning 26.6 hours, with audited source-clean pairs and target masks. Building on PanoVOR, we propose PANOERASE, a unified model that incorporates spherical geometry into panoramic object and effect removal. PanoErase first learns panorama-aware completion from large-scale panoramic videos using SC-RoPE, which models spherical positional relationships. It is then refined with paired removal supervision through SphericalControl, which combines target-specific local evidence with panorama-wide context for effect-aware removal while preserving unrelated content. Experiments on PanoVOR-Eval and PanoVOR-Wild demonstrate consistent improvements over existing object removal methods, particularly in seam-crossing and polar regions. The dataset, code and models will be released publicly. The anonymous project website is available at https://panoerase-anonymous.pages.dev/."},"_bibtex":{"value":"@inproceedings{\nanonymous2026panoerasespherical,\ntitle={PanoErase:Spherical Geometry-Aware Panoramic Video Object Removal},\nauthor={Anonymous},\nbooktitle={Submitted to The Fifteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=F9WNs3dn5t},\nnote={under review}\n}"},"title":{"value":"PanoErase:Spherical Geometry-Aware Panoramic Video Object Removal"},"pdf":{"value":"/pdf/3628366bd7d1ad6f3213f655334f5aca4f75241b.pdf"},"venueid":{"value":"ICLR.cc/2027/Conference/Submission"}},"tmdate":1791228110567,"tcdate":1789370049389,"writers":["ICLR.cc/2027/Conference","ICLR.cc/2027/Conference/Submission17238/Authors"],"signatures":["ICLR.cc/2027/Conference/Submission17238/Authors"],"forum":"F9WNs3dn5t","license":"CC BY 4.0","number":17238,"cdate":1789370049389,"readers":["everyone"],"invitations":["ICLR.cc/2027/Conference/-/Submission","ICLR.cc/2027/Conference/Submission17238/-/Full_Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Bidding","ICLR.cc/2027/Conference/-/Submission_Change_Before_Reviewing","ICLR.cc/2027/Conference/-/Edit"],"mdate":1791228110567,"odate":1791031328286,"domain":"ICLR.cc/2027/Conference","id":"F9WNs3dn5t","version":2},{"content":{"summary":{"value":"The manuscript presents an Action Dynamics Benchmark (ActionBench) with three new evaluation metrics for existing video-language models. Through ActionBench, the authors find that existing video-language models essentially rely on recognizing objects to recognize actions. Hence, a parameter-efficient component named Knowledge Patcher is connected in parallel to the video-language model and is trained with Discriminative Video Dynamics Modeling to improve the action understanding ability of the model. Finally, the authors present a knowledge fuser to infuse the knowledge learned by the Knowledge Patcher into the video-language model for downstream tasks. Empirical results show that training a knowledge patcher for use is effective in improving the video-text retrieval and temporal VQA tasks. \n"},"soundness":{"value":"3 good"},"confidence":{"value":"4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work."},"questions":{"value":"- To help object-centric understanding, why don't you use the pre-trained representations from the backbone as well (e.g., concatenate or add them to the KP during fine-tuning)?"},"rating":{"value":"7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations."},"code_of_conduct":{"value":"Yes"},"presentation":{"value":"4 excellent"},"contribution":{"value":"3 good"},"strengths":{"value":"- Although it is proposed before that current video models may rely on spatial information to recognize actions in videos, the manuscript is the first to present such a benchmark with proper evaluation metrics to show the problem. This can inspire further work to solve the problem in video understanding. \n\n- The Discriminative Video Dynamics Modeling is straightforward and effective in training the video-language model to be aware of the action and the temporal direction of the videos. \n\n- The approach that the manuscript presents is similar to post-pre-training, and it is shown in Table 2 and Table 3 that such a post-pre-training strategy can benefit downstream retrieval tasks. \n\n- The writing is good and the organization of the manuscript is clear. "},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The evaluation in the downstream task is limited to video-text retrieval on SSV2 and VQA on NExT-QA, which makes it difficult to compare the presented approach with existing methods. Since the evaluation is based on InternVideo backbone, it is possible to provide some other comparisons where InternVideo has some published results. \n\n- Further on the above note, I am not quite familiar with VQA but on SSV2, the performance seems oddly low. Especially for video-to-text retrieval, which is essentially a text-driven classification for SSv2 videos, the performance has reached 61% for EVL-B/16 [1] when 8 frames are used, which also connect a side network to the pre-trained encoder. I didn't find details in both the manuscript and the supplemental material why this is the case. Could you provide an explanation for this? \n\n[1] Frozen CLIP Models are Efficient Video Learners\n"},"limitations":{"value":"See weaknesses."}},"nonreaders":[],"tmdate":1702411250203,"tcdate":1688654925807,"writers":["NeurIPS.cc/2023/Conference","NeurIPS.cc/2023/Conference/Submission9752/Reviewer_SP6c"],"signatures":["NeurIPS.cc/2023/Conference/Submission9752/Reviewer_SP6c"],"forum":"blm1pqiOXe","number":3,"license":"CC BY 4.0","cdate":1688654925807,"mdate":1702411250203,"readers":["everyone"],"invitations":["NeurIPS.cc/2023/Conference/Submission9752/-/Official_Review","NeurIPS.cc/2023/Conference/-/Edit"],"domain":"NeurIPS.cc/2023/Conference","replyto":"blm1pqiOXe","id":"RlLL4yAL5C","forumContent":{"venue":{"value":"NeurIPS 2023 spotlight"},"keywords":{"value":["video-language model","action knowledge benchmarking","action understanding","temporal understanding"]},"supplementary_material":{"value":"/attachment/3e1ca2e010d52351ce0d1c23fe50b309c0627091.zip"},"_bibtex":{"value":"@inproceedings{\nwang2023paxion,\ntitle={Paxion: Patching Action Knowledge in Video-Language Foundation Models},\nauthor={Zhenhailong Wang and Ansel Blume and Sha Li and Genglin Liu and Jaemin Cho and Zineng Tang and Mohit Bansal and Heng Ji},\nbooktitle={Thirty-seventh Conference on Neural Information Processing Systems},\nyear={2023},\nurl={https://openreview.net/forum?id=blm1pqiOXe}\n}"},"title":{"value":"Paxion: Patching Action Knowledge in Video-Language Foundation Models"},"paperhash":{"value":"wang|paxion_patching_action_knowledge_in_videolanguage_foundation_models"},"TLDR":{"value":"Benchmarking and enhancing video-language foundation models with action knowledge"},"abstract":{"value":"Action knowledge involves the understanding of textual, visual, and temporal aspects of actions. We introduce the **Action Dynamics Benchmark (ActionBench)** containing two carefully designed probing tasks: Action Antonym and Video Reversal, which targets multimodal alignment capabilities and temporal understanding skills of the model, respectively. Despite recent video-language models’ (VidLM) impressive performance on various benchmark tasks, our diagnostic tasks reveal their surprising deficiency (near-random performance) in action knowledge, suggesting that current models rely on object recognition abilities as a shortcut for action understanding. To remedy this, we propose a novel framework, **Paxion**, along with a new **Discriminative Video Dynamics Modeling (DVDM)** objective. The Paxion framework utilizes a **Knowledge Patcher** network to encode new action knowledge and a **Knowledge Fuser** component to integrate the Patcher into frozen VidLMs without compromising their existing capabilities. Due to limitations of the widely-used Video-Text Contrastive (VTC) loss for learning action knowledge, we introduce the DVDM objective to train the Knowledge Patcher. DVDM forces the model to encode the correlation between the action text and the correct ordering of video frames. Our extensive analyses show that Paxion and DVDM together effectively fill the gap in action knowledge understanding (~50% → 80%), while maintaining or improving performance on a wide spectrum of both object- and action-centric downstream tasks."},"pdf":{"value":"/pdf/75f70ba6a193d347accace7069e781ae098ccee9.pdf"},"venueid":{"value":"NeurIPS.cc/2023/Conference"},"authorids":{"value":["~Zhenhailong_Wang1","~Ansel_Blume1","~Sha_Li1","~Genglin_Liu1","~Jaemin_Cho1","~Zineng_Tang1","~Mohit_Bansal2","~Heng_Ji3"]},"authors":{"value":["Zhenhailong Wang","Ansel Blume","Sha Li","Genglin Liu","Jaemin Cho","Zineng Tang","Mohit Bansal","Heng Ji"]}},"version":2},{"content":{"summary":{"value":"The authors of the paper propose a multimodal large language model (MLLM) called SafeWatch, designed to follow customized safety policies and provide multi-label video guardrail categorical outputs with answer explanations in a zero-shot manner. They also introduce SafeWatch-Bench, a large-scale video guardrail benchmark containing over 2 million videos spanning 6 safety broad categories and covering over 30 finer-grained risk categories to ensure comprehensive coverage of potential safety scenarios.\n\nThe technical contributions include:\n- Model Design: The authors introduce two key plug-and-play modules: Parallel Equivalent Policy Encoding (PEPE) and Policy-Aware Adaptive Pruning (PAP). \n  - PEPE mitigates high latency from extensive input contexts and policy positional bias by dividing lengthy safety guidelines into independent chunks encoded in parallel with equal importance. \n  - PAP, on the other hand, reduces latency by selecting the most relevant visual tokens for each policy while discarding those with low relevance.\n\n- Data: Each instance in SafeWatch-Bench is annotated with multi-label guardrail categories and detailed explanations. The dataset includes 2 million videos—both real-world and generative from various SOTA models—comprising an instruction-tuning set and a test set of 1K hand-selected, high-quality annotated videos across subcategories.\n\n- Training strategy: The authors fine-tune InternVL2-8B with their modeling changes on this new data via three stages, i.e., multi-task training, adaptive-pruning training, and preference post-tuning. \n  - Stage 1: Only PEPE is trained during this stage on a large corpus of unsafe videos, as well as traditional VQA and captioning tasks on normal videos. \n  - Stage 2: Both PEPE and PAP are fine-tuned on guardrail tasks. \n  - Stage 3: Preference pairs are curated to enable the preference post-tuning."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. Could you provide some evaluation results for Safety-aware Event Sampling? Currently, its effectiveness or limitations are unclear.\n2. Is SafeWatch-Bench truly a video benchmark? Specifically, does it truly require reasoning across multiple frames for a model to achieve high performance?\n3. It appears that humans do not directly provide annotations for SafeWatch-Bench; instead, annotations are model-generated, with human reviewers checking if re-annotation is necessary. In the caption for Figure 2, you mention that the 1K test set has high-quality annotations. Were these test set annotations created directly by humans, or were they produced via the multi-agent propose-discuss-consensus pipeline?\n4. Could you describe how the preference pairs were curated for the preference post-tuning stage? Additionally, how were the challenging benign examples—those easily identified by humans but likely to mislead guardrail models—selected? A detailed explanation of these curation processes would be helpful.\n5. What is the quality of the synthetic videos generated by the GenAI models? Are they accurately aligned with the unsafe prompts? Was any video filtering applied to filter out misaligned videos?\n6. Could you clarify your model training data recipe? Are the unsafe videos used in stage-1 training and the guardrail tasks in stage-2 training drawn from the instruction-tuning dataset within SafeWatch-Bench? Additionally, how many samples are used for each stage, task, and dataset?\n7. Which layers are tuned in the preference post-tuning stage?\n8. In Table 3 of Appendix A.1, there is a column named \"Temporal location\". What does it mean?\n9. What is the SFT Baseline mentioned in Figure 5 and Table 5?\n10. It was claimed that prior datasets lack detailed descriptions of the videos, suggesting that SafeWatch-Bench offers a detailed description for each video. Is that correct?\n11. There is unsafe content in SafeWatch-Bench. Will SafeWatch and SafeWatch-Bench be released? If so, how do you plan to ensure their proper use?\n\nMinor comments:\nLine 749 has placeholder text."},"rating":{"value":8},"details_of_ethics_concerns":{"value":"The authors claim data contributions, but there is unsafe content in SafeWatch-Bench."},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":3},"strengths":{"value":"1. Their model, SafeWatch, outperforms SOTA video guardrails on SafeWatch-Bench by 19.6% and on existing benchmarks by 15.4%, while reducing inference costs by an average of 25%. SafeWatch also demonstrates strong policy-following abilities, outperforming baselines by 20% in zero-shot adaptability to new policies. Additionally, both LLM-as-a-judge and human evaluators confirm the high quality of the explanations provided by SafeWatch.\n2. The design choices are well-founded, following best practices for efficient MLLM construction.\n3. This is an important area of study, with meaningful contributions (if these contributions are reproducible)."},"flag_for_ethics_review":{"value":["Yes, Discrimination / bias / fairness concerns","Yes, Privacy, security and safety","Yes, Responsible research practice (e.g., human subjects, data release)"]},"weaknesses":{"value":"1. Missing Evaluation: The evaluation of the Safety-aware Event Sampling step is absent. This step should be crucial for model performance, as the authors used TransnetV2 to segment videos into safety-aware events, sampling a single frame per event for further MLLM processing.\n2. Dataset Clarity: The dataset’s specifics and its exact use in model training remain unclear. For instance, there is no information on the average video length or the typical length of an explanation in SafeWatch-Bench. Additionally, the quality of the SafeWatch-Bench test set is not fully addressed, which is particularly important for an evaluation dataset.\n3. Reproducibility Concerns: Reproducibility in data collection and model training is questionable. For instance, Section 4.2 on \"multi-agent consensus video annotation\" provides a basic idea of the processes but lacks sufficient detail for replication (e.g., missing the prompts used, configurations such as the number of frames used for each model, etc.). Additional issues are noted in the \"Questions\" section."}},"nonreaders":[],"tmdate":1732681377590,"tcdate":1730346723979,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission9949/Reviewer_YtKs"],"signatures":["ICLR.cc/2025/Conference/Submission9949/Reviewer_YtKs"],"forum":"xjKz6IxgCX","number":2,"license":"CC BY 4.0","cdate":1730346723979,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission9949/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732681377590,"domain":"ICLR.cc/2025/Conference","replyto":"xjKz6IxgCX","id":"8sFuEVadMA","forumContent":{"TLDR":{"value":"We propose an efficient MLLM-based video guardrail model and a large-scale video safety benchmark dataset to enforce safety policy compliance and transparent explanations."},"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video Guardrail Model","Safe Foundation Models","Efficient LLMs Inference","LLM Safety","Multimodal Foundation Models"]},"primary_area":{"value":"alignment, fairness, safety, privacy, and societal considerations"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"With the rise of generative AI and rapid growth of high-quality video generation, video guardrails have become more crucial than ever to ensure safety and security across platforms. Current video guardrails, however, are either overly simplistic, relying on pure classification models trained on simple policies with limited unsafe categories, which lack detailed explanations, or prompting multimodal large language models (MLLMs) with long safety guidelines, which are inefficient and impractical for guardrailing real-world content. To bridge this gap, we propose SafeWatch, an efficient MLLM-based video guardrail model designed to follow customized safety policies and provide multi-label video guardrail outputs with content-specific explanations in a zero-shot manner. In particular, unlike traditional MLLM-based guardrails that encode all safety policies autoregressively, causing inefficiency and bias, SafeWatch uniquely encodes each policy chunk in parallel and eliminates their position bias such that all policies are attended simultaneously with equal importance. In addition, to improve efficiency and accuracy, SafeWatch incorporates a policy-aware visual token pruning algorithm that adaptively selects the most relevant video tokens for each policy, discarding noisy or irrelevant information. This allows for more focused, policy-compliant guardrail with significantly reduced computational overhead. Considering the limitations of existing video guardrail benchmarks, we propose SafeWatch-Bench, a large-scale video guardrail benchmark comprising over 2M videos spanning six safety categories which covers over 30 tasks to ensure a comprehensive coverage of all potential safety scenarios. We have conducted extensive experiments, showing that SafeWatch outperforms all SOTA video guardrails on SafeWatch-Bench by 28.2%, and achieves a 13.6% improvement on existing benchmarks, all while reducing inference costs by an average of 10%. SafeWatch also demonstrates strong policy-following abilities and outperforms previous SOTAs by 5.6% and 15.6% in zero-shot generalizability to new policies and new prompting tasks. Additionally, both LLM-as-a-judge and human evaluators confirm the high quality of the explanations provided by SafeWatch. Our project is open-sourced at https://safewatch-aiguard.github.io."},"_bibtex":{"value":"@inproceedings{\nchen2025safewatch,\ntitle={SafeWatch: An Efficient Safety-Policy Following Video Guardrail Model with Transparent Explanations},\nauthor={Zhaorun Chen and Francesco Pinto and Minzhou Pan and Bo Li},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=xjKz6IxgCX}\n}"},"title":{"value":"SafeWatch: An Efficient Safety-Policy Following Video Guardrail Model with Transparent Explanations"},"pdf":{"value":"/pdf/bebab0da9d02c81b7ccbdfc1c7db6ab2f3a21c53.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"chen|safewatch_an_efficient_safetypolicy_following_video_guardrail_model_with_transparent_explanations"},"authorids":{"value":["~Zhaorun_Chen1","~Francesco_Pinto1","~Minzhou_Pan1","~Bo_Li19"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Zhaorun Chen","Francesco Pinto","Minzhou Pan","Bo Li"]}},"version":2},{"content":{"summary":{"value":"This work focuses on generating high-quality, controllable open-world game videos that feature game engine traits. It emphasizes interactive controllability to simulate gameplay effectively. Notably, the authors collected a large-scale Open-World Video Game Dataset (OGameData), which consists of over one million diverse gameplay video clips from more than 150 games, along with informative captions generated by GPT-4o. Methodologically, they introduce a diffusion transformer model as the foundation model and a specially designed network called InstructNet for interactive control. The model is trained on the large-scale OGameData dataset using a two-stage process involving pre-training of the foundation model and instruction tuning for InstructNet."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1.  Please provide data on the time required to generate a video segment at different resolutions or for different types of content. A section to analyze the trade-offs between generation quality and speed would be better.\n\n2.  The training details of InstructNet lack specificity regarding the acquisition of video data corresponding to keyboard bindings. It would be beneficial to include more comprehensive information on the data collection process and the training methodology employed."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. This work collects a substantial number of open-world game videos from over 150 games, ultimately constructing more than 1,000,000 text-video pairs with highly detailed annotations. Its scale and diversity of annotations make it stand out, and the release of this dataset is expected to advance the field of game video generation.\n2. It produces high-quality, more general realistic game video content. Previous works on game video generation often focused on specific game types, primarily 2D games or limited early 3D games. This work offers a more diverse and high-definition range of scene types for game video generation."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. This work attempts to address the interactive control of open-world game video generation for gameplay simulation. However, to fully tackle the interactive issue, the generation speed needs to be considered,  as interactive experiences demand stringent timing requirements, which poses significant challenges. For instance, Google’s [1] achieves real-time rendering, even making it a viable game engine. While this work focuses on higher-resolution video generation, exploring the relationship between speed and performance would be beneficial, along with providing data on rendering time and speed. \n\n[1] Dani Valevski, Yaniv Leviathan, Moab Arar, and Shlomi Fruchter. Diffusion models are real-time game engines. arXiv preprint arXiv:2408.14837, 2024.\n\n2. The paper claims to simulate game engine features like diverse events, yet the examples provided offer quite limited dynamic event simulation, primarily addressing environmental changes like weather and lighting. There remains a gap to true gameplay simulation, such as incorporating NPC interactions or triggering more game-like special events."}},"nonreaders":[],"tmdate":1731427472603,"tcdate":1730170832100,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission1714/Reviewer_1M8F"],"signatures":["ICLR.cc/2025/Conference/Submission1714/Reviewer_1M8F"],"forum":"8VG8tpPZhe","number":1,"license":"CC BY 4.0","cdate":1730170832100,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission1714/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427472603,"domain":"ICLR.cc/2025/Conference","replyto":"8VG8tpPZhe","id":"DiahPF6aVi","forumContent":{"venue":{"value":"ICLR 2025 Poster"},"TLDR":{"value":"We introduce GameGen-$\\mathbb{X}$, the first diffusion transformer model tailored for the generation and controllable interaction of open-world game videos, unifying multi-modal game-related control signals."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Open-world Game Video Generation","Interactive Control","Diffusion Transformers"]},"primary_area":{"value":"generative models"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"We introduce GameGen-$\\mathbb{X}$, the first diffusion transformer model specifically designed for both generating and interactively controlling open-world game videos. \n    This model facilitates high-quality, open-domain generation by approximating various game elements, such as innovative characters, dynamic environments, complex actions, and diverse events. \n    Additionally, it provides interactive controllability, predicting and altering future content based on the current clip, thus allowing for gameplay simulation.\n    To realize this vision, we first collected and built an Open-World Video Game Dataset (OGameData) from scratch. \n    It is the first and largest dataset for open-world game video generation and control, which comprises over one million diverse gameplay video clips with informative captions.\n    GameGen-$\\mathbb{X}$ undergoes a two-stage training process, consisting of pre-training and instruction tuning. \n    Firstly, the model was pre-trained via text-to-video generation and video continuation, enabling long-sequence open-domain game video generation with improved fidelity and coherence.\n    Further, to achieve interactive controllability, we designed InstructNet to incorporate game-related multi-modal control signal experts.\n    This allows the model to adjust latent representations based on user inputs, advancing the integration of character interaction and scene content control in video generation.\n    During instruction tuning, only the InstructNet is updated while the pre-trained foundation model is frozen, enabling the integration of interactive controllability without loss of diversity and quality of generated content. \n    GameGen-$\\mathbb{X}$ contributes to advancements in open-world game design using generative models. \n    It demonstrates the potential of generative models to serve as auxiliary tools to traditional rendering techniques, demonstrating the potential for merging creative generation with interactive capabilities.\n    The project will be available at https://github.com/GameGen-X/GameGen-X."},"_bibtex":{"value":"@inproceedings{\nche2025gamegenx,\ntitle={GameGen-X: Interactive Open-world Game Video Generation},\nauthor={Haoxuan Che and Xuanhua He and Quande Liu and Cheng Jin and Hao Chen},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=8VG8tpPZhe}\n}"},"title":{"value":"GameGen-X: Interactive Open-world Game Video Generation"},"pdf":{"value":"/pdf/5a4c3701e3960ca9c924f406799bec4693421252.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"che|gamegenx_interactive_openworld_game_video_generation"},"authorids":{"value":["~Haoxuan_Che1","~Xuanhua_He1","~Quande_Liu1","~Cheng_Jin3","~Hao_Chen1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Haoxuan Che","Xuanhua He","Quande Liu","Cheng Jin","Hao Chen"]}},"version":2},{"content":{"summary":{"value":"This paper proposes FAME (Formal Abstract Minimal Explanations), a framework for generating formally verified minimal explanations for neural network predictions. The method combines abstract interpretation with LiRPA (Linear Relaxation based Perturbation Analysis) to efficiently identify the minimal subset of input features responsible for a model’s decision. FAME removes the traversal-order dependency in prior formal explanation methods through an abstraction-based parallel mechanism that eliminates multiple irrelevant features simultaneously. Experiments on MNIST and GTSRB show that FAME produces more compact explanations and up to 25× faster runtime compared to VERIX+, supporting its theoretical claims."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. I do not fully understand the validation setup described in Section 7.2, where the authors claim that FAME achieves formally minimal explanations but provide no direct comparison against exhaustive search or exact verification results. What metric was used to assess proximity to the true minimal set, and how do the authors justify that this indirect evaluation sufficiently demonstrates minimality?\n2. The paper reports in Section 6 that FAME relies on LiRPA-derived abstract bounds to identify irrelevant features. However, recent work has shown that local linear relaxations may miss globally relevant feature interactions (see Lu et al., 2024, ICML — EiG-Search: Generating Edge-Induced Subgraphs for GNN Explanation in Linear Time). Could such limitations affect the completeness of FAME’s explanations, and have the authors considered adaptive or hierarchical abstraction domains to mitigate this risk?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The paper introduces a novel and theoretically grounded framework, FAME (Formal Abstract Minimal Explanations), that advances formal explainable AI by combining abstract interpretation with LiRPA for efficient and provably minimal explanations.\n2. It effectively removes the traversal-order dependency that limits prior formal XAI methods and achieves substantial improvements in explanation compactness and runtime on benchmark datasets."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The paper introduces formal minimal explanations as a novel interpretability construct, yet the theoretical operationalization of “interpretability” remains vague. The framework lacks a precise articulation of how its formal minimality relates to established interpretive criteria, such as faithfulness or epistemic transparency, which weakens the clarity of its conceptual contribution.\n2. The evaluation is limited to MNIST and GTSRB, which are too simple to validate FAME’s scalability or general effectiveness. More representative benchmarks such as CIFAR-10 for visual interpretability or COMPAS for tabular reasoning would better demonstrate the framework’s robustness across domains.\n3. Although the paper asserts improved explanation quality, the experiments report only efficiency-related metrics (runtime and explanation size). Without fidelity- or stability-based evaluation, the claimed enhancement in explanation quality is not empirically supported.\n4. The iterative optimization in Abstract Batch Freeing is described as a key component of the framework, but its contribution has not been empirically isolated. No ablation or comparative results are provided to verify whether this step improves efficiency or explanation compactness. Without such analysis, the practical impact of this mechanism remains speculative."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762941992979,"tcdate":1761792079708,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission21947/Reviewer_CJQr"],"signatures":["ICLR.cc/2026/Conference/Submission21947/Reviewer_CJQr"],"forum":"VJkNqJJAhV","number":1,"license":"CC BY 4.0","cdate":1761792079708,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission21947/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762941992979,"domain":"ICLR.cc/2026/Conference","replyto":"VJkNqJJAhV","id":"PFq276FBPo","forumContent":{"TLDR":{"value":"We introduce FAME, a novel method grounded in abstract interpretation that efficiently generates formal, minimal explanations for large neural networks by leveraging dedicated perturbation domains."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["abductive explanations","abstract interpretation","robustness","NN verification"]},"supplementary_material":{"value":"/attachment/f7d8bfbc5ca153ce22c33a7039615c2613acbf00.pdf"},"primary_area":{"value":"interpretability and explainable AI"},"abstract":{"value":"We propose $\\textbf{FAME}$ (Formal Abstract Minimal Explanations), a new class of abductive explanations grounded in abstract interpretation. FAME is the first method to scale to large neural networks while reducing explanation size. Our main contribution is the design of dedicated perturbation domains that eliminate the need for traversal order. FAME progressively shrinks these domains and leverages LiRPA-based bounds to discard irrelevant features, ultimately converging to a $\\textbf{formal abstract minimal explanation}$.  To assess explanation quality, we introduce a procedure that measures the worst-case distance between an abstract minimal explanation and a true minimal explanation. This procedure combines adversarial attacks with an optional $VERI{\\large X}+$ refinement step. We benchmark FAME against $VERI{\\large X}+$ and demonstrate consistent gains in both explanation size and runtime on medium- to large-scale neural networks."},"_bibtex":{"value":"@inproceedings{\nboumazouza2026fame,\ntitle={{FAME}: Formal Abstract Minimal Explanation for Neural Networks},\nauthor={Ryma Boumazouza and Raya Elsaleh and Melanie Ducoffe and Shahaf Bassan and Guy Katz},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=VJkNqJJAhV}\n}"},"title":{"value":"FAME: Formal Abstract Minimal Explanation for Neural Networks"},"pdf":{"value":"/pdf/d189862bd3e17f18cb525885463d83a8f31454ff.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"boumazouza|fame_formal_abstract_minimal_explanation_for_neural_networks"},"authorids":{"value":["~Ryma_Boumazouza1","~Raya_Elsaleh1","~Melanie_Ducoffe1","~Shahaf_Bassan1","~Guy_Katz1"]},"authors":{"value":["Ryma Boumazouza","Raya Elsaleh","Melanie Ducoffe","Shahaf Bassan","Guy Katz"]}},"version":2},{"content":{"summary":{"value":"This paper introduces VideoAlchemy, a video customization method that generates personalized videos with multiple subjects and specified backgrounds, eliminating the need for test-time fine-tuning. The approach leverages a DiT architecture with an additional cross-attention layer and linear projection. Furthermore, this paper curates a large-scale training dataset, while employing data augmentation and conditional subject sampling strategy for training. For evaluation, this work develops a multi-subject video customization benchmark with four metrics. Experimental results demonstrate that VideoAlchemy surpasses existing methods both quantitatively and qualitatively."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"1. Please refer to the weaknesses section. If my concerns are well addressed, I am willing to modify my score.\n2. Typo: \"issue\" in line 60.\n3. Typo: line 176 has two periods.\n4. In line 339, the paper mentions using six metrics but only introduces four metrics."},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. This paper studies the multi-subject open-set video customization task, which is interesting and meaningful for practical video generation.\n2. The paper is well-written and easy to understand.\n3. This work curates a large multi-subject video dataset, containing 37.8M videos with subjects, objects, and backgrounds. It also introduces a multi-subject video customization evaluation benchmark.\n4. The generated videos exhibit relatively high quality by using the DiT model and abundant data."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. One of the important contributions of this paper is the curated dataset. However, I have concerns about:\n    - The quality and diversity of the training and evaluation datasets. From Fig. 8's word cloud, humans and static objects/scenes (laptop, branch, house, sunset, building, etc.) make up a larger proportion of the dataset. Videos with static objects may have minimal motion. Also, based on my experience, the videos in MSRVTT primarily feature humans and scenes, with fewer animals and objects, potentially limiting test set diversity.\n    - Availability of training dataset. While the paper promises to release the test dataset, it does not confirm making the training dataset publicly available. I believe releasing the training dataset is of great significance to the open-source community, especially since this paper uses the curated dataset as one of its main contributions.\n2. The method's innovation is relatively limited. Although the differences from the IP-Adapter are mentioned in the method part, the similar idea of using separate cross-attention layers for text and image embeddings has been explored in previous work [1]. It would be better to discuss the differences.\n3. The introduced evaluation benchmark mostly uses existing metrics and lacks comprehensiveness. It would be better to introduce new metrics, like temporal consistency or subject motion intensity as in VBench [2], to enhance the benchmark.\n4. Another concern I have is about experiments:\n    - Comparison Fairness. The method uses a DiT architecture, while comparison methods are mostly based on UNet, raising concerns that improvements might stem from a more robust video base model rather than the method itself.\n    - Baseline Selection. It would be better to apply and reproduce image customization methods to video generative models, as in VideoBooth [3], rather than animation. On the other hand, there are some face customization works for video generation, such as Magic-Me[4] and ID-Animator [5]. It is best to compare with these methods to verify the effectiveness of the proposed method.\n    - Lack of comparisons with multi-subject video customization methods. Since the paper studies the multi-subject customization task, existing methods like VideoDreamer [6], CustomVideo [7], and DisenStudio [8] should be discussed and chosen two or more for comparison.\n    - Ablation Study. It would be better to conduct an ablation study on excluding text embeddings from image embeddings in image encoder.\n5. I carefully viewed each video provided in the supplementary material. The video examples provided are not diverse enough for multi-subject video customization. There are only two categories, human and dog, and only one motion prompt \"a ... is petting a dog on ...\". It would be better to provide examples with different categories and prompts to demonstrate the effectiveness of the method on multi-subject customization.\n\n[1] Li X, Hou X, Loy C C. When stylegan meets stable diffusion: a w+ adapter for personalized image generation.\n\n[2] Huang Z, He Y, Yu J, et al. Vbench: Comprehensive benchmark suite for video generative models.\n\n[3] Jiang Y, Wu T, Yang S, et al. VideoBooth: Diffusion-based video generation with image prompts.\n\n[4] Ma Z, Zhou D, Yeh C H, et al. Magic-me: Identity-specific video customized diffusion.\n\n[5] He X, Liu Q, Qian S, et al. ID-Animator: Zero-Shot Identity-Preserving Human Video Generation.\n\n[6] Chen H, Wang X, Zeng G, et al. Videodreamer: Customized multi-subject text-to-video generation with disen-mix finetuning.\n\n[7] Wang Z, Li A, Xie E, et al. Customvideo: Customizing text-to-video generation with multiple subjects.\n\n[8] Chen H, Wang X, Zhang Y, et al. DisenStudio: Customized Multi-subject Text-to-Video Generation with Disentangled Spatial Control."}},"nonreaders":[],"tmdate":1731427359756,"tcdate":1730006211754,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission1025/Reviewer_5xgj"],"signatures":["ICLR.cc/2025/Conference/Submission1025/Reviewer_5xgj"],"forum":"popKM1zAYa","number":1,"license":"CC BY 4.0","cdate":1730006211754,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission1025/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427359756,"domain":"ICLR.cc/2025/Conference","replyto":"popKM1zAYa","id":"eSeDBigI0O","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["generative models","video generation","content personalization","content customization"]},"supplementary_material":{"value":"/attachment/90a1dc46fc4c5d3b08046482b13f8b24441c9fe7.zip"},"primary_area":{"value":"generative models"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Video personalization methods allow us to synthesize videos with specific concepts such as people, pets, and places. However, existing methods often focus on limited domains, require time-consuming optimization per subject, or support only a single subject. We present $VideoAlchemy~-$ a video model equipped with built-in multi-subject, open-set personalization capabilities for both foreground objects and backgrounds, eliminating the need for time-consuming test-time optimization. Our model is built on a new Diffusion Transformer module that fuses each reference image conditioning and its corresponding subject-level text prompt with cross-attention layers. Developing such a large model presents two main challenges: $dataset$ and $evaluation$. First, as paired datasets of reference images and videos are extremely hard to collect, we opt to sample video frames as reference images and synthesize entire videos. This approach, however, introduces data biases issue, where models can easily denoise training videos but fail to generalize to new contexts during inference. To mitigate these issue, we carefully design a new automatic data construction pipeline with extensive image augmentation and sampling techniques. Second, evaluating open-set video personalization is a challenge in itself. To address this, we introduce a new personalization benchmark with evaluation protocols focusing on accurate subject fidelity assessment and accommodating different types of personalization conditioning. Finally, our extensive experiments show that our method significantly outperforms existing personalization methods, regarding quantitative and qualitative evaluations."},"_bibtex":{"value":"@misc{\nchen2024videoalchemy,\ntitle={VideoAlchemy: Open-set Personalization in Video Generation},\nauthor={Tsai-Shien Chen and Aliaksandr Siarohin and Willi Menapace and Yuwei Fang and Ivan Skorokhodov and Jun-Yan Zhu and Kfir Aberman and Ming-Hsuan Yang and Sergey Tulyakov},\nyear={2024},\nurl={https://openreview.net/forum?id=popKM1zAYa}\n}"},"title":{"value":"VideoAlchemy: Open-set Personalization in Video Generation"},"pdf":{"value":"/pdf/0e2fed221a1c5db8ab777a4ecaeaf7ffc16ae3ba.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"chen|videoalchemy_openset_personalization_in_video_generation"},"authorids":{"value":["~Tsai-Shien_Chen1","~Aliaksandr_Siarohin1","~Willi_Menapace1","~Yuwei_Fang1","~Ivan_Skorokhodov1","~Jun-Yan_Zhu1","~Kfir_Aberman1","~Ming-Hsuan_Yang1","~Sergey_Tulyakov1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Tsai-Shien Chen","Aliaksandr Siarohin","Willi Menapace","Yuwei Fang","Ivan Skorokhodov","Jun-Yan Zhu","Kfir Aberman","Ming-Hsuan Yang","Sergey Tulyakov"]}},"version":2},{"content":{"summary":{"value":"The paper introduces LLaVA-Surg, the first multimodal surgical assistant capable of understanding surgical videos and engaging in open-ended conversations about them. The authors create Surg-QA, a dataset of 102,000 surgical video-instruction pairs, using a novel two-stage question-answer generation pipeline. This approach reduces LLM hallucinations and costs by breaking down the generation process. The resulting model demonstrates superior performance in surgical video question-answering compared to previous general-domain models."},"soundness":{"value":4},"confidence":{"value":4},"questions":{"value":"Please address the weaknesses mentioned above."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":3},"strengths":{"value":"1. The pipeline is comprehensive: A two-stage question-answer generation process minimizes hallucinations by extracting information prior to generating pairs, which enhances data quality and reliability compared to Quilt-1M[1], which has a similar approach.\n\n2. Integrating surgical visual concept alignment data through action triplets improves text-visual alignment, enhancing the model’s grasp of surgical concepts.\n\n3. The idea is interesting: using the Spearman rank correlation between expert and GPT scores effectively validates the reliability of large-scale GPT evaluation.\n\n[1] Ikezogwo, Wisdom, et al. \"Quilt-1m: One million image-text pairs for histopathology.\" Advances in neural information processing systems 36 (2024)."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Could you provide results for the three existing surgical-domain datasets (EndoVis-18-VQA, Cholec80-VQA, and SSG-VQA) trained on Surg-QA? These results could demonstrate Surg-QA's potential as a foundational dataset in the surgical domain.\n\n2. Maybe considering to use other video VLM models, which provides a more sophisticated approach to temporal fusion than simple average pooling."}},"nonreaders":[],"tmdate":1731428252107,"tcdate":1730361702171,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission4896/Reviewer_zRU3"],"signatures":["ICLR.cc/2025/Conference/Submission4896/Reviewer_zRU3"],"forum":"063FuFYQQd","number":1,"license":"CC BY 4.0","cdate":1730361702171,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission4896/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428252107,"domain":"ICLR.cc/2025/Conference","replyto":"063FuFYQQd","id":"MxvlHdZ7WE","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Multimodal assistant","surgical","multimodal instruction-following data","dataset"]},"primary_area":{"value":"foundation or frontier models, including LLMs"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Multimodal large language models (LLMs) have achieved notable success across various domains, while research in the medical field has largely focused on unimodal images. Meanwhile, current general-domain multimodal models for videos still lack the capabilities to understand and engage in conversations about surgical videos. One major contributing factor is the absence of datasets in the surgical field. In this paper, we create a new dataset, Surg-QA, consisting of 102,000 surgical video-instruction pairs, the largest of its kind so far. To build such a dataset, we propose a novel two-stage question-answer generation pipeline with LLM to learn surgical knowledge in a structured manner from the publicly available surgical lecture videos. The pipeline breaks down the generation process into two stages to significantly reduce the task complexity, allowing us to use a more affordable, locally deployed open-source LLM than the premium paid LLM services. It also mitigates the risk of LLM hallucinations during question-answer generation, thereby enhancing the overall quality of the generated data. We further train LLaVA-Surg, a novel vision-language conversational assistant capable of answering open-ended questions about surgical videos, on this Surg-QA dataset, and conduct comprehensive evaluations on zero-shot surgical video question-answering tasks. We show that LLaVA-Surg significantly outperforms all previous general-domain models, demonstrating exceptional multimodal conversational skills in answering open-ended questions about surgical videos. We will release our code, model, and the instruction-tuning dataset."},"_bibtex":{"value":"@misc{\nli2025llavasurg,\ntitle={{LL}a{VA}-Surg: Towards Multimodal Surgical Assistant via Structured Lecture Learning},\nauthor={Jiajie Li and Garrett Skinner and Brian R Quaranto and Gene Yang and Steven D Schwaitzberg and Peter C W Kim and Jinjun Xiong},\nyear={2025},\nurl={https://openreview.net/forum?id=063FuFYQQd}\n}"},"title":{"value":"LLaVA-Surg: Towards Multimodal Surgical Assistant via Structured Lecture Learning"},"pdf":{"value":"/pdf/04d73daf100581d96e3a971dd358d0aad68ebdd1.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"li|llavasurg_towards_multimodal_surgical_assistant_via_structured_lecture_learning"},"authorids":{"value":["~Jiajie_Li2","~Garrett_Skinner1","~Brian_R_Quaranto1","~Gene_Yang1","~Steven_D_Schwaitzberg1","~Peter_C_W_Kim1","~Jinjun_Xiong1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Jiajie Li","Garrett Skinner","Brian R Quaranto","Gene Yang","Steven D Schwaitzberg","Peter C W Kim","Jinjun Xiong"]}},"version":2},{"content":{"summary":{"value":"Syncphony demonstrates meaningful progress in the audio-to-video generation domain. Its main contributions include: (1) a motion-aware loss that re-weights the training objective based on optical flow to focus the model’s learning on motion-intensive regions; (2) a training-free guidance technique that enhances the injection of audio information during inference without additional training; and (3) a more intuitive and reasonable evaluation metric for measuring audio-visual synchronization through reconstructed audio from generated videos. Experiments on two public datasets and qualitative case studies further validate the method’s effectiveness in improving both synchronization and visual quality."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"- Could the authors elaborate on how they plan to address the challenges posed by more realistic (or longer) audio-to-video scenarios, where optical flow estimation can be severely affected by background motion, camera movement, or scene transitions beyond the main subject?\n\n- Could the authors show how the skip-layer behavior manifests in other video generation models and whether similar phenomena can be consistently observed? In addition, if temporal alignment is indeed a crucial aspect of the task, why not consider using energy-based audio features as a more direct form of control or guidance?\n\n- Given that modern multimodal large language models (e.g., Gemini 2.5) already demonstrate strong capabilities in understanding audio-visual information, and that existing video-to-audio alignment metrics (such as DeSync, which can be applied similarity in a2v field) provide reliable proxies for spatiotemporal correspondence, could the authors clarify or compare what specific advantages their proposed metric offers over these established approaches?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"- The proposed training-free guidance is novel and effectively points out the core challenge of achieving precise spatiotemporal alignment in audio-to-video generation.\n\n- The proposed evaluation metric is intuitively reasonable and addresses the limitation of conventional metrics like FVD, which fail to effectively measure spatiotemporal alignment.\n\n- The use of RoPE for spatiotemporal positional encoding enhances temporal consistency and spatial coherence in video generation, contributing to smoother and more structured motion representation.\n\n- The demo results demonstrate good temporal alignment consistency."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Although the work is technically sound, there are several issues that I feel must be discussed.\n\n- Regarding the motion-aware loss, it relies on a strong assumption that the visual content remains static and confined to a single scene. This assumption holds almost perfectly under the authors’ setting where short video clips of around two seconds. However, in more realistic video generation scenarios, background changes, camera motion, and scene transitions can occur without producing strong audio cues, but may catastrophically distort optical flow estimation and thus limit the scalability of the proposed approach.\n\n- The proposed guidance technique randomly drops audio cross-attention layers during inference; however, this inference structure (unlike CFG or Autoguidance) is never encountered during training, making the resulting “unconditional” outputs less predictable. Moreover, although the authors provide some visual demonstrations, the underlying motivation may primarily apply to relatively smaller models where audio conditioning tends to be weaker. In larger-scale video generation frameworks such as Wan or Hunyuan-Video, audio conditions may not be as easily ignored, and models could exhibit different skip-layer behaviors. Rather than validating only on additional datasets, I would prefer the authors to verify this idea across multiple video generation baselines to strengthen the generality of their claims.\n\n- Although the proposed new metric is intuitively reasonable, current video-to-audio (V2A) models also suffer from (or are still addressing) spatiotemporal alignment issues, which means they may not serve as a fully reliable ground-truth proxy. Moreover, the approach fundamentally increases evaluation time, potentially limiting scalability to larger experiments."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762926189685,"tcdate":1761960454109,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission15982/Reviewer_4rJg"],"signatures":["ICLR.cc/2026/Conference/Submission15982/Reviewer_4rJg"],"forum":"sG8dGZMaub","number":3,"license":"CC BY 4.0","cdate":1761960454109,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission15982/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762926189685,"domain":"ICLR.cc/2026/Conference","replyto":"sG8dGZMaub","id":"dtxrS4k8YS","forumContent":{"TLDR":{"value":"We propose improved audio-aligned video generation by leveraging a pretrained video generation model, while preserving its original performance."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Audio-to-Video Generation","Multimodal Synthesis","Temporal Synchronization","Diffusion Transformer","Video Generation","Audio-Conditioned Generation"]},"supplementary_material":{"value":"/attachment/c567cd1d3273111f54b832167575f6a1e472f1e3.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Text-to-video and image-to-video generation have made rapid progress in visual quality, but they remain limited in controlling the precise timing of motion. \nIn contrast, audio provides temporal cues aligned with video motion, making it a promising condition for temporally controlled video generation. \nHowever, existing audio-to-video (A2V) models struggle with fine-grained synchronization due to indirect conditioning mechanisms or limited temporal modeling capacity.\nWe present Syncphony, which generates 380×640 resolution, 24fps videos synchronized with diverse audio inputs. Our approach builds upon a pre-trained video backbone and incorporates two key components to improve synchronization: \n(1) Motion-aware Loss, which emphasizes learning at high-motion regions; \n(2) Audio Sync Guidance, which guides the full model using a visually aligned off-sync model without audio layers to better exploit audio cues at inference while maintaining visual quality.\nTo evaluate synchronization, we propose CycleSync, a video-to-audio-based metric that measures the amount of motion cues in the generated video to reconstruct the original audio. Experiments on AVSync15 and The Greatest Hits datasets demonstrate that Syncphony outperforms existing methods in both synchronization accuracy and visual quality."},"_bibtex":{"value":"@inproceedings{\nsong2026syncphony,\ntitle={Syncphony: Synchronized Audio-to-Video Generation with Diffusion Transformers},\nauthor={Jibin Song and Mingi Kwon and Jaeseok Jeong and Youngjung Uh},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=sG8dGZMaub}\n}"},"title":{"value":"Syncphony: Synchronized Audio-to-Video Generation with Diffusion Transformers"},"pdf":{"value":"/pdf/f6962ed3e6685ba3f57c38238cbece80684fe7b8.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"song|syncphony_synchronized_audiotovideo_generation_with_diffusion_transformers"},"authorids":{"value":["~Jibin_Song1","~Mingi_Kwon1","~Jaeseok_Jeong2","~Youngjung_Uh2"]},"authors":{"value":["Jibin Song","Mingi Kwon","Jaeseok Jeong","Youngjung Uh"]}},"version":2},{"content":{"summary":{"value":"This paper introduces CAT-Video, a training framework that improves video generation models' resilience to imperfect inputs. The authors address a key limitation of standard noise-injection techniques, which often disrupt video coherence. Their solution involves two novel methods: Batch-Centered Noise Injection (BCNI) to maintain semantic consistency within a batch, and Spectrum-Aware Contextual Noise (SACN) to preserve smooth temporal dynamics."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"Please refer to the strengths and weaknesses"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"​​+  a New Problem​​: It is one of the first to systematically address how to make video AI models robust to imperfect instructions, focusing on a key weakness in existing methods that break video coherence.\n\n​​+  Its two new techniques, BCNI and SACN, are simple and add almost no extra cost, yet outperform much larger and more expensive models.\n\n​​+ Extensive testing across different datasets and metrics proves the methods reliably enhance video quality and motion."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Is this problem meaningful in the research field?  In what kind of scenerios, noise will introduced into the input? Does the model robust for  adversarial attack? \n- Does the proposed model work for large video models?"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762931542070,"tcdate":1761910789362,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission19699/Reviewer_5D4F"],"signatures":["ICLR.cc/2026/Conference/Submission19699/Reviewer_5D4F"],"forum":"unZhwukf0T","number":2,"license":"CC BY 4.0","cdate":1761910789362,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission19699/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762931542070,"domain":"ICLR.cc/2026/Conference","replyto":"unZhwukf0T","id":"Y7n5Y6heSR","forumContent":{"TLDR":{"value":"We introduce CAT-Video, a corruption-aware training framework that improves robustness and temporal coherence in video diffusion models through structured, data-aligned noise injection."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["video diffusion","corruption-aware training","robust video generation","structured noise injection","multimodal robustness","temporal coherence"]},"supplementary_material":{"value":"/attachment/c99cdbeea034fc1e74ad38310569a2906228cc13.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Latent Video Diffusion Models (LVDMs) have achieved state-of-the-art generative quality for image and video generation; however, they remain brittle under noisy conditioning, where small perturbations in text or multimodal embeddings can cascade over timesteps and cause semantic drift. Existing corruption strategies from image diffusion (Gaussian, Uniform) fail in video settings because static noise disrupts temporal fidelity. In this paper, we propose **CAT-Video**, a corruption-aware training framework with structured, data-aligned noise injection tailored for video diffusion. Our two operators—*Batch-Centered Noise Injection (BCNI)* and *Spectrum-Aware Contextual Noise (SACN)*—align perturbations with batch semantics or spectral dynamics to preserve coherence. CAT-Video yields substantial gains: BCNI reduces FVD by **31.9%** on WebVid-2M, MSR-VTT, and MSVD, while SACN improves UCF-101 by **12.3%**, outperforming Gaussian, Uniform, and even large diffusion baselines like DEMO (2.3B) and Lavie (3B) despite training on $\\mathbf{5}\\times$ less data. Ablations confirm the unique value of low-rank, data-aligned noise, and theory establishes why these operators tighten robustness and generalization bounds. CAT-Video thus sets a new framework for robust video diffusion, and our experiments show that it can also be extended to autoregressive generation and multimodal video understanding LLMs."},"_bibtex":{"value":"@misc{\nmaduabuchi2026catvideo,\ntitle={{CAT}-{VIDEO}: {CORRUPTION}-{AWARE} {TRAINING} {FOR} {ROBUST} {VIDEO} {DIFFUSION} {MODELS}},\nauthor={Chika Maduabuchi and Hao Chen and Yujin Han and Jindong Wang},\nyear={2026},\nurl={https://openreview.net/forum?id=unZhwukf0T}\n}"},"title":{"value":"CAT-VIDEO: CORRUPTION-AWARE TRAINING FOR ROBUST VIDEO DIFFUSION MODELS"},"pdf":{"value":"/pdf/36b779e3109b1503084b84e3c1c30d2bc4985918.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"maduabuchi|catvideo_corruptionaware_training_for_robust_video_diffusion_models"},"authorids":{"value":["~Chika_Maduabuchi1","~Hao_Chen15","~Yujin_Han1","~Jindong_Wang4"]},"authors":{"value":["Chika Maduabuchi","Hao Chen","Yujin Han","Jindong Wang"]}},"version":2},{"content":{"summary":{"value":"This paper addresses the challenge of offline reinforcement learning (offline RL), particularly in overcoming the issue of overestimating values for out-of-distribution (OOD) actions when training from a fixed dataset. It introduces the Action-Space Pseudo-Labeling (ASPL) method, which assigns pseudo Q-values to randomly sampled actions in the state–action space, providing training signals for actions outside the behavior support. This approach mitigates the overestimation problem by using distance-aware pseudo-labels that decay with the distance from the behavior support. The method is shown to outperform existing offline RL methods with minimal tuning on various benchmarks, including D4RL Gym–MuJoCo tasks.\n\nThe paper proposes a novel and practical solution to the overestimation problem in offline RL through ASPL, an effective method that augments training with pseudo-labels for OOD actions. The experimental results on D4RL benchmarks demonstrate that ASPL performs\ncompetitively with minimal hyperparameter tuning. However, while the method is shown to work well empirically, the theoretical foundations could benefit from further elaboration, particularly regarding the behavior-aware weighting mechanism. Additionally, a more thorough analysis of the method's generalization across diverse RL tasks would strengthen the claims. The paper provides solid empirical evidence but lacks sufficient theoretical grounding in some areas."},"soundness":{"value":2},"confidence":{"value":2},"questions":{"value":"Provide deeper theoretical analysis of the behavior-aware weight:\n1. Elaborate on the formal properties of the behavior-aware weight mechanism and its effects on training stability and performance. A more thorough mathematical analysis could clarify the behavior-aware weight's role and its relationship with other regularization techniques (Sec. 4.2).\n2. Consider formalizing the convergence properties of ASPL to help justify its effectiveness and robustness.\n\nExpand baseline comparisons:\n1. Include additional baselines that use ensemble critics, model-based approaches, or other advanced techniques in offline RL. This will provide a clearer understanding of how ASPL compares to a broader range of methods (Sec. 5.1).\n2. Also, consider evaluating ASPL on tasks with more complex action spaces, such as discrete action tasks or multi-agent settings, to explore its generalization beyond continuous action spaces (Sec. 5.1).\n\nAddress generalization to other action spaces:\n1. Discuss and provide potential extensions of ASPL for discrete or mixed action spaces. A clear discussion on how the method can be adapted for other action representations, such as categorical or structured actions, would help broaden its applicability (Sec. 6).\n\nClarify the pseudo-labeling process:\n1. Provide a more detailed explanation of the pseudo-labeling process, particularly the distance metric used to calculate the pseudo-Q targets. Clarifying how the pseudolabeling interacts with the Bellman backup would help in understanding its impact on training dynamics (Sec. 4.1)."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"Simple and effective solution to overestimation in offline RL:\n1. The ASPL method tackles the challenge of sparse coverage in the action space, effectively providing training signals for OOD actions (Sec. 4.1; Eq. 6). This addresses a critical issue in offline RL and improves performance without adding complexity to existing actor–critic pipelines.\n2. The integration of pseudo-labeling into the critic update is minimalistic yet effective, requiring no additional actor-side constraints or auxiliary networks (Sec. 4.3). This simplicity enhances the method's appeal for real-world applications where computational efficiency is important.\n\nEmpirical validation on D4RL benchmarks:\n1. ASPL consistently outperforms several strong offline RL baselines, including TD3+BC, CQL, and IQL, across various D4RL Gym–MuJoCo tasks (Sec. 5.1; Table 1). These results demonstrate the method's robustness and its ability to achieve stable training and competitive performance with minimal hyperparameter tuning.\n2. The paper includes detailed experiments showing the sensitivity of ASPL to key hyperparameters like the number of random actions sampled (N) and the pseudolabeling coefficient (α), highlighting its resilience to hyperparameter changes (Sec.5.2; Fig. 3, Fig. 4).\n\nBehavior-aware weighting mechanism:\n1. The dynamic adjustment of pseudo-label weight based on behavior coverage (Eq. 9) is a unique and promising feature. This ensures that updates are Bellman-dominated in well-supported regions and pseudo-label-dominated in unsupported regions, improving stability and mitigating extrapolation errors (Sec. 4.2).\n2. The simplicity of the behavior-aware weight mechanism—decreasing the weight with increasing behavior coverage—makes the method both effective and easy to integrate into existing RL pipelines (Sec. 4.2).\n\nMinimal tuning burden:\n1. ASPL reduces the sensitivity to hyperparameters, particularly the pseudo-label weight α, which simplifies the tuning process compared to other methods (Sec. 5.2). This is a significant advantage for practitioners who need reliable methods with minimal configuration effort."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Limited theoretical analysis of the behavior-aware weight mechanism:\n1. While the paper introduces a behavior-aware weight (Eq. 9), the theoretical explanation and formal analysis of this weight mechanism are not fully developed.\n2. The implications of this dynamic weighting, especially in complex environments, could be better justified with more rigorous theoretical analysis (Sec. 4.2). No direct evidence found in the manuscript.\nThe paper could benefit from a more detailed explanation of why the specific form of distance-aware pseudo-labeling works well across various tasks and datasets. Whilethe empirical results are strong, further formalism in the explanation would add to the method's credibility (Sec. 4.2).\n\nComparative analysis with more baselines:\n1. Although the paper compares ASPL against several state-of-the-art offline RL algorithms, the set of baselines could be expanded to include more diverse approaches, especially methods that use ensemble critics or model-based planning (Sec. 5.1). This would provide a more comprehensive comparison and further demonstrate ASPL's advantages.\n2. The paper mainly focuses on tasks in the Gym–MuJoCo environment. Evaluating ASPL on a broader range of tasks, including those with more complex action spaces or other domains like NLP or robotics, would offer insights into the generalizability of\nthe method (Sec. 5.1). No direct evidence found in the manuscript.\n\nInsufficient discussion of generalization to other action spaces:\n1. While the method performs well in continuous action spaces, its applicability to discrete or mixed action spaces is not fully addressed (Sec. 6). A more detailed discussion on how ASPL could be adapted to different action spaces or other types of\ndecision-making environments would be valuable, particularly for applications beyond Gym–MuJoCo tasks.\n\nClarifications needed for pseudo-labeling process:\n1. The pseudo-labeling process involves sampling random actions from the action space and using distance-based Q-values as pseudo-targets. However, the handling of these pseudo-labels could be explained more clearly, particularly regarding the distance metric and its effects on training dynamics (Eq. 7). Further clarification of how these pseudo-labels interact with the Bellman backup process would improve understanding and reproducibility (Sec. 4.1)."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762921680917,"tcdate":1761934068567,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission10354/Reviewer_iaNB"],"signatures":["ICLR.cc/2026/Conference/Submission10354/Reviewer_iaNB"],"forum":"wzX6hi5QGj","number":4,"license":"CC BY 4.0","cdate":1761934068567,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission10354/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762921680917,"domain":"ICLR.cc/2026/Conference","replyto":"wzX6hi5QGj","id":"3Pu1o4vvqL","forumContent":{"TLDR":{"value":"We propose Action-Space Pseudo-Labeling (ASPL), a simple method that regularizes offline RL with pseudo-labels for unseen actions, yielding stable and consistent gains without fragile tuning."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["offline reinforcement learning","pseudo-labeling"]},"supplementary_material":{"value":"/attachment/ba2d4f9059d89a0d4b6ee8347b9b0b6769c411c2.zip"},"primary_area":{"value":"reinforcement learning"},"abstract":{"value":"he critical challenge of offline reinforcement learning (offline RL) is improving\nfrom a fixed dataset while avoiding overestimation on out-of-distribution (OOD)\nactions. Existing methods typically regularize the learned policy to avoid choosing overestimated OOD actions. However, we argue that this often over-constrains policy improvement or requires sensitive hyperparameter tuning. We restate this challenge as the absence of explicit training signals for the\nvalue function in parts of the state–action space. A more effective approach is to provide explicit training signals across the entire action space to eliminate overestimation. We introduce a surprisingly simple yet effective method: $\\textbf{Action-Space Pseudo-Labeling (ASPL)}$ to resolve this challenge. It completes the value-function’s\nmissing signals by assigning pseudo Q-targets that decrease with distance from\nthe behavior support (i.e., the support of the behavior policy). In practice, ASPL achieves an implicit behavior-aware regularization that strengthens as behavior likelihood decreases. On D4RL datasets,\nwe observe stable training and consistent improvements over strong offline baselines with minor tuning burden. Code for reproducing the experiments is provided in the supplementary material."},"_bibtex":{"value":"@misc{\nzhou2026offline,\ntitle={Offline Reinforcement Learning via Action-Space Pseudo-Labeling},\nauthor={Yunfan Zhou and Xijun Li and Jianguo Yao and Haibing Guan},\nyear={2026},\nurl={https://openreview.net/forum?id=wzX6hi5QGj}\n}"},"title":{"value":"Offline Reinforcement Learning via Action-Space Pseudo-Labeling"},"pdf":{"value":"/pdf/f6ebaa0dd1a392d50b7d2e901596fa2a81b47f3b.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"zhou|offline_reinforcement_learning_via_actionspace_pseudolabeling"},"authorids":{"value":["~Yunfan_Zhou2","~Xijun_Li1","~Jianguo_Yao1","~Haibing_Guan1"]},"authors":{"value":["Yunfan Zhou","Xijun Li","Jianguo Yao","Haibing Guan"]}},"version":2},{"content":{"summary":{"value":"This paper proposed a new dataset, CofCA for evaluating LLM’s counterfactual multi-hop reasoning ability, which focuses on both final answer evaluation and reasoning step evaluation. The dataset construction involves 2 steps, 1) the sampled Wikipedia paragraphs are sent to GPT-4 for entity replacement and paraphrase, and then back-translation is done to further modify the paragraphsm and then these paragraphs are manually checked 2) GPT-4 is then used to generate QA pairs based on the newly generate paragraphs. Here GPT-4 is asked to generate both overall multi-hop QA pair and subquestion answer pairs, and the resulting data is also manually verified. \nThe authors then conducted experiments on both newly generated data and HotpotQA, 2WikiHop and Musique, using both closed-sourced models and open-source models. The results show that all LLMs suffer significant performance drop on counterfactual questions and the joint performance which considers both final answer and reasoning process correctness can be even worse, exposing the weakness of current LLMs."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"See weaknesses"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. CofCA could be a valueable resource for studying LLM’s true reasoning ability. Due to its counterfactual nature, the LLMs can no longer use their internal knowledge for finding reasoning shortcut, hence the model have to reason over provided context. Also, the annotation of intermediate sub question answer pairs can also enable examination of the reasoning process. \n\n2. The experiments and analysis are very thorough. The results show a clear trend of performance vs question complexity and the analysis on the reasoning process further exposes the weaknesses of the SOTA LLMs."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"My main concern is the scale of the passages set when constructing the data. For previous datasets such as. HotpotQA and 2WikiHop, they usually construct questions using multiple passages to ensure that the answer have to be deducted from multiple pieces of evidence. However, it seems that for CofCA, the questions are generate using only a single passage? If this is the case, how can you make sure the generated multi-hop QA pairs are fully grounded to the passage?"}},"nonreaders":[],"tmdate":1731428026362,"tcdate":1730706904738,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission7256/Reviewer_6xmP"],"signatures":["ICLR.cc/2025/Conference/Submission7256/Reviewer_6xmP"],"forum":"q2DmkZ1wVe","number":4,"license":"CC BY 4.0","cdate":1730706904738,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission7256/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428026362,"domain":"ICLR.cc/2025/Conference","replyto":"q2DmkZ1wVe","id":"lX1Pi0YX5r","forumContent":{"TLDR":{"value":"A novel evaluation method that comprehensively and objectively reveal LLMs' real multi-step reasoning performance without data contamination."},"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["LLM evaluation","Multi-hop QA evaluation"]},"primary_area":{"value":"datasets and benchmarks"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"While Large Language Models (LLMs) excel in question-answering (QA) tasks, their real reasoning abilities on multiple evidence retrieval and integration on Multi-hop QA tasks remain less explored. Firstly, LLMs sometimes generate answers that rely on internal memory rather than retrieving evidence and reasoning in the given context, which brings concerns about the evaluation quality of real reasoning abilities. Although previous counterfactual QA benchmarks can separate the internal memory of LLMs, they focus solely on final QA performance, which is insufficient for reporting LLMs' real reasoning abilities. Because LLMs are expected to engage in intricate reasoning processes that involve evidence retrieval and answering a series of sub-questions from given passages. Moreover, current factual Multi-hop QA (MHQA) benchmarks are annotated on open-source corpora such as Wikipedia, although useful for multi-step reasoning evaluation, they show limitations due to the potential data contamination in LLMs' pre-training stage. To address these issues, we introduce the Step-wise and Counterfactual benchmark (CofCA), a novel evaluation benchmark consisting of factual data and counterfactual data that reveals LLMs' real reasoning abilities on multi-step reasoning and reasoning chain evaluation. Our experimental results reveal a significant performance gap of several LLMs between Wikipedia-based factual data and counterfactual data, deeming data contamination issues in existing benchmarks. Moreover, we observe that LLMs usually bypass the correct reasoning chain, showing an inflated multi-step reasoning performance. We believe that our CofCA benchmark will enhance and facilitate the evaluations of trustworthy LLMs."},"_bibtex":{"value":"@inproceedings{\nwu2025cofca,\ntitle={Cof{CA}: A {STEP}-{WISE} Counterfactual Multi-hop {QA} benchmark},\nauthor={Jian Wu and Linyi Yang and Zhen Wang and Manabu Okumura and Yue Zhang},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=q2DmkZ1wVe}\n}"},"title":{"value":"CofCA: A STEP-WISE Counterfactual Multi-hop QA benchmark"},"pdf":{"value":"/pdf/8bcb310c5052589c38968512150bf5fa37abb5c4.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"wu|cofca_a_stepwise_counterfactual_multihop_qa_benchmark"},"authorids":{"value":["~Jian_Wu8","~Linyi_Yang1","~Zhen_Wang13","~Manabu_Okumura2","~Yue_Zhang7"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Jian Wu","Linyi Yang","Zhen Wang","Manabu Okumura","Yue Zhang"]}},"version":2},{"content":{"venue":{"value":"Video-Langauge Models Poster"},"TLDR":{"value":"We present MOTIFY, a novel method using Vision-Language Models to predict scene transitions and actions in mobile OS task videos. It requires no manual annotation, outperforms baselines, and aims to enable scalable mobile agent development."},"keywords":{"value":["video understanding","vision and language","large language model","vision language model","LLM","VLM","operating system","mobile","Mobile OS","procedure understanding","video summarization"]},"supplementary_material":{"value":"/attachment/35858d3a33699bcc57fd8daefbcd7a76b171cf3f.zip"},"abstract":{"value":"We present MOTIFY, a novel approach for predicting scene transitions and actions from mobile operating system (OS) task videos. By leveraging pretrained Vision-Language Models (VLMs), MOTIFY extract the task sequences from real-world YouTube videos without manual annotation. Our method addresses the limitations of existing approaches, which rely on manual data annotation or simulation environments. We demonstrate MOTIFY's effectiveness on a diverse set of mobile OS tasks across multiple platforms, outperforming baseline methods in scene transition detection and action prediction. This approach opens new possibilities for scalable, real-world mobile agent development and video understanding research."},"_bibtex":{"value":"@inproceedings{\njang2025mobile,\ntitle={Mobile {OS} Task Procedure Extraction from YouTube},\nauthor={Yunseok Jang and Yeda Song and Sungryull Sohn and Lajanugen Logeswaran and Tiange Luo and Honglak Lee},\nbooktitle={Workshop on Video-Language Models @ NeurIPS 2024},\nyear={2025},\nurl={https://openreview.net/forum?id=PcwaP4o7vk}\n}"},"title":{"value":"Mobile OS Task Procedure Extraction from YouTube"},"pdf":{"value":"/pdf/e55e9426cdb01ac10d959ede98aaeba686ba9c85.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Workshop/Video-Langauge_Models"},"paperhash":{"value":"jang|mobile_os_task_procedure_extraction_from_youtube"},"authorids":{"value":["~Yunseok_Jang1","~Yeda_Song1","~Sungryull_Sohn1","~Lajanugen_Logeswaran1","~Tiange_Luo1","~Honglak_Lee2"]},"track":{"value":"Short Paper Track (up to 3 pages)"},"authors":{"value":["Yunseok Jang","Yeda Song","Sungryull Sohn","Lajanugen Logeswaran","Tiange Luo","Honglak Lee"]}},"tmdate":1736861080959,"pdate":1730081753356,"tcdate":1726152278805,"writers":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission48/Authors"],"signatures":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission48/Authors"],"forum":"PcwaP4o7vk","license":"CC BY 4.0","number":48,"cdate":1726152278805,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Post_Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission48/-/Full_Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Edit","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission48/-/Camera-Ready_Revision"],"mdate":1736861080959,"odate":1736861080944,"domain":"NeurIPS.cc/2024/Workshop/Video-Langauge_Models","id":"PcwaP4o7vk","version":2},{"content":{"summary":{"value":"This paper introduces LONGVTG-R1, a RL framework for long-video temporal grounding. It proposes Token-aware KL regularization to balance exploration and exploitation, a dense CenDist reward to mitigate sparse supervision, and an automatically constructed SceneTG dataset with sharp boundaries and unambiguous queries. LONGVTG-R1 provide  evidence that precise temporal localization is a foundational skill for broader video understanding."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1.  Please report an “SFT-only” curve trained on the same 40 k SceneTG samples. If that curve already closes most of the gap with LONGVTG-R1, it would show that data quality, not RL, drives the gains.\n2. Weaknesses1 and Weaknesses2"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. CenDist reward: The paper gives an inexpensive, model-agnostic auxiliary signal that densifies the naturally sparse IoU reward without extra annotation.\n2. The paper purposed  an automatic “propose-then-annotate” pipeline that uses scene-boundary detection plus uniqueness filtering to generate sharp-boundary, unambiguous training pairs, scalable to 20 k unlabeled videos.\n3. Technical sections derive Token-aware KL from first principles and give closed-form properties; pseudocode and illustrative figures accompany data-construction pipeline."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Compared to the baseline model VideoChat-R1, there was no significant performance improvement on several mainstream TG datasets, Charades(61.5 vs 60.8), ANet(36.8 vs 36.6), QVHighLight(57.3 vs 57.5), The model shows significant improvements over the baseline model on some longer datasets, but it cannot be proven whether the improvement is due to the introduction of more long-term training data (40K vs 16K) or the result of the proposed method.\n2. Missing qualitative error analysis. Failure cases are only shown for QA (Fig. 3). Provide grounding failure modes: drifting beyond scene cuts, duplicate predictions, or confusion between “first” vs. “last” event. Categorizing 50 errors would guide future improvements.\n3. Token-aware KL generality claim is overstated. Theory assumes discrete, finite vocabulary; vision-language models often use sub-word tokens where timestamp digits are fragmented (e.g., “1”, “7”, “:” are separate). The paper does not show how often fragmented digits appear or whether relaxing KL on all digit tokens accidentally releases semantically important numeric tokens unrelated to time.\n\nMinor formatting error: in Table2, QVHigh column, two values, 57.5 and 57.3, are both bolded."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762925943881,"tcdate":1761819294698,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission15694/Reviewer_CjPm"],"signatures":["ICLR.cc/2026/Conference/Submission15694/Reviewer_CjPm"],"forum":"8H1HmGH8ua","number":4,"license":"CC BY 4.0","cdate":1761819294698,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission15694/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762925943881,"domain":"ICLR.cc/2026/Conference","replyto":"8H1HmGH8ua","id":"l6AUZ9kVG3","forumContent":{"TLDR":{"value":"We present LongVTG-R1, the first RL-based framework for long-video temporal grounding, which leverages novel regularization, reward design, and dataset construction to achieve state-of-the-art performance and generalize to QA tasks."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["video understanding","temporal grounding","multimodal large language model (MLLM)"]},"supplementary_material":{"value":"/attachment/a7526f2447a3b01b5d091f2498a198a2aae0ee49.zip"},"primary_area":{"value":"foundation or frontier models, including LLMs"},"abstract":{"value":"We present the first Reinforcement Learning (RL)-based framework that equips Multimodal Large Language Models (MLLMs) with long-video temporal grounding skills, and demonstrate that this approach also generalizes to improve performance on general question-answering (QA) tasks. Unlike dominant supervised fine-tuning (SFT) methods, RL enables models to acquire temporal grounding abilities without risking catastrophic forgetting of their core understanding. However, adopting RL for long-video temporal grounding reveals a challenge in balancing exploitation of pre-trained knowledge with exploration of new localization skills. To address this, we propose Token-aware KL Regularization, which selectively relaxes the KL-divergence regularization on timestamp-related tokens to guide exploration. Moreover, effective optimization requires a learning signal that alleviates the sparsity of key events in long videos, for which we introduce a denser reward, the Center Distance Reward (CenDist). To further mitigate grounding ambiguity between language queries and visually similar content, and to facilitate effective RL training, we propose an automatic data construction method and construct a small but high-quality dataset, SceneTG. Our resulting model, QwenLongTG, delivers substantial improvements across three long-video temporal grounding datasets among efficiently fine-tuned MLLMs, and even approaches the performance of densely pre-trained or continually trained models. Beyond temporal grounding, we further verify its generalization to long-video QA: under a “Ground-then-Answer” strategy, QwenLongTG consistently enhances downstream QA performance, serving as an effective first-stage grounding module."},"_bibtex":{"value":"@misc{\nzhang2026longvtgr,\ntitle={Long{VTG}-R1: Reinforcement Learning for Robust Long-Video Temporal Grounding},\nauthor={Zheyu Aqa Zhang and Shixing Chen and Ziqi Pang and Xiang Hao and Kushan Thakkar and Yu-Xiong Wang},\nyear={2026},\nurl={https://openreview.net/forum?id=8H1HmGH8ua}\n}"},"title":{"value":"LongVTG-R1: Reinforcement Learning for Robust Long-Video Temporal Grounding"},"pdf":{"value":"/pdf/1a44d3a0d532c7e585afb40db0defe29d293c3cd.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"zhang|longvtgr1_reinforcement_learning_for_robust_longvideo_temporal_grounding"},"authorids":{"value":["~Zheyu_Aqa_Zhang1","~Shixing_Chen1","~Ziqi_Pang1","~Xiang_Hao1","~Kushan_Thakkar1","~Yu-Xiong_Wang1"]},"authors":{"value":["Zheyu Aqa Zhang","Shixing Chen","Ziqi Pang","Xiang Hao","Kushan Thakkar","Yu-Xiong Wang"]}},"version":2},{"content":{"venue":{"value":"CoRR 2025"},"pdf":{"value":"https://arxiv.org/pdf/2511.19525v2"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"pal|shortcut_invariance_targeted_jacobian_regularization_in_disentangled_latent_space"},"html":{"value":"https://doi.org/10.48550/arXiv.2511.19525"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2511-19525,\n  publtype={informal},\n  author={Shivam Pal and Sakshi Varshney and Piyush Rai},\n  title={Shortcut Invariance: Targeted Jacobian Regularization in Disentangled Latent Space},\n  year={2025},\n  month={November},\n  cdate={1761955200000},\n  journal={CoRR},\n  volume={abs/2511.19525},\n  url={https://doi.org/10.48550/arXiv.2511.19525}\n}\n"},"abstract":{"value":"Deep neural networks are prone to learning shortcuts, spurious correlations present in the training data that undermine out-of-distribution (OOD) generalization. Most prior work mitigates shortcut learning through input-space reweighting, either relying on explicit shortcut labels or inferring shortcut structure from heuristics such as per-sample loss. Moreover, these approaches typically assume the presence of some shortcut-conflicting examples in the training set, an assumption that is often violated in practice, particularly in medical imaging where data is aggregated across institutions with different acquisition protocols. We propose a latent-space method that views shortcut learning as over-reliance on shortcut-aligned axes. In a disentangled latent space, we identify candidate shortcut-aligned axes via their strong correlation with labels and reduce classifier reliance on them by injecting targeted anisotropic noise during training. Unlike prior latent-space based approaches that remove, project out, or adversarially suppress shortcut features, our method preserves the full representation and instead impose functional invariance by regularizing the classifier's sensitivity along those axes. We show that injecting anisotropic noise induces targeted Jacobian and curvature regularization, effectively flattening the decision boundary along shortcut axes while leaving core feature dimensions largely unaffected. Our method achieves state-of-the-art OOD performance across standard shortcut-learning benchmarks without requiring shortcut labels or shortcut-conflicting samples."},"title":{"value":"Shortcut Invariance: Targeted Jacobian Regularization in Disentangled Latent Space"},"authors":{"value":[{"fullname":"Shivam Pal","username":"~Shivam_Pal1"},{"fullname":"Sakshi Varshney","username":""},{"fullname":"Piyush Rai","username":""}]}},"tmdate":1778004922790,"pdate":1767139200000,"externalIds":["dblp:journals/corr/abs-2511-19525"],"tcdate":1778004919307,"writers":["~","OpenReview.net/Public_Article/DBLP.org","OpenReview.net/Support"],"signatures":["~Shivam_Pal1"],"forum":"WMK6tmUkGy","license":"CC BY-SA 4.0","number":6062,"cdate":1761955200000,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/DBLP.org/-/Record","OpenReview.net/Public_Article/-/Edit"],"mdate":1778004922790,"domain":"OpenReview.net/Public_Article","id":"WMK6tmUkGy","version":2},{"content":{"summary":{"value":"This paper proposes MuVi, a new method for generating music that aligns with video content, focusing on both semantic alignment and rhythmic synchronization. MuVi's design includes a \"visual adaptor\" that extracts relevant visual features from videos, which helps guide music generation to match the mood and rhythm of the video. To improve synchronization between visual events and musical beats, the authors use a pre-training technique that contrasts synchronized and unsynchronized video-music pairs, helping the model learn rhythmic alignment."},"soundness":{"value":2},"confidence":{"value":5},"questions":{"value":"In addition to the weaknesses, here are some points that raise further confusion or seem inconsistent in the paper:\n\n1.Claim on Previous V2M Methods (lines 39-41): The authors claim that \"Previous V2M methods focus on global features,\" presenting this as a limitation of past approaches. However, this appears inconsistent with prior work, as several existing methods focus on local clip features for training. For instance, V2Meow [1] and VMAS [2] emphasize local clip-based features, while VidMuse [3] captures both local and global features through long-short-term modeling. The authors should clarify and provide evidence to support their assertion about the emphasis on global features in previous V2M approaches.\n\n2.Choice of Beat Synchronization Metrics and Exclusion of Dance Video Music Generation for Comparison: The authors select Beats Coverage Score (BCS) and Beats Hit Score (BHS) as metrics to evaluate beat synchronization, following the approach in [4] (line 346), which specifically targets music generation for dance videos. However, the authors then claim in line 364 that \"D2M-GAN are not considered for comparison because their scope of application differs from ours.\" If dance-related videos are outside MuVi’s intended scope, it is unclear why dance-specific metrics are being applied for evaluation. This raises a need for clarification.\n\n3.Choice of MuVi(beta) Setting for Comparison: The paper claims \"use CLIP-ViT(base-patch16) and the attention pooling adaptor as the visual encoder\" for MuVi(beta) (lines 366-367). However, Table 1 shows that the VideoMAE V2 with a Softmax adaptor yields better results for this setting. It is unclear why a suboptimal setting was selected for MuVi(beta), as this choice could impact the fairness and interpretability of the comparisons. An explanation from the authors on the rationale for this choice would provide more clarity.\n\n[1] Su K, Li J Y, Huang Q, et al. V2Meow: Meowing to the Visual Beat via Video-to-Music Generation[C]//Proceedings of the AAAI Conference on Artificial Intelligence. 2024, 38(5): 4952-4960.\n\n[2] Lin Y B, Tian Y, Yang L, et al. VMAS: Video-to-Music Generation via Semantic Alignment in Web Music Videos[J]. arXiv preprint arXiv:2409.07450, 2024.\n\n[3] Tian Z, Liu Z, Yuan R, et al. VidMuse: A Simple Video-to-Music Generation Framework with Long-Short-Term Modeling[J]. arXiv preprint arXiv:2406.04321, 2024.\n\n[4] Zhu Y, Olszewski K, Wu Y, et al. Quantized gan for complex music generation from dance videos[C]//European Conference on Computer Vision. Cham: Springer Nature Switzerland, 2022: 182-199."},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. This paper introduces a new method for generating music that aligns with video content, focusing on both semantic alignment and rhythmic synchronization within a generative video-music Diffusion Transformer framework.\n2. The model employs a joint encoder-decoder architecture that integrates a contrastive pre-training scheme for improved synchronization. The inclusion of a \"visual adaptor\" enhances the model’s ability to compress and process high-frame-rate visual inputs, capturing video cues for music generation.\n3. The paper is well-organized, presenting MUVI's methodology alongside a series of experiments. The framework demonstrates superior performance over the baseline on the test dataset across various evaluation metrics, showcasing its effectiveness in video-to-music generation."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1.Novelty and Contribution: The paper presents its main contributions as a visual adaptor and a contrastive training scheme, but visual adaptor techniques and contrastive learning have already been used in video-to-music generation tasks [1, 2] and are commonly employed in multi-modal learning [3, 4]. The design of the visual adaptor lacks unique innovation, primarily involving a selection of common aggregation and pooling methods, which appears more as an ablation study to find the best setting. Overall, the proposed method lacks novelty, and the results in Table 2 indicate that the proposed method does not outperform the baseline across all metrics.\n\n2.Lack of Justification and Explanation: Another weakness is the lack of clear justification and explanation across different sections, from design choices to metric selection.\n\n2.1 The adaptor design section lacks a clear justification. For instance, why were these three adaptor methods chosen, instead of exploring alternative multi-modal adaptors [1, 3]? Why is CLS set as the query instead of the key-value pair?\n\n2.2 The paper introduces various metrics for evaluating model performance, but lacks explanations for each metric. For instance, in lines 446-449: “resulting in lower sound quality, diversity, and poorer synchronization,” it is unclear which metrics specifically measure sound quality, diversity, or synchronization. Additionally, the statement “the channel-wise fusion of the visual conditioning still aids in synchronization” lacks experimental evidence to substantiate this claim.\n\n2.3 Ambiguous phrases like “we believe” (lines 222, 483) and “might lead to” (lines 77, 483) appear multiple times in the paper. Clear support or reasoning should be provided for these assertions.\n\n3. Presentation and Writing: There are some presentation and writing issues within the paper.\n\n3.1 The introduction (line 63) highlights tackling “Integration of foley and sound \teffects,” yet no further details or experiments addressing this topic are provided in the rest of the paper.\n\n3.2 The temporal shift method introduced in Section 3.2 is motivated as a significant contribution, but it lacks a clear explanation. Additionally, some symbols, like \"m,\" are redefined multiple times, \"C'\" is not defined, which may cause confusion for readers.\n\n\n4. Experimental Comparison: A main weakness of the paper is lack of the experimental comparisons, which include only one baseline method.\n\n4.1 A simple baseline could have been constructed by combining an existing video understanding model with a music generation model, similar to the approach in [2, 6].\n\n4.2 The experiments omit comparisons with several relevant state-of-the-art methods, such as Diff-BGM [5], VidMuse [6], and Dance2Music-Diffusion [7].\n\n4.3 The M^2Ugen method shows comparable or superior results in terms of audio quality (Table 2). Fine-tuning this method on the dataset used in this paper could provide additional insight into its performance.\n\nReferences: \n[1] Liu S, Hussain A S, Sun C, et al. M$^{2}$UGen: Multi-modal Music Understanding and Generation with the Power of Large Language Models[J]. arXiv preprint arXiv:2311.11255, 2023.\n\n[2] Lin Y B, Tian Y, Yang L, et al. VMAS: Video-to-Music Generation via Semantic Alignment in Web Music Videos[J]. arXiv preprint arXiv:2409.07450, 2024.\n\n[3] Zhang R, Han J, Liu C, et al. Llama-adapter: Efficient fine-tuning of language models with zero-init attention[J]. arXiv preprint arXiv:2303.16199, 2023.\n\n[4] Radford A, Kim J W, Hallacy C, et al. Learning transferable visual models from natural language supervision[C]//International conference on machine learning. PMLR, 2021: 8748-8763.\n\n[5] Li S, Qin Y, Zheng M, et al. Diff-BGM: A Diffusion Model for Video Background Music Generation[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2024: 27348-27357.\n\n[6] Tian Z, Liu Z, Yuan R, et al. VidMuse: A Simple Video-to-Music Generation Framework with Long-Short-Term Modeling[J]. arXiv preprint arXiv:2406.04321, 2024.\n\n[7] Zhang C, Hua Y. Dance2Music-Diffusion: leveraging latent diffusion models for music generation from dance videos[J]. EURASIP Journal on Audio, Speech, and Music Processing, 2024, 2024(1): 48."}},"nonreaders":[],"tmdate":1731429369873,"tcdate":1730661059717,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission10758/Reviewer_yVkt"],"signatures":["ICLR.cc/2025/Conference/Submission10758/Reviewer_yVkt"],"forum":"E040QmNETN","number":1,"license":"CC BY 4.0","cdate":1730661059717,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission10758/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731429369873,"domain":"ICLR.cc/2025/Conference","replyto":"E040QmNETN","id":"9PqbdkTZSs","forumContent":{"TLDR":{"value":"This paper proposes a novel approach to generate music conditioned on video, focusing on the semantic alignment and rhythmic synchronization between music and video."},"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video-to-music generation","music generation"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Generating music that aligns with the visual content of a video has been a challenging task, as it requires a deep understanding of visual semantics and involves generating music whose melody, rhythm, and dynamics harmonize with the visual narratives. This paper presents MuVi, a novel framework that effectively addresses these challenges to enhance the cohesion and immersive experience of audio-visual content. MuVi analyzes video content through a specially designed visual adaptor to extract contextually and temporally relevant features. These features are used to generate music that not only matches the video’s mood and theme but also its rhythm and pacing. We also introduce a contrastive music-visual pre-training scheme to ensure synchronization, based on the periodicity nature of music phrases. In addition, we demonstrate that our flow-matching-based music generator has in-context learning ability, allowing us to control the style and genre of the generated music. Experimental results show that MuVi demonstrates superior performance in both audio quality and temporal synchronization. The generated music video samples are available at muvi-v2m.github.io."},"_bibtex":{"value":"@misc{\nli2025muvi,\ntitle={MuVi: Video-to-Music Generation with Semantic Alignment and Rhythmic Synchronization},\nauthor={Ruiqi Li and Siqi Zheng and Xize Cheng and Ziang Zhang and Shengpeng Ji and Zhou Zhao},\nyear={2025},\nurl={https://openreview.net/forum?id=E040QmNETN}\n}"},"title":{"value":"MuVi: Video-to-Music Generation with Semantic Alignment and Rhythmic Synchronization"},"pdf":{"value":"/pdf/e5daca4235827fa95cbb8a80969b6ca4c588722c.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"li|muvi_videotomusic_generation_with_semantic_alignment_and_rhythmic_synchronization"},"authorids":{"value":["~Ruiqi_Li2","~Siqi_Zheng1","~Xize_Cheng1","~Ziang_Zhang1","~Shengpeng_Ji1","~Zhou_Zhao3"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Ruiqi Li","Siqi Zheng","Xize Cheng","Ziang Zhang","Shengpeng Ji","Zhou Zhao"]}},"version":2},{"content":{"summary":{"value":"The paper proposes Pusa V1.0, a lightweight adaptation for large pretrained text-to-video (T2V) diffusion models that replaces the scalar timestep with a vectorized timestep (one per frame) and learns with a frame-aware flow-matching objective. The key contribution is Vectorized Timestep Adaptation, a non-destructive modification that enhance the base model’s generation ability. With minimal fine-tuning (LoRA or full FT) on ~3.9k Wan-T2V generated clips, Pusa achieves SOTA-level image-to-video (I2V) results on VBench-I2V and further shows zero-shot behavior for other downstream tasks."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"See the weakness."},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. Simple, general idea with clear motivation. Moving from a scalar to a vectorized timestep directly attacks the synchronized-frame limitation of standard VDMs; the frame-aware flow-matching formulation is clean and well explained.\n2. High efficiency. Comparable I2V performance to Wan-I2V with a tiny fraction of the compute/data, plus favorable results with only 10 sampling steps. The ablations (LoRA vs full FT, timestep sampling) are helpful.\n3. Unified capability. The same model handles I2V, start–end, and extension without task-specific heads or destructive finetuning."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Lack of quantitative evaluation for non-I2V tasks. While the author mentions zero-shot, start–end and other applications of its method, it lacks quantitative evaluation for these applications; most quantitative focus is I2V on VBench-I2V. Concrete metrics are missing from the main text. It would be better for the auther to add quantitative results for applications.\n2. Data source and generalization. Fine-tuning uses videos generated by the same base video model. This could limit diversity and introduce bias toward the base model’s distribution; it’s unclear how well Pusa generalizes to challenging real photos as I2V inputs beyond curated benchmarks. Could you provide some details of the ~3,860 Wan-T2V samples? And are there any data filtering you do?"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762918491646,"tcdate":1761573533360,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission6133/Reviewer_TcSx"],"signatures":["ICLR.cc/2026/Conference/Submission6133/Reviewer_TcSx"],"forum":"4adY8FepXg","number":1,"license":"CC BY 4.0","cdate":1761573533360,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission6133/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762918491646,"domain":"ICLR.cc/2026/Conference","replyto":"4adY8FepXg","id":"wGMJuznhcA","forumContent":{"TLDR":{"value":"Pusa proposes a new fintuing paradigm that leverages vectorized timestep adaptation (VTA) to enable fine-grained temporal control within a unified video diffusion framework, achieving SOTA level performance with unprecedented training efficiency"},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Vectorized Timesteps","Flow Matching","Temporal Modeling","Video Generation"]},"supplementary_material":{"value":"/attachment/28f2b8c62646d383693ac91b727ad6e77b71cfc7.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"The rapid advancement of video diffusion models has been hindered by fundamental limitations in temporal modeling, particularly the rigid synchronization of frame evolution imposed by conventional scalar timestep variables. While task-specific adaptations and autoregressive models have sought to address these challenges, they remain constrained by computational inefficiency, catastrophic forgetting, or narrow applicability. In this work, we present \\textbf{Pusa} V1.0, a versatile model that leverages \\textbf{vectorized timestep adaptation (VTA)} to enable fine-grained temporal control within a unified video diffusion framework. Note that VTA is a non-destructive adaptation, which means that it fully preserves the capabilities of the base model.\n\\textbf{Unlike conventional methods like Wan-I2V, which finetune a base text-to-video (T2V) model with abundant resources to do image-to-video (I2V), we achieve comparable results in a zero-shot manner after an ultra-efficient finetuning process based on VTA. Moreover, this method also unlocks many other zero-shot capabilities simultaneously, such as start-end frames and video extension ---all without task-specific training. Meanwhile, it keeps the T2V capability from the base model.} Mechanistic analyses also reveal that our approach preserves the foundation model's generative priors while surgically injecting temporal dynamics, avoiding the combinatorial explosion inherent to the vectorized timestep. This work establishes a scalable, efficient, and versatile paradigm for next-generation video synthesis, democratizing high-fidelity video generation for research and industry alike."},"_bibtex":{"value":"@inproceedings{\nliu2026pusa,\ntitle={Pusa V1.0: Unlocking Temporal Control in Pretrained Video Diffusion Models via Vectorized Timestep Adaptation},\nauthor={Yaofang Liu and Yumeng REN and Aitor Artola and Yuxuan Hu and Xiaodong Cun and Xiaotong Zhao and Alan Zhao and Raymond H. Chan and Suiyun Zhang and Rui Liu and Dandan Tu and Jean-michel Morel},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=4adY8FepXg}\n}"},"title":{"value":"Pusa V1.0: Unlocking Temporal Control in Pretrained Video Diffusion Models via Vectorized Timestep Adaptation"},"pdf":{"value":"/pdf/8d42fc9215f3f407e6a4a5dd31a5cf579ea8be78.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"liu|pusa_v10_unlocking_temporal_control_in_pretrained_video_diffusion_models_via_vectorized_timestep_adaptation"},"authorids":{"value":["~Yaofang_Liu1","~Yumeng_REN1","~Aitor_Artola1","~Yuxuan_Hu3","~Xiaodong_Cun1","~Xiaotong_Zhao1","~Alan_Zhao1","~Raymond_H._Chan1","~Suiyun_Zhang1","~Rui_Liu17","~Dandan_Tu1","~Jean-michel_Morel1"]},"authors":{"value":["Yaofang Liu","Yumeng REN","Aitor Artola","Yuxuan Hu","Xiaodong Cun","Xiaotong Zhao","Alan Zhao","Raymond H. Chan","Suiyun Zhang","Rui Liu","Dandan Tu","Jean-michel Morel"]}},"version":2},{"content":{"summary":{"value":"This paper proposes MotionWeaver, a multi-humanoid image-to-video framework that (1) learns identity-agnostic 4D motion tokens via the Unified-Choreography Core (UCC), (2) fuses motion and video latents in a shared 4D space using the Hyper-Scene Integrator (HSI) with depth-aware attention and Dynamic Cross-RoPE, and (3) trains with Hierarchical-4D Supervision (H4S) that adds occlusion supervision at high noise and motion-level supervision at low noise. The authors also curate MultiHuman46 (46 hours of multi-human video) and release DualDynamics (300 two-character clips) to stress-test interactions/occlusions, reporting consistent SOTA on that benchmark."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Please see the weaknesses."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The paper clearly argues that 2D control entangles appearance and motion and lacks explicit depth reasoning in multi-person scenes, and responds with a fully 4D-anchored pipeline (UCC, HSI and H4S) that separates motion from morphology, injects depth ordering, and supervises motion explicitly.\n\n2. UCC standardizes SMPL joints to strip appearance cues and binds per-person motion with group attention. HSI adds depth-aware cross-attention plus Dynamic Cross-RoPE to align (t, x, y) between camera and image planes. H4S schedules occlusion vs. motion-level supervision by noise step. These pieces are concrete and novel in combination.\n\n3. This paper provides extensive experiments to support its claim. On DualDynamics, MotionWeaver tops all baselines on every metric, supporting the claim that explicit depth/occlusion handling and identity-motion binding help in multi-humanoid settings. Module-wise ablations and attention visualizations show each component (motion normalization, group attention, depth-aware attention, Dynamic Cross-RoPE, timestep-aware training) matters. Qualitative figures highlight fewer identity swaps and cleaner occlusion ordering than baselines."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"My major concern about this paper is the detailed comparison over existing methods to show the contribution clearly. The multi-person interaction and 4D motion tokens are adopted by previous methods already. I hope the authors could clearly clarify the difference with existing methods.\n\n1. MTVCrafter (Ding et al., 2025) also models raw 4D motion via a tokenizer (4DMoT) and conditions a DiT with 4D positional encodings, reporting large gains on open-world human animation. MotionWeaver likewise trains a 4D tokenizer, uses 4D positional cues (Dynamic Cross-RoPE), and claims generalization beyond humans, so the novelty boundary feels blurred. The paper would benefit from a direct experimental comparison (same control signals/base model), or at least an ablation contrasting quantized motion tokens (MTVCrafter) versus the authors’ standardized-skeleton motion units and group-attention binding.\n\n2. DanceTogether (Chen et al., 2025) introduces a MaskPoseAdapter that fuses tracking masks with noisy pose heatmaps at every denoising step to suppress identity drift when actors swap positions and interact over long horizons. MotionWeaver’s identity binding relies on group attention and occlusion-aware depth cues but evaluates on 49-frame clips. It doesn’t stress identity under frequent cross-overs or long sequences. A head-to-head on DanceTogether’s scenarios/metrics, or adding long-horizon tests with frequent position exchanges, would strengthen claims. Besides, Structural Video Diffusion [a] in ICCV 2025 proposes identity-specific embeddings plus structural learning with depth + surface normals. MotionWeaver also models depth (via depth-aware attention and occlusion loss). A comparison on the [a]'s setup or cross-evaluation across datasets would clarify when explicit identity embeddings vs. identity-agnostic binding are preferable.\n\n[a] Zhenzhi Wang, Yixuan Li, Yanhong Zeng, Yuwei Guo, Dahua Lin, Tianfan Xue, Bo Dai. Multi-identity Human Image Animation with Structural Video Diffusion, ICCV 2025. https://openaccess.thecvf.com/content/ICCV2025/papers/Wang_Multi-identity_Human_Image_Animation_with_Structural_Video_Diffusion_ICCV_2025_paper.pdf\n\n3. DualDynamics emphasizes two-character, 49-frame interactions crafted by an animation team. MultiHuman46 includes AI-generated material and web-crawled clips. While appropriate for controlled studies, this may underestimate long-range identity drift and real-world messiness. Evaluating on longer, real-capture sequences (or adopting external multi-human sets) would improve external validity."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762916148588,"tcdate":1761847034462,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission2220/Reviewer_Yd3q"],"signatures":["ICLR.cc/2026/Conference/Submission2220/Reviewer_Yd3q"],"forum":"KjlLwRsiUE","number":4,"license":"CC BY 4.0","cdate":1761847034462,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission2220/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762916148588,"domain":"ICLR.cc/2026/Conference","replyto":"KjlLwRsiUE","id":"GHKNw7Zz94","forumContent":{"TLDR":{"value":"MotionWeaver is a thoroughly 4D-anchored framework for multi-humanoid animation, enabling robust synthesis across diverse body forms, complex interactions, and severe occlusions."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Character Animation","Diffusion Model","Video Generation"]},"supplementary_material":{"value":"/attachment/c3bd1322ae619af123937e86e128cbee71e70eb0.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Character image animation, which synthesizes videos of reference characters driven by pose sequences, has advanced rapidly but remains largely limited to single-human settings. Existing methods struggle to generalize to multi-humanoid scenarios, which involve diverse humanoid forms, complex interactions, and frequent occlusions. We address this gap with two key innovations. First, we introduce unified motion representations that extract identity-agnostic motions and explicitly bind them to corresponding characters, enabling generalization across diverse humanoid forms and seamless extension to multi-humanoid scenarios. Second, we propose a holistic 4D-anchored paradigm that constructs a shared 4D space to fuse motion representations with video latents, and further reinforces this process with hierarchical 4D-level supervision to better handle interactions and occlusions. We instantiate these ideas in MotionWeaver, an end-to-end framework for multi-humanoid image animation. To support this setting, we curate a 46-hour dataset of multi-human videos with rich interactions, and construct a 300-video benchmark featuring paired humanoid characters. Quantitative and qualitative experiments demonstrate that MotionWeaver not only achieves state-of-the-art results on our benchmark but also generalizes effectively across diverse humanoid forms, complex interactions, and challenging multi-humanoid scenarios."},"_bibtex":{"value":"@inproceedings{\nhu2026motionweaver,\ntitle={MotionWeaver: Holistic 4D-Anchored Framework for Multi-Humanoid Image Animation},\nauthor={Xirui Hu and Yanbo Ding and Jiahao Wang and Tingting Shi and Yali Wang and Guo Zhi Zhi and Weizhan Zhang},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=KjlLwRsiUE}\n}"},"title":{"value":"MotionWeaver: Holistic 4D-Anchored Framework for Multi-Humanoid Image Animation"},"pdf":{"value":"/pdf/d1a321ec84cdba29d5159ac53da30aa5885592f8.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"hu|motionweaver_holistic_4danchored_framework_for_multihumanoid_image_animation"},"authorids":{"value":["~Xirui_Hu1","~Yanbo_Ding2","~Jiahao_Wang14","~Tingting_Shi2","~Yali_Wang1","~Guo_Zhi_Zhi1","~Weizhan_Zhang1"]},"authors":{"value":["Xirui Hu","Yanbo Ding","Jiahao Wang","Tingting Shi","Yali Wang","Guo Zhi Zhi","Weizhan Zhang"]}},"version":2},{"content":{"venue":{"value":"CoRR 2024"},"venueid":{"value":"dblp.org/journals/CORR/2024"},"paperhash":{"value":"torres|davide_depthaware_video_deblurring"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:German_F._Torres:","https://dblp.org/search/pid/api?q=author:Jussi_Kalliola:","https://dblp.org/search/pid/api?q=author:Soumya_Tripathy:","~Erman_Acar4","https://dblp.org/search/pid/api?q=author:Joni-Kristian_Kämäräinen:"]},"html":{"value":"https://doi.org/10.48550/arXiv.2409.01274"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2409-01274,\n  publtype={informal},\n  author={German F. Torres and Jussi Kalliola and Soumya Tripathy and Erman Acar and Joni-Kristian Kämäräinen},\n  title={DAVIDE: Depth-Aware Video Deblurring},\n  year={2024},\n  cdate={1704067200000},\n  journal={CoRR},\n  volume={abs/2409.01274},\n  url={https://doi.org/10.48550/arXiv.2409.01274}\n}\n"},"abstract":{"value":"Video deblurring aims at recovering sharp details from a sequence of blurry frames. Despite the proliferation of depth sensors in mobile phones and the potential of depth information to guide deblurring, depth-aware deblurring has received only limited attention. In this work, we introduce the 'Depth-Aware VIdeo DEblurring' (DAVIDE) dataset to study the impact of depth information in video deblurring. The dataset comprises synchronized blurred, sharp, and depth videos. We investigate how the depth information should be injected into the existing deep RGB video deblurring models, and propose a strong baseline for depth-aware video deblurring. Our findings reveal the significance of depth information in video deblurring and provide insights into the use cases where depth cues are beneficial. In addition, our results demonstrate that while the depth improves deblurring performance, this effect diminishes when models are provided with a longer temporal context. Project page: https://germanftv.github.io/DAVIDE.github.io/ ."},"title":{"value":"DAVIDE: Depth-Aware Video Deblurring"},"authors":{"value":["German F. Torres","Jussi Kalliola","Soumya Tripathy","Erman Acar","Joni-Kristian Kämäräinen"]}},"tmdate":1762941743334,"pdate":1704067200000,"externalIds":["dblp:journals/corr/abs-2409-01274"],"tcdate":1762937978642,"writers":["~"],"signatures":["~Erman_Acar4"],"forum":"FLpGeHRTb7","license":"CC BY-SA 4.0","number":692695,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1762941743334,"domain":"DBLP.org","id":"FLpGeHRTb7","version":2},{"content":{"venue":{"value":"ECCV Workshops (9) 2024"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-031-91838-4_10.pdf"},"venueid":{"value":"dblp.org/conf/ECCV/2024"},"paperhash":{"value":"torres|davide_depthaware_video_deblurring"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:German_F._Torres:","https://dblp.org/search/pid/api?q=author:Jussi_Kalliola:","https://dblp.org/search/pid/api?q=author:Soumya_Tripathy:","~Erman_Acar4","https://dblp.org/search/pid/api?q=author:Joni-Kristian_Kämäräinen:"]},"html":{"value":"https://doi.org/10.1007/978-3-031-91838-4_10"},"_bibtex":{"value":"@inproceedings{DBLP:conf/eccv/TorresKTAK24,\n  author={German F. Torres and Jussi Kalliola and Soumya Tripathy and Erman Acar and Joni-Kristian Kämäräinen},\n  title={DAVIDE: Depth-Aware Video Deblurring},\n  year={2024},\n  cdate={1704067200000},\n  pages={161-179},\n  url={https://doi.org/10.1007/978-3-031-91838-4_10},\n  booktitle={ECCV Workshops (9)},\n  crossref={conf/eccv/2024-w9}\n}\n"},"abstract":{"value":"Video deblurring aims at recovering sharp details from a sequence of blurry frames. Despite the proliferation of depth sensors in mobile phones and the potential of depth information to guide deblurring, depth-aware deblurring has received only limited attention. In this work, we introduce the ‘Depth-Aware VIdeo DEblurring’ (DAVIDE) dataset to study the impact of depth information in video deblurring. The dataset comprises synchronized blurred, sharp, and depth videos. We investigate how the depth information should be injected into the existing deep RGB video deblurring models, and propose a strong baseline for depth-aware video deblurring. Our findings reveal the significance of depth information in video deblurring and provide insights into the use cases where depth cues are beneficial. In addition, our results demonstrate that while the depth improves deblurring performance, this effect diminishes when models are provided with a longer temporal context. Project page: https://germanftv.github.io/DAVIDE.github.io/."},"title":{"value":"DAVIDE: Depth-Aware Video Deblurring"},"authors":{"value":["German F. Torres","Jussi Kalliola","Soumya Tripathy","Erman Acar","Joni-Kristian Kämäräinen"]}},"tmdate":1762941740335,"pdate":1704067200000,"externalIds":["dblp:conf/eccv/TorresKTAK24"],"tcdate":1762937978505,"writers":["~"],"signatures":["~Erman_Acar4"],"forum":"qH9oG3rSGI","license":"CC BY-SA 4.0","number":692693,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1762941740335,"domain":"DBLP.org","id":"qH9oG3rSGI","version":2},{"content":{"venue":{"value":"MICCAI (8) 2024"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-031-72111-3_59.pdf"},"venueid":{"value":"dblp.org/conf/MICCAI/2024"},"paperhash":{"value":"lin|shortcut_learning_in_medical_image_segmentation"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Manxi_Lin:","~Nina_Weng1","https://dblp.org/search/pid/api?q=author:Kamil_Mikolaj:","https://dblp.org/search/pid/api?q=author:Zahra_Bashir:","https://dblp.org/search/pid/api?q=author:Morten_Bo_Søndergaard_Svendsen:","https://dblp.org/search/pid/api?q=author:Martin_Grønnebæk_Tolsgaard:","https://dblp.org/search/pid/api?q=author:Anders_Nymark_Christensen:","https://dblp.org/search/pid/api?q=author:Aasa_Feragen:"]},"html":{"value":"https://doi.org/10.1007/978-3-031-72111-3_59"},"_bibtex":{"value":"@inproceedings{DBLP:conf/miccai/LinWMBSTCF24,\n  author={Manxi Lin and Nina Weng and Kamil Mikolaj and Zahra Bashir and Morten Bo Søndergaard Svendsen and Martin Grønnebæk Tolsgaard and Anders Nymark Christensen and Aasa Feragen},\n  title={Shortcut Learning in Medical Image Segmentation},\n  year={2024},\n  cdate={1704067200000},\n  pages={623-633},\n  url={https://doi.org/10.1007/978-3-031-72111-3_59},\n  booktitle={MICCAI (8)},\n  crossref={conf/miccai/2024-8}\n}\n"},"abstract":{"value":"Shortcut learning is a phenomenon where machine learning models prioritize learning simple, potentially misleading cues from data that do not generalize well beyond the training set. While existing research primarily investigates this in the realm of image classification, this study extends the exploration of shortcut learning into medical image segmentation. We demonstrate that clinical annotations such as calipers, and the combination of zero-padded convolutions and center-cropped training sets in the dataset can inadvertently serve as shortcuts, impacting segmentation accuracy. We identify and evaluate the shortcut learning on two different but common medical image segmentation tasks. In addition, we suggest strategies to mitigate the influence of shortcut learning and improve the generalizability of the segmentation models. By uncovering the presence and implications of shortcuts in medical image segmentation, we provide insights and methodologies for evaluating and overcoming this pervasive challenge and call for attention in the community for shortcuts in segmentation. Our code is public at https://github.com/nina-weng/shortcut_skinseg."},"title":{"value":"Shortcut Learning in Medical Image Segmentation"},"authors":{"value":["Manxi Lin","Nina Weng","Kamil Mikolaj","Zahra Bashir","Morten Bo Søndergaard Svendsen","Martin Grønnebæk Tolsgaard","Anders Nymark Christensen","Aasa Feragen"]}},"tmdate":1741155072443,"pdate":1704067200000,"tcdate":1741155069047,"writers":["~"],"signatures":["~Nina_Weng1"],"forum":"LXOpCwYMKM","license":"CC BY-SA 4.0","number":352654,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1741155072443,"domain":"DBLP.org","id":"LXOpCwYMKM","version":2},{"content":{"venue":{"value":"ICLR 2025 Oral"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["diffusion","flow-matching","fast inference","distillation"]},"primary_area":{"value":"generative models"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Diffusion models and flow matching models have enabled generating diverse and realistic images by learning to transfer noise to data. However, sampling from these models involves iterative denoising over many neural network passes, making generation slow and expensive. Previous approaches for speeding up sampling require complex training regimes, such as multiple training phases, multiple networks, or fragile scheduling. We introduce Shortcut Models, a family of generative models that use a single network and training phase to produce high-quality samples in a single or multiple sampling steps. Shortcut models condition the network not only on the current noise level but also on the desired step size, allowing the model to skip ahead in the generation process. Across a wide range of sampling step budgets, shortcut models consistently produce higher quality samples than previous approaches, such as consistency models and reflow. Compared to distillation, shortcut models reduce complexity to a single network and training phase and additionally allow varying step budgets at inference time."},"_bibtex":{"value":"@inproceedings{\nfrans2025one,\ntitle={One Step Diffusion via Shortcut Models},\nauthor={Kevin Frans and Danijar Hafner and Sergey Levine and Pieter Abbeel},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=OlzB6LnXcS}\n}"},"title":{"value":"One Step Diffusion via Shortcut Models"},"pdf":{"value":"/pdf/834505d749dfe267a983986c87c26443a58835cb.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"frans|one_step_diffusion_via_shortcut_models"},"authorids":{"value":["~Kevin_Frans1","~Danijar_Hafner1","~Sergey_Levine1","~Pieter_Abbeel2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Kevin Frans","Danijar Hafner","Sergey Levine","Pieter Abbeel"]}},"tmdate":1740869755195,"pdate":1737562391781,"tcdate":1727296983696,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission5115/Authors"],"signatures":["ICLR.cc/2025/Conference/Submission5115/Authors"],"forum":"OlzB6LnXcS","license":"CC BY 4.0","number":5115,"cdate":1727296983696,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/-/Submission","ICLR.cc/2025/Conference/-/Post_Submission","ICLR.cc/2025/Conference/Submission5115/-/Full_Submission","ICLR.cc/2025/Conference/-/Edit","ICLR.cc/2025/Conference/Submission5115/-/Camera_Ready_Revision"],"mdate":1740869755195,"odate":1728008565725,"domain":"ICLR.cc/2025/Conference","id":"OlzB6LnXcS","version":2},{"content":{"summary":{"value":"The authors propose a two-stage approach for behaviour cloning via video generation. In the first stage, a video diffusion model generates aa video of a rollout of a policy conditioned on the image of the scene and a task description. In the second stage an inverse dynamics model is used to control the robot given the video. The video model is pre-trained on a large dataset of robot demonstrations before fine-tuning on a few demonstrations from a target domain. Experiments demonstrate that this approach demonstrates strong results compared to end2end video+action generation models like VPP in a very low data regime in the real world. In simulation it performs on par with Pi0."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"Please provide a major revision of the positioning of the paper (a rewrite of the abstract and introduction) to address the positioning issues raised above.\n\nPlease provide more convincing evidence of the benefits of the propped approach compared to prior work."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"The paper is relatively well written and easy to follow.\n\nThe proposed approach is sound. The masked inverse dynamics model has some non-trivial novelty.\n\nStrong results are reported in a real-world evaluation in a very low data regime.\n\nA minimal ablation study is reported."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The positioning of the paper is inaccurate in several ways:\n\n1. The abstracts opens with this statement: “pixel-to-action VLA pipelines typically degenerate under background and view-point shifts”, which is simply incorrect for the latest pixel-to-action pipelines (as demonstrated later in the paper where the prosed approach performs worse than pixel-to-action Pi0 out of domain on the RoboTwin benchmark). \n\n2. While the authors do discuss existing works on video generation + robot control in related work, the introduction is written in such a way as if they are the first to propose this idea. \n\n3. In the related work section the authors mischaracterize the multi-view, 2-stage Dreamitate model as being end2end. \n\n4. Introduction claims that “Empirically, Vidar achieves state-of-the-art performance on the RoboTwin”, which is not true, it underperforms compared to Pi0 in an out-of-distribution evaluation.\n\nOverall, both the novelty and the contributions are over-claimed. The only (minimal) novelty I can see is in the design of the masked inverse dynamics model and the benefits of the proposed approach are not convincingly demonstrated. Real world evaluation is not reproducible and on RoboTwin the proposed approach performs on par with prior work (slightly better in-domain, slightly worse out-of-domain, which is supposed to be the key strength of the proposed approach). \n\nFor some reason the authors only use 40% of the demonstrations in RoboTwin. Results with 100% of the demonstrations need to be reported for reference. \n\nIt would be even better to report results on the more widely used RoboCasa or LIBERO benchmarks instead, using the standard protocol for these benchmarks. RoboCasa specifically targets out-of-domain generalization in a low data regime (50 demonstrations per task). \n\nIt's unclear why different video generation models are used for real and sim experiments. Results should be reported under a uniform model (or with both models in both setting for completeness).\n\nThe proposed approach is primarily tested on short-horizon pick-and-place and simple bimanual grasping tasks, while baselines such as VPP or UniPi were originally demonstrated on more complex or diverse scenarios."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762923321664,"tcdate":1761849905226,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission12434/Reviewer_pjgR"],"signatures":["ICLR.cc/2026/Conference/Submission12434/Reviewer_pjgR"],"forum":"CFuNu8dK4s","number":2,"license":"CC BY 4.0","cdate":1761849905226,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission12434/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762923321664,"domain":"ICLR.cc/2026/Conference","replyto":"CFuNu8dK4s","id":"rT45LFMouk","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Robotics","Video Diffusion Model","Computer Vision"]},"supplementary_material":{"value":"/attachment/800ae62e2c85f1f1900121ca4f82206892dd63d0.zip"},"primary_area":{"value":"applications to robotics, autonomy, planning"},"abstract":{"value":"Scaling general-purpose manipulation to new robot embodiments remains challenging: each platform typically needs large, homogeneous demonstrations, and end-to-end pixel-to-action pipelines may degenerate under background and viewpoint shifts. Based on previous advances in video-based robot control, we present Vidar, consisting of an embodied video diffusion model as the generalizable prior and a masked inverse dynamics model (MIDM) as the adapter. We leverage a video diffusion model pre-trained at Internet scale, and further continuously pre-train it for the embodied domain using 750K multi-view trajectories collected from three real-world robot platforms. For this embodied pre-training, we introduce a unified observation space that jointly encodes robot, camera, task, and scene contexts. The MIDM module learns action-relevant pixel masks without dense labels, grounding the prior into the target embodiment’s action space while suppressing distractors. With only ∼20 minutes of human demonstrations on an unseen robot (∼1% of typical data), Vidar outperforms state-of-the-art baselines and generalizes to unseen tasks, backgrounds, and camera layouts. Our results suggest a scalable recipe for “one prior, many embodiments”: strong, inexpensive video priors together with minimal on-robot alignment."},"_bibtex":{"value":"@misc{\nfeng2026vidar,\ntitle={Vidar: Embodied Video Diffusion Model for Generalist Manipulation},\nauthor={Yao Feng and Hengkai Tan and Xinyi Mao and Chendong Xiang and Guodong Liu and Shuhe Huang and Hang Su and Jun Zhu},\nyear={2026},\nurl={https://openreview.net/forum?id=CFuNu8dK4s}\n}"},"title":{"value":"Vidar: Embodied Video Diffusion Model for Generalist Manipulation"},"pdf":{"value":"/pdf/b45cafef389fd0f6dabb96177b698f3894696ba9.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"feng|vidar_embodied_video_diffusion_model_for_generalist_manipulation"},"authorids":{"value":["~Yao_Feng2","~Hengkai_Tan1","~Xinyi_Mao1","~Chendong_Xiang1","~Guodong_Liu8","~Shuhe_Huang1","~Hang_Su3","~Jun_Zhu2"]},"authors":{"value":["Yao Feng","Hengkai Tan","Xinyi Mao","Chendong Xiang","Guodong Liu","Shuhe Huang","Hang Su","Jun Zhu"]}},"version":2},{"content":{"summary":{"value":"This paper introduces UniCon, a unified framework for handling diverse conditional generation tasks by learning a joint distribution over correlated image pairs using diffusion models. UniCon employs a simple architecture that adds minimal learned parameters (15% of base model) and uses a single efficient training stage while maintaining the standard model input. Through extensive experiments, the authors demonstrate that their single model can produce comparable or better results than specialized methods and prior unified approaches, while also showing that multiple models can be effectively combined for multi-signal conditional generation."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. What is the inference cost of depth prediction compared to other traditional and recent approaches?\n\n2. How does your method compare with ZoeDepth, Depth Anything, and Depth Anything v2?\n\n3. What are the specific advantages of joint distribution modeling over other approaches discussed in the paper, particularly in terms of inference, accuracy, and other metrics?"},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"- Unified framework for handling diverse conditional generation tasks through a joint distribution approach\n- Lightweight adaptation of existing diffusion models with minimal parameter overhead\n- Clear empirical demonstration of the proposed framework"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. **Missing Comparisons/References**\n\n* The paper lacks comparisons with several important recent methods in depth estimation, and this limits our understanding of where the method stands in relation to the current state-of-the-art in depth estimation.\n   - ZoeDepth\n   - Depth Anything\n   - Depth Anything v2 (which is only used as an annotator in this work)\n\n* In addition, it will be helpful to discuss \"DAG: Depth-Aware Guidance with Denoising Diffusion Probabilistic Models\" [Kim et al.], which appears to be a relevant line of work for conditional generation with depth information.\n\n2. **Technical Ambiguities**\n* Section 4.2 should explicitly state the model architecture, e.g., whether the approach uses two pre-trained diffusion models and requires multiple feed-forward passes.\n* This paper does not provide sufficient analysis of the inference cost. Based on current understanding, while it uses LoRA, it requires twice as many feed-forward passes, making the inference step complex and resource-intensive.\n\n3. **Joint Distribution Modeling**\n\n* While the paper presents joint distribution modeling as a key contribution, it should better justify why this particular formulation is advantageous over alternatives. The theoretical benefits of learning the joint distribution p(x,y) versus other approaches could be more thoroughly explained and experimented. For instance, their joint modeling might mutually enhance depth-aware image generation through improved representation."}},"nonreaders":[],"tmdate":1731427911852,"tcdate":1730687149765,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission4930/Reviewer_xQEt"],"signatures":["ICLR.cc/2025/Conference/Submission4930/Reviewer_xQEt"],"forum":"tAGmxz1TUi","number":4,"license":"CC BY 4.0","cdate":1730687149765,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission4930/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427911852,"domain":"ICLR.cc/2025/Conference","replyto":"tAGmxz1TUi","id":"149jjD1ACn","forumContent":{"TLDR":{"value":"We propose a simple architecture and training scheme that produces a single diffusion model that can be used for flexible conditional generation, estimation, and joint diffusion."},"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["image generation","controllability","estimation"]},"primary_area":{"value":"generative models"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Recent progress in image generation has sparked research into controlling these models through condition signals, with various methods addressing specific challenges in conditional generation. Instead of proposing another specialized technique, we introduce a simple, unified framework to handle diverse conditional generation tasks involving a specific image-condition correlation. By learning a joint distribution over a correlated image pair (e.g. image and depth) with a diffusion model, our approach enables versatile capabilities via different inference-time sampling schemes, including controllable image generation (e.g. depth to image), estimation (e.g. image to depth), signal guidance, joint generation (image \\& depth), and coarse control. Previous attempts at unification often introduce complexity through multi-stage training, architectural modification, or increased parameter counts. In contrast, our simplified formulation requires a single, computationally efficient training stage, maintains the standard model input, and adds minimal learned parameters (15% of the base model). Moreover, our model supports additional capabilities like non-spatially aligned and coarse conditioning. Extensive results show that our single model can produce comparable results with specialized methods and better results than prior unified methods. We also demonstrate that multiple models can be effectively combined for multi-signal conditional generation."},"_bibtex":{"value":"@inproceedings{\nli2025a,\ntitle={A Simple Approach to Unifying Diffusion-based Conditional Generation},\nauthor={Xirui Li and Charles Herrmann and Kelvin C.K. Chan and Yinxiao Li and Deqing Sun and Chao Ma and Ming-Hsuan Yang},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=tAGmxz1TUi}\n}"},"title":{"value":"A Simple Approach to Unifying Diffusion-based Conditional Generation"},"pdf":{"value":"/pdf/1eae9881e546a4bb5d913b4f39a62a8b85c3eb65.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"li|a_simple_approach_to_unifying_diffusionbased_conditional_generation"},"authorids":{"value":["~Xirui_Li1","~Charles_Herrmann1","~Kelvin_C.K._Chan1","~Yinxiao_Li2","~Deqing_Sun2","~Chao_Ma3","~Ming-Hsuan_Yang1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Xirui Li","Charles Herrmann","Kelvin C.K. Chan","Yinxiao Li","Deqing Sun","Chao Ma","Ming-Hsuan Yang"]}},"version":2},{"content":{"summary":{"value":"The paper proposes CoSeLECT, a training-free adaptive frame selection method that efficiently selects the most informative frames from a large pool by combining query relevance and temporal continuity. This approach achieves better performance than existing training-free methods on multiple video understanding benchmarks."},"soundness":{"value":2},"confidence":{"value":5},"questions":{"value":"- I am confused about the avoidance of redundant calculations mentioned in the article in line 204. Although  LLaVA-OneVision uses SigLIP-So400M-patch14-384 as the visual encoder, it fine-tunes SigLIP during the training process, which results in them having the same structure but different parameters. So is it really possible to avoid redundant calculations? \n> Since SigLIP is also used as the vision encoder in LLaVA-OneVision (Li et al., 2024a), these embeddings can be directly reused, avoiding redundant computation.\n\n- Does pre-$\\textbf{E}_{im}$  in section 4 mean $N$ in section 3, and post-$\\textbf{E}_{im}$  means  $K$? If so, please use consistent expressions to improve reading; if not, I hope the author can further elaborate on the difference.\n\n- Lack of in-depth analysis of Table 1. With the increase of pre-$\\textbf{E}_{im}$, there is no consistent performance improvement across all benchmarks. Does this contradict the motivation of the paper?\n> Crucially, these heavier methods are typically limited to sparsely pre-sampled frame pools in order to remain computationally feasible—risking the permanent loss of “needle-in-a-haystack” moments before the selection algorithm can even evaluate them, a limitation that becomes particularly acute under resource constraints.\n\n- I am confused about the experimental results in Table 2. Is this comparison meaningful?\n  - Frame-Voyager is similar to the proposed CoSeLECT, and is also a plug-and-play model. But its LLM Size does not seem to be 7B. \n  - LongVA and VideoChat2 are Video-LLMs. How to compare them with CoSeLECT ? \n\n- Supplement CoSeLECT compares the results of the Qwen2-VL [1] experiment with AKS and Q-Frame [2], which will provide a more comprehensive evaluation. \n\n[1] Wang P, Bai S, Tan S, et al. Qwen2-vl: Enhancing vision-language model's perception of the world at any resolution[J]. arXiv preprint arXiv:2409.12191, 2024.\n\n[2] Zhang S, Yang J, Yin J, et al. Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs[J]. arXiv preprint arXiv:2506.22139, 2025."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"The paper presents a clear and practical training-free frame selection method that combines query relevance and visual continuity in a straightforward manner. While the individual components are not novel, their combination into an adaptive, query-aware selection pipeline is sensibly designed and effectively executed. The method is well-described, easy to reproduce, and evaluated across multiple benchmarks with ablations that support key design choices. Its significance lies in offering a lightweight, plug-and-play solution that improves efficiency and performance for video understanding with MLLMs without requiring model retraining. This is useful for real-world deployment, though not theoretically groundbreaking."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The novelty of the paper is limited. Query-aware frame selection for Videl-LLM is not innovative, as discussed in KeyVideoLLM [1], AKS [2], and Q-Frame [3]. The paper pointed out that `these heavier methods are typically limited to sparsely pre-sampled frame pools` is interesting, but the experiment did not support the solution of this problem.\n\n- It is not clear whether the comparison of the experimental results in the paper with other training-based methods is fair.\n\n- The paper claims the proposed CoSeLECT is lightweight. However, there is a lack of systematic evaluation of latency and computing consumption, which is crucial for actual deployment.\n> but in its lightweight, principled fusion of two readily available ones—frame–text similarity for semantic relevance and inter-frame similarity for temporal continuity.\n\n- Lack of discussion of limitations.\n\n- Minor weaknesses\n  - The first equation in Section 3.4 is missing a number\n  - $\\sqrt{D_i}$ in equation (4) lacks definition\n\n[1] Liang H, Li J, Bai T, et al. Keyvideollm: Towards large-scale video keyframe selection[J]. arXiv preprint arXiv:2407.03104, 2024.\n[2] Tang X, Qiu J, Xie L, et al. Adaptive keyframe sampling for long video understanding[C]//Proceedings of the Computer Vision and Pattern Recognition Conference. 2025: 29118-29128.\n[3] Zhang S, Yang J, Yin J, et al. Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs[J]. arXiv preprint arXiv:2506.22139, 2025."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762923541807,"tcdate":1761110132110,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission12714/Reviewer_Lg9P"],"signatures":["ICLR.cc/2026/Conference/Submission12714/Reviewer_Lg9P"],"forum":"Pr3I3ewBFU","number":1,"license":"CC BY 4.0","cdate":1761110132110,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission12714/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762923541807,"domain":"ICLR.cc/2026/Conference","replyto":"Pr3I3ewBFU","id":"Nyp52AoR6W","forumContent":{"TLDR":{"value":"Training-Free Adaptive Frame Selection to improve efficiency and accuracy of Video Large Language Models"},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Video Language Understanding","Efficient Multimodal Large Language Models"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Multimodal Large Language Models (MLLMs) have shown strong performance on image understanding tasks, but video comprehension remains a significant challenge due to the high computational cost of processing long frame sequences and the limited token capacity of underlying Large Language Models (LLMs). Prior approaches to address this often rely on uniform frame sampling, query-agnostic pruning, or require costly training of dedicated compression modules. In this work, we introduce CoSeLECT, a training-free, plug-and-play, query-guided frame selection method that intelligently subsamples video frames for efficient use in MLLMs. CoSeLECT leverages two key signals: temporal redundancy, which identifies similar frame clusters, and query relevance, which selects frames based on their semantic alignment with the input query. By combining these signals through an adaptive frame selection strategy, CoSeLECT selects frames that are both diverse and highly relevant to the query, without requiring any model-specific tuning. Our results on various base MLLMs show that CoSeLECT consistently outperforms state-of-the-art methods, including trained methods like LongVU by +3.8\\% on MVBench and +0.8\\% on VideoMME."},"_bibtex":{"value":"@misc{\ndevnani2026trainingfree,\ntitle={Training-Free Adaptive Frame Selection for Video-Language Understanding},\nauthor={Bhavika Suresh Devnani and Jitesh Jain and Humphrey Shi and Judy Hoffman},\nyear={2026},\nurl={https://openreview.net/forum?id=Pr3I3ewBFU}\n}"},"title":{"value":"Training-Free Adaptive Frame Selection for Video-Language Understanding"},"pdf":{"value":"/pdf/85bab755c7254aed0b86d31707b4fac0a92d777d.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"devnani|trainingfree_adaptive_frame_selection_for_videolanguage_understanding"},"authorids":{"value":["~Bhavika_Suresh_Devnani1","~Jitesh_Jain1","~Humphrey_Shi1","~Judy_Hoffman1"]},"authors":{"value":["Bhavika Suresh Devnani","Jitesh Jain","Humphrey Shi","Judy Hoffman"]}},"version":2},{"content":{"summary":{"value":"This paper presents HAWK, a framework that uses large Visual Language Models (VLMs) to accurately interpret video anomalies. HAWK integrates motion and video modalities through a dual-branch framework, enhanced by an auxiliary consistency loss to focus on motion-related features. The authors annotated over 8,000 anomaly videos with language descriptions and created 8,000 question-answer pairs to train the model across diverse scenarios. HAWK demonstrates state-of-the-art performance."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"See the weaknesses."},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The framework benefits from extensive annotations and question-answer pairs, improving training quality. \n2. HAWK achieves superior results in video description generation and question-answering tasks, outperforming existing baselines.\n3. The paper is well-written and easy to follow."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. In Table 1, it would be better to indicate the LLM backbone of the methods for a fair comparison.\n2. From the ablation study in Table 3, it appears that even without the motion modality, the baseline model achieves comparable performance to other methods without motion modality. This may be because the high-quality dataset and strong LLM backbone contribute more to the performance, which weakens the perceived technical contribution of the motion modality."},"limitations":{"value":"See the weaknesses."}},"nonreaders":[],"tmdate":1730878748530,"tcdate":1721147781767,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission2096/Reviewer_b4dX"],"signatures":["NeurIPS.cc/2024/Conference/Submission2096/Reviewer_b4dX"],"forum":"vBKoEZ1PG3","number":4,"license":"CC BY 4.0","cdate":1721147781767,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission2096/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730878748530,"domain":"NeurIPS.cc/2024/Conference","replyto":"vBKoEZ1PG3","id":"KnFAQEs5XF","forumContent":{"venue":{"value":"NeurIPS 2024 poster"},"TLDR":{"value":"A large vision-language model for understanding open-world video anomalies."},"keywords":{"value":["Video Anomalies Understanding"]},"primary_area":{"value":"machine_vision"},"abstract":{"value":"Video Anomaly Detection (VAD) systems can autonomously monitor and identify disturbances, reducing the need for manual labor and associated costs. However, current VAD systems are often limited by their superficial semantic understanding of scenes and minimal user interaction. Additionally, the prevalent data scarcity in existing datasets restricts their applicability in open-world scenarios.\nIn this paper, we introduce HAWK, a novel framework that leverages interactive large Visual Language Models (VLM) to interpret video anomalies precisely. Recognizing the difference in motion information between abnormal and normal videos, HAWK explicitly integrates motion modality to enhance anomaly identification. To reinforce motion attention, we construct an auxiliary consistency loss within the motion and video space, guiding the video branch to focus on the motion modality. Moreover, to improve the interpretation of motion-to-language, we establish a clear supervisory relationship between motion and its linguistic representation. Furthermore, we have annotated over 8,000 anomaly videos with language descriptions, enabling effective training across diverse open-world scenarios, and also created 8,000 question-answering pairs for users' open-world questions. The final results demonstrate that HAWK achieves SOTA performance, surpassing existing baselines in both video description generation and question-answering. Our codes/dataset/demo will be released at https://github.com/jqtangust/hawk."},"_bibtex":{"value":"@inproceedings{\ntang2024hawk,\ntitle={{HAWK}: Learning to Understand Open-World Video Anomalies},\nauthor={Jiaqi Tang and Hao LU and RUIZHENG WU and Xiaogang Xu and Ke Ma and Cheng Fang and Bin Guo and Jiangbo Lu and Qifeng Chen and Ying-Cong Chen},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=vBKoEZ1PG3}\n}"},"title":{"value":"HAWK: Learning to Understand Open-World Video Anomalies"},"pdf":{"value":"/pdf/72f5b57d7e449d765338f2a60df13977e0717488.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"tang|hawk_learning_to_understand_openworld_video_anomalies"},"authorids":{"value":["~Jiaqi_Tang1","~Hao_LU8","~RUIZHENG_WU1","~Xiaogang_Xu2","~Ke_Ma11","~Cheng_Fang6","~Bin_Guo3","~Jiangbo_Lu1","~Qifeng_Chen1","~Ying-Cong_Chen1"]},"authors":{"value":["Jiaqi Tang","Hao LU","RUIZHENG WU","Xiaogang Xu","Ke Ma","Cheng Fang","Bin Guo","Jiangbo Lu","Qifeng Chen","Ying-Cong Chen"]}},"version":2},{"content":{"research_area_keywords":{"value":"contrastive learning"},"keywords":{"value":["shortcut learning","spurious correlations","deployment-time","contrastive learning"]},"B_use_or_create_scientific_artifacts":{"value":"Yes"},"languages_studied":{"value":"English"},"D4_ethics_review_board_approval":{"value":"N/A"},"C4_parameters_for_packages":{"value":"Yes"},"B6_elaboration":{"value":"Appendix Section A.1"},"B1_cite_creators_of_artifacts":{"value":"Yes"},"B2_discuss_the_license_for_artifacts":{"value":"Yes"},"B4_elaboration":{"value":"Appendix Section D"},"B5_elaboration":{"value":"Section 6.1 and Appendix A.1"},"D_human_subjects_including_annotators":{"value":"No"},"A2_elaboration":{"value":"Appendix Section E"},"C3_elaboration":{"value":"Section 4,6, Appendix A.5"},"B2_elaboration":{"value":"Appendix D"},"B3_elaboration":{"value":"Appendix D"},"C2_elaboration":{"value":"Appendix A.5"},"C4_elaboration":{"value":"Appendix A.5 and A.6"},"D3_elaboration":{"value":"Appendix Section D"},"E_ai_assistants_in_research_or_writing":{"value":"Yes"},"C1_elaboration":{"value":"Appendix A.5"},"B1_elaboration":{"value":"Section6.1 discussed the datasets, and Appendix A.6 discussed the baselines compared in this manuscript."},"C2_experimental_setup_and_hyperparameters":{"value":"Yes"},"E1_elaboration":{"value":"AI assistants were used for grammar, writing assistance, code debugging, literature search. All research content, experimental design, and conclusions are the work of the authors."},"C1_model_size_and_budget":{"value":"Yes"},"venue":{"value":"ACL ARR 2026 May Submission"},"D3_data_consent":{"value":"Yes"},"_bibtex":{"value":"@inproceedings{\nanonymous2026models,\ntitle={Models Know Their Shortcuts: Deployment-Time Shortcut Mitigation},\nauthor={Anonymous},\nbooktitle={Submitted to ACL Rolling Review - May 2026},\nyear={2026},\nurl={https://openreview.net/forum?id=B73oQIFoLk},\nnote={under review}\n}"},"title":{"value":"Models Know Their Shortcuts: Deployment-Time Shortcut Mitigation"},"C3_descriptive_statistics":{"value":"Yes"},"contribution_types":{"value":["NLP engineering experiment"]},"A1_limitations_section":{"value":"This paper has a limitations section."},"B6_statistics_for_data":{"value":"Yes"},"B4_data_contains_personally_identifying_info_or_offensive_content":{"value":"Yes"},"B3_artifact_use_consistent_with_intended_use":{"value":"Yes"},"A2_potential_risks":{"value":"Yes"},"E1_information_about_use_of_ai_assistants":{"value":"Yes"},"abstract":{"value":"Pretrained text encoders are prone to shortcut learning, relying on token-label correlations that fail once the distribution shifts in deployment. Existing shortcut mitigation methods mainly operate at training time and assume access to training data, training dynamics, or shortcut annotations, which are hardly available during deployment, where only the converged model remains. We show that this model alone suffices to mitigate shortcuts during deployment: a biased model internalizes a signal of its learned shortcuts that can be captured via unsupervised gradient-based attribution. We further prove that deployment-time mitigation is information-theoretically upper-bounded by training-time mitigation. Nevertheless, exploiting this gradient signal, our proposed unsupervised deployment-time shortcut mitigation framework for pretrained text encoders, Shortcut Guardrail, recovers substantial performance under shortcut distribution shift, matching or outperforming training-time baselines across sentiment classification, toxicity detection, and natural language inference."},"D1_instructions_given_to_participants":{"value":"N/A"},"paper_type":{"value":"Long"},"C_computational_experiments":{"value":"Yes"},"pdf":{"value":"/pdf/cfd9e5d5e84577e719842abe6cc74031877c79dd.pdf"},"research_area":{"value":"Machine Learning for NLP"},"B5_documentation_of_artifacts":{"value":"Yes"},"EMNLP_2026_AI_Reviewing_Experiment":{"value":"no"},"D2_recruitment_and_payment":{"value":"N/A"},"venueid":{"value":"aclweb.org/ACL/ARR/2026/May/Submission"}},"tmdate":1790696951167,"tcdate":1779678008333,"writers":["aclweb.org/ACL/ARR/2026/May","aclweb.org/ACL/ARR/2026/May/Submission6296/Authors"],"signatures":["aclweb.org/ACL/ARR/2026/May/Submission6296/Authors"],"forum":"B73oQIFoLk","license":"CC BY 4.0","number":6296,"cdate":1779678008333,"readers":["everyone"],"invitations":["aclweb.org/ACL/ARR/2026/May/-/Submission","aclweb.org/ACL/ARR/2026/May/-/Edit","aclweb.org/ACL/ARR/2026/May/-/Post_Submission","aclweb.org/ACL/ARR/2026/May/-/Preprint_Post_Submission","aclweb.org/ACL/ARR/2026/May/Submission6296/-/Blind_Submission_License_Agreement"],"mdate":1790696951167,"odate":1780385598983,"domain":"aclweb.org/ACL/ARR/2026/May","id":"B73oQIFoLk","version":2},{"content":{"venue":{"value":"IEEE Trans. Circuits Syst. Video Technol. 2024"},"pdf":{"value":"https://ieeexplore.ieee.org/iel7/76/10422872/10167696.pdf"},"venueid":{"value":"dblp.org/journals/TCSV/2024"},"paperhash":{"value":"cao|unsupervised_hdr_image_and_video_tone_mapping_via_contrastive_learning"},"authorids":{"value":["~Cong_Cao1","~Huanjing_Yue2","~Xin_Liu13","https://dblp.org/search/pid/api?q=author:Jing-Yu_Yang_0002:"]},"html":{"value":"https://doi.org/10.1109/TCSVT.2023.3290351"},"_bibtex":{"value":"@article{DBLP:journals/tcsv/CaoYLY24,\n  author={Cong Cao and Huanjing Yue and Xin Liu and Jing-Yu Yang},\n  title={Unsupervised HDR Image and Video Tone Mapping via Contrastive Learning},\n  year={2024},\n  month={February},\n  cdate={1706745600000},\n  journal={IEEE Trans. Circuits Syst. Video Technol.},\n  volume={34},\n  number={2},\n  pages={786-798},\n  url={https://doi.org/10.1109/TCSVT.2023.3290351}\n}\n"},"abstract":{"value":"Capturing high dynamic range (HDR) images (videos) is attractive because it can reveal the details in both dark and bright regions. Since the mainstream screens only support low dynamic range (LDR) content, tone mapping algorithm is required to compress the dynamic range of HDR images (videos). Although image tone mapping has been widely explored, video tone mapping is lagging behind, especially for the deep-learning-based methods, due to the lack of HDR-LDR video pairs. In this work, we propose a unified framework (IVTMNet) for unsupervised image and video tone mapping. To improve unsupervised training, we propose domain and instance based contrastive learning loss. Instead of using a universal feature extractor, such as VGG to extract the features for similarity measurement, we propose a novel latent code, which is an aggregation of the brightness and contrast of extracted features, to measure the similarity of different pairs. We totally construct two negative pairs and three positive pairs to constrain the latent codes of tone mapped results. For the network structure, we propose a spatial-feature-enhanced (SFE) module to enable information exchange and transformation of nonlocal regions. For video tone mapping, we propose a temporal-feature-replaced (TFR) module to efficiently utilize the temporal correlation and improve the temporal consistency of video tone-mapped results. We construct a large-scale unpaired HDR-LDR video dataset to facilitate the unsupervised training process for video tone mapping. Experimental results demonstrate that our method outperforms state-of-the-art image and video tone mapping methods. Our code and dataset are available at https://github.com/cao-cong/UnCLTMO ."},"title":{"value":"Unsupervised HDR Image and Video Tone Mapping via Contrastive Learning"},"authors":{"value":["Cong Cao","Huanjing Yue","Xin Liu","Jing-Yu Yang"]}},"tmdate":1744192427052,"pdate":1704067200000,"tcdate":1727460471248,"writers":["~"],"signatures":["~Xin_Liu25"],"forum":"y0l0bxIwPW","license":"CC BY-SA 4.0","number":92341,"cdate":1706745600000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1744192427052,"domain":"DBLP.org","id":"y0l0bxIwPW","version":2},{"content":{"summary":{"value":"This paper introduces a novel framework called Pre-training and Fine-tuning Video Super-Resolution (PFVSR) for efficiently handling the video super-resolution (SR) task. The authors address the challenges of expensive computational resources, high dimensionality of video data, and scarcity of high-quality video SR datasets. They propose a two-step approach where a pre-trained image SR model is first fine-tuned for spatial adaptation and then adapted for temporal modeling. Through extensive experiments, the authors demonstrate that PFVSR achieves significant improvements in efficiency without compromising output quality compared to existing methods."},"presentation":{"value":"2 fair"},"contribution":{"value":"1 poor"},"soundness":{"value":"3 good"},"strengths":{"value":"- It proposes a two-step approach that adapts pre-trained image SR models for spatial and temporal modeling, addressing the challenges in training video SR models.\n- This paper utilizes pre-trained image SR models and adapts them for video SR tasks. It incorporates adapter modules for spatial and temporal adaptation, enhancing the model's capability to handle video sequences effectively.\n- The paper conducts extensive experiments on public datasets to validate the proposed approach. It compares the performance of PFVSR with existing methods and showcases the significant improvements in efficiency without compromising output quality."},"flag_for_ethics_review":{"value":["Yes, Research integrity issues (e.g., plagiarism, dual submission)"]},"weaknesses":{"value":"- Insights and analysis are rather limited.\n    - This work simply reshapes the video patch embedding and finetunes the adapter without any investigation and analysis of the specific challenges of video SR. As it is believed that temporal propagation is the key element for video SR, the authors did not provide any analysis on why the temporal adaptation works and how well it works.\n    - In Table 3, B1 + Spatial Adaptation (B2) is still an image-based model and generates video on a per-frame basis. Why simply adding an adapter can bring 1.1 dB PSNR gain?\n    - In Table 3, B4 achieves 32.46 dB on REDS4, which is on par with BasicVSR++ and poorer than RSRT and RVRT. Compared to those SOTA methods, It seems the most performance gain stems from changing backbone (HAT), rather than the proposed Spatial-Temporal Adaptation.\n\n- Technical contributions are limited.\n    - The detailed architecture of Figure 1 looks almost the same as Figure 2 in AIM [1].\n    - Section 3.3 is similar to Section 3.3 in [1].\n\n- It is unclear how the authors adopt the trajectory-aware attention into the frozen MSA. The authors should provide more technical details if this is the default setting of the proposed PFVSR.\n- In Section 4.1, the authors mentioned they initialize the SpyNet with pre-trained weights. However, there is no alignment module adopted for the adapter. Why SpyNet is used in implementation?\n\nOverall the idea of repurposing pre-trained image SR models for video SR is interesting. However, the lack of significant analysis and unclear results in Table 3 could be misleading to the community. Given that, this paper is not publication-ready in its current form.\n\n[1] AIM: Adapting Image Models for Efficient Video Action Recognition. ICLR'23"},"confidence":{"value":"5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully."},"questions":{"value":"See weaknesses"},"rating":{"value":"1: strong reject"},"details_of_ethics_concerns":{"value":"This paper exhibits significant overlap with [1], raising concerns about research integrity. \n\n[1] AIM: Adapting Image Models for Efficient Video Action Recognition. ICLR'23"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636181658,"tcdate":1699211171408,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission2457/Reviewer_cshF"],"signatures":["ICLR.cc/2024/Conference/Submission2457/Reviewer_cshF"],"forum":"GDNo5oLpMx","number":4,"license":"CC BY 4.0","cdate":1699211171408,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission2457/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636181658,"domain":"ICLR.cc/2024/Conference","replyto":"GDNo5oLpMx","id":"fIwdd8Qj2M","forumContent":{"venue":{"value":"ICLR 2024 Conference Withdrawn Submission"},"TLDR":{"value":"We propose a novel framework for adapting pre-trained image super-resolution (SR) models to tackle the challenging task of efficient video super-resolution"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Image Super-Resolution; Efficient Video Super-Resolution; Pre-Training and Fine-Tuning"]},"supplementary_material":{"value":"/attachment/b9bd7d5e65449be23e4f2cd68f4930d861b1369c.zip"},"primary_area":{"value":"representation learning for computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"In this paper, we propose a novel framework for adapting pre-trained image super-resolution (SR) models to tackle the challenging task of efficient video SR. This is achieved by freezing the pre-trained image SR model and fine-tuning it with the addition of several lightweight adapter modules. These adapters facilitate spatial and temporal learning, progressively equipping the image SR model with spatiotemporal reasoning capabilities for video SR. Also, these Adapters are compact and extendable, embedding only a few trainable parameters for each video dataset. Moreover, the parameters of the image SR model remain unchanged, resulting in substantial parameter sharing. This allows us to train video SR models quickly and efficiently. Remarkably, despite having significantly fewer parameters, our proposed method achieves competitive or even superior performance compared to existing video SR methods across multiple benchmarks."},"_bibtex":{"value":"@misc{\ntang2024pretraining,\ntitle={Pre-Training and Fine-Tuning Image Super-Resolution Models for Efficient Video Super-Resolution},\nauthor={Hao Tang and Ling Shao and Luc Van Gool},\nyear={2024},\nurl={https://openreview.net/forum?id=GDNo5oLpMx}\n}"},"title":{"value":"Pre-Training and Fine-Tuning Image Super-Resolution Models for Efficient Video Super-Resolution"},"pdf":{"value":"/pdf/bbf22217a6ae61508269df2fdc157cc31eccffbe.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Withdrawn_Submission"},"paperhash":{"value":"tang|pretraining_and_finetuning_image_superresolution_models_for_efficient_video_superresolution"},"authorids":{"value":["~Hao_Tang6","~Ling_Shao1","~Luc_Van_Gool1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Hao Tang","Ling Shao","Luc Van Gool"]}},"version":2},{"content":{"summary":{"value":"This paper introduces Tri-Factor Saliency, a novel framework for efficient, training-free token pruning in LVLMs applied to long-form video. The core problem addressed is the quadratic computational cost of self-attention, which makes processing long videos prohibitively expensive. The authors challenge the assumption that preserving visual diversity during pruning requires expensive computations in the original high-dimensional feature space. Their central hypothesis is that a compact, low-dimensional representation can guide an equally effective pruning process with a fraction of the overhead. To validate this, they propose a method that first projects high-dimensional token features into a 3D \"saliency-space.\" This projection is composed of three interpretable and largely orthogonal factors: Dynamic Saliency, Regional Saliency, and Focal Saliency. This low-dimensional representation is then used to drive a sophisticated pruning pipeline that includes: adaptive video segmentation based on a \"pace signal,\" location-aware clustering to identify spatio-temporally coherent entities, and a diversity-preserving stratified sampling mechanism within each cluster. Experiments across multiple LVLMs and benchmarks show that the method can prune up to 75% of tokens while retaining over 95% of the original model's performance, demonstrating a highly effective balance between efficiency and diversity preservation."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"See the weakness."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"- The framework is thoughtfully designed. The Tri-Factor Saliency model is interpretable and comprehensive. The subsequent pruning pipeline, particularly the location-aware clustering and diversity-preserving stratified sampling, is sophisticated and directly aligned with the stated goals.\n- The experiments are extensive, covering multiple state-of-the-art LVLMs and a wide array of video understanding benchmarks. The results are strong and consistently demonstrate the method's effectiveness."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The paper introduces too many hyperparameters. However, there is no discussion of how these values were chosen or how sensitive the model's performance is to them.\n- The framework is a complex multi-stage pipeline involving three separate saliency calculations, adaptive segmentation, location-aware clustering, stratified sampling, and a final temporal coalescing step. Many design choices and hyperparameters appear to be hand-tuned and based on heuristics. While effective, this intricate, engineered approach lacks simplicity and may be \"brittle,\" potentially struggling with video types or scenarios that fall outside the distribution of the test sets.\n- In an era where end-to-end learning prevails, the entire TFS-Pruner is a fixed, non-learnable module. Although its \"plug-and-play\" nature is a strength, it also means its definitions of saliency and its pruning rules are static. They cannot adapt to different data distributions (e.g., animation vs. real-world footage) or to the specific architectural biases of the LVLM they are paired with. \n- The paper presents the strong performance of the full system but lacks a detailed ablation study to dissect the contribution of each component. This makes it difficult to ascertain the precise impact of each part."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762943221038,"tcdate":1760599326442,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission24848/Reviewer_nDbx"],"signatures":["ICLR.cc/2026/Conference/Submission24848/Reviewer_nDbx"],"forum":"pAgiqavopA","number":1,"license":"CC BY 4.0","cdate":1760599326442,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission24848/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762943221038,"domain":"ICLR.cc/2026/Conference","replyto":"pAgiqavopA","id":"8aHYvIBPM7","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["video token compression","video understanding"]},"primary_area":{"value":"optimization"},"abstract":{"value":"The quadratic computational overhead of self-attention severely limits the application of Large Vision-Language Models (LVLMs) to long-form video. While training-free token pruning offers a promising avenue for acceleration, current methods still struggle for balancing the token diversity and pruning efficiency. Query-based approaches prune tokens irrelevant to a specific prompt, but consequently sacrifice the intrinsic diversity of the video content. Conversely, methods that preserve diversity by clustering or matching based on the raw, high-dimensional token features incur prohibitive computational costs, making them impractical for long video inputs.\nIn this work, we challenge the assumption that preserving diversity necessitates expensive computations in the original high-dimensional feature space. We hypothesize that a low-dimensional yet informative representation engineered for pruning can achieve comparable results with a fraction of the overhead. To validate this, we propose a framework that first projects the original token features into a highly informative 3D \"saliency-space.\" This projection is achieved via our Tri-Factor Saliency (TFS) model, which computes three largely orthogonal sub-features from a local spatio-temporal neighborhood: (1) Dynamic Saliency, which captures the magnitude of movement; (2) Regional Saliency, which identifies coherent objects that stand out from their background; and (3) Focal Saliency, which pinpoints unpredictable, fine-grained details.\nThis low-dimensional representation enables subsequent entity-aware clustering and diversity-preserving stratified sampling to be performed with minimal computational cost. Our experiments show that this approach allows for the pruning of up to 75% of tokens while retaining 95% of the original model's performance on video understanding benchmarks. Our work demonstrates that a well-designed, low-dimensional perceptual projection can effectively replace expensive high-dimensional feature matching for video token pruning, charting a new course that achieves both high efficiency and strong diversity preservation."},"_bibtex":{"value":"@misc{\nhuang2025trifactor,\ntitle={Tri-Factor Saliency: A Low-Dimensional Representation for Efficient and Diversity-Aware Video Token Pruning},\nauthor={Zhuangqiu Huang and Minxin Lai and Shuo Liu and Yu Zhang and Jiaqi Wang},\nyear={2025},\nurl={https://openreview.net/forum?id=pAgiqavopA}\n}"},"title":{"value":"Tri-Factor Saliency: A Low-Dimensional Representation for Efficient and Diversity-Aware Video Token Pruning"},"pdf":{"value":"/pdf/13c0e5f612e17f492890cfafbbdbecdaeea68f41.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"huang|trifactor_saliency_a_lowdimensional_representation_for_efficient_and_diversityaware_video_token_pruning"},"authorids":{"value":["~Zhuangqiu_Huang1","~Minxin_Lai1","~Shuo_Liu14","~Yu_Zhang51","~Jiaqi_Wang1"]},"authors":{"value":["Zhuangqiu Huang","Minxin Lai","Shuo Liu","Yu Zhang","Jiaqi Wang"]}},"version":2},{"content":{"TLDR":{"value":"We identify calendar shortcuts in text-assisted forecasting and propose SPRA, a shortcut controlled, state conditioned low-rank residual adapter that improves five forecasting backbones across nine domains."},"venue":{"value":"ICLR 2027 Conference Submission"},"keywords":{"value":["time series forecasting","text-assisted forecasting","shortcut learning"]},"supplementary_material":{"value":"/attachment/5d49c94ce7f8acc2351b24d54cdcaf1834f64a22.zip"},"primary_area":{"value":"learning on time series and dynamical systems"},"abstract":{"value":"Text-assisted forecasting pairs numerical windows with contemporaneous documents to capture event semantics unavailable from the series alone. Evidence for this premise rests largely on modality ablation (adding text helps, removing it hurts), which establishes predictive signal but does not identify its source. We design controlled interventions on trained forecasters to localize that signal: reversing event semantics alters MAE by less than 1\\%, whereas masking only dates and timestamp expressions reproduces at least 74\\% of the full text removal penalty in five of the six settings. We term this mode the calendar shortcut: forecasters recover periodic phase from textual calendar cues while largely bypassing event content. To overcome it, we propose State Conditioned Paired Residual Adaptation (SPRA), a lightweight, backbone agnostic framework for state-conditioned residual learning. Shortcut controlled document pairs preserve event content while suppressing calendar cues, and a paired consistency objective reduces the adapter's reliance on such cues. A low-rank adapter then estimates what the text contributes beyond the observed trajectory and maps it to a target specific multi-horizon residual correction of a frozen backbone. Across nine domains and five backbones, SPRA ranks first in 79 of 90 comparisons, improves upon its unimodal backbone in 44 of 45 settings, and retains its gains under the same shortcut interventions, supporting attribution to forecast-relevant event content rather than incidental calendar identity."},"_bibtex":{"value":"@inproceedings{\nanonymous2026reading,\ntitle={{READING} {THE} {DATE}, {NOT} {THE} {EVENT}: {THE} {CALENDAR} {SHORTCUT} {IN} {TEXT}-{ASSISTED} {TIME} {SERIES} {FORECASTING}},\nauthor={Anonymous},\nbooktitle={Submitted to The Fifteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=XKFGWcNoFF},\nnote={under review}\n}"},"title":{"value":"READING THE DATE, NOT THE EVENT: THE CALENDAR SHORTCUT IN TEXT-ASSISTED TIME SERIES FORECASTING"},"pdf":{"value":"/pdf/f4e8cbf22e6363d8d0969355c5eec6eb835760a6.pdf"},"venueid":{"value":"ICLR.cc/2027/Conference/Submission"}},"tmdate":1791222848738,"tcdate":1788499296195,"writers":["ICLR.cc/2027/Conference","ICLR.cc/2027/Conference/Submission8973/Authors"],"signatures":["ICLR.cc/2027/Conference/Submission8973/Authors"],"forum":"XKFGWcNoFF","license":"CC BY 4.0","number":8973,"cdate":1788499296195,"readers":["everyone"],"invitations":["ICLR.cc/2027/Conference/-/Submission","ICLR.cc/2027/Conference/-/Edit","ICLR.cc/2027/Conference/-/Submission_Change_Before_Bidding","ICLR.cc/2027/Conference/Submission8973/-/Full_Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Reviewing"],"mdate":1791222848738,"odate":1791031328286,"domain":"ICLR.cc/2027/Conference","id":"XKFGWcNoFF","version":2},{"content":{"summary":{"value":"This work proposes ViPO, visual preference optimization at scale. It includes two contributions: (a) Poly-DPO which proposes including a polynomial term alongside the standard diffusion loss to dynamically weight preferences in the face of uncertainty in the preference datasets, (b) ViPO dataset, a massive-scale preference dataset with 1M image (1024resolution) pairs across five categories  and 300K video pairs (720P+ resolution) across three categories. Numerical experiments have been done to evaluate the proposed method on preference optimization compared to the baselines."},"soundness":{"value":2},"confidence":{"value":5},"questions":{"value":"- New insights regarding existing datasets is not that new. i.e., These datasets also suffer from low resolution (512-768resolution), limited prompt diversity, outdated generation models. Many existing works have demonstrated the same thing. For instance, see RankDPO.\n- Why not evaluate SDXL and SD1.5 on DPG-Bench benchmark like in Tab. 5? Why only show the GenEval scores?\n- What’s the recommended value of alpha in poly-loss? How does one tune this hyper-parameter for a new preference dataset?\n- For the proposed new dataset, since the recommended value of alpha=0 ==> its standard diffusion DPO. Does it mean that Poly-Loss does not help at all? Because the illustrated examples in which alpha!=0, are very old datasets and would no longer be used due to their outdated generation models and low-res images.\n- In Tab.5, why does SFT not improve SD3.5-medium and Flux.1 Dev performance? Technically, the ViPO-1M image dataset is better in quality. \n- How robust is the proposed Poly-DPO is to model collapse which is observed very frequently with DPO formulation with longer training?\n\nMissing References:\n- DSPO — Direct Score Preference Optimization for Diffusion Model Alignment https://openreview.net/forum?id=xyfb9HHvMe \n- RankDPO — Scalable Ranked Preference Optimization for Text-to-Image Generation:  https://arxiv.org/abs/2410.18013  \n- Flow-GRPO— Training Flow Matching Models via Online RL : https://arxiv.org/pdf/2505.05470 \n- Bridging SFT and DPO for Diffusion Model Alignment with Self-Sampling Preference Optimization https://arxiv.org/abs/2410.05255"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- Large scale preference optimization dataset for image and video generation \n- Evaluation on a range of image and video generation models to show the efficacy of the proposed method"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- It is unclear what’s the impact of Polynomial component in Poly-DPO on a decent DPO formulation. Most of the datasets referred in the paper are quite old and would not be used by any modern SOTA T2I models for preference optimization. Besides the alpha becomes yet another hyper-parameter in the DPO formulation which is already quite prone to collapse\n- One concerning issue with the VIPO dataset is that SFT with this dataset does not seem to improve models like SD3.5-medium and FLUX.1 dev on DPG bench which is harder benchmark compared to GenEval. This could point to biases in the dataset construction. \n- Claims the new insights regarding existing datasets (low resolution images, outdated models, conflicting win/lose pairs) as a contribution, but this has been observed earlier in the literature."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762917229388,"tcdate":1761786806755,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission4207/Reviewer_Cn5U"],"signatures":["ICLR.cc/2026/Conference/Submission4207/Reviewer_Cn5U"],"forum":"x5zP3k64Nl","number":2,"license":"CC BY 4.0","cdate":1761786806755,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission4207/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762917229388,"domain":"ICLR.cc/2026/Conference","replyto":"x5zP3k64Nl","id":"4kVcBqYFIW","forumContent":{"TLDR":{"value":"Scaling preference optimization in visual generation through our proposed large-scale datasets and algorithmic improvements"},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Diffusion Model","Image Generation","Video Generation","Visual Generation","DPO"]},"primary_area":{"value":"generative models"},"abstract":{"value":"While preference optimization is crucial for improving visual generative models, how to effectively scale this paradigm for visual generation remains largely unexplored.\nCurrent open-source preference datasets typically contain substantial conflicting preference patterns, where winners excel in some dimensions but underperform in others. Naively optimizing on such noisy datasets fails to learn meaningful preferences, fundamentally hindering effective scaling. To enhance the robustness of preference algorithms against noise, we propose Poly-DPO, which extends the DPO objective with an additional polynomial term that dynamically adjusts model confidence during training based on dataset characteristics, enabling effective learning across diverse data distributions from noisy to trivially simple patterns.\nBeyond biased patterns, existing datasets suffer from low resolution, limited prompt diversity, and imbalanced distributions. To facilitate large-scale visual preference optimization by tackling key data bottlenecks, we construct ViPO, a massive-scale preference dataset with 1M image pairs (1024px) across five categories and 300K video pairs (720p+) across three categories. Leveraging state-of-the-art generative models and diverse prompts ensures consistent, reliable preference signals with balanced distributions.\nRemarkably, when applying Poly-DPO to our high-quality dataset, the optimal configuration converges to standard DPO. This convergence validates both our dataset quality and Poly-DPO's adaptive nature: sophisticated optimization becomes unnecessary with sufficient data quality, yet remains valuable for imperfect datasets.\nWe comprehensively validate our approach across various visual generation models. On noisy datasets like Pick-a-Pic V2, Poly-DPO achieves 6.87 and 2.32 gains over Diffusion-DPO on GenEval for SD1.5 and SDXL, respectively. For our high-quality ViPO dataset, models achieve performance far exceeding those trained on existing open-source preference datasets. These results confirm that addressing both algorithmic adaptability and data quality is essential for scaling visual preference optimization. Code, models and open-source datasets will be released at: https://github.com/liming-ai/ViPO"},"_bibtex":{"value":"@inproceedings{\nli2026vipo,\ntitle={Vi{PO}: Visual Preference Optimization at Scale},\nauthor={Ming Li and Jie Wu and Justin Cui and Xiaojie Li and Rui Wang and Chen Chen},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=x5zP3k64Nl}\n}"},"title":{"value":"ViPO: Visual Preference Optimization at Scale"},"pdf":{"value":"/pdf/cc33a36e8a16848dd99f08763eb9d28783c0f5cc.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"li|vipo_visual_preference_optimization_at_scale"},"authorids":{"value":["~Ming_Li19","~Jie_Wu8","~Justin_Cui1","~Xiaojie_Li14","~Rui_Wang32","~Chen_Chen18"]},"authors":{"value":["Ming Li","Jie Wu","Justin Cui","Xiaojie Li","Rui Wang","Chen Chen"]}},"version":2},{"content":{"summary":{"value":"This paper introduces Model-aware Counterfactual Data based Contrastive Decoding, a training-free inference method designed to mitigate hallucination in Video-LLMs. The core idea is to improve Contrastive Decoding by replacing generic random perturbations with model-aware counterfactual data. This data is constructed by leveraging the Video-LLM's internal feedback, specifically, by performing gradient ascent on a soft mask applied to the video. The optimization goal is to maximize the query reconstruction loss, thus generating an adversarial perturbation that selectively erases the critical visual cues necessary for the correct prediction. Experiments on diverse benchmarks and various Video-LLMs demonstrate the effectiveness of the proposed method in consistently reducing hallucination while maintaining or improving accuracy."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. The optimization maximizes the loss to construct the counterfactual. Is it possible for a strong perturbation (high loss) to destroy too much information, leading to an overly weak contrastive view that is unhelpful for decoding? Did the authors experiment with alternative loss functions or constraints that target minimal information removal to achieve a threshold loss (making the perturbation just strong enough)?\n\n2. The paper focuses on comparing GeWu to VCD and SID. Could GeWu's model-aware counterfactual data generation module be integrated with other advanced decoding strategies (e.g., self-consistency methods) to achieve further performance gains? A brief discussion on this potential synergy would be valuable."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The primary motivation for generating counterfactual data is clear. By utilizing a gradient-ascent approach to learn adversarial perturbations, the method can directionally erase the critical information required by the Video-LLM. This results in a superior contrastive view compared to those generated via simple random noise, directly addressing a key limitation of prior CD methods.\n\n2. The overall framework, including counterfactual data augmentation, generation, and contrastive decoding, is thoughtfully adapted from image-based methods to the temporal and spatial complexities of video and Video-LLMs. The integration of object detection/tracking, soft mask generation, and joint object-level/frame-level masking effectively controls the visual cues across the spatio-temporal domain.\n\n3. The method is evaluated across multiple modern Video-LLMs and three distinct, challenging benchmarks, providing strong evidence of its effectiveness in mitigating hallucination."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Major Concerns\n\n1. The method, while \"training-free,\" involves a gradient optimization step (gradient ascent) for every counterfactual sample during inference. This process, which requires backpropagation through the visual encoder and potentially the full LLM, introduces a significant computational overhead compared to standard decoding or simple random perturbations. It would be better to provide an explicitly quantification of this latency cost and a stronger justification for why this overhead is acceptable for real-time video processing. The \"training-free\" claim must be balanced against the increased inference time.\n\n2. The optimization seeks to find an optimal mask $r^*$ that maximizes the loss for a specific query $q$ and video $V$. It is a local optimization problem. Authors need to discuss or investigate the robustness of this optimization: Are the resultant counterfactuals truly the best at isolating the hallucination source, or could the optimization converge to spurious local maxima?"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762942726375,"tcdate":1761989594393,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission23595/Reviewer_Jp5V"],"signatures":["ICLR.cc/2026/Conference/Submission23595/Reviewer_Jp5V"],"forum":"ioGQhr1lhZ","number":4,"license":"CC BY 4.0","cdate":1761989594393,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission23595/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762942726375,"domain":"ICLR.cc/2026/Conference","replyto":"ioGQhr1lhZ","id":"zfGDiH5UQK","forumContent":{"venue":{"value":"ICLR 2026 Conference Desk Rejected Submission"},"keywords":{"value":["Video-language models","Hallucination mitigation","Contrastive decoding","Object-level data augmentation","Counterfactual inputs"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Video language models (Video-LLMs) are prone to hallucinations, often generating plausible but ungrounded content when visual evidence is weak, ambiguous, or biased. Existing decoding methods, such as contrastive decoding (CD), rely on random perturbations to construct contrastive data for mitigating hallucination patterns. However, such a way is hard to control the visual cues that drive hallucination or well align with model weaknesses. \nWe propose Model-aware Counterfactual Data based Contrastive Decoding (GeWu), a new inference strategy that combines model-guided counterfactual construction with decoding. Our approach uses the Video-LLM’s own feedback to identify object regions most responsible for hallucination, generating targeted counterfactual inputs at the object level rather than arbitrary frame or temporal modifications. These model-aware counterfactual data is then integrated into CD to enforce evidence-grounded token selection during decoding. Experiments on EventHallusion, MVBench, and Perception-test show that GeWu consistently reduces hallucination while maintaining or improving task accuracy across diverse Video-LLMs, including InternVL3, Qwen2.5-VL and Qwen2-VL families. The method is especially effective in challenging scenarios involving small, occluded, or co-occurring objects. Our code and data will be publicly released."},"_bibtex":{"value":"@misc{\nanonymous2026modelaware,\ntitle={Model-aware Counterfactual Data based Contrastive Decoding for Video-{LLM}},\nauthor={Anonymous},\nyear={2026},\nurl={https://openreview.net/forum?id=ioGQhr1lhZ}\n}"},"title":{"value":"Model-aware Counterfactual Data based Contrastive Decoding for Video-LLM"},"pdf":{"value":"/pdf/0b0047cec22626f1aa96ce3bb46d12bde4818579.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Desk_Rejected_Submission"},"paperhash":{"value":"xiao|modelaware_counterfactual_data_based_contrastive_decoding_for_videollm"},"authorids":{"value":["~Qixin_Xiao1","~Kun_Zhou2"]},"authors":{"value":["Qixin Xiao","Kun Zhou"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a training-free method to leverage pre-trained image restoration diffusion models for zero-shot video restoration. Specifically, the proposed method employs hierarchical latent warping and an enhanced token merging strategy to maintain temporal consistency and restore details across video frames. Experimental results demonstrate the versatility of the method across various tasks, including denoising, super-resolution, and depth estimation."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"1. Is there a comparison of inference efficiency with baselines?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The proposed method is pioneering in achieving zero-shot video restoration by leveraging pre-trained image restoration diffusion models, enabling multiple video restoration tasks without additional training.\n2. The method presents strong quantitative and visual results compared with state-of-the-art methods, balancing temporal consistency and detail generation."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. In line 267, the authors claim that the combination of hierarchical latent warping and hybrid flow-guided spatial-aware token merging could achieve adaptation to various degradation types. However, it is not sufficiently discussed that why the combination could handle various degradation types.\n2. While the use of optical flow guidance aligns with previous video restoration works [1,2], the paper introduces similarity-based guidance to capture correspondences distinct from optical flow. However, the specific benefits of similarity-based guidance over optical flow are not thoroughly discussed.\n3. Although recent video restoration methods, such as Shift-Net and FMA-Net, are included as baselines, some classic methods like BasicVSR++ and RVRT are not compared in the experiments.\n\n[1] Chan K C K, Zhou S, Xu X, et al. Basicvsr++: Improving video super-resolution with enhanced propagation and alignment[C]//Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2022: 5972-5981.\n\n[2] Liang J, Fan Y, Xiang X, et al. Recurrent video restoration transformer with guided deformable attention[J]. Advances in Neural Information Processing Systems, 2022, 35: 378-393."}},"nonreaders":[],"tmdate":1731427336371,"tcdate":1730339258279,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission893/Reviewer_ttiq"],"signatures":["ICLR.cc/2025/Conference/Submission893/Reviewer_ttiq"],"forum":"qpDqO7qa3R","number":1,"license":"CC BY 4.0","cdate":1730339258279,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission893/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427336371,"domain":"ICLR.cc/2025/Conference","replyto":"qpDqO7qa3R","id":"rDdX5oCT4l","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video restoration","Zero-shot","Training-free","Diffusion models"]},"supplementary_material":{"value":"/attachment/ece9b098ba89dfe1e82311bd70fc3f703077883c.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"This paper introduces a method for zero-shot video restoration using pre-trained image restoration diffusion models. Traditional video restoration methods often need retraining for different settings and struggle with limited generalization across various degradation types and datasets. Our approach uses a hierarchical latent warping strategy for keyframes and local frames, combined with token merging that uses a hybrid correspondence mechanism that integrates spatial information, optical flow, and feature-based matching. We show that our method not only achieves top performance in zero-shot video restoration but also significantly surpasses trained models in generalization across diverse datasets and extreme degradations (8$\\times$ super-resolution and high-standard deviation video denoising). We present evidence through quantitative metrics and visual comparisons on various challenging datasets. Additionally, our technique works with any 2D restoration diffusion model, offering a versatile and powerful tool for video enhancement tasks without extensive retraining."},"_bibtex":{"value":"@misc{\nyeh2025diffirvrzero,\ntitle={Diff{IR}2{VR}-Zero: Zero-Shot Video Restoration with Diffusion-based Image Restoration Models},\nauthor={Changhan Yeh and Chin-Yang Lin and Zhixiang Wang and Chi-Wei Hsiao and Ting-Hsuan Chen and Hau-Shiang Shiu and Yu-Lun Liu},\nyear={2025},\nurl={https://openreview.net/forum?id=qpDqO7qa3R}\n}"},"title":{"value":"DiffIR2VR-Zero: Zero-Shot Video Restoration with Diffusion-based Image Restoration Models"},"pdf":{"value":"/pdf/c2dae8a064bbe66eaf052c701fcf55e6452ab4dc.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"yeh|diffir2vrzero_zeroshot_video_restoration_with_diffusionbased_image_restoration_models"},"authorids":{"value":["~Changhan_Yeh1","~Chin-Yang_Lin1","~Zhixiang_Wang1","~Chi-Wei_Hsiao2","~Ting-Hsuan_Chen1","~Hau-Shiang_Shiu1","~Yu-Lun_Liu2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Changhan Yeh","Chin-Yang Lin","Zhixiang Wang","Chi-Wei Hsiao","Ting-Hsuan Chen","Hau-Shiang Shiu","Yu-Lun Liu"]}},"version":2},{"content":{"venue":{"value":"EMNLP 2024"},"pdf":{"value":"https://aclanthology.org/2024.emnlp-main.679.pdf"},"venueid":{"value":"dblp.org/conf/EMNLP/2024"},"paperhash":{"value":"yuan|do_llms_overcome_shortcut_learning_an_evaluation_of_shortcut_challenges_in_large_language_models"},"authorids":{"value":["~Yu_Yuan5","~Lili_Zhao3","~Kai_Zhang12","~Guangting_Zheng2","~Qi_Liu3"]},"html":{"value":"https://doi.org/10.18653/v1/2024.emnlp-main.679"},"_bibtex":{"value":"@inproceedings{DBLP:conf/emnlp/YuanZZZL24,\n  author={Yu Yuan and Lili Zhao and Kai Zhang and Guangting Zheng and Qi Liu},\n  title={Do LLMs Overcome Shortcut Learning? An Evaluation of Shortcut Challenges in Large Language Models},\n  year={2024},\n  cdate={1704067200000},\n  pages={12188-12200},\n  url={https://doi.org/10.18653/v1/2024.emnlp-main.679},\n  booktitle={EMNLP},\n  crossref={conf/emnlp/2024}\n}\n"},"abstract":{"value":"Large Language Models (LLMs) have shown remarkable capabilities in various natural language processing tasks. However, LLMs may rely on dataset biases as shortcuts for prediction, which can significantly impair their robustness and generalization capabilities. This paper presents Shortcut Suite, a comprehensive test suite designed to evaluate the impact of shortcuts on LLMs’ performance, incorporating six shortcut types, five evaluation metrics, and four prompting strategies. Our extensive experiments yield several key findings: 1) LLMs demonstrate varying reliance on shortcuts for downstream tasks, which significantly impairs their performance. 2) Larger LLMs are more likely to utilize shortcuts under zero-shot and few-shot in-context learning prompts. 3) Chain-of-thought prompting notably reduces shortcut reliance and outperforms other prompting strategies, while few-shot prompts generally underperform compared to zero-shot prompts. 4) LLMs often exhibit overconfidence in their predictions, especially when dealing with datasets that contain shortcuts. 5) LLMs generally have a lower explanation quality in shortcut-laden datasets, with errors falling into three types: distraction, disguised comprehension, and logical fallacy. Our findings offer new insights for evaluating robustness and generalization in LLMs and suggest potential directions for mitigating the reliance on shortcuts."},"title":{"value":"Do LLMs Overcome Shortcut Learning? An Evaluation of Shortcut Challenges in Large Language Models"},"authors":{"value":["Yu Yuan","Lili Zhao","Kai Zhang","Guangting Zheng","Qi Liu"]}},"tmdate":1785325748638,"pdate":1704067200000,"externalIds":["dblp:conf/emnlp/YuanZZZL24"],"tcdate":1768439537145,"writers":["~"],"signatures":["~Kai_Zhang12"],"forum":"JKzRjqPp28","license":"CC BY-SA 4.0","number":753277,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1785325748638,"domain":"DBLP.org","id":"JKzRjqPp28","version":2},{"content":{"summary":{"value":"This paper proposes Data-augmented Phrase-level Alignment (DPA), a novel loss function designed to reduce object hallucinations in multimodal large language models (MLLMs) while preserving their general vision-language capabilities. By generating pairs of hallucinated and correct responses through data augmentation, DPA trains MLLMs to distinguish hallucinated phrases from correct ones. Experimental results show that MLLMs fine-tuned with DPA achieve significant improvements, reducing hallucination rates and enhancing performance on visual question-answering and image description tasks."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Please see the weakness part."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"* The DPA loss function is an innovative approach that mitigates object hallucinations by targeting phrase-level distinctions, offering a focused solution for multimodal hallucination issues.\n\n* The data augmentation method is straightforward yet effective, generating hallucinated-correct response pairs that enable the model to learn nuanced differences with minimal complexity.\n\n* DPA demonstrates significant performance gains"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The core idea of the paper is to generate correct-hallucinated data pairs through data augmentation. However, I have three questions about this process.\n1. In hallucination-related datasets, object hallucination does not frequently occur, raising questions about the validity of replacing every possible object and attribute with hallucinations. Since models seldom generate such hallucinations, this augmentation strategy might introduce excessive \"non-realistic\" hallucination cases, leading to a mismatch between training and real-world distributions, potentially impacting the model's generalization.\n2. The method’s effectiveness may be limited by the diversity of data augmentation. Since hallucinated data generation relies on a finite set of replacements, it may not fully cover the types of hallucinations that could appear in practical applications, limiting the model’s ability to handle unseen hallucinations.\n3. The data augmentation strategy itself lacks independent experimental evaluation. The experiments mainly focus on improvements in model performance across different benchmarks, without assessing the augmentation strategy’s generalization effect and impact on model training stability across tasks."}},"nonreaders":[],"tmdate":1731427280042,"tcdate":1730554036439,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission568/Reviewer_Qf9P"],"signatures":["ICLR.cc/2025/Conference/Submission568/Reviewer_Qf9P"],"forum":"yG1fW8igzP","number":4,"license":"CC BY 4.0","cdate":1730554036439,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission568/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427280042,"domain":"ICLR.cc/2025/Conference","replyto":"yG1fW8igzP","id":"98jGzQopnw","forumContent":{"TLDR":{"value":"We introduce phrase-level alignment method that can be applied to off-the-shelf MLLMs for mitigating hallucinations, while preserving their general vision-language capabilities."},"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Multimodal LLMs","Object Hallucination","Vision-language Models"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Despite their significant advancements, Multimodal Large Language Models\n(MLLMs) often generate factually inaccurate information, referred to as hallucination.\nIn this work, we address object hallucinations in MLLMs, where information\nis generated about an object not present in the input image. We introduce Data-augmented\nPhrase-level Alignment (DPA), a novel loss which can be applied to\ninstruction-tuned off-the-shelf MLLMs to mitigate hallucinations, while preserving\ntheir general vision-language capabilities. To fine-tune MLLMs with DPA, we first\ngenerate a set of 'hallucinated' and 'correct' response pairs through generative data\naugmentation by selectively altering the ground-truth information of the correct\nresponses at a phrase level. The DPA loss is then used to train MLLMs to reduce\nthe likelihood of hallucinated phrases compared to the correct ones. Our thorough\nevaluation on various benchmarks confirms the effectiveness of DPA in mitigating\nhallucination while retaining the out-of-the-box performance of the MLLMs on\ngeneral tasks. For instance, MLLMs finetuned with DPA, which we refer to as Hallucination\nAttenuated Language and Vision Assistant (HALVA), improve F1 by up\nto 13.4% on hallucination visual question-answering and reduce the hallucination\nrate by up to 4.2% on image description tasks."},"_bibtex":{"value":"@inproceedings{\nsarkar2025mitigating,\ntitle={Mitigating Object Hallucination in {MLLM}s via Data-augmented Phrase-level Alignment},\nauthor={Pritam Sarkar and Sayna Ebrahimi and Ali Etemad and Ahmad Beirami and Sercan O Arik and Tomas Pfister},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=yG1fW8igzP}\n}"},"title":{"value":"Mitigating Object Hallucination in MLLMs via Data-augmented Phrase-level Alignment"},"pdf":{"value":"/pdf/16ddb11aaac076fdbb13977aeb28540015cc32db.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"sarkar|mitigating_object_hallucination_in_mllms_via_dataaugmented_phraselevel_alignment"},"authorids":{"value":["~Pritam_Sarkar1","~Sayna_Ebrahimi1","~Ali_Etemad1","~Ahmad_Beirami1","~Sercan_O_Arik1","~Tomas_Pfister1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Pritam Sarkar","Sayna Ebrahimi","Ali Etemad","Ahmad Beirami","Sercan O Arik","Tomas Pfister"]}},"version":2},{"content":{"venue":{"value":"IEEE Trans. Circuits Syst. Video Technol. 2024"},"pdf":{"value":"https://ieeexplore.ieee.org/iel7/76/10492617/10234439.pdf"},"venueid":{"value":"dblp.org/journals/TCSV/2024"},"paperhash":{"value":"liu|semanticaware_contrastive_learning_with_proposal_suppression_for_video_semantic_role_grounding"},"authorids":{"value":["","https://dblp.org/search/pid/api?q=author:Di_Zhou:","https://dblp.org/search/pid/api?q=author:Jie_Guo:","https://dblp.org/search/pid/api?q=author:Xin_Luo_0006:","~Zan_Gao1",""]},"html":{"value":"https://doi.org/10.1109/TCSVT.2023.3310296"},"_bibtex":{"value":"@article{DBLP:journals/tcsv/LiuZGLGN24,\n  author={Meng Liu and Di Zhou and Jie Guo and Xin Luo and Zan Gao and Liqiang Nie},\n  title={Semantic-Aware Contrastive Learning With Proposal Suppression for Video Semantic Role Grounding},\n  year={2024},\n  month={April},\n  cdate={1711929600000},\n  journal={IEEE Trans. Circuits Syst. Video Technol.},\n  volume={34},\n  number={4},\n  pages={3003-3016},\n  url={https://doi.org/10.1109/TCSVT.2023.3310296}\n}\n"},"abstract":{"value":"Video semantic role grounding has gained substantial interest from both the academic and industrial communities. While existing methods have demonstrated considerable performance improvements, the influence of noisy and intra-object proposals, referring to proposals with the same object label, has yet to be explored in video semantic role grounding. In this study, we propose a semantic-aware contrastive learning network with proposal suppression to enhance the accuracy of grounding referenced objects. To fully exploit the semantic information in each semantic role, we introduce a novel semantic role encoding module that allows for precise representations of each semantic role. We also design a semantic-aware proposal suppression network to reduce the impact of noisy proposals on object representation learning. Additionally, we propose a proposal contrastive loss to improve cross-modal alignment and reduce the effect of irrelevant intra-object proposals. Extensive experiments on four datasets demonstrate that our model achieves significant improvements over state-of-the-art methods."},"title":{"value":"Semantic-Aware Contrastive Learning With Proposal Suppression for Video Semantic Role Grounding"},"authors":{"value":["Meng Liu","Di Zhou","Jie Guo","Xin Luo","Zan Gao","Liqiang Nie"]}},"tmdate":1774413088972,"pdate":1704067200000,"tcdate":1727651927907,"writers":["~"],"signatures":["~Meng_Liu4"],"forum":"ONIz9LywHX","license":"CC BY-SA 4.0","number":107484,"cdate":1711929600000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1774413088972,"domain":"DBLP.org","id":"ONIz9LywHX","version":2},{"content":{"summary":{"value":"This work proposes a benchmark called StreamingBench to evaluate video LLM capabilities in streaming settings. StreamingBench introduces several tasks tailored to streaming scenarios, including real-time visual understanding, omni-source understanding, and contextual understanding."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"As mentioned in the weakness"},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- This work introduces a new benchmark designed to evaluate video models in streaming scenarios.\n- It conducts insightful experiments, such as \"Does Redundant Information Affect Contextual Understanding?\", which provide valuable perspectives in this area."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Although this benchmark focuses on streaming scenarios, a standard video LLM can handle it effectively with simple preprocessing. For instance, whenever a question arises, the model can process all frames up to that timestamp. With this approach, the benchmark may not differ significantly from traditional video benchmarks. Therefore, it is essential for this benchmark to identify scenarios that cannot be simplified to an offline setting.\n\n- While handling redundant information is indeed critical for video LLMs, this challenge is not exclusive to streaming scenarios; it is a general issue for any long-video task. As a result, the insights from this paper may be overshadowed by findings from benchmarks specifically focused on long-video understanding.\n\n- The annotation process lacks clarity. Specifically, how do human annotators manually label QA pairs for omni-source understanding and other contextual understanding tasks? What measures are in place to ensure the quality of each question, and what specific strategies were employed?"}},"nonreaders":[],"tmdate":1731427999979,"tcdate":1730798581333,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission14232/Reviewer_7TRD"],"signatures":["ICLR.cc/2025/Conference/Submission14232/Reviewer_7TRD"],"forum":"qnAZqlMGTB","number":4,"license":"CC BY 4.0","cdate":1730798581333,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission14232/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427999979,"domain":"ICLR.cc/2025/Conference","replyto":"qnAZqlMGTB","id":"z3AD6FocQ6","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Benchmark","Streaming Video Understanding","Multimodal Large Language Models","Video Benchmark","Evaluation"]},"primary_area":{"value":"datasets and benchmarks"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"The rapid development of Multimodal Large Language Models (MLLMs) has expanded their capabilities from image comprehension to video understanding. However, most of these MLLMs focus primarily on ofﬂine video comprehension, necessitating extensive processing of all video frames before any queries can be made. This presents a signiﬁcant gap compared to the human ability to watch, listen, think, and respond to streaming inputs in real time, highlighting the limitations of current MLLMs. In this paper, we introduce StreamingBench, the ﬁrst comprehensive benchmark designed to evaluate the streaming video understanding capabilities of MLLMs. StreamingBench assesses three core aspects of streaming video understanding: (1) real-time visual understanding, (2) omni-source understanding and (3) contextual understanding. The benchmark consists of 18 tasks, featuring 900 videos and 4,500 human-curated QA pairs. Each video features ﬁve questions presented at different time points to simulate a continuous streaming scenario. We conduct experiments on StreamingBench with 15 open-source and proprietary MLLMs and ﬁnd that even the most advanced proprietary MLLMs like Gemini 1.5 Pro and GPT-4o perform signiﬁcantly below human-level streaming video understanding capabilities. We hope our work can facilitate further advancements for MLLMs, empowering them to approach human-level video comprehension and interaction in more realistic scenarios."},"_bibtex":{"value":"@misc{\nlin2025streamingbench,\ntitle={StreamingBench: Assessing the Gap for {MLLM}s to Achieve Streaming Video Understanding},\nauthor={Junming Lin and Zheng Fang and Zihao Wan and Fuwen Luo and Chi Chen and Peng Li and Yang Liu and Maosong Sun},\nyear={2025},\nurl={https://openreview.net/forum?id=qnAZqlMGTB}\n}"},"title":{"value":"StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding"},"pdf":{"value":"/pdf/510ba84e9bb94cbbd1bcb5060d0bf89b4703ddb7.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"lin|streamingbench_assessing_the_gap_for_mllms_to_achieve_streaming_video_understanding"},"authorids":{"value":["~Junming_Lin1","~Zheng_Fang14","~Zihao_Wan2","~Fuwen_Luo1","~Chi_Chen1","~Peng_Li2","~Yang_Liu19","~Maosong_Sun1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Junming Lin","Zheng Fang","Zihao Wan","Fuwen Luo","Chi Chen","Peng Li","Yang Liu","Maosong Sun"]}},"version":2},{"content":{"summary":{"value":"The paper introduces a shortcut design (self-consistency objective) in flow models for efficient and high-quality sampling. Unlike complex methods, this shortcut leverages existing networks and requires a single training phase.\nThe shortcut design involves incorporating an additional step size condition into the generative models, enabling them to learn the objective while considering the step size. This approach showcases improved performance in image and policy generation domains."},"soundness":{"value":4},"confidence":{"value":4},"questions":{"value":"* In Equation 5, Flow-Matching and Self-Consistency are intended to correspond to different units of time. Flow-Matching is represented by d=0, whereas Self-Consistency is associated with d > 0. Does it affect the training data samples or procedures?  \n\n* In Algorithm1, why stopgrad is applied in self-consistency target? Will it disable the gradient update?  \n\n* Table1 suggests two-phase training can give better results than all end-to-end methods in one step sampling: progressive distillation in 1-step:14.8 and 35.6 for Celeb and ImageNet. Does the two-phase method has higher performance ceiling than end-to-end? \n\n* In Figure 7, does increasing the sampling steps lead to a higher success rate for Short-cut Policy, similar to the diffusion policy? In policy planning, the primary focus is on achieving multi-mode and high precision performance."},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"The paper introduces a shortcut design (self-consistency objective) in flow models to maintain consistency in intermediate sampling trajectories. This shortcut involves adding a self-consistency objective during training without significant additional burden. Empirically, it consistently outperforms other end-to-end methods in terms of generation quality."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Limited comparison with one-step generator GAN is provided. In Appendix Table 2, StyleGAN-X demonstrates the best performance (2.3). Does this suggest that GAN is currently the optimal choice for one-step sampling?"}},"nonreaders":[],"tmdate":1732012892438,"tcdate":1729328450601,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission5115/Reviewer_NSZo"],"signatures":["ICLR.cc/2025/Conference/Submission5115/Reviewer_NSZo"],"forum":"OlzB6LnXcS","number":1,"license":"CC BY 4.0","cdate":1729328450601,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission5115/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732012892438,"domain":"ICLR.cc/2025/Conference","replyto":"OlzB6LnXcS","id":"BWskyFRmVk","forumContent":{"venue":{"value":"ICLR 2025 Oral"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["diffusion","flow-matching","fast inference","distillation"]},"primary_area":{"value":"generative models"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Diffusion models and flow matching models have enabled generating diverse and realistic images by learning to transfer noise to data. However, sampling from these models involves iterative denoising over many neural network passes, making generation slow and expensive. Previous approaches for speeding up sampling require complex training regimes, such as multiple training phases, multiple networks, or fragile scheduling. We introduce Shortcut Models, a family of generative models that use a single network and training phase to produce high-quality samples in a single or multiple sampling steps. Shortcut models condition the network not only on the current noise level but also on the desired step size, allowing the model to skip ahead in the generation process. Across a wide range of sampling step budgets, shortcut models consistently produce higher quality samples than previous approaches, such as consistency models and reflow. Compared to distillation, shortcut models reduce complexity to a single network and training phase and additionally allow varying step budgets at inference time."},"_bibtex":{"value":"@inproceedings{\nfrans2025one,\ntitle={One Step Diffusion via Shortcut Models},\nauthor={Kevin Frans and Danijar Hafner and Sergey Levine and Pieter Abbeel},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=OlzB6LnXcS}\n}"},"title":{"value":"One Step Diffusion via Shortcut Models"},"pdf":{"value":"/pdf/834505d749dfe267a983986c87c26443a58835cb.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"frans|one_step_diffusion_via_shortcut_models"},"authorids":{"value":["~Kevin_Frans1","~Danijar_Hafner1","~Sergey_Levine1","~Pieter_Abbeel2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Kevin Frans","Danijar Hafner","Sergey Levine","Pieter Abbeel"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a joint audio-video generation (JAVG) model built on the pretrained text-to-video model Wan2.1.\nTo enable synchronized audio and video generation, three mechanisms are introduced: a model design called MS-MOE for joint generation, a new RoPE variant (TA-RoPE) that manages both cross-modal interaction and single-modal structure, and an audio-video direct preference optimization method (AV-DPO) that aligns the model with human preference. \nExperimental results demonstrate that the proposed method outperforms existing baselines."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"- Did the authors evaluate the model performance after the Audio-Video SFT stage, prior to AV-DPO finetuning?\n- For TA-RoPE, could the authors visualize or analyze the similarity structure between RoPE embeddings used for audio and video?\n- For the training data used in AV-DPO, how did the authors sort the videos based on three scores?\n- Is AV-DPO applicable to other existing JAVG models?\n- Did the authors conduct a user study to compare the proposed method with baselines?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"- The proposed model design (MS-MOE with LoRA finetuning) is simple yet effective, achieving both a high cross-modal alignment and a high single-modal generation quality. It is also computationally efficient, as inference cost per token remains constant while the parameter size increases.\n- The proposed TA-RoPE requires minimal engineering efforts. It provides a natural extension of Wan's RoPE to support both inter- and intra-modal interactions of audio and video.\n- AV-DPO sounds novel, as applying DPO for aligning JAVG models with human preference is underexplored."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The primary concern lies in the **fairness and clarity of the experimental evaluation**.\n\nFor Table 1:\n- It is unclear which T2A models are used for T2A+A2V and which T2V models are used for T2A+V2A. Also, many results of T2A+A2V and T2V+V2A baselines are missing. \n- The authors re-evaluate JavisDiT, but the reported scores deviate substantially from those in the original paper. For instance, TA-IB = 0.197 (original) vs. 0.151 (this paper), CLIP = 0.325 vs. 0.308, AV-IB = 0.201 vs. 0.197, AVHScore =  0.183 vs. 0.179, and JavisScore = 0.158 vs. 0.154. These metrics consistently decrease, and in some cases, the original scores are higher than those of the proposed method.\n- Only the final model (finetuned using AV-DPO) is reported. It would also be important to show results after the Audio-Video SFT stage to isolate the contributions of the model architecture and AV-DPO optimization.\n\nFor Fig.2:\n- All evaluation metrics appear identical to those used in AV-DPO training, relying on the same reward models. Since only the proposed model is optimized for these metrics, this evaluation is biased and not directly comparable to other baselines.\n\nFor Fig.7:\n- The definitions of A-LoRA, A-noLoRA, AV-AttnLoRA, and AV-LoRA are missing, making the figure difficult to interpret. A more detailed explanation is needed.\n\n\n**Lack of justification for the TA-RoPE design.**  \nWhile the reviewer agrees that TA-RoPE is a reasonable engineering extension, the motivation behind the specific positional correspondence between audio and video modalities is not fully explained.\nIn particular, the mapping of video height and audio time, and video width and audio frequency, may introduce implicit correlations that lack perceptual grounding.\nA brief visualization or analysis of the positional-similarity structure would help clarify how the TA-RoPE affects the interaction between audio and video modalities.\n\n\n**Lack of subjective evaluation.**  \nAlthough the paper claims that AV-DPO improves perceptual alignment with human preference, no user study is conducted to support this claim.\nIncluding even a small-scale user study would make the evaluation more convincing."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1764358791747,"tcdate":1761822218241,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission19605/Reviewer_qjFx"],"signatures":["ICLR.cc/2026/Conference/Submission19605/Reviewer_qjFx"],"forum":"hRRWfFpKRp","number":3,"license":"CC BY 4.0","cdate":1761822218241,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission19605/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1764358791747,"domain":"ICLR.cc/2026/Conference","replyto":"hRRWfFpKRp","id":"6SfFRmA5rJ","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"TLDR":{"value":"We introduce JavisDiT++, a concise yet powerful DiT model to generate semantically and temporally aligned sounding videos with textual conditions."},"keywords":{"value":["AIGC","Diffusion Model","Sounding Video Generation","Text-to-Audio-Video Generation","Joint Audio-Video Generation","Video Generation"]},"supplementary_material":{"value":"/attachment/937fbbba3ea309a5d017dc53f9f14cc1ebcf67b2.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Recent AIGC advances have rapidly expanded from text-to-image generation toward high-quality multimodal synthesis across video and audio. Within this context, joint audio-video generation (JAVG) has emerged as a fundamental task that produces synchronized and semantically aligned sound and vision from textual descriptions. However, compared with advanced commercial models such as Veo3, existing open-source methods still suffer from limitations in generation quality, temporal synchrony, and alignment with human preferences. To bridge the gap, this paper presents JavisDiT++, a concise yet powerful framework for efficient and effective JAVG. First, we introduce a modality-specific mixture-of-experts (MS-MoE) design that enables cross-modal interaction efficacy while enhancing single-modal generation quality. Then, we propose a temporal-aligned RoPE (TA-RoPE) strategy to achieve explicit, frame-level synchronization between audio and video tokens. Besides, we develop an audio-video direct preference optimization (AV-DPO) method to align model outputs with human preference across quality, consistency, and synchrony dimensions. Built upon Wan2.1-1.3B-T2V, our model achieves state-of-the-art performance merely with around 1M public training entries, significantly outperforming prior approaches in both qualitative and quantitative evaluations. Comprehensive ablation studies have been conducted to validate the effectiveness of our proposed modules. All the code, model, and dataset are released at https://JavisVerse.github.io/JavisDiT2-page."},"_bibtex":{"value":"@inproceedings{\nliu2026javisdit,\ntitle={JavisDiT++: Unified Modeling and Optimization for Joint Audio-Video Generation},\nauthor={Kai Liu and Yanhao Zheng and Kai Wang and Shengqiong Wu and Rongjunchen Zhang and Jiebo Luo and Dimitrios Hatzinakos and Ziwei Liu and Hao Fei and Tat-Seng Chua},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=hRRWfFpKRp}\n}"},"title":{"value":"JavisDiT++: Unified Modeling and Optimization for Joint Audio-Video Generation"},"pdf":{"value":"/pdf/58e60f5d85c34b49e00fd9bb4bc03f2fd1003993.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"liu|javisdit_unified_modeling_and_optimization_for_joint_audiovideo_generation"},"authorids":{"value":["~Kai_Liu8","~Yanhao_Zheng1","~Kai_Wang39","~Shengqiong_Wu2","~Rongjunchen_Zhang1","~Jiebo_Luo1","~Dimitrios_Hatzinakos1","~Ziwei_Liu1","~Hao_Fei1","~Tat-Seng_Chua2"]},"authors":{"value":["Kai Liu","Yanhao Zheng","Kai Wang","Shengqiong Wu","Rongjunchen Zhang","Jiebo Luo","Dimitrios Hatzinakos","Ziwei Liu","Hao Fei","Tat-Seng Chua"]}},"version":2},{"content":{"summary":{"value":"This paper proposes VELR (Efficient Video Reward Feedback via Ensemble Latent Reward Models), an innovative framework to overcome the prohibitive memory cost of applying large-scale video Reward Models (RMs) in Reward Feedback Learning (ReFL) for Text-to-Video (T2V) generation. By training an Ensemble Latent Reward Model (LRM) to predict rewards directly in the latent space, the framework successfully bypasses expensive backpropagation through the VAE decoder and the large video RM. The method achieves a substantial memory reduction (up to 150GB ) while maintaining comparable performance to standard ReFL, making previously infeasible Video RM-based fine-tuning a reality for large T2V models. The work is highly relevant and addresses a critical practical limitation in aligning video diffusion models."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Please refer to Weaknesses for more details. If the concerns are solved, I will be glad to raise my score rate."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1.\tThe proposed reward model enables the effective use of powerful, temporal-aware Video RMs (like UnifiedReward and CausalVQA) that were previously computationally inaccessible.\n2.\tThe introduction of the Ensemble LRM is technically sound. \n3.\tExperimental results indicate the effectiveness of the proposed method."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1.\tThe primary benefit of moving from Image RMs to Video RMs is the improvement in temporal coherence. The paper critically lacks supplementary video material for the generated samples. Without visual evidence, it is impossible for the reviewer to verify the claimed improvements in temporal consistency and to judge the subjective quality, flicker, and artifacts of the generated videos. This omission severely undermines the empirical claims of the paper.\n2.\tWhile the paper compares against LRM-adapted baselines, there is no direct comparison against a dedicated, high-fidelity Image Reward model (running in its original pixel space). Additionally, Since Video RMs primarily prioritize temporal objectives, a convincing comparison is necessary to demonstrate that the VELR framework, despite its efficiency and temporal gains, fully retains the high perceptual visual quality boost provided by state-of-the-art Image RMs. The current results focus heavily on memory and speed, but not enough on the visual quality trade-off (if any) compared to the best image-focused alignment methods.\n3.\tWhile motivated, \"Truncated Mid Step Setting\" introduces a model-specific heuristic for selecting the \"mid-step regime\" that is dependent on the velocity prediction and noise schedule of the base T2V model (e.g., Wan-2.1 ). This reliance limits the general applicability and plug-and-play nature of the ReFL solution across different diffusion architectures."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762918883099,"tcdate":1761655559305,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission6523/Reviewer_u2cT"],"signatures":["ICLR.cc/2026/Conference/Submission6523/Reviewer_u2cT"],"forum":"TctJWv7Suz","number":1,"license":"CC BY 4.0","cdate":1761655559305,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission6523/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762918883099,"domain":"ICLR.cc/2026/Conference","replyto":"TctJWv7Suz","id":"vGc4NTIz68","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Diffusion Model","Text-to-Video Generation","Generative Models","Reward Feedback Learning"]},"primary_area":{"value":"generative models"},"abstract":{"value":"Reward feedback learning (ReFL) is effective for both text-to-image (T2I) and text-to-video (T2V) generation with image reward models (RMs). However, image RMs are misaligned with temporal objectives of T2V, motivating ReFL with video reward models. Nevertheless, directly deploying video RMs is impractical due to their large parameter size and the prohibitive cost of memory. To address this, we propose VELR: an efficient framework that employs ensemble latent reward models (LRMs) to predict rewards directly in latent space, bypassing expensive backpropagation through VAE decoders and video RMs. Specifically, we introduce the ensemble technique for the LRM, which enhances capacity, quantifies uncertainty, and mitigates reward hacking. VELR achieves a reduction of up to 150GB in memory, requiring as little as 12.4\\% of the memory compared to standard ReFL. Experiments on OpenSora, CogVideoX-1.5, and Wan-2.1 with large-scale video RMs demonstrate that VELR achieves comparable performance as standard ReFL and enables efficient and robust video RM-based ReFL at scales previously unattainable."},"_bibtex":{"value":"@misc{\nzhang2026velr,\ntitle={{VELR}: Efficient Video Reward Feedback via Ensemble Latent Reward Models},\nauthor={Liyu Zhang and Kehan Li and Tao Zhou and Zeyi Huang and Chao Li and Jiming Chen},\nyear={2026},\nurl={https://openreview.net/forum?id=TctJWv7Suz}\n}"},"title":{"value":"VELR: Efficient Video Reward Feedback via Ensemble Latent Reward Models"},"pdf":{"value":"/pdf/4afe2ce4f2156ac4450ff58713720bc800da2230.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"zhang|velr_efficient_video_reward_feedback_via_ensemble_latent_reward_models"},"authorids":{"value":["~Liyu_Zhang2","~Kehan_Li2","~Tao_Zhou8","~Zeyi_Huang4","~Chao_Li24","~Jiming_Chen1"]},"authors":{"value":["Liyu Zhang","Kehan Li","Tao Zhou","Zeyi Huang","Chao Li","Jiming Chen"]}},"version":2},{"content":{"summary":{"value":"Previous work struggles with hour-long videos because existing datasets fail to capture long-range context, and current models are not designed to handle such extensive temporal dependencies.\nTo address this limitation, the authors introduce a new task called Hierarchical Dense Video Captioning (HDVC) for long-form video understanding. HDVC requires generating both scene-level captions and a global narrative caption for the entire video. They also release HourHDVC, a new dataset that supports this task.\nIn addition, the authors propose LOCO (LOng COntext memory-based hierarchical dense video captioning), a model that uses a two-tier memory system, Context-aware Memory and Long-term Context Memory, to maintain coherent narratives over extended time spans."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"1. It appears that the context-aware memory is not updated across scenes, and instead each scene has its own separate context-aware memory. Is this correct? If so, if the video gets long, wouldn't it need large storage?\n2. In line 290, when you mention parameter sharing, does this mean that the same LoRA weights are shared, or that it is the same base model with different LoRA adapters for each component? The figure only shows LoRA applied to the video-narrative captioning module, so clarification would be helpful.\n3. In Equation (2), I assume M represents memory, but it would improve readability if all symbols were explicitly defined in the text."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"1. The authors introduce a new task and dataset with multi-level annotations, providing scene-level and video-level narrative captions. This dataset is likely to be valuable for future research on long-form video understanding.\n2. They present a new model, LOCO, which outperforms existing methods on the HourHDVC dataset.\n3. They also propose a new evaluation metric, ConSim, designed to measure how well models capture contextual information."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The paper lacks sufficient detail regarding the LOCO architecture. In particular, it is unclear which backbone models are used for each component (e.g., the temporal encoder), whether the framework generalizes to alternative model choices, and what hyperparameters are used during training. More implementation details would help clarify the method and improve reproducibility.\n2. It is also unclear why the performance of Scene Dense Video Captioning decreases when context-aware memory is added in Table 4. Since this component is intended to enhance the modeling of long-range context, additional explanation is needed to understand why it does not consistently improve performance. \n3. It seems like the comparison setup is unfair for Table 2: the proposed method is trained in-domain, whereas some baselines may not be. \n4. More detail is needed about the human refinement of the evaluation set. For example, it is unclear what specific changes annotators made, how frequently edits were applied, who the annotators were, and what procedures were used to ensure annotation quality.\n5. The description of the ConSim metric is also incomplete. How ConSim is computed? Is it a form of LLM-as-a-judge–style evaluation?\n6. The paper does not report results from other methods on the Ego4D-HCap dataset (Table 5). Could you also add the results of different methods?"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762922082729,"tcdate":1761864340104,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission10868/Reviewer_1GSZ"],"signatures":["ICLR.cc/2026/Conference/Submission10868/Reviewer_1GSZ"],"forum":"gr0Z4kWUdC","number":2,"license":"CC BY 4.0","cdate":1761864340104,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission10868/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762922082729,"domain":"ICLR.cc/2026/Conference","replyto":"gr0Z4kWUdC","id":"Ob6OwAzYNd","forumContent":{"TLDR":{"value":"We present HourHDVC, a benchmark and model for hour-long dense video captioning that leverage scene-to-narrative structure and long-context memory, setting a new standard for coherent, paragraph-level video descriptions."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Dense Video Captioning","Hour-long Video Understanding","Paragraph Captioning"]},"supplementary_material":{"value":"/attachment/9d114e2fc76ee32754d2ba066d8a8b30c06f61b8.zip"},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"While existing Dense Video Captioning (DVC) research has shown promise for short video clips, current approaches struggle with hour-long videos due to a critical lack of datasets that capture long-term context and models capable of managing extensive temporal dependencies. To address these challenges, we introduce Hierarchical Dense Video Captioning (HDVC), a novel task designed for long-form videos that involves both scene-level and video-level narrative captioning. For this task, we propose HourHDVC, a new dataset providing comprehensive annotations for hour-long videos. We also present LOng COntext memory-based hierarchical dense video captioning (LOCO), an end-to-end model explicitly designed to manage extensive temporal dependencies by modeling the scene-to-narrative structure inherent in HDVC. LOCO leverages a two-tier memory system, Context-aware Memory and Long-term Context Memory, to maintain narrative coherence across extended durations. Experiments on HourHDVC demonstrate that LOCO establishes strong baselines for the HDVC task, while highlighting the remaining challenges of modeling long-form video narratives."},"_bibtex":{"value":"@misc{\nkim2026how,\ntitle={How Do You Watch a Movie? Hour{HDVC}: Hour-Long Hierarchical Dense Video Captioning},\nauthor={Minkuk Kim and Heedong Kim and Jinyoung Moon and Jinwoo Choi and Seong Tae Kim},\nyear={2026},\nurl={https://openreview.net/forum?id=gr0Z4kWUdC}\n}"},"title":{"value":"How Do You Watch a Movie? HourHDVC: Hour-Long Hierarchical Dense Video Captioning"},"pdf":{"value":"/pdf/650b338229244f601a70a5776028598fa1a3d5dd.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"kim|how_do_you_watch_a_movie_hourhdvc_hourlong_hierarchical_dense_video_captioning"},"authorids":{"value":["~Minkuk_Kim1","~Heedong_Kim2","~Jinyoung_Moon1","~Jinwoo_Choi1","~Seong_Tae_Kim1"]},"authors":{"value":["Minkuk Kim","Heedong Kim","Jinyoung Moon","Jinwoo Choi","Seong Tae Kim"]}},"version":2},{"content":{"venue":{"value":"CoRR 2026"},"pdf":{"value":"https://arxiv.org/pdf/2602.07358v2"},"venueid":{"value":"dblp.org/journals/CORR/2026"},"paperhash":{"value":"he|utopia_unlearnable_tabular_data_via_decoupled_shortcut_embedding"},"authorids":{"value":["~Jiaming_He2","","","","","","",""]},"html":{"value":"https://doi.org/10.48550/arXiv.2602.07358"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2602-07358,\n  publtype={informal},\n  author={Jiaming He and Fuming Luo and Hongwei Li and Wenbo Jiang and Wenshu Fan and Zhenbo Shi and Xudong Jiang and Yi Yu},\n  title={UTOPIA: Unlearnable Tabular Data via Decoupled Shortcut Embedding},\n  year={2026},\n  month={February},\n  cdate={1769904000000},\n  journal={CoRR},\n  volume={abs/2602.07358},\n  url={https://doi.org/10.48550/arXiv.2602.07358}\n}\n"},"abstract":{"value":"Unlearnable examples (UE) have emerged as a practical mechanism to prevent unauthorized model training on private vision data, while extending this protection to tabular data is nontrivial. Tabular data in finance and healthcare is highly sensitive, yet existing UE methods transfer poorly because tabular features mix numerical and categorical constraints and exhibit saliency sparsity, with learning dominated by a few dimensions. Under a Spectral Dominance condition, we show certified unlearnability is feasible when the poison spectrum overwhelms the clean semantic spectrum. Guided by this, we propose Unlearnable Tabular Data via DecOuPled Shortcut EmbeddIng (UTOPIA), which exploits feature redundancy to decouple optimization into two channels: high saliency features for semantic obfuscation and low saliency redundant features for embedding a hyper correlated shortcut, yielding constraint-aware dominant shortcuts while preserving tabular validity. Extensive experiments across tabular datasets and models show UTOPIA drives unauthorized training toward near random performance, outperforming strong UE baselines and transferring well across architectures."},"title":{"value":"UTOPIA: Unlearnable Tabular Data via Decoupled Shortcut Embedding"},"authors":{"value":["Jiaming He","Fuming Luo","Hongwei Li","Wenbo Jiang","Wenshu Fan","Zhenbo Shi","Xudong Jiang","Yi Yu"]}},"tmdate":1777180082521,"pdate":1798675200000,"tcdate":1777180070892,"writers":["~"],"signatures":["~Jiaming_He2"],"forum":"XMruEMUnRj","license":"CC BY-SA 4.0","number":870425,"cdate":1769904000000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1777180082521,"domain":"DBLP.org","id":"XMruEMUnRj","version":2},{"content":{"venue":{"value":"AAAI 2024"},"pdf":{"value":"https://ojs.aaai.org/index.php/AAAI/article/download/29211/30285"},"venueid":{"value":"dblp.org/conf/AAAI/2024"},"paperhash":{"value":"kim|adaptive_shortcut_debiasing_for_online_continual_learning"},"authorids":{"value":["~Doyoung_Kim2","https://dblp.org/search/pid/api?q=author:Dongmin_Park:","https://dblp.org/search/pid/api?q=author:Yooju_Shin:","https://dblp.org/search/pid/api?q=author:Jihwan_Bang:","~Hwanjun_Song1","~Jae-Gil_Lee1"]},"html":{"value":"https://doi.org/10.1609/aaai.v38i12.29211"},"_bibtex":{"value":"@inproceedings{DBLP:conf/aaai/KimPSBS024,\n  author={Doyoung Kim and Dongmin Park and Yooju Shin and Jihwan Bang and Hwanjun Song and Jae-Gil Lee},\n  title={Adaptive Shortcut Debiasing for Online Continual Learning},\n  year={2024},\n  cdate={1704067200000},\n  pages={13122-13131},\n  url={https://doi.org/10.1609/aaai.v38i12.29211},\n  booktitle={AAAI},\n  crossref={conf/aaai/2024}\n}\n"},"abstract":{"value":"We propose a novel framework DropTop that suppresses the shortcut bias in online continual learning (OCL) while being adaptive to the varying degree of the shortcut bias incurred by continuously changing environment. By the observed high-attention property of the shortcut bias, highly-activated features are considered candidates for debiasing. More importantly, resolving the limitation of the online environment where prior knowledge and auxiliary data are not ready, two novel techniques---feature map fusion and adaptive intensity shifting---enable us to automatically determine the appropriate level and proportion of the candidate shortcut features to be dropped. Extensive experiments on five benchmark datasets demonstrate that, when combined with various OCL algorithms, DropTop increases the average accuracy by up to 10.4% and decreases the forgetting by up to 63.2%."},"title":{"value":"Adaptive Shortcut Debiasing for Online Continual Learning"},"authors":{"value":["Doyoung Kim","Dongmin Park","Yooju Shin","Jihwan Bang","Hwanjun Song","Jae-Gil Lee"]}},"tmdate":1769340959665,"pdate":1704067200000,"tcdate":1727656743758,"writers":["~"],"signatures":["~Jae-Gil_Lee1"],"forum":"WKbpbV6zeB","license":"CC BY-SA 4.0","number":107711,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1769340959665,"domain":"DBLP.org","id":"WKbpbV6zeB","version":2},{"content":{"venue":{"value":"CoRR 2023"},"pdf":{"value":"http://arxiv.org/pdf/2312.08677v1"},"venueid":{"value":"dblp.org/journals/CORR/2023"},"paperhash":{"value":"kim|adaptive_shortcut_debiasing_for_online_continual_learning"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Doyoung_Kim:","https://dblp.org/search/pid/api?q=author:Dongmin_Park:","https://dblp.org/search/pid/api?q=author:Yooju_Shin:","https://dblp.org/search/pid/api?q=author:Jihwan_Bang:","~Hwanjun_Song1","~Jae-Gil_Lee1"]},"html":{"value":"https://doi.org/10.48550/arXiv.2312.08677"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2312-08677,\n  publtype={informal},\n  author={Doyoung Kim and Dongmin Park and Yooju Shin and Jihwan Bang and Hwanjun Song and Jae-Gil Lee},\n  title={Adaptive Shortcut Debiasing for Online Continual Learning},\n  year={2023},\n  cdate={1672531200000},\n  journal={CoRR},\n  volume={abs/2312.08677},\n  url={https://doi.org/10.48550/arXiv.2312.08677}\n}\n"},"abstract":{"value":"We propose a novel framework DropTop that suppresses the shortcut bias in online continual learning (OCL) while being adaptive to the varying degree of the shortcut bias incurred by continuously changing environment. By the observed high-attention property of the shortcut bias, highly-activated features are considered candidates for debiasing. More importantly, resolving the limitation of the online environment where prior knowledge and auxiliary data are not ready, two novel techniques -- feature map fusion and adaptive intensity shifting -- enable us to automatically determine the appropriate level and proportion of the candidate shortcut features to be dropped. Extensive experiments on five benchmark datasets demonstrate that, when combined with various OCL algorithms, DropTop increases the average accuracy by up to 10.4% and decreases the forgetting by up to 63.2%."},"title":{"value":"Adaptive Shortcut Debiasing for Online Continual Learning"},"authors":{"value":["Doyoung Kim","Dongmin Park","Yooju Shin","Jihwan Bang","Hwanjun Song","Jae-Gil Lee"]}},"tmdate":1768563554159,"pdate":1672531200000,"tcdate":1727656744072,"writers":["~"],"signatures":["~Jae-Gil_Lee1"],"forum":"KBxHBxQZL1","license":"CC BY-SA 4.0","number":107724,"cdate":1672531200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1768563554159,"domain":"DBLP.org","id":"KBxHBxQZL1","version":2},{"content":{"summary":{"value":"The goal of this work is to tackle video virtual try-on task. To achieve this purpose, this paper builds a new virtual try-on dataset, which consists of a total of 5,350 video-image pairs. In addition, this work constructs the focus tunnel to achieve better performance, and uses Kalman filter, tunnel position embedding and environment context to further enhance the authenticity of try-on results."},"suitability":{"value":3},"strengths":{"value":"1. This work proposes the first diffusion-based video virtual try-on model and has better performance than previous work.\n2. This work collects a new virtual try-on dataset which consists of a total of 5,350 video-image pairs.\n3. This work constructs the focus tunnel to emphasize the clothing region in videos and utilizes Kalman filter, tunnel position embedding and environment context to further enhance model's performance."},"confidence":{"value":3},"rating":{"value":2},"limitations":{"value":"1. The contribution of tunnel extraction is not novel enough. It zooms in the agnostic person but does not deal with the clothing. So it does not seem to be significantly helpful in enhancing the details of the original clothing. In addition, tunnel extraction may destroy the coherent motion in the original video.\n2. It may be more reasonable to provide a detailed comparison between the dataset proposed by this work and the VVT dataset, as well as give some visual comparisons. What is the resolution of the proposed dataset?\n3. The experiments in this work seem somewhat insufficient. The paper seems to only provide quantitative comparison experiments and ablation study on the VVT dataset. It might be better to add more experiments.\n4. It might be better to explain the implementation details of the temporal aggregation technique to combine different video clips in the testing phase.\n5. For the video virtual try-on task, is it possible to use the image-based virtual try-on method to generate a certain frame, and then use the video pose transfer method to generate the video try-on result? It might be better to conduct such experiments in contrast to the approach of this work."}},"nonreaders":[],"tmdate":1721657143523,"tcdate":1715583553022,"writers":["acmmm.org/ACMMM/2024/Conference","acmmm.org/ACMMM/2024/Conference/Submission573/Reviewer_8zFG"],"signatures":["acmmm.org/ACMMM/2024/Conference/Submission573/Reviewer_8zFG"],"forum":"yoPXKdyctQ","number":1,"license":"CC BY 4.0","cdate":1715583553022,"readers":["everyone"],"invitations":["acmmm.org/ACMMM/2024/Conference/Submission573/-/Official_Review","acmmm.org/ACMMM/2024/Conference/-/Edit","acmmm.org/ACMMM/2024/Conference/Submission573/Official_Review1/-/Final_Rating"],"mdate":1721657143523,"domain":"acmmm.org/ACMMM/2024/Conference","replyto":"yoPXKdyctQ","id":"DQ9BC0KZOC","forumContent":{"venue":{"value":"MM2024 Poster"},"supplementary_material":{"value":"/attachment/4e40152847be001eb5aea3e3b9cbe4631d9e889c.zip"},"abstract":{"value":"Video try-on is challenging and has not been well tackled in previous works. The main obstacle lies in preserving the clothing details and modeling the coherent motions simultaneously. Faced with those difficulties, we address video try-on by proposing a diffusion-based framework named \"Tunnel Try-on.\" The core idea is excavating a ``focus tunnel'' in the input video that gives close-up shots around the clothing regions. We zoom in on the region in the tunnel to better preserve the fine details of the clothing. To generate coherent motions, we leverage the Kalman filter to smooth the tunnel and inject its position embedding into attention layers to improve the continuity of the generated videos. In addition, we develop an environment encoder to extract the context information outside the tunnels. Equipped with these techniques, Tunnel Try-on keeps fine clothing details and synthesizes stable and smooth videos. Demonstrating significant advancements, Tunnel Try-on could be regarded as the first attempt toward the commercial-level application of virtual try-on in videos. The project page is https://mengtingchen.github.io/tunnel-try-on-page/."},"relevance_to_conference":{"value":"Our paper aligns well with the sub-theme of \"Generative Multimedia\" under the broader theme of \"Multimedia in the Generative AI Era\". Specifically, we have developed a video virtual try-on system based on diffusion models, which takes both clothing images and user videos as input and generates high-fidelity try-on videos. Compared to image-based try-on systems that output single still images, our video virtual try-on model offers users a more immersive and realistic try-on experience. Additionally, as the first diffusion-based video virtual try-on model, our system supports various types of tops and bottoms, as well as complex backgrounds and diverse movements in real-world scenarios. This allows users to input more personalized videos for virtual try-on, thereby enhancing the overall user experience. To conclude, our work contributes to the application of video generation models in the fashion domain, making a impact in the multimedia field."},"_bibtex":{"value":"@inproceedings{\nxu2024tunnel,\ntitle={Tunnel Try-on: Excavating Spatial-temporal Tunnels for High-quality Virtual Try-on in Videos},\nauthor={Zhengze Xu and Mengting Chen and Zhao Wang and Linyu XING and Zhonghua Zhai and Nong Sang and Jinsong Lan and Shuai Xiao and Changxin Gao},\nbooktitle={ACM Multimedia 2024},\nyear={2024},\nurl={https://openreview.net/forum?id=yoPXKdyctQ}\n}"},"title":{"value":"Tunnel Try-on: Excavating Spatial-temporal Tunnels for High-quality Virtual Try-on in Videos"},"secondary_subject_area":{"value":["[Generation] Generative Multimedia"]},"pdf":{"value":"/pdf/e4e4634b5c3c99189fa47448e35ae91fe03f6b3b.pdf"},"venueid":{"value":"acmmm.org/ACMMM/2024/Conference"},"paperhash":{"value":"xu|tunnel_tryon_excavating_spatialtemporal_tunnels_for_highquality_virtual_tryon_in_videos"},"primary_subject_area":{"value":"[Generation] Generative Multimedia"},"authorids":{"value":["~Zhengze_Xu1","~Mengting_Chen1","~Zhao_Wang9","~Linyu_XING1","~Zhonghua_Zhai1","~Nong_Sang1","~Jinsong_Lan1","~Shuai_Xiao1","~Changxin_Gao1"]},"authors":{"value":["Zhengze Xu","Mengting Chen","Zhao Wang","Linyu XING","Zhonghua Zhai","Nong Sang","Jinsong Lan","Shuai Xiao","Changxin Gao"]}},"version":2},{"content":{"summary":{"value":"The authors present an analysis of shortcut learning in image classification. Their study includes experiments on bird, chest x-ray pneumothorax, and skin lesion classification. The experiments were conducted with three different neural network architectures, three learning rates, and two optimizers. Shortcuts were either naturally present (e.g., background in water bird classification or drains on chest x-rays) or added in a controlled setup.\nThe study shows that through PD and KL-divergence analysis, one can locate the layers most prone to a given shortcut, having important implications for mitigation strategies. The authors demonstrate that more complex shortcuts can be located deeper inside the network than simple confounders (i.e., red squares in fixed locations)."},"final_rating":{"value":5,"readers":["everyone"]},"justification_of_final_rating":{"value":"I am fully satisfied with the replies and edits by the authors and recommend accepting this contribution. The revision has further improved the text and readability and how has added relevant references.","readers":["everyone"]},"justification_of_the_preliminary_rating":{"value":"The study is motivated by a pressing issue in medical AI and puts the effort into the context of prior research. The study is carefully designed and analyzed. The authors conducted experiments on various architectures, learning strategies, and datasets and included shortcuts with different complexities, making the study very relevant for shortcut/confounder mitigation techniques."},"strengths":{"value":"- The paper is written concisely, guiding the reader through the text.\n- I appreciate the clear experimental design, thoughtful selection of shortcuts with different complexities, and precise analysis.\n- The figures are prepared well and contribute to understanding the method and support the findings described in the text.\n- The effort the authors put into benchmarking with three datasets, networks, learning rates, and two optimizers is commendable\n- The study contributes to a pressing issue in medical AI."},"weaknesses":{"value":"- No significant weaknesses in my opinion.\n- Since you tested many different setups, I would appreciate it if you could elaborate when describing the findings if you observed common patterns for all tested networks and optimizers (e.g., on the PD vs. shortcut complexity). I am aware of the length constraint; consider, e.g., adding a table showing the results for all experiments or plots for the network architectures you do not showcase in the main text as a supplementary."},"confidence":{"value":4},"detailed_comments":{"value":"I appreciate the very thorough experiments and the clear experimental setup. You are able to show the connection between the shortcut complexity and PD nicely.\nSince you did many more tests than you can present in the tables/text, it would be good if you could elaborate if and how your findings applied to the tested neural networks, optimizers, etc. (you already mention the influence of the LR on mitigating shortcuts).\nI the abstract, you mention that you explore shortcut mitigation strategies. I agree that your work has important implications for mitigation strategies, but do not see the exploration part in the text."},"special_issue":{"value":"Yes"},"recommendation":{"value":"Oral"},"questions_to_address_in_the_rebuttal":{"value":"- Were the findings the same for the two tested optimizers and for the three tested neural networks?"},"preliminary_rating":{"value":5}},"nonreaders":[],"tmdate":1712410980295,"tcdate":1709196608405,"writers":["MIDL.io/2024/Conference","MIDL.io/2024/Conference/Submission270/Reviewer_N54m"],"signatures":["MIDL.io/2024/Conference/Submission270/Reviewer_N54m"],"forum":"3UMzxqDcpY","number":3,"license":"CC BY 4.0","cdate":1709196608405,"readers":["everyone"],"invitations":["MIDL.io/2024/Conference/Submission270/-/Official_Review","MIDL.io/2024/Conference/-/Edit","MIDL.io/2024/Conference/Submission270/Official_Review3/-/Review_Revision"],"mdate":1712410980295,"domain":"MIDL.io/2024/Conference","replyto":"3UMzxqDcpY","id":"AWNXMRMZyN","forumContent":{"venue":{"value":"MIDL 2024 Poster"},"keywords":{"value":["shortcut learning","bias","prediction depth","model interpretation","clinical machine learning","spurious correlations","model robustness","generalization"]},"abstract":{"value":"Many studies have reported human-level accuracy (or better) for AI-powered algorithms performing a specific clinical task, such as detecting pathology. However, these results often fail to generalize to other scanners or populations. Several mechanisms have been identified that confound generalization. One such is shortcut learning, where a network erroneously learns to depend on a fragile spurious feature, such as a text label added to the image, rather than scrutinizing the genuinely useful regions of the image. In this way, systems can exhibit misleadingly high test-set results while the labels are present but fail badly elsewhere where the relationship between the label and the spurious feature breaks down. In this paper, we investigate whether it is possible to detect shortcut learning and locate where the shortcut is happening in a neural network. We propose a novel methodology utilizing the sample difficulty metric Prediction Depth (PD) and KL divergence to identify specific layers of a neural network model where the learned features of a shortcut manifest. We demonstrate that our approach can effectively isolate these layers across several shortcuts, model architectures, and datasets. Using this, we show a correlation between the visual complexity of a shortcut, the depth of its feature manifestation within the model, and the extent to which a model relies on it. Finally, we highlight the nuanced relationship between learning rate and shortcut learning."},"_bibtex":{"value":"@inproceedings{\nboland2024there,\ntitle={There Are No Shortcuts to Anywhere Worth Going: Identifying Shortcuts in Deep Learning Models for Medical Image Analysis},\nauthor={Christopher Boland and Keith A Goatman and Sotirios A. Tsaftaris and Sonia Dahdouh},\nbooktitle={Medical Imaging with Deep Learning},\nyear={2024},\nurl={https://openreview.net/forum?id=3UMzxqDcpY}\n}"},"title":{"value":"There Are No Shortcuts to Anywhere Worth Going: Identifying Shortcuts in Deep Learning Models for Medical Image Analysis"},"latex_code":{"value":"/attachment/6430a253dcfaddf41854115954ca18c9fc001a76.zip"},"pdf":{"value":"/pdf/3c6888b96f569bccf12925981f91ed4bc7d9a6d0.pdf"},"copyright_form":{"value":"/attachment/27380095359df7fa78ed39bf9f36960888ca219b.pdf"},"venueid":{"value":"MIDL.io/2024/Conference"},"paperhash":{"value":"boland|there_are_no_shortcuts_to_anywhere_worth_going_identifying_shortcuts_in_deep_learning_models_for_medical_image_analysis"},"authorids":{"value":["~Christopher_Boland1","~Keith_A_Goatman1","~Sotirios_A._Tsaftaris1","sonia.dahdouh@mre.medical.canon"]},"authors":{"value":["Christopher Boland","Keith A Goatman","Sotirios A. Tsaftaris","Sonia Dahdouh"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a novel video generation guidance method capable of producing videos with drastically different cross-frame descriptions, such as texture changes, morph transformations, and large motions. The core innovation of the method is the state triplet, which decomposes the video into different phases, each with its own distinct description. The state triplet can be generated either from a large language model (LLM) or manually. The proposed method is training-free and has been evaluated on a new video benchmark, achieving performance comparable to methods that require task-specific training."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"In addition to the key information missing mentioned above, I have several questions related to the details of the paper:\n\n1. How to handle different combinations of text and image prompt? For example, what if we have the triplet description and only the end frame of the generated image? How is this case different from the case where we have the triplet description and only the first frame of the generated image?\n\n2. How is the method able to handle the morph transform even although the pretrained model is rarely trained on videos with morphism since it is not common? More discussion on this would be appreciated.\n\n3. Have the authors try a more detailed caption vs. the simple ones used in the paper? Will that lead to better motion?"},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"The proposed method is practically useful given that:\n\n* It enables the transformations that would be very difficult to model with pretrained models, such as morph transformation, drastic texture changes over frames. \n\n* The method is training-free and does not require text-video pairs with the target motion patterns, which is hard to collect in scale. The training free method achieves comparable performance as the compared training-based method."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The T2V limitation mentioned in L202 is not convincing enough. The prompts used in the paper are relatively simple, consisting of only one or two sentences. Some related works, such as CogVideoX [1], utilize DiT-based structures and demonstrate that detailed prompts significantly improve video generation quality, both in appearance and motion. Therefore, if we consider a scenario with both detailed video captions and a DiT-based model that relies less on explicit per-frame modeling due to (1) temporal compression in the tokenizer, and (2) stronger spatial-temporal modeling capabilities (i.e., cross-frame modeling rather than per-frame), the limitation highlighted for the T2V model becomes less relevant. This is because we would not be restricted by a limited T2I model (point 2) and would benefit from enhanced spatial-temporal modeling (point 1).  \n\n- The paper lacks key information on how the transition order is maintained. While Eq. 1 models the joint conditional distribution given the prompt triplet, it does not specify how the generated images are constrained to follow the prompt order: initial -> transition -> final stage.  Ensuring this sequential alignment is crucial for achieving controllability and realism in the generated video.\n\n- As mentioned in the limitation section in the supplementary, the method introduces additional hyper-parameters, such as the guidance scale at for the triplet states. Tweaking those hyper-parameter would be a case-specific effort and paper does not propose a principled approach for estimating/optimizing those hyper-parameters. \n\n- As mentioned in the first item, there are strong models taking much more descriptive prompt as input for video generation. However, the paper does not include the comparison with those methods. The lack of this comparison makes the claim about the T2V limitation and the proposed method less convincing.\n\n[1] CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer. Yang, Zhuoyi and Teng, Jiayan and Zheng, Wendi and Ding, Ming and Huang, Shiyu and Xu, Jiazheng and Yang, Yuanming and Hong, Wenyi and Zhang, Xiaohan and Feng, Guanyu and others. arXiv preprint arXiv:2408.06072"}},"nonreaders":[],"tmdate":1731427675314,"tcdate":1730698535673,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission7611/Reviewer_EnJp"],"signatures":["ICLR.cc/2025/Conference/Submission7611/Reviewer_EnJp"],"forum":"zkGxROm7D3","number":2,"license":"CC BY 4.0","cdate":1730698535673,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission7611/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427675314,"domain":"ICLR.cc/2025/Conference","replyto":"zkGxROm7D3","id":"6AkcCFBiNh","forumContent":{"TLDR":{"value":"We present two new sampling methods for Text-to-Video (T2V) diffusion models that enhance pre-trained models, allowing for dynamic scene generation and zero-shot image-to-video and image-image-to-video generation (based on the first and last frames)."},"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Text-to-Video Generation","Diffusion Models","Diffusion Guidance","Zero-shot Image-to-Video Generation"]},"supplementary_material":{"value":"/attachment/4bc8d4947b6af20c7b7dbf7a4b9558133f61039d.zip"},"primary_area":{"value":"generative models"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Current text-to-video (T2V) models have made significant progress in generating high-quality video. However, these models are limited when it comes to generating dynamic video scenes where the description per frame can vary dramatically. Changing the color, shape, position and state of objects in the scene is a challenge that current video models cannot handle. In addition, the lack of a cheap image-based conditioning mechanism limits their creative application. To address these challenges and extend the applicability of T2V models, we propose two innovative approaches: **State Guidance** and **Image Guidance**. **State Guidance** uses advanced guidance mechanisms to control motion dynamics and scene transformation smoothness by navigating the diffusion process between a state triplet <initial state, transition state, final state>. This mechanism enables the generation of dynamic video scenes (Dynamic Scene T2V) and allows to control the speed and the expressiveness of the scene transformation by introducing temporal dynamics via a guidance weight schedule across video frames. **Image Guidance** enables Zero-Shot Image-to-Video generation (Zero-Shot I2V) by injecting reference image into the initial diffusion steps noise predictions. Furthermore, the combination of **State Guidance** and **Image Guidance** allows for zero-shot transitions between two input reference frames of a video (Zero-Shot II2V). Finally, we introduce the novel **Dynamic Scene Benchmark** to evaluate the ability of the models to generate dynamic video scenes. Extensive experiments show that **State Guidance** and **Image Guidance** successfully address the aforementioned challenges and significantly improve the generation capabilities of existing T2V architectures."},"_bibtex":{"value":"@misc{\nsobolev2025state,\ntitle={State \\& Image Guidance: Teaching Old Text-to-Video Diffusion Models New Tricks},\nauthor={Konstantin Sobolev and Mikhail Zhirnov and Arsen Kuzhamuratov and Denis Dimitrov and Andrey Kuznetsov and Anton Konushin},\nyear={2025},\nurl={https://openreview.net/forum?id=zkGxROm7D3}\n}"},"title":{"value":"State & Image Guidance: Teaching Old Text-to-Video Diffusion Models New Tricks"},"pdf":{"value":"/pdf/e09748202d19f56b06015b45722d82b51021a14c.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"sobolev|state_image_guidance_teaching_old_texttovideo_diffusion_models_new_tricks"},"authorids":{"value":["~Konstantin_Sobolev2","~Mikhail_Zhirnov1","~Arsen_Kuzhamuratov1","~Denis_Dimitrov2","~Andrey_Kuznetsov2","~Anton_Konushin1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Konstantin Sobolev","Mikhail Zhirnov","Arsen Kuzhamuratov","Denis Dimitrov","Andrey Kuznetsov","Anton Konushin"]}},"version":2},{"content":{"summary":{"value":"This paper proposes Cumulative Self-Consistency Loss (CSL) as an extension of the shortcut training loss (SL) used in one-step and few-step diffusion models. The authors reinterpret shortcut training through an optimal-control lens. CSL generalizes this by integrating consistency along the remaining trajectory, encouraging alignment not just locally but cumulatively over time. The paper derives a continuous formulation of this loss, motivates it with a “cumulative gradient” analysis, and implements a discrete estimator involving R rollouts (typically R=2). Empirically, CSL yields improved FID scores on CIFAR-10 and CelebA-256 for 1-, 2-, and 4-step sampling, with minimal additional compute cost."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. How is the ∇ₓJ term in Eq. 13 realized in practice? Are gradients ever propagated through rollout steps, or is CSL purely a multi-point supervision objective?  \n2. Why does Algorithm 1 advance the 2d step from the midpoint state rather than from the starting point? Was this empirically motivated?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"- Conceptually clear reinterpretation of shortcut loss using an optimal-control framework, identifying why standard self-consistency only optimizes instantaneous alignment.  \n- The cumulative loss is simple to implement and shows consistent, measurable FID improvements with small computational overhead (~5–10%).  \n- Thorough experimental evaluation with ablations (number of rollout terms, backbone scale, training cost) and comparisons across DiT backbones."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The “cumulative gradient” claimed in Eq. 13 does not seem to match the implementation. Algorithm 1 detaches rollout targets with `stopgrad` and does not backpropagate through future steps, so it seems that the optimization reduces to multiple local MSEs rather than true cumulative assignment. \n- The midpoint update rule used during 2d rollouts is unconventional and unexplained. \n- Evaluation scope is a bit narrow, authos only evaluate on CIFAR-10 and CelebA-256, with large baselines borrowed from prior work under different setups."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762931627588,"tcdate":1761954219548,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission19781/Reviewer_z676"],"signatures":["ICLR.cc/2026/Conference/Submission19781/Reviewer_z676"],"forum":"cZqAk87Lu4","number":3,"license":"CC BY 4.0","cdate":1761954219548,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission19781/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762931627588,"domain":"ICLR.cc/2026/Conference","replyto":"cZqAk87Lu4","id":"ffsnnT3iO4","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["One-Step Diffusion","Optimal Control","Shortcut Diffusion Models"]},"primary_area":{"value":"generative models"},"abstract":{"value":"Although iterative denoising (i.e., diffusion/flow) methods offer strong generative performance, they suffer from low generation efficiency, requiring hundreds of steps of network forward passes to simulate a single sample. Mitigating this requires taking larger step-sizes during simulation, thereby allowing one- or few-step generation. Recently proposed shortcut model learns larger step-sizes by enforcing alignment between its direction and the path defined by a base many-step flow-matching model through a self-consistency loss. However, its generation quality is significantly lower than the base model. In this paper, we formulate few-step generation as a controlled base generative process, and show that self-consistency loss can be understood through the lens of optimal control. This perspective naturally motivates its generalization to the proposed cumulative self-consistency loss that cumulatively penalizes misalignment along the entire trajectory. This encourages larger step-sizes that not only align with the base model at the current time step but also support alignment in the subsequent steps, facilitating high-quality generation. Furthermore, we draw a connection between our approach and reinforcement learning, potentially opening the door to a new set of approaches for few-step generation. Experiments show that we significantly improve one- and few-step generation quality under the same training budget. Implementation is available at: [https://github.com/paribeshregmi/Shortcut-CSL](https://github.com/paribeshregmi/Shortcut-CSL)"},"_bibtex":{"value":"@inproceedings{\nregmi2026shortcut,\ntitle={Shortcut Diffusion Training with Cumulative Consistency Loss: An Optimal Control View},\nauthor={Paribesh Regmi and Sandesh Ghimire and Rui Li},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=cZqAk87Lu4}\n}"},"title":{"value":"Shortcut Diffusion Training with Cumulative Consistency Loss: An Optimal Control View"},"pdf":{"value":"/pdf/1b88b1373e822c88adbcbd6921499915ee5a276b.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"regmi|shortcut_diffusion_training_with_cumulative_consistency_loss_an_optimal_control_view"},"authorids":{"value":["~Paribesh_Regmi1","~Sandesh_Ghimire2","~Rui_Li3"]},"authors":{"value":["Paribesh Regmi","Sandesh Ghimire","Rui Li"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a novel method for storytelling video generation that can do seamless transitions between generated video segments. The main contributions are three aspects: dynamics-informed prompt weighting, temporally-aware blending with bidirectional constraints, and structured semantic action representation. Dynamics-informed prompt weighting can balance the contributions of two prompts. The second contribution designs a time-weighted blending mechanism that dynamically balances past and future frames. The third contribution encodes high-level action semantics into the blending process using a pre-trained text encoder."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Please see above. I think the authors should compare with more baseline methods. In addition, the video demo shown in the supplementary material is not very impressive."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. The paper is well organized.  \n\n2. The topic of storytelling long video generation is worth exploring in the research community. This paper targets this important problem.\n\n3. The paper provides video demonstrations to help reviewers better evaluate the performance of the proposed method."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"There are some concerns and questions about this paper:\n\n1.\tThe description of the Dynamics-Informed Prompt Weighting method seems complicated. Could the author provide diagrams to help readers understand it?\n\n2.\tAccording to Formula 3, the generation of the next video clip depends on the generation of the previous video clip. Won't such an operation introduce error accumulation?\n\n3.\tWhere is alpha^’ applied in Formula 7? How exactly is it used?\n\n4.\tThe long video demo shown in the supplementary materials looks more like a series of video clips spliced together, with very obvious transitions between each video clip, making it difficult for viewers to have an immersive experience.\n\n5.\tWith only two comparison methods, could the authors compare it with more methods? For example, One-Minute Video Generation with Test-Time Training (CVPR 2025, open-source).\n\n6.   The three innovations proposed in the paper lack a clear connection. For example, when introducing SAR, the authors only mention that \"Preserving temporal smoothness alone is insufficient,\" but why? Is this problem caused by the introduction of the first two innovations? It would be helpful if the authors could more clearly explain the connection and motivation between each innovation."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762918621776,"tcdate":1761726019660,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission6326/Reviewer_eX6B"],"signatures":["ICLR.cc/2026/Conference/Submission6326/Reviewer_eX6B"],"forum":"pSwlegpXZ0","number":1,"license":"CC BY 4.0","cdate":1761726019660,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission6326/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762918621776,"domain":"ICLR.cc/2026/Conference","replyto":"pSwlegpXZ0","id":"rmqR6LNK9N","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Long-form Story Generation","Dynamics-Informed Prompt Weighting","Time-Weighted Blending","Semantic Action Representation"]},"supplementary_material":{"value":"/attachment/7e442395fc8a654b728aa132857caffae74e95c3.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Generating coherent long-form video sequences from discrete input using only text prompts is a critical task in content creation. While diffusion-based models excel at short video synthesis, long-form storytelling from text remains largely unexplored and a challenge due to difficulties in temporal coherency, preserving semantic meaning, and maintaining both scene context and action continuity across the video. We introduce a novel storytelling framework that achieves this by integrating scene and action prompts through dynamics-inspired prompt mixing. Specifically, we first present a bidirectional time-weighted latent blending strategy to ensure temporal consistency between segments of the long-form video being generated. We then propose a dynamics-informed prompt weighting (DIPW) mechanism that adaptively balances the influence of scene and action prompts at each diffusion timestep by jointly considering CLIP-based alignment, narrative continuity, and temporal smoothness. To further enhance motion continuity, we incorporate a semantic action representation to encode high-level action semantics into the blending process, dynamically adjusting transitions based on action similarity and ensuring smooth yet adaptable motion changes. Latent space blending maintains spatial coherence between objects in a scene, while time-weighted blending enforces bidirectional constraints for temporal consistency. The resulting integrative system prevents abrupt transitions while ensuring fluid storytelling that faithfully reflects both scene and action cues. Extensive experiments demonstrate significant improvements over baselines, achieving temporally consistent and visually compelling video narratives without any additional training. This approach bridges the gap between short clips and extended video to establish a new paradigm in GenAI-driven video synthesis from text."},"_bibtex":{"value":"@misc{\nkang2025dynamicsinspired,\ntitle={Dynamics-Inspired Text-Guided Video Storytelling},\nauthor={Taewon Kang and Divya Kothandaraman and Ming Lin},\nyear={2025},\nurl={https://openreview.net/forum?id=pSwlegpXZ0}\n}"},"title":{"value":"Dynamics-Inspired Text-Guided Video Storytelling"},"pdf":{"value":"/pdf/df8f20e018ec433d2d3ea26b9f59db8b88179bf3.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"kang|dynamicsinspired_textguided_video_storytelling"},"authorids":{"value":["~Taewon_Kang2","~Divya_Kothandaraman1","~Ming_Lin2"]},"authors":{"value":["Taewon Kang","Divya Kothandaraman","Ming Lin"]}},"version":2},{"content":{"summary":{"value":"The paper introduces VGrounding-QA, a framework that constructs a training dataset based on an event-centric knowledge graph. Long videos are divided into uniformly sized short segments, each paired with corresponding descriptions. These segment-description pairs are then used to train multimodal large language models (MLLMs) for temporal grounding within videos. The trained ViTL model outperforms existing multimodal large language models (MLLMs) on both long video question answering and temporal grounding benchmarks, demonstrating its effectiveness in handling extended temporal contexts."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"- How does the proposed method compare in terms of computational efficiency? It would be helpful to report training and inference time, as well as memory usage, to better understand the method’s practicality and scalability.\n- Compared to existing approaches, which component or design choice of the proposed method contributes most significantly to the observed performance improvements? A detailed analysis or ablation would help clarify this."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"- The paper proposes a novel training dataset constructed from an event-based knowledge graph, providing structured and semantically rich supervision for long video temporal grounding.\n- It introduces a two-stage GRPO framework that leverages grounded spans, improving the alignment between language queries and temporally localized video segments.\n- The proposed method achieves notable performance improvements over existing MLLMs on both long video question answering and temporal grounding benchmarks."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- It is unclear which dataset was used for training. \n- The set of baselines appears incomplete. For instance, it is unclear whether models like LLaVA-Video or InternVL3 utilize a greater number of video frames, which could impact performance. The comparison table should also include computational metrics—such as the number of input frames, inference time, and memory usage—for each method to ensure a fair and comprehensive evaluation.\n- How does the proposed data construction pipeline differ from existing approaches in weakly supervised video learning or synthetic data generation? A more explicit comparison would help clarify the novelty and advantages of the method.\n- Since the dataset is specifically designed for temporal grounding and question answering, the improved performance on these tasks is expected. However, the generalization ability of the proposed method to other video understanding tasks remains unclear and is not thoroughly evaluated."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762927411333,"tcdate":1761912311676,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission17539/Reviewer_qPGy"],"signatures":["ICLR.cc/2026/Conference/Submission17539/Reviewer_qPGy"],"forum":"BOFzC3xndr","number":3,"license":"CC BY 4.0","cdate":1761912311676,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission17539/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762927411333,"domain":"ICLR.cc/2026/Conference","replyto":"BOFzC3xndr","id":"iamTTCUlv2","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Long Video Understanding","RL","Video Grounding"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"We present $\\textit{Video-in-the-Loop}$ (ViTL), a two-stage long-video QA framework that preserves a fixed token budget by first $\\textit{localizing}$ question-relevant interval(s) with a low-fps skim and then $\\textit{answering}$ via span-aware reallocation of visual tokens at higher effective frame rate, emitting an interleaved output with both spans and the final option for direct attribution. We also introduce $\\textit{VGrounding-QA}$, which converts description based event graphs into $\\textit{span-grounded}$ multiple-choice QA by pairing each question with $\\textit{ground-truth}$ time span(s) and related reasoning. ViTL is trained end-to-end with an interleaved group-relative objective that couples temporal IoU for localization with answer correctness, allowing credit to flow from answers back to spans without increasing compute. Under fixed token budgets, ViTL attains up to 8.6\\% with 50\\% less frame input on long-video QA and temporal grounding (e.g., Charades-STA, ActivityNet-Captions) and ablations show that span-aware token reallocation consistently surpasses uniform sampling. Together, $\\textit{VGrounding-QA}$ and ViTL provide an interpretable, compute-efficient recipe for scalable long-video QA."},"_bibtex":{"value":"@misc{\nwang2026videointheloop,\ntitle={Video-in-the-Loop: Span-Grounded Long Video {QA} with Interleaved Reasoning},\nauthor={Chendong Wang and Donglin Bai and Yifan Yang and Xiao Jin and Anlan Zhang and Rui Wang and Shiqi Jiang and Yuqing Yang and Hao Wu and Qi Dai and Chong Luo and Ting Cao and Lili Qiu and Suman Banerjee},\nyear={2026},\nurl={https://openreview.net/forum?id=BOFzC3xndr}\n}"},"title":{"value":"Video-in-the-Loop: Span-Grounded Long Video QA with Interleaved Reasoning"},"pdf":{"value":"/pdf/c64609aeecd8f51a41a6a79d52911ee293b4435c.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"wang|videointheloop_spangrounded_long_video_qa_with_interleaved_reasoning"},"authorids":{"value":["~Chendong_Wang1","~Donglin_Bai1","~Yifan_Yang9","~Xiao_Jin2","~Anlan_Zhang1","~Rui_Wang18","~Shiqi_Jiang4","~Yuqing_Yang1","~Hao_Wu36","~Qi_Dai4","~Chong_Luo1","~Ting_Cao1","~Lili_Qiu1","~Suman_Banerjee3"]},"authors":{"value":["Chendong Wang","Donglin Bai","Yifan Yang","Xiao Jin","Anlan Zhang","Rui Wang","Shiqi Jiang","Yuqing Yang","Hao Wu","Qi Dai","Chong Luo","Ting Cao","Lili Qiu","Suman Banerjee"]}},"version":2},{"content":{"summary":{"value":"To address unreliable rankings and high costs in MLLM evaluation caused by shortcut questions answerable from a single modality, this paper introduces the M³-IRT and M²-IRT frameworks. These extend classical Item Response Theory (IRT) by decomposing both model abilities and question difficulties into image-only, text-only, and cross-modal components. This allows for quantifying a model's cross-modal reasoning and an item's cross-modal demand. Experiments show the framework effectively prioritizes genuine cross-modal questions, faithfully reproducing full-benchmark rankings with as little as a 10% subset. It remains robust even when 50% of the benchmark is contaminated with low-quality items, significantly reducing evaluation costs while enhancing reliability."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"Have you tested any non-linear variants? If not, how sensitive are your conclusions to the linearity assumption?\n\nIn addition, beyond swapping-based corruption, does your filter still identify low-quality items under subtler artifacts?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. The paper makes a contribution by addressing the critical challenge of \"shortcut questions\" in multimodal benchmarks. It provides a systematic solution that enhances the reliability of evaluations while simultaneously reducing computational costs.\n2. The paper effectively targets two major pain points in current multimodal evaluation: unreliable model rankings caused by shortcut question contamination, and the high computational cost associated with large-scale benchmarks. The proposed M³-IRT framework offers an efficient and effective solution to both issues.\n3. By decomposing model abilities into image-only, text-only, and cross-modal components, it offers valuable and interpretable insights into the specific strengths and weaknesses of different MLLMs, moving beyond a single, monolithic accuracy score."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The model's core assumption of linear decomposition for abilities and difficulties might oversimplify the complex, potentially non-linear interactions that occur during cross-modal reasoning. \n2. The interpretability of the estimated parameters such as cross-modal difficulty is derived purely from model performance patterns and lacks external validation against human cognitive judgments of what constitutes a cross-modal task.\n3. It is noted that the paper generates artificial low-quality questions through full swapping of images or text, which effectively simulates extreme scenarios involving complete modal mismatch. That said, it may be worth considering whether this approach fully covers the more prevalent and subtle shortcut features commonly observed in real-world multimodal benchmarks. In practical contexts, shortcuts often exhibit characteristics of concealment and diversity. Given that the framework’s ability to filter such subtle shortcuts has not yet been evaluated, it might be valuable to further explore its generalizability when applied to real-world benchmarks."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762921753768,"tcdate":1761630810259,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission10457/Reviewer_akTA"],"signatures":["ICLR.cc/2026/Conference/Submission10457/Reviewer_akTA"],"forum":"jMbyMp5DCh","number":3,"license":"CC BY 4.0","cdate":1761630810259,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission10457/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762921753768,"domain":"ICLR.cc/2026/Conference","replyto":"jMbyMp5DCh","id":"O1DhVGo957","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"TLDR":{"value":"M3IRT decomposes IRT into single-modal and cross-modal components to identify and filter out shortcut questions in MLLM benchmarks, enabling compact and efficient evaluation of the multimodal reasoning skill of MLLMs."},"keywords":{"value":["VLM","Evaluation","IRT"]},"supplementary_material":{"value":"/attachment/4cc46f02ed5d1fdcbb6d58ebafed2a04a5dabc5b.zip"},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Multimodal Large Language Models (MLLMs) have recently emerged as general architectures capable of reasoning over diverse modalities.\nBenchmarks for MLLMs should measure their ability for cross‑modal integration. However, current benchmarks are filled with shortcut questions, which can be solved using only single modality, and thereby yielding unreliable rankings. For example, in vision-language cases, we can find the correct answer without either the image or the text. These low-quality questions unnecessarily increase the size and computational requirements of benchmarks. We introduce a multi-modal and multidimensional item response theory framework (M3IRT) that extends classical IRT by decomposing both model ability and item difficulty into image‑only, text‑only, and cross‑modal components. M3IRT estimates cross‑modal ability of MLLMs and each question’s cross‑modal difficulty, enabling compact, high‑quality subsets that better reflect multimodal reasoning. Across 24 VLMs on three benchmarks, M3IRT prioritizes genuinely cross‑modal questions over shortcuts and preserves ranking fidelity even when 50\\% of items are artificially generated low‑quality questions, thereby reducing evaluation cost while improving reliability. M3IRT thus offers a practical tool for assessing cross‑modal reasoning and refining multimodal benchmarks."},"_bibtex":{"value":"@inproceedings{\nuebayashi2026evaluating,\ntitle={Evaluating Cross-Modal Reasoning Ability and Problem Characteristics with Multimodal Item Response Theory},\nauthor={Shunki Uebayashi and Kento Masui and Kyohei Atarashi and Han Bao and Hisashi Kashima and Naoto Inoue and Mayu Otani and Koh Takeuchi},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=jMbyMp5DCh}\n}"},"title":{"value":"Evaluating Cross-Modal Reasoning Ability and Problem Characteristics with Multimodal Item Response Theory"},"pdf":{"value":"/pdf/5af32a3eb819a4ead9222e437104d91a206b5e95.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"uebayashi|evaluating_crossmodal_reasoning_ability_and_problem_characteristics_with_multimodal_item_response_theory"},"authorids":{"value":["~Shunki_Uebayashi1","~Kento_Masui1","~Kyohei_Atarashi2","~Han_Bao2","~Hisashi_Kashima2","~Naoto_Inoue1","~Mayu_Otani1","~Koh_Takeuchi1"]},"authors":{"value":["Shunki Uebayashi","Kento Masui","Kyohei Atarashi","Han Bao","Hisashi Kashima","Naoto Inoue","Mayu Otani","Koh Takeuchi"]}},"version":2},{"content":{"summary":{"value":"This paper delves into the long-term video-text learning task and reveals a new problem named multi-granularity noisy correspondence (MNC), referring to both course-grained clip-caption misalignment and fine-grained frame-word misalignment. Clearly, such a problem would hinder temporal learning and video understanding. To address the MNC problem, the authors propose NOise Robust Temporal Optimal traNsport (Norton), which formulates the solutions to both course- and fine-grained NC into a unified optimal transport (OT) framework. On the one hand, Norton filters the irrelevant clips and captions using an alignable prompt bucket and realigns the asynchronous clip-caption pairs based on transport distance, contributing to robustness against course-grained NC. On the other hand, a soft-maximum operator is used to identify crucial words and keyframes so that the negative impact of fine-grained NC can be alleviated. The effectiveness of Norton is validated through extensive experiments on commonly used video-text tasks, including video retrieval, video QA, and action segmentation."},"presentation":{"value":"4 excellent"},"contribution":{"value":"4 excellent"},"soundness":{"value":"4 excellent"},"strengths":{"value":"**Revealing a new problem**. This paper studies a new and practical challenge in the context of long-term video-text representation learning, namely, multi-granularity noisy correspondence (MNC). MNC encompasses both coarse-grained clip-caption misalignment and fine-grained frame-word misalignment, both of which hinder temporal learning and video comprehension. While some studies have been concentrated on addressing the coarse-grained clip-caption misalignment, as far as I know, there are no formal studies on the fine-grained NC for video-text learning. From this perspective, I think this paper would bring some new insights to the community.\n\n**Novel approach**. To handle MNC and achieve robust long-term video-text learning, this paper proposes Norton, which formulates the solutions to both course- and fine-grained NC into a unified optimal transport (OT) framework. Norton first incorporates a token-wise soft-maximum operator to identify crucial words and keyframes within each clip-caption pair, so that the fine-grained NC could be eliminated. After that, Norton filters the irrelevant clips and captions using an alignable prompt bucket, and realigns the asynchronous clip-caption pairs based on transport distance, leading to robustness against course-grained NC. \n\n**Good shape**. This paper is well-written and structured. Besides, the experiment designs are interesting and sufficient. Extensive experimental results validated the effectiveness of the proposed methods and the necessity of solving MNC problems."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- While the authors effectively illustrate the motivation in Fig. 1, the advantages of the proposed OT method over DTW require more elaboration. It is advisable to include further discussions to expound upon and clarify the claims made regarding the superiority of OT over DTW.\n- The use of the stop-gradient operation in transport assignment Q is highlighted by the authors as a means to enhance the efficiency of their video-paragraph contrastive loss. However, the rationale behind this operation's efficacy is not entirely clear. To remedy this, additional discussions should be included to provide a more comprehensive explanation of why this operation is meaningful and how it improves efficiency.\n- To enhance clarity, it is suggested that the authors highlight the second-best results alongside the primary results in each table. This practice can provide a useful point of reference for readers and facilitate a more comprehensive understanding of the findings.\n- The results of DTW should be included as a baseline for comparison. This can help demonstrate the advancements made in the proposed methodology and provide a clearer context for the contributions of this work."},"confidence":{"value":"5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully."},"questions":{"value":"The major concern is the lack of some claims, as highlighted in the weaknesses."},"rating":{"value":"8: accept, good paper"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699635927984,"tcdate":1698636165645,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission39/Reviewer_pP85"],"signatures":["ICLR.cc/2024/Conference/Submission39/Reviewer_pP85"],"forum":"9Cu8MRmhq2","number":2,"license":"CC BY 4.0","cdate":1698636165645,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission39/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699635927984,"domain":"ICLR.cc/2024/Conference","replyto":"9Cu8MRmhq2","id":"tHlhMgjMW5","forumContent":{"venue":{"value":"ICLR 2024 oral"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video-language pre-training","Noisy correspondence"]},"primary_area":{"value":"representation learning for computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Existing video-language studies mainly focus on learning short video clips, leaving long-term temporal dependencies rarely explored due to over-high computational cost of modeling long videos. To address this issue, one feasible solution is learning the correspondence between video clips and captions, which however inevitably encounters the multi-granularity noisy correspondence (MNC) problem. To be specific, MNC refers to the clip-caption misalignment (coarse-grained) and frame-word misalignment (fine-grained), hindering temporal learning and video understanding. In this paper, we propose NOise Robust Temporal Optimal traNsport (Norton) that addresses MNC in a unified optimal transport (OT) framework. In brief, Norton employs video-paragraph and clip-caption contrastive losses to capture long-term dependencies based on OT. To address coarse-grained misalignment in video-paragraph contrast, Norton filters out the irrelevant clips and captions through an alignable prompt bucket and realigns asynchronous clip-caption pairs based on transport distance. To address the fine-grained misalignment, Norton incorporates a soft-maximum operator to identify crucial words and key frames. Additionally, Norton exploits the potential faulty negative samples in clip-caption contrast by rectifying the alignment target with OT assignment to ensure precise temporal modeling. Extensive experiments on video retrieval, videoQA, and action segmentation verify the effectiveness of our method. \nCode is available at https://lin-yijie.github.io/projects/Norton."},"_bibtex":{"value":"@inproceedings{\nlin2024multigranularity,\ntitle={Multi-granularity Correspondence Learning from Long-term Noisy Videos},\nauthor={Yijie Lin and Jie Zhang and Zhenyu Huang and Jia Liu and zujie wen and Xi Peng},\nbooktitle={The Twelfth International Conference on Learning Representations},\nyear={2024},\nurl={https://openreview.net/forum?id=9Cu8MRmhq2}\n}"},"title":{"value":"Multi-granularity Correspondence Learning from Long-term Noisy Videos"},"pdf":{"value":"/pdf/578b0930059c165430921bc67cd65b6a0657e518.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference"},"paperhash":{"value":"lin|multigranularity_correspondence_learning_from_longterm_noisy_videos"},"authorids":{"value":["~Yijie_Lin1","~Jie_Zhang42","~Zhenyu_Huang1","~Jia_Liu4","~zujie_wen1","~Xi_Peng3"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Yijie Lin","Jie Zhang","Zhenyu Huang","Jia Liu","zujie wen","Xi Peng"]}},"version":2},{"content":{"venue":{"value":"IEEE Trans. Circuits Syst. Video Technol. 2024"},"pdf":{"value":"https://ieeexplore.ieee.org/iel7/76/10550083/10329331.pdf"},"venueid":{"value":"dblp.org/journals/TCSV/2024"},"paperhash":{"value":"fu|multimodal_imbalanceaware_gradient_modulation_for_weaklysupervised_audiovisual_video_parsing"},"authorids":{"value":["~Jie_Fu10","https://dblp.org/search/pid/api?q=author:Junyu_Gao_0002:","https://dblp.org/search/pid/api?q=author:Bing-Kun_Bao:","https://dblp.org/search/pid/api?q=author:Changsheng_Xu:"]},"html":{"value":"https://doi.org/10.1109/TCSVT.2023.3337134"},"_bibtex":{"value":"@article{DBLP:journals/tcsv/FuGBX24,\n  author={Jie Fu and Junyu Gao and Bing-Kun Bao and Changsheng Xu},\n  title={Multimodal Imbalance-Aware Gradient Modulation for Weakly-Supervised Audio-Visual Video Parsing},\n  year={2024},\n  month={June},\n  cdate={1717200000000},\n  journal={IEEE Trans. Circuits Syst. Video Technol.},\n  volume={34},\n  number={6},\n  pages={4843-4856},\n  url={https://doi.org/10.1109/TCSVT.2023.3337134}\n}\n"},"abstract":{"value":"Weakly-supervised audio-visual video parsing (WS-AVVP) aims to localize the temporal extents of audio, visual and audio-visual event instances as well as identify the corresponding event categories with only video-level category labels for training. Most previous efforts have been devoted to refining the supervision for each modality or extracting fruitful cross-modality information for more reliable feature learning. None of them have noticed the imbalanced feature learning between different modalities in the task. In this paper, to balance the feature learning processes of different modalities, a dynamic gradient modulation (DGM) mechanism is explored, where a novel and effective metric function is designed to measure the imbalanced feature learning between audio and visual modalities. Furthermore, by going in depth into the principle of traditional WS-AVVP pipelines, two additional challenges are identified: confusing multimodal calculation will hamper the precise measurement of audio-visual imbalanced feature learning, as well as the global supervision provided by video-level labels can not provide explicit guidance for robust semantic feature learning in each action subspace. To cope with the above issues, the modality-separated decision unit (MSDU) and semantic-aware feature extractor (SAFE) are designed for precise measurement of imbalanced feature learning and unambiguous semantic-aware feature extraction separately. Comprehensive experiments are conducted on public benchmarks and the corresponding experimental results demonstrate the effectiveness of our proposed method."},"title":{"value":"Multimodal Imbalance-Aware Gradient Modulation for Weakly-Supervised Audio-Visual Video Parsing"},"authors":{"value":["Jie Fu","Junyu Gao","Bing-Kun Bao","Changsheng Xu"]}},"tmdate":1744360108131,"pdate":1704067200000,"tcdate":1744360028490,"writers":["~"],"signatures":["~Jie_Fu10"],"forum":"Wvyi760gqx","license":"CC BY-SA 4.0","number":389713,"cdate":1717200000000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1744360108131,"domain":"DBLP.org","id":"Wvyi760gqx","version":2},{"content":{"venue":{"value":"ECCV (86) 2024"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-031-73016-0_20.pdf"},"venueid":{"value":"dblp.org/conf/ECCV/2024"},"paperhash":{"value":"weng|fast_diffusionbased_counterfactuals_for_shortcut_removal_and_generation"},"authorids":{"value":["~Nina_Weng1","~Paraskevas_Pegios1","~Eike_Petersen1","https://dblp.org/search/pid/api?q=author:Aasa_Feragen:","https://dblp.org/search/pid/api?q=author:Siavash_Arjomand_Bigdeli:"]},"html":{"value":"https://doi.org/10.1007/978-3-031-73016-0_20"},"_bibtex":{"value":"@inproceedings{DBLP:conf/eccv/WengPPFB24,\n  author={Nina Weng and Paraskevas Pegios and Eike Petersen and Aasa Feragen and Siavash Arjomand Bigdeli},\n  title={Fast Diffusion-Based Counterfactuals for Shortcut Removal and Generation},\n  year={2024},\n  cdate={1704067200000},\n  pages={338-357},\n  url={https://doi.org/10.1007/978-3-031-73016-0_20},\n  booktitle={ECCV (86)},\n  crossref={conf/eccv/2024-86}\n}\n"},"abstract":{"value":"Shortcut learning is when a model – e.g. a cardiac disease classifier – exploits correlations between the target label and a spurious shortcut feature, e.g. a pacemaker, to predict the target label based on the shortcut rather than real discriminative features. This is common in medical imaging, where treatment and clinical annotations correlate with disease labels, making them easy shortcuts to predict disease. We propose a novel detection and quantification of the impact of potential shortcut features via a fast diffusion-based counterfactual image generation that can synthetically remove or add shortcuts. Via a novel self-optimized masking scheme we spatially limit the changes made with no extra inference step, encouraging the removal of spatially constrained shortcut features while ensuring that the shortcut-free counterfactuals preserve their remaining image features to a high degree. Using these, we assess how shortcut features influence model predictions."},"title":{"value":"Fast Diffusion-Based Counterfactuals for Shortcut Removal and Generation"},"authors":{"value":["Nina Weng","Paraskevas Pegios","Eike Petersen","Aasa Feragen","Siavash Arjomand Bigdeli"]}},"tmdate":1766126376782,"pdate":1704067200000,"tcdate":1741037308160,"writers":["~"],"signatures":["~Paraskevas_Pegios1"],"forum":"JEkzlbJcGE","license":"CC BY-SA 4.0","number":346811,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1766126376782,"domain":"DBLP.org","id":"JEkzlbJcGE","version":2},{"content":{"venue":{"value":"IJCAI 2025"},"pdf":{"value":"https://www.ijcai.org/proceedings/2025/1133.pdf"},"venueid":{"value":"dblp.org/conf/IJCAI/2025"},"paperhash":{"value":"wang|hallucinationaware_prompt_optimization_for_texttovideo_synthesis"},"authorids":{"value":["","","","~Lianwen_Jin1"]},"html":{"value":"https://doi.org/10.24963/ijcai.2025/1133"},"_bibtex":{"value":"@inproceedings{DBLP:conf/ijcai/Wang0HJ25,\n  author={Jiapeng Wang and Chengyu Wang and Jun Huang and Lianwen Jin},\n  title={Hallucination-Aware Prompt Optimization for Text-to-Video Synthesis},\n  year={2025},\n  cdate={1735689600000},\n  pages={10198-10206},\n  url={https://doi.org/10.24963/ijcai.2025/1133},\n  booktitle={IJCAI},\n  crossref={conf/ijcai/2025}\n}\n"},"abstract":{"value":"The rapid advancements in AI-generated content (AIGC) have led to extensive research and application of deep text-to-video (T2V) synthesis models, such as OpenAI's Sora. These models typically rely on high-quality prompt-video pairs and detailed text prompts for model training in order to produce high-quality videos. To boost the effectiveness of Sora-like T2V models, we introduce VidPrompter, an innovative large multi-modal model supporting T2V applications with three key functionalities: (1) generating detailed prompts from raw videos, (2) enhancing prompts from videos grounded with short descriptions, and (3) refining simple user-provided prompts to elevate T2V video quality. We train VidPrompter using a hybrid multi-task paradigm and propose the hallucination-aware direct preference optimization (HDPO) technique to improve the multi-modal, multi-task prompt optimization process. Experiments on various tasks show our method surpasses strong baselines and other competitors."},"title":{"value":"Hallucination-Aware Prompt Optimization for Text-to-Video Synthesis"},"authors":{"value":["Jiapeng Wang","Chengyu Wang","Jun Huang","Lianwen Jin"]}},"tmdate":1772674727985,"pdate":1767139200000,"externalIds":["dblp:conf/ijcai/Wang0HJ25"],"tcdate":1772674716847,"writers":["~"],"signatures":["~Lianwen_Jin1"],"forum":"9YXznndGlZ","license":"CC BY-SA 4.0","number":848144,"cdate":1735689600000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1772674727985,"domain":"DBLP.org","id":"9YXznndGlZ","version":2},{"content":{"submission_length":{"value":"Regular submission (no more than 12 pages of main content)"},"venue":{"value":"Accepted by TMLR"},"code":{"value":"https://github.com/sherwinbahmani/3dvideogeneration/"},"supplementary_material":{"value":"/attachment/d2bcac778f56a086ed43b12e00f5f98b3c627238.zip"},"abstract":{"value":"Generative models have emerged as an essential building block for many image synthesis and editing tasks. Recent advances in this field have also enabled high-quality 3D or video content to be generated that exhibits either multi-view or temporal consistency. With our work, we explore 4D generative adversarial networks (GANs) that learn unconditional generation of 3D-aware videos. By combining neural implicit representations with time-aware discriminator, we develop a GAN framework that synthesizes 3D video supervised only with monocular videos. We show that our method learns a rich embedding of decomposable 3D structures and motions that enables new visual effects of spatio-temporal renderings while producing imagery with quality comparable to that of existing 3D or video GANs."},"_bibtex":{"value":"@article{\nbahmani2023daware,\ntitle={3D-Aware Video Generation},\nauthor={Sherwin Bahmani and Jeong Joon Park and Despoina Paschalidou and Hao Tang and Gordon Wetzstein and Leonidas Guibas and Luc Van Gool and Radu Timofte},\njournal={Transactions on Machine Learning Research},\nissn={2835-8856},\nyear={2023},\nurl={https://openreview.net/forum?id=SwlfyDq6B3},\nnote={}\n}"},"title":{"value":"3D-Aware Video Generation"},"certifications":{"value":[]},"changes_since_last_submission":{"value":"- Acknowledged limitation of 2D upsampling layers decreasing 3D consistency\n- Added geometry evaluation results\n- Visualized TaiChi depth maps\n- Clarified hyperparameter selection\n- Clarified inductive bias for motion/content disentanglement\n- Clarified missing ACD and CPBD metrics\n- Updated Fig. 2\n- Added missing references\n- Fixed equations\n- Fixed citations"},"license":{"value":"Creative Commons Attribution 4.0 International (CC BY 4.0)"},"pdf":{"value":"/pdf/b92e981a4818749888cd238c1745546ed0931fbf.pdf"},"venueid":{"value":"TMLR"},"paperhash":{"value":"bahmani|3daware_video_generation"},"authorids":{"value":["~Sherwin_Bahmani1","~Jeong_Joon_Park2","~Despoina_Paschalidou1","~Hao_Tang6","~Gordon_Wetzstein3","~Leonidas_Guibas1","~Luc_Van_Gool1","~Radu_Timofte1"]},"assigned_action_editor":{"value":"~Mathieu_Salzmann1"},"authors":{"value":["Sherwin Bahmani","Jeong Joon Park","Despoina Paschalidou","Hao Tang","Gordon Wetzstein","Leonidas Guibas","Luc Van Gool","Radu Timofte"]}},"tmdate":1726599538528,"pdate":1686120177765,"tcdate":1674281679265,"writers":["TMLR"],"signatures":["TMLR/Paper791/Authors"],"forum":"SwlfyDq6B3","number":791,"license":"CC BY 4.0","cdate":1674281679265,"mdate":1726599538528,"readers":["everyone"],"invitations":["TMLR/-/Submission","TMLR/Paper791/-/Revision","TMLR/-/Under_Review","TMLR/-/Edit","TMLR/Paper791/-/Camera_Ready_Revision","TMLR/-/Accepted"],"odate":1674759363914,"domain":"TMLR","id":"SwlfyDq6B3","version":2},{"content":{"venue":{"value":"CoRR 2024"},"pdf":{"value":"https://arxiv.org/pdf/2410.13343v1"},"venueid":{"value":"dblp.org/journals/CORR/2024"},"paperhash":{"value":"yuan|do_llms_overcome_shortcut_learning_an_evaluation_of_shortcut_challenges_in_large_language_models"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Yu_Yuan:","","~Kai_Zhang12","","~Qi_Liu3"]},"html":{"value":"https://doi.org/10.48550/arXiv.2410.13343"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2410-13343,\n  publtype={informal},\n  author={Yu Yuan and Lili Zhao and Kai Zhang and Guangting Zheng and Qi Liu},\n  title={Do LLMs Overcome Shortcut Learning? An Evaluation of Shortcut Challenges in Large Language Models},\n  year={2024},\n  cdate={1704067200000},\n  journal={CoRR},\n  volume={abs/2410.13343},\n  url={https://doi.org/10.48550/arXiv.2410.13343}\n}\n"},"abstract":{"value":"Large Language Models (LLMs) have shown remarkable capabilities in various natural language processing tasks. However, LLMs may rely on dataset biases as shortcuts for prediction, which can significantly impair their robustness and generalization capabilities. This paper presents Shortcut Suite, a comprehensive test suite designed to evaluate the impact of shortcuts on LLMs' performance, incorporating six shortcut types, five evaluation metrics, and four prompting strategies. Our extensive experiments yield several key findings: 1) LLMs demonstrate varying reliance on shortcuts for downstream tasks, significantly impairing their performance. 2) Larger LLMs are more likely to utilize shortcuts under zero-shot and few-shot in-context learning prompts. 3) Chain-of-thought prompting notably reduces shortcut reliance and outperforms other prompting strategies, while few-shot prompts generally underperform compared to zero-shot prompts. 4) LLMs often exhibit overconfidence in their predictions, especially when dealing with datasets that contain shortcuts. 5) LLMs generally have a lower explanation quality in shortcut-laden datasets, with errors falling into three types: distraction, disguised comprehension, and logical fallacy. Our findings offer new insights for evaluating robustness and generalization in LLMs and suggest potential directions for mitigating the reliance on shortcuts. The code is available at \\url {https://github.com/yyhappier/ShortcutSuite.git}."},"title":{"value":"Do LLMs Overcome Shortcut Learning? An Evaluation of Shortcut Challenges in Large Language Models"},"authors":{"value":["Yu Yuan","Lili Zhao","Kai Zhang","Guangting Zheng","Qi Liu"]}},"tmdate":1785325782181,"pdate":1704067200000,"externalIds":["dblp:journals/corr/abs-2410-13343"],"tcdate":1768439538195,"writers":["~"],"signatures":["~Kai_Zhang12"],"forum":"xuxNVx6D0m","license":"CC BY-SA 4.0","number":753283,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1785325782181,"domain":"DBLP.org","id":"xuxNVx6D0m","version":2},{"content":{"summary":{"value":"This paper introduces **ReWatch-R1**, which tackles the challenge of video reasoning from two aspects: \n1. Data. A large synthetic dataset called **ReWatch** is collected, which includes detailed video captions, challenging question-answer pairs, and high-quality, video-grounded reasoning traces (COT data). These are generated using a multi-agent ReAct system. \n2. Model. Authors post-train QwenVL by SFT and GRPO. They introduce a new Observation & Reasoning (O&R) reward, which evaluates the accuracy of video observations and the validity of reasoning process. \nReWatch-R1 achieves promising performance compared with other 7B models.\nAnalysis shows that high-quality reasoning data is crucial for RL and RL on \"thinking\" mode improves the reasoning efficiency."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. Equation (20): the design of non-format rewards lacks motivation. For example, why not use the (weighted) summation of all rewards. The design has no explanation or experimental results.\n2. RL on 7B model shows promising improvements. However, the performance still lags behind the larger 32B model. Therefore, whether the same RL on 32B leads to improvement is questionable.\n\nSome typos:\n1. L255: \"a novel O&R reward mechanism we propos\" -> \"propose\"\n2. L771: \"Tabale 4\" -> \"Table 4\""},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The proposed multi-agent COT data synthesis pipeline is scalable to curate large-scale video-grounding reasoning data.\n2. The reward design shows the emphasis on explicit observation and reasoning is beneficial for video reasoning.\n3. The data and model design achieves state-of-the-art performance on video reasoning and understanding benchmarks in 7B-scale models. The extensive analysis shows insights on the role and importance of SFT and RL."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The information source for data synthesis is semantic segmentation and detailed video description from Gemini. The accuracy of this multi-step hierarchical captioning is not validated. Therefore there is no direct quality assessment of the synthetic data.\n2. All results are based on a 7B model. The benefits of high-quality COT data and O&R reward mechanism are not validated on larger-scale models."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762927833401,"tcdate":1761989929871,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission18045/Reviewer_UAaJ"],"signatures":["ICLR.cc/2026/Conference/Submission18045/Reviewer_UAaJ"],"forum":"xindJJLSr1","number":4,"license":"CC BY 4.0","cdate":1761989929871,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission18045/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762927833401,"domain":"ICLR.cc/2026/Conference","replyto":"xindJJLSr1","id":"kFWMHLn917","forumContent":{"TLDR":{"value":"We introduce an agent-based pipeline to synthesize a high-quality video reasoning dataset (ReWatch) and a novel reinforcement learning reward (O&R) to train LVLMs, achieving state-of-the-art performance."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Video Reasoning","Large Vision-Language Models (LVLMs)","Agentic Data Synthesis","Multi-Agent ReAct","Reinforcement Learning with Verifiable Reward (RLVR)","Chain-of-Thought (CoT)"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"While Reinforcement Learning with Verifiable Reward (RLVR) significantly advances image reasoning in Large Vision-Language Models (LVLMs), its application to complex video reasoning remains underdeveloped. This gap stems primarily from a critical data bottleneck: existing datasets lack the challenging, multi-hop questions and high-quality, video-grounded Chain-of-Thought (CoT) data necessary to effectively bootstrap RLVR. To address this, we introduce ReWatch, a large-scale dataset built to foster advanced video reasoning. We propose a novel multi-stage synthesis pipeline to synthesize its three components: ReWatch-Caption, ReWatch-QA, and ReWatch-CoT. A core innovation is our Multi-Agent ReAct framework for CoT synthesis, which simulates a human-like \"re-watching\" process to generate video-grounded reasoning traces by explicitly modeling information retrieval and verification. Building on this dataset, we develop ReWatch-R1 by post-training a strong baseline LVLM with Supervised Fine-Tuning (SFT) and our RLVR framework. This framework incorporates a novel Observation \\& Reasoning (O\\&R) reward mechanism that evaluates both the final answer's correctness and the reasoning's alignment with video content, directly penalizing hallucination. Our experiments show that ReWatch-R1 achieves state-of-the-art average performance on five challenging video reasoning benchmarks."},"_bibtex":{"value":"@inproceedings{\nzhang2026rewatchr,\ntitle={ReWatch-R1: Boosting Complex Video Reasoning in Large Vision-Language Models through Agentic Data Synthesis},\nauthor={Congzhi Zhang and Zhibin Wang and Yinchao Ma and Jiawei Peng and Yihan Wang and Qiang Zhou and Jun Song and Bo Zheng},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=xindJJLSr1}\n}"},"title":{"value":"ReWatch-R1: Boosting Complex Video Reasoning in Large Vision-Language Models through Agentic Data Synthesis"},"pdf":{"value":"/pdf/188764f9319a82af607db76b20fa18a29c9fcf00.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"zhang|rewatchr1_boosting_complex_video_reasoning_in_large_visionlanguage_models_through_agentic_data_synthesis"},"authorids":{"value":["~Congzhi_Zhang1","~Zhibin_Wang1","~Yinchao_Ma2","~Jiawei_Peng2","~Yihan_Wang29","~Qiang_Zhou8","~Jun_Song5","~Bo_Zheng5"]},"authors":{"value":["Congzhi Zhang","Zhibin Wang","Yinchao Ma","Jiawei Peng","Yihan Wang","Qiang Zhou","Jun Song","Bo Zheng"]}},"version":2},{"content":{"venue":{"value":"PRICAI (4) 2024"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-981-96-0125-7_16.pdf"},"venueid":{"value":"dblp.org/conf/PRICAI/2024"},"paperhash":{"value":"zheng|rsanet_relationshipaware_symmetric_alignment_network_for_finegrained_videotext_retrieval"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Min_Zheng:","~Chunpeng_Wu1","https://dblp.org/search/pid/api?q=author:Ke_Chang:","https://dblp.org/search/pid/api?q=author:Weiwei_Liu:","https://dblp.org/search/pid/api?q=author:Dongsheng_Wang:","https://dblp.org/search/pid/api?q=author:Yue_Wang:","https://dblp.org/search/pid/api?q=author:Qinghe_Ye:","https://dblp.org/search/pid/api?q=author:Zhi_Yu:"]},"html":{"value":"https://doi.org/10.1007/978-981-96-0125-7_16"},"_bibtex":{"value":"@inproceedings{DBLP:conf/pricai/ZhengWCLWWYY24,\n  author={Min Zheng and Chunpeng Wu and Ke Chang and Weiwei Liu and Dongsheng Wang and Yue Wang and Qinghe Ye and Zhi Yu},\n  title={RSANet: Relationship-Aware Symmetric Alignment Network for Fine-Grained Video-Text Retrieval},\n  year={2024},\n  cdate={1704067200000},\n  pages={190-201},\n  url={https://doi.org/10.1007/978-981-96-0125-7_16},\n  booktitle={PRICAI (4)},\n  crossref={conf/pricai/2024-4}\n}\n"},"abstract":{"value":"Video-text retrieval aims to retrieve the most similar samples from the database in another modality, given a query of one modality (e.g. text, video). The primary challenge lies in capturing the fine-grained relationship, including individual objects features and interactions among diverse objects. Existing approaches predominantly prioritize text parsing independently, neglecting video parsing and thereby not effectively capturing the complementary relations across video-text pairs. In this paper, we introduce a novel Relationship-aware Symmetric Alignment Network (RSANet) for Fine-grained Video-Text Retrieval. Specifically, we first adaptively recalibrate frame-wise features of videos and extract relationship-aware fine-grained features from the recalibrated frames. Subsequently, a tailored heterogeneous graph convolution network is formulated to encode each recalibrated frame. Correspondingly, we parse texts into relationship-aware nodes and employ a pre-trained model to extracts the contextual features of nodes. As a result, relationship-aware cross-modality features can be obtained, which enables the alignment in a more plausible fine-grained manner. In addition, a negative sample enhanced ranking loss is proposed to optimize the RSANet, which promotes the model output with a larger inter-class variation and a smaller intra-class variation. Extensive experiments on three public datasets, namely MSR-VTT, VATEX, and PKU FG-XMedia, show the effectiveness of RSANet surpasses state-of-the-art methods."},"title":{"value":"RSANet: Relationship-Aware Symmetric Alignment Network for Fine-Grained Video-Text Retrieval"},"authors":{"value":["Min Zheng","Chunpeng Wu","Ke Chang","Weiwei Liu","Dongsheng Wang","Yue Wang","Qinghe Ye","Zhi Yu"]}},"tmdate":1768974724041,"pdate":1735603200000,"externalIds":["dblp:conf/pricai/ZhengWCLWWYY24"],"tcdate":1768974719426,"writers":["~"],"signatures":["~Chunpeng_Wu1"],"forum":"LU52s4zbR5","license":"CC BY-SA 4.0","number":778594,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1768974724041,"domain":"DBLP.org","id":"LU52s4zbR5","version":2},{"content":{"summary":{"value":"This paper presents a novel method for camera-controllable multi-view video generation. Leveraging a pre-trained I2V model, the method introduces camera-conditioned Plücker coordinates in the latent space for conditional control. Additionally, it extends spatial attention into cross-view attention and promotes 1D temporal attention into cross-frame 3D attention. These enhancements have been evaluated and found beneficial for achieving cross-view and cross-frame consistency, as well as precise camera control.\n\nTo address the challenge of limited multi-view video data, an effective joint training strategy is proposed, utilizing a curated mixture of static, monocular dynamic, and multi-view dynamic videos. Experiments demonstrate the superiority of the proposed method over existing competitors in terms of multi-view consistency and camera control precision. However, it is worth noting that most of the demonstrated examples pertain to specific domains, involving minimal dynamics and simple backgrounds, which raises questions about the method's performance in more general scenarios."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. In Table 3, I'm curious why the impact of \"w/o cross-view\" is even smaller than that of \"w/o cross-frame.\" Intuitively, cross-view should be an essential component for multi-view setups (without it, it's challenging to maintain multi-view consistency and motion synchronization), whereas cross-frame primarily enhances the accuracy. Does this result imply that most of the outcomes are actually static, and similar results could be achieved through independent monocular camera control?\n\n2. Is your model fully tuned? What is the parameter count?\n\n3. Are the results of the ablation study in Table 3 based on monocular camera control or multi-view camera control? It would be advisable to evaluate these two settings separately."},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. From the qualitative and quantitative results, it is evident that this method indeed offers superior camera trajectory control accuracy. This advantage is attributed to the benefits brought by 3D temporal attention (maybe also fulltuning) and the carefully collected training dataset.\n\n2. To overcome the scarcity of multi-view video dataset, a curated mixture of existing available static, monocular dynamic, and multi-view dynamic videos is exploited via joint training scheme. This provides a solution for addressing data issues in future research on this task.\n\n3. The experimental evaluation and comparison are comprehensive, evaluating the geometric accuracy of camera control from multiple perspectives. It demonstrates noticeable superiority of the proposed method."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The technical novelty is mild. Adopting plucker embedding for camera control is not new. Extending spatial attention for cross-view attention is also straightforward. Anyway, it is inspiring that 3D attention instead of 1D temporal attention is highly desired to help camera control accuracy, and this is well validated by multiple experiments.\n\n\n2. One of the major weakness is the less convincing results. \nThe most important aspects of multi-view video generation are motion synchronization and geometric consistency. However, most of the examples presented in the paper involve minimal object movement (either a single object with a simple, weakly moving background or entirely static scenes), and the camera angles do not vary significantly, with most shots taken from the same side of the object. This makes it difficult to evaluate the proposed method's capabilities in this task.\n\nFor example, I would like to see examples like this: \"on a street with vehicles and pedestrians, a horse is running, with one camera capturing a global view from behind the horse and another camera tracking from the side\". Showing more examples like this would be more convincing and demonstrate the value of controllable multi-view video generation.\n\n\n3. In my understanding, making use of monocular videos could promote the model's generalization. But most of the cases in the paper, are close to the style of Objeverse dataset. Is it because the fulltuning destroyed the capabiliy of SVD in generating realistic style and scene? If so, this should be claimed as a limitation, especially compared to those adapter plugin methods like MotionCtrl and CameraCtrl.\n\n\n4. From the results in Fig. 8, it appears that while the introduction of monocular video increases generalization ability, such as maintaining the content of the input images, it also affects the accuracy of multi-view camera control. In the second row of (b), the shooting direction for the fish and the fox did not turn as expected."}},"nonreaders":[],"tmdate":1731427558639,"tcdate":1730276207130,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission2215/Reviewer_FTwe"],"signatures":["ICLR.cc/2025/Conference/Submission2215/Reviewer_FTwe"],"forum":"sNntRFmn72","number":2,"license":"CC BY 4.0","cdate":1730276207130,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission2215/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427558639,"domain":"ICLR.cc/2025/Conference","replyto":"sNntRFmn72","id":"XQSucH9s42","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video Generation"]},"supplementary_material":{"value":"/attachment/59d1a5baedc574ecafa96bd45b05b5c57e7e564e.zip"},"primary_area":{"value":"generative models"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"In recent years there have been remarkable breakthroughs in image-to-video generation.\nHowever, the 3D consistency and camera controllability of generated frames have remained unsolved. Recent studies have attempted to incorporate camera control into the generation process, but their results are often limited to simple trajectories or lack the ability to generate consistent videos from multiple distinct camera paths for the same scene. To address these limitations, we introduce **Cavia**, a novel framework for camera-controllable, multi-view video generation, capable of converting an input image into multiple spatiotemporally consistent videos. Our framework extends the spatial and temporal attention modules into view-integrated attention modules, improving both viewpoint and temporal consistency. This flexible design allows for joint training with diverse curated data sources, including scene-level static videos, object-level synthetic multi-view dynamic videos, and real-world monocular dynamic videos. To our best knowledge, Cavia is the first of its kind that allows the user to precisely specify camera motion while obtaining object motion. To the best of our knowledge, Cavia is the first framework that enables users to generate multiple videos of the same scene with precise control over camera motion, while simultaneously preserving object motion. Extensive experiments demonstrate that Cavia surpasses state-of-the-art methods in terms of geometric consistency and perceptual quality."},"_bibtex":{"value":"@misc{\nxu2025cavia,\ntitle={Cavia: Camera-controllable Multi-view Video Diffusion with View-Integrated Attention},\nauthor={Dejia Xu and Yifan Jiang and Chen Huang and Liangchen Song and Thorsten Gernoth and Liangliang Cao and Zhangyang Wang and Hao Tang},\nyear={2025},\nurl={https://openreview.net/forum?id=sNntRFmn72}\n}"},"title":{"value":"Cavia: Camera-controllable Multi-view Video Diffusion with View-Integrated Attention"},"pdf":{"value":"/pdf/b8d2c40ed3d51a3f5f683b69537e25d39c0a948b.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"xu|cavia_cameracontrollable_multiview_video_diffusion_with_viewintegrated_attention"},"authorids":{"value":["~Dejia_Xu1","~Yifan_Jiang2","~Chen_Huang1","~Liangchen_Song1","~Thorsten_Gernoth1","~Liangliang_Cao1","~Zhangyang_Wang1","~Hao_Tang16"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Dejia Xu","Yifan Jiang","Chen Huang","Liangchen Song","Thorsten Gernoth","Liangliang Cao","Zhangyang Wang","Hao Tang"]}},"version":2},{"content":{"summary":{"value":"This paper proposes Slerp+, a method for zero-shot Composed Visual Retrieval (CVR). CVR unifies composed image retrieval and composed video retrieval into a single framework. The proposed method uses a vision-language model (based on BLIP) trained on both image-caption and video-caption pairs concurrently. It then utilizes the Spherical Linear Interpolation (i.e. Slerp) to compose visual and textual embeddings during inference. The methods is evaluated on popular composed image retrieval and composed video retrieval  datasets and introduce a new, more benchmark called Activitynet-CoVR for the video retrieval part. Results show that the proposed method outperforms existing methods in both image and video retrieval tasks."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"- the Slerp fusion seems to be the most important reason of the resulting performance. Comparison with the application of this component to other methods is limited to that of Slerp + TAT. It would be interesting to see other ablation studies with Slerp.\n\n- are the method and Activitynet-CoVR going to be released?"},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":3},"strengths":{"value":"- the proposal to build a unified composed retrieval system with both image video and text is novel.\n\n- the paper is well written, well structured and experiments are well carried out. \n\n- the use of Slerp makes the proposed approach simple and effective, achieving state of the art results on both composed image retrieval and composed video retrieval datasets."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The extension to video is straightforward since the method does not propose any particular strategy that goes beyond taking frames from the videos (e.g. for instance no temporal aspect is explicitly considered in the model).\n\n- the use of Slerp seems to be the most important aspect of the resulting performance. The ablation regarding the use of this component is limited.\n\n- Limited discussion of limitations or weaknesses of the proposed method.\n\nTypos:\n- Check the use of CoVR vs CVR terminology"}},"nonreaders":[],"tmdate":1731427508511,"tcdate":1730655330307,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission1907/Reviewer_abA2"],"signatures":["ICLR.cc/2025/Conference/Submission1907/Reviewer_abA2"],"forum":"YCOVTlMFIG","number":4,"license":"CC BY 4.0","cdate":1730655330307,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission1907/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427508511,"domain":"ICLR.cc/2025/Conference","replyto":"YCOVTlMFIG","id":"9SqsyR2DhZ","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["multi-modal representation learning","composed retrieval"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Zero-shot composed image/video retrieval is a challenging task that involves using a combination of a reference visual input and a relative caption as a query to search for target visual data. Earlier studies have treated composed image retrieval and composed video retrieval methods separately, potentially neglecting the benefits of integrating image-video-text representation learning.  In this paper, we consolidate these tasks into a single Composed \\emph{Visual} Retrieval (CVR) task, which requires the composition of image and video samples with textual modifications using a unified retrieval model. Our principal insight is that the video modality can be effectively added to existing vision-language pretrained models. When integrated with the Spherical Linear Interpolation (Slerp) method previously proposed for Composed Image Retrieval (CoIR), we found that it results in an effective approach for solving the CVR task, which we called $\\text{Slerp}^{+}$. Extensive experiments demonstrate $\\text{Slerp}^{+}$'s superiority across various composed image and video retrieval benchmarks, including our newly proposed video benchmark. Notably, $\\text{Slerp}^{+}$ mutually enhances image and video retrieval performance over single-modality models, underscoring its potential to transform the field of compositional visual retrieval."},"_bibtex":{"value":"@misc{\njang2024textslerp,\ntitle={\\${\\textbackslash}text\\{Slerp\\}{\\textasciicircum}\\{+\\}\\$: Spherical Linear Interpolation for Unified Compositional Retrieval},\nauthor={Young Kyun Jang and Donghyun Kim and Bo He and Zihang Meng and Ser-Nam Lim},\nyear={2024},\nurl={https://openreview.net/forum?id=YCOVTlMFIG}\n}"},"title":{"value":"$\\text{Slerp}^{+}$: Spherical Linear Interpolation for Unified Compositional Retrieval"},"pdf":{"value":"/pdf/25ed87289cb3c609d20358983e2e562ad83f3404.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"jang|\\textslerp^_spherical_linear_interpolation_for_unified_compositional_retrieval"},"authorids":{"value":["~Young_Kyun_Jang1","~Donghyun_Kim2","~Bo_He1","~Zihang_Meng1","~Ser-Nam_Lim3"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Young Kyun Jang","Donghyun Kim","Bo He","Zihang Meng","Ser-Nam Lim"]}},"version":2},{"content":{"summary":{"value":"The paper proposes a new pipeline for preparing video features for LLMs. Beyond simple keyframe retrieval, it introduces an agentic flow designed to capture temporally ordered events and reconstruct the underlying narrative. Based on the extracted components, the authors employ a CoT prompting strategy to enhance reasoning and improve understanding. Experiments are conducted across several benchmark datasets."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"The numbers reported in Table 1 (for Qwen2.5-VL) show a large discrepancy compared to the original paper. For instance, LVBench should report 45.3 for Qwen2.5-VL-7B. This inconsistency raises concerns about the results’ reliability. Although the relative improvement over the baseline is significant, the absolute performance values are not aligned with prior reports."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The paper attempts to construct scene graphs to decompose video content, which is an interesting idea.\n\n2. The proposed pipeline is reasonable, and the implementation details are concrete and easy to understand.\n\n3. The experiments are comprehensive, covering most mainstream long-video benchmarks currently available."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The performance does not reach state-of-the-art results. For example, it is notably inferior to Video-XL-2 [1]. Additionally, some training-free retrieval methods (e.g., BOLT [2]) are missing from the comparison table, which weakens the technical contribution.\n\n2. The main contribution lies in pipeline design rather than technical innovation. The approach feels closer to a text-based agent framework, so the title’s emphasis on “representation” may be misleading—it seems more like an engineering effort.\n\n\n\n[1] Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification\n\n[2] BOLT: Boost Large Vision-Language Model Without Training for Long-form Video Understanding"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762920278515,"tcdate":1761837639692,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission8370/Reviewer_6jm1"],"signatures":["ICLR.cc/2026/Conference/Submission8370/Reviewer_6jm1"],"forum":"aLQsPnVNnk","number":3,"license":"CC BY 4.0","cdate":1761837639692,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission8370/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762920278515,"domain":"ICLR.cc/2026/Conference","replyto":"aLQsPnVNnk","id":"KPpgSMypbV","forumContent":{"TLDR":{"value":"We introduce Video-EM, a training‑free framework that treats long video question answering as an episodic memory retrieval-and‑reasoning problem inspired by human cognitive psychology."},"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Multi-modal Vision","Video Understanding"]},"primary_area":{"value":"foundation or frontier models, including LLMs"},"abstract":{"value":"Video Large Language Models (Video-LLMs) excel at general video understanding but struggle with long-form videos due to context-window limits. Consequently, recent approaches focus on keyframe retrieval, condensing lengthy videos into a small set of informative frames. Despite their practicality, these methods simplify the problem to static text-image matching, overlooking spatio-temporal relationships crucial for capturing scene transitions and contextual continuity, and may yield redundant keyframes with limited information, diluting salient cues essential for accurate video question answering. To address these limitations, we introduce Video-EM, a training-free framework inspired by the principles of human episodic memory, designed to facilitate robust and contextually grounded reasoning. Rather than treating keyframes as isolated visual entities, Video-EM explicitly models them as temporally ordered episodic events, capturing both spatial relationships and temporal dynamics necessary for accurately reconstructing the underlying narrative. Furthermore, the framework leverages chain-of-thought (CoT) thinking with LLMs to iteratively identify a minimal yet highly informative subset of episodic memories, enabling efficient and accurate question answering by Video-LLMs. Extensive evaluations on multiple mainstream long-video benchmarks demonstrate the superiority of Video-EM, which achieves highly competitive results while using fewer frames."},"_bibtex":{"value":"@misc{\nwang2026episodic,\ntitle={Episodic Memory Representation for Long Video Understanding},\nauthor={Yun wang and Long Zhang and Jingren Liu and Jiaqi Yan and Zhanjie Zhang and Jiahao Zheng and Ao Ma and Xun Yang and Dapeng Wu and Xiangyu Chen and Xuelong Li},\nyear={2026},\nurl={https://openreview.net/forum?id=aLQsPnVNnk}\n}"},"title":{"value":"Episodic Memory Representation for Long Video Understanding"},"pdf":{"value":"/pdf/f020aad44e91a19ec0c3c30fd66c852f9da5374f.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"wang|episodic_memory_representation_for_long_video_understanding"},"authorids":{"value":["~Yun_wang13","~Long_Zhang4","~Jingren_Liu1","~Jiaqi_Yan3","~Zhanjie_Zhang2","~Jiahao_Zheng1","~Ao_Ma2","~Xun_Yang1","~Dapeng_Wu1","~Xiangyu_Chen5","~Xuelong_Li2"]},"authors":{"value":["Yun wang","Long Zhang","Jingren Liu","Jiaqi Yan","Zhanjie Zhang","Jiahao Zheng","Ao Ma","Xun Yang","Dapeng Wu","Xiangyu Chen","Xuelong Li"]}},"version":2},{"content":{"summary":{"value":"This paper presents AccidentBench, a large-scale benchmark designed to evaluate video reasoning, anticipation, and explanation in the context of traffic accidents. The dataset comprises over 32,000 video clips collected from dashcams, driving simulations, and surveillance footage, covering 26 types of traffic incidents.\n\nEach video is annotated with multi-level multimodal information — including accident type, temporal boundaries, causal descriptions, predicted outcomes, and responsibility attribution — forming a comprehensive platform for both visual prediction and language-based causal reasoning.\n\nThe benchmark defines three core tasks:\n1/ Accident Anticipation — predicting if and when an accident will occur;\n2/ Causal Reasoning — explaining why the accident happens;\n3/ Responsibility Attribution — identifying who is responsible.\n\nExtensive experiments evaluate a range of baselines, from video transformers (VideoMAE, TimeSformer) to multimodal large language models (Video-LLaVA, GPT-4V, Qwen2-VL, Gemini-1.5-Pro). The results show that while foundation models achieve strong descriptive capability, they still struggle with temporal alignment, causal inference, and grounded reasoning — highlighting the need for specialized architectures for video understanding in safety-critical domains."},"soundness":{"value":4},"confidence":{"value":5},"questions":{"value":"Please see the weakness section."},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":3},"strengths":{"value":"1/ AccidentBench fills a clear gap in multimodal research by unifying video anticipation, causal reasoning, and attribution tasks under a single dataset. The data design — with dense annotations and temporally aligned textual explanations — is impressive and likely to become a valuable community resource.\n\n2/ The use of real-world dashcam footage alongside simulation and surveillance videos improves domain diversity. Covering 26 accident categories ensures the benchmark captures a wide spectrum of risky interactions and visual conditions.\n\n3/ The benchmark tasks are well defined, with metrics that encourage both early anticipation (Time-to-Accident) and high-quality explanations (BLEU, CIDEr, human consistency). This structured task decomposition provides clarity and reproducibility.\n\n4/ AccidentBench provides a foundation for developing causal-aware video reasoning models and could influence areas like embodied AI, self-driving perception, and video safety analysis."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1/ The related work section overlooks several recent 2025 works on accident video reasoning, such as [1]. Further literature review shall be encouraged.\n\n2/ The paper could better discuss potential biases (e.g., geographic or weather distribution) and provide details on how labeling consistency was ensured across different video domains.\n\n3/ Most results focus on quantitative metrics; there is limited qualitative analysis on failure cases, particularly where LLMs produce plausible but factually wrong explanations.\n\n[1] AVD2: Accident Video Diffusion for Accident Video Description (ICRA 2025)"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762924067640,"tcdate":1761904784793,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission13442/Reviewer_5oHn"],"signatures":["ICLR.cc/2026/Conference/Submission13442/Reviewer_5oHn"],"forum":"f5lIozG83H","number":3,"license":"CC BY 4.0","cdate":1761904784793,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission13442/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762924067640,"domain":"ICLR.cc/2026/Conference","replyto":"f5lIozG83H","id":"tJTwZ9En4V","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Multimodal Understanding and Reasoning","Large-Scale Dataset","Traffic Accident","Land Space","Airplane Navigation","Ship Motion"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Rapid advances in multimodal models demand benchmarks that rigorously evaluate understanding and reasoning in safety-critical, dynamic real-world settings. We present AccidentBench, a large-scale benchmark that combines vehicle accident scenarios with Beyond domains, safety-critical settings in air and water that emphasize spatial and temporal reasoning (e.g., navigation, orientation, multi-vehicle motion). The benchmark contains approximately 2000 videos and over 19000 human-annotated question--answer pairs spanning multiple video lengths (short/medium/long) and difficulty levels (easy/medium/hard). Tasks systematically probe core capabilities: temporal, spatial, and intent understanding and reasoning.  By unifying accident-centric traffic scenes with broader safety-critical scenarios in air and water, AccidentBench offers a comprehensive, physically grounded testbed for evaluating models under real-world variability. Evaluations of state-of-the-art models (e.g., Gemini-2.5 Pro and GPT-5) show that even the strongest models achieve only about 18% accuracy on the hardest tasks and longest videos, revealing substantial gaps in real-world temporal, spatial, and intent reasoning. AccidentBench is designed to expose these critical gaps and drive the development of multimodal models that are safer, more robust, and better aligned with real-world safety-critical challenges. The code and dataset are available at: http://accident-bench.site"},"_bibtex":{"value":"@misc{\ngu2026accidentbench,\ntitle={AccidentBench: Benchmarking Multimodal Understanding and Reasoning in Vehicle Accidents and Beyond},\nauthor={Shangding Gu and Xiaohan Wang and Donghao Ying and Haoyu Zhao and Runing Yang and Ming Jin and Boyi Li and Marco Pavone and Serena Yeung-Levy and Jun Wang and Dawn Song and Costas Spanos},\nyear={2026},\nurl={https://openreview.net/forum?id=f5lIozG83H}\n}"},"title":{"value":"AccidentBench: Benchmarking Multimodal Understanding and Reasoning in Vehicle Accidents and Beyond"},"pdf":{"value":"/pdf/f92d19d32838720da5d3cb35e4ad61301871e9ba.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"gu|accidentbench_benchmarking_multimodal_understanding_and_reasoning_in_vehicle_accidents_and_beyond"},"authorids":{"value":["~Shangding_Gu1","~Xiaohan_Wang2","~Donghao_Ying1","~Haoyu_Zhao3","~Runing_Yang1","~Ming_Jin2","~Boyi_Li1","~Marco_Pavone1","~Serena_Yeung-Levy1","~Jun_Wang2","~Dawn_Song1","~Costas_Spanos1"]},"authors":{"value":["Shangding Gu","Xiaohan Wang","Donghao Ying","Haoyu Zhao","Runing Yang","Ming Jin","Boyi Li","Marco Pavone","Serena Yeung-Levy","Jun Wang","Dawn Song","Costas Spanos"]}},"version":2},{"content":{"summary":{"value":"1.This paper proposes AURORACAP, a video captioner based on a large multimodal model. AURORACAP follows a simple architecture design without additional parameters for temporal modeling. To address the overhead caused by lengthy video sequences, it implements a token merging strategy, reducing the number of input visual tokens with little performance loss. AURORACAP shows superior performance on various video and image captioning benchmarks. \n2.Existing video caption benchmarks only include simple descriptions, so the authors develop VDC, a video detailed captioning benchmark with over one thousand carefully annotated structured captions. \n3.The authors also propose a new LLM-assisted metric VDCSCORE for better evaluation, which transforms long caption evaluation into multiple short question-answer pairs."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"the questions are listed above."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1.the proposed AURORACAP shows good performance. Using bipartite soft matching method to merge similar token is novel and show potential on token reduction.  \n2.VDC with detailed video caption is good contribution for video captioning research community, and authors provide the detailed description of dataset curation. \n3.the proposed metric VDCscore is a good improvement for captioning task evaluation with novelty. \n4.Paper is in good writing for understanding."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1.AURORACAP: there is no description of the AURORACAP's 1.3M pretraining data. We could't conclude the reseasons for the good peroformance of AUROREACAP. Could authors provide details on:\n\na. The sources and types of data included in the 1.3M pretraining dataset\n\nb. Any preprocessing or filtering steps applied to this data\n\nc. How this pretraining data compares to datasets used by other models? \n\nThis information would help readers better understand the factors contributing to AURORACAP's performance.\n\n2.VDCSCORE: the potential stability issue of the metric.\n\na. Is there a fixed number of question-answer pairs generated for each caption, or does it vary?\n\nb. If it varies, how do the authors ensure consistency in the metric across captions of different lengths or complexities?\n\nc. Did the authors conduct any experiments to test the stability of the metric with varying numbers of question-answer pairs?\n\n\n3.VDCSCORE: miss detailed information in the Elo Ranking. Could authors provide details on:\n\na. The specific dataset used for the Elo Ranking comparison, including its size and composition\n\nb. The criteria used for selecting this dataset for the comparison\n\nc. How this dataset relates to or differs from the VDC benchmark?\n\n\n4.VDCSCORE：some quality problems of caption are about the repetition or grammar errors. Does this metric could evaluate these types of quality?\n\na.How does VDCSCORE handle linguistic aspects of caption quality, such as repetition and grammatical errors?\n\nb.Were any specific measures incorporated into the metric to detect and penalize such issues?\n\nc.Could the authors provide examples or experiments demonstrating how the metric performs on captions with these types of linguistic problems compared to human evaluation?"}},"nonreaders":[],"tmdate":1731427292272,"tcdate":1730543386684,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission640/Reviewer_BpiW"],"signatures":["ICLR.cc/2025/Conference/Submission640/Reviewer_BpiW"],"forum":"tTDUrseRRU","number":3,"license":"CC BY 4.0","cdate":1730543386684,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission640/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427292272,"domain":"ICLR.cc/2025/Conference","replyto":"tTDUrseRRU","id":"Lo86RABi04","forumContent":{"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video Captioning","Benchmark","Multimodel Large Language Model"]},"primary_area":{"value":"datasets and benchmarks"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Video detailed captioning is a key task which aims to generate comprehensive and coherent textual descriptions of video content, benefiting both video understanding and generation. In this paper, we propose AuroraCap, a video captioner based on a large multimodal model. We follow the simplest architecture design without additional parameters for temporal modeling. To address the overhead caused by lengthy video sequences, we implement the token merging strategy, reducing the number of input visual tokens. Surprisingly, we found that this strategy results in little performance loss. AuroraCap shows superior performance on various video and image captioning benchmarks, for example, obtaining a CIDEr of 88.9 on Flickr30k, beating GPT-4V (55.3) and Gemini-1.5 Pro (82.2). However, existing video caption benchmarks only include simple descriptions, consisting of a few dozen words, which limits research in this field. Therefore, we develop VDC, a video detailed captioning benchmark with over one thousand carefully annotated structured captions. In addition, we propose a new LLM-assisted metric VDCscore for bettering evaluation, which adopts a divide-and-conquer strategy to transform long caption evaluation into multiple short question-answer pairs. With the help of human Elo ranking, our experiments show that this benchmark better correlates with human judgments of video detailed captioning quality."},"_bibtex":{"value":"@inproceedings{\nchai2025auroracap,\ntitle={AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark},\nauthor={Wenhao Chai and Enxin Song and Yilun Du and Chenlin Meng and Vashisht Madhavan and Omer Bar-Tal and Jenq-Neng Hwang and Saining Xie and Christopher D Manning},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=tTDUrseRRU}\n}"},"title":{"value":"AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark"},"pdf":{"value":"/pdf/54378994740e36257bfa53e767f50f28a4e36483.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"chai|auroracap_efficient_performant_video_detailed_captioning_and_a_new_benchmark"},"authorids":{"value":["~Wenhao_Chai1","~Enxin_Song2","~Yilun_Du1","~Chenlin_Meng1","~Vashisht_Madhavan3","~Omer_Bar-Tal2","~Jenq-Neng_Hwang1","~Saining_Xie2","~Christopher_D_Manning1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Wenhao Chai","Enxin Song","Yilun Du","Chenlin Meng","Vashisht Madhavan","Omer Bar-Tal","Jenq-Neng Hwang","Saining Xie","Christopher D Manning"]}},"version":2},{"content":{"venue":{"value":"EMNLP 2023 Findings"},"TLDR":{"value":"Automatically discovereing irrationale reasoning process of models with minimal predefinition."},"keywords":{"value":["Shortcut Reasoning","Inference","Robustness"]},"abstract":{"value":"Shortcut reasoning is an irrational process of inference, which degrades the robustness of an NLP model.\nWhile a number of previous work has tackled the identification of shortcut reasoning, there are still two major limitations: (i) a method for quantifying the severity of the discovered shortcut reasoning is not provided; (ii) certain types of shortcut reasoning may be missed.\nTo address these issues, we propose a novel method for identifying shortcut reasoning.\nThe proposed method quantifies the severity of the shortcut reasoning by leveraging out-of-distribution data and does not make any assumptions about the type of tokens triggering the shortcut reasoning.\nOur experiments on Natural Language Inference and Sentiment Analysis demonstrate that our framework successfully discovers known and unknown shortcut reasoning in the previous work."},"_bibtex":{"value":"@inproceedings{\nharaguchi2023discovering,\ntitle={Discovering Highly Influential Shortcut Reasoning: An Automated Template-Free Approach},\nauthor={Daichi Haraguchi and Kiyoaki Shirai and Naoya Inoue and Natthawut Kertkeidkachorn},\nbooktitle={The 2023 Conference on Empirical Methods in Natural Language Processing},\nyear={2023},\nurl={https://openreview.net/forum?id=czxX6jjpVJ}\n}"},"title":{"value":"Discovering Highly Influential Shortcut Reasoning: An Automated Template-Free Approach"},"Submission_Type":{"value":"Regular Short Paper"},"pdf":{"value":"/attachment/fbbdab7aacfff56e033687adc285e8c1fe2663bc.pdf"},"Submission_Track":{"value":"Interpretability, Interactivity, and Analysis of Models for NLP"},"venueid":{"value":"EMNLP/2023/Conference"},"paperhash":{"value":"haraguchi|discovering_highly_influential_shortcut_reasoning_an_automated_templatefree_approach"},"authorids":{"value":["~Daichi_Haraguchi2","~Kiyoaki_Shirai1","~Naoya_Inoue1","~Natthawut_Kertkeidkachorn1"]},"authors":{"value":["Daichi Haraguchi","Kiyoaki Shirai","Naoya Inoue","Natthawut Kertkeidkachorn"]}},"tmdate":1701459730575,"pdate":1696709253198,"tcdate":1686835437392,"writers":["EMNLP/2023/Conference","EMNLP/2023/Conference/Submission1692/Authors"],"signatures":["EMNLP/2023/Conference/Submission1692/Authors"],"forum":"czxX6jjpVJ","number":1692,"cdate":1686835437392,"mdate":1701459730575,"readers":["everyone"],"invitations":["EMNLP/2023/Conference/-/Submission","EMNLP/2023/Conference/-/Post_Submission","EMNLP/2023/Conference/Submission1692/-/Revision","EMNLP/2023/Conference/-/Edit","EMNLP/2023/Conference/Submission1692/-/Camera_Ready_Revision"],"odate":1701459730560,"domain":"EMNLP/2023/Conference","id":"czxX6jjpVJ","version":2},{"content":{"summary":{"value":"The authors present a compression or optimization technique applicable to video transformers for both training and inference paradigms. Through empirical evaluation, the work showcases efficiency gains achieved for fine-tuning video transformer models and also showcases inference time efficiency without any training at all, with minimal quality degradation. The code has been made available for reproducibility."},"soundness":{"value":4},"confidence":{"value":4},"questions":{"value":"Out of curiosity, how did the formulation of RLT come up?"},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":4},"strengths":{"value":"The paper introduces run length tokenization (RLT) as a mechanism to reduce static tokens in training time dynamically. RLT takes into account spatial changes with respect to the temporal aspect and effectively condense static tokens adding run length information in the input embedding. RLT draws inspiration from widely used video compression techniques such as HEVC and AVC which are content aware, making RLT also content aware as part of the input embedding of the video transformer. \n\nThe originality of the work is commendable and the presentation quality of the work is exceptional. The efficiency gains achieved using this tokenization mechanism are state of the art, in both training and inference time with no performance degradation makes this work exceptionally useful to the extended community to train and deploy video transformer models more efficiently than previous literature in the field."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"No additional weaknesses apart from the limitations pointed out by the authors in Section 5."},"limitations":{"value":"NA"}},"nonreaders":[],"tmdate":1730879691174,"tcdate":1720850130195,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission14182/Reviewer_B8qa"],"signatures":["NeurIPS.cc/2024/Conference/Submission14182/Reviewer_B8qa"],"forum":"b1ggjW00NI","number":3,"license":"CC BY 4.0","cdate":1720850130195,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission14182/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730879691174,"domain":"NeurIPS.cc/2024/Conference","replyto":"b1ggjW00NI","id":"P6L4g9YuVc","forumContent":{"TLDR":{"value":"We make transformers 40% faster on video, with no performance drop, by identifying consecutive runs of tokens repeated in time, and treating them as a single token with variable length."},"venue":{"value":"NeurIPS 2024 spotlight"},"keywords":{"value":["video understanding","vision transformers","efficient transformers"]},"primary_area":{"value":"machine_vision"},"abstract":{"value":"Video transformers are slow to train due to extremely large numbers of input tokens, even though many video tokens are repeated over time. Existing methods to remove uninformative tokens either have significant overhead, negating any speedup, or require tuning for different datasets and examples. We present Run-Length Tokenization (RLT), a simple approach to speed up video transformers inspired by run-length encoding for data compression. RLT efficiently finds and removes `runs' of patches that are repeated over time before model inference, then replaces them with a single patch and a positional encoding to represent the resulting token's new length. \nOur method is content-aware, requiring no tuning for different datasets, and fast, incurring negligible overhead. \nRLT yields a large speedup in training, reducing the wall-clock time to fine-tune a video transformer by 30% while matching baseline model performance. RLT also works without training, increasing model throughput by 35% with only 0.1% drop in accuracy.\nRLT speeds up training at 30 FPS by more than 100%, and on longer video datasets, can reduce the token count by up to 80\\%. Our project page is at  rccchoudhury.github.io/projects/rlt."},"_bibtex":{"value":"@inproceedings{\nchoudhury2024dont,\ntitle={Don't Look Twice: Faster Video Transformers with Run-Length Tokenization},\nauthor={Rohan Choudhury and Guanglei Zhu and Sihan Liu and Koichiro Niinuma and Kris M. Kitani and Laszlo Attila Jeni},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=b1ggjW00NI}\n}"},"title":{"value":"Don't Look Twice: Faster Video Transformers with Run-Length Tokenization"},"pdf":{"value":"/pdf/dd99432c8f2dff551850e041e3841a2d75932e8c.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"choudhury|dont_look_twice_faster_video_transformers_with_runlength_tokenization"},"authorids":{"value":["~Rohan_Choudhury1","~Guanglei_Zhu1","~Sihan_Liu3","~Koichiro_Niinuma1","~Kris_M._Kitani1","~Laszlo_Attila_Jeni1"]},"authors":{"value":["Rohan Choudhury","Guanglei Zhu","Sihan Liu","Koichiro Niinuma","Kris M. Kitani","Laszlo Attila Jeni"]}},"version":2},{"content":{"venue":{"value":"CoRR 2025"},"pdf":{"value":"https://arxiv.org/pdf/2502.09150v2"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"suhail|shortcut_learning_susceptibility_in_vision_classifiers"},"html":{"value":"https://doi.org/10.48550/arXiv.2502.09150"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2502-09150,\n  publtype={informal},\n  author={Pirzada Suhail and Amit Sethi},\n  title={Shortcut Learning Susceptibility in Vision Classifiers},\n  year={2025},\n  month={February},\n  cdate={1738368000000},\n  journal={CoRR},\n  volume={abs/2502.09150},\n  url={https://doi.org/10.48550/arXiv.2502.09150}\n}\n"},"abstract":{"value":"Shortcut learning, where machine learning models exploit spurious correlations in data instead of capturing meaningful features, poses a significant challenge to building robust and generalizable models. This phenomenon is prevalent across various machine learning applications, including vision, natural language processing, and speech recognition, where models may find unintended cues that minimize training loss but fail to capture the underlying structure of the data. Vision classifiers based on Convolutional Neural Networks (CNNs), Multi-Layer Perceptrons (MLPs), and Vision Transformers (ViTs) leverage distinct architectural principles to process spatial and structural information, making them differently susceptible to shortcut learning. In this study, we systematically evaluate these architectures by introducing deliberate shortcuts into the dataset that are correlated with class labels both positionally and via intensity, creating a controlled setup to assess whether models rely on these artificial cues or learn actual distinguishing features. We perform both quantitative evaluation by training on the shortcut-modified dataset and testing on two different test sets-one containing the same shortcuts and another without them-to determine the extent of reliance on shortcuts. Additionally, qualitative evaluation is performed using network inversion-based reconstruction techniques to analyze what the models internalize in their weights, aiming to reconstruct the training data as perceived by the classifiers. Further, we evaluate susceptibility to shortcut learning across different learning rates. Our analysis reveals that CNNs at lower learning rates tend to be more reserved against entirely picking up shortcut features, while ViTs, particularly those without positional encodings, almost entirely ignore the distinctive image features in the presence of shortcuts."},"title":{"value":"Shortcut Learning Susceptibility in Vision Classifiers"},"authors":{"value":[{"fullname":"Pirzada Suhail","username":"~Pirzada_Suhail1"},{"fullname":"Amit Sethi","username":""}]}},"tmdate":1785934430457,"pdate":1767139200000,"externalIds":["dblp:journals/corr/abs-2502-09150"],"tcdate":1785934426741,"writers":["~","OpenReview.net/Public_Article/DBLP.org","OpenReview.net/Support"],"signatures":["~Pirzada_Suhail1"],"forum":"bMwCWibzxh","license":"CC BY-SA 4.0","number":123121,"cdate":1738368000000,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/DBLP.org/-/Record","OpenReview.net/Public_Article/-/Edit"],"mdate":1785934430457,"domain":"OpenReview.net/Public_Article","id":"bMwCWibzxh","version":2},{"content":{"venue":{"value":"SCSL @ ICLR 2025"},"TLDR":{"value":"This paper evaluates shortcut learning susceptibility in MLPs, CNNs, and ViTs across multiple datasets, showing that ViTs are the most vulnerable to shortcuts while CNNs are the least and higher learning rates reinforce shortcut reliance."},"keywords":{"value":["Shortcut Learning","Spurious Correlations","Network Inversion","Reconstructions"]},"abstract":{"value":"Shortcut learning, where machine learning models exploit spurious correlations in data instead of capturing meaningful features, poses a significant challenge to building robust and generalizable models. This phenomenon is prevalent across various machine learning applications, including vision, natural language processing, and speech recognition, where models may find unintended cues that minimize training loss but fail to capture the underlying structure of the data. Vision classifiers based on Convolutional Neural Networks (CNNs), Multi-Layer Perceptrons (MLPs), and Vision Transformers (ViTs) leverage distinct architectural principles to process spatial and structural information, making them differently susceptible to shortcut learning. In this study, we systematically evaluate these architectures by introducing deliberate shortcuts into the dataset that are correlated with class labels both positionally and via intensity, creating a controlled setup to assess whether models rely on these artificial cues or learn actual distinguishing features. We perform both quantitative evaluation by training on the shortcut-modified dataset and testing on two different test sets—one containing the same shortcuts and another without them—to determine the extent of reliance on shortcuts. Additionally, qualitative evaluation is performed using network inversion-based reconstruction techniques to analyze what the models internalize in their weights, aiming to reconstruct the training data as perceived by the classifiers. Further, we evaluate susceptibility to shortcut learning across different learning rates. Our analysis reveals that CNNs at lower learning rates tend to be more reserved against entirely picking up shortcut features, while ViTs, particularly those without positional encodings, almost entirely ignore the distinctive image features in the presence of shortcuts."},"_bibtex":{"value":"@inproceedings{\nsuhail2025shortcut,\ntitle={Shortcut Learning Susceptibility in Vision Classifiers},\nauthor={Pirzada Suhail and Amit Sethi},\nbooktitle={Workshop on Spurious Correlation and Shortcut Learning: Foundations and Solutions},\nyear={2025},\nurl={https://openreview.net/forum?id=dvafjL2zXP}\n}"},"title":{"value":"Shortcut Learning Susceptibility in Vision Classifiers"},"Anonymization":{"value":"This submission has been anonymized for double-blind review via the removal of identifying information such as names, affiliations, and identifying URLs."},"pdf":{"value":"/pdf/d6702ebe749e0bd43dbb679f502050a899ed45c4.pdf"},"venueid":{"value":"ICLR.cc/2025/Workshop/SCSL"},"paperhash":{"value":"suhail|shortcut_learning_susceptibility_in_vision_classifiers"},"authorids":{"value":["~Pirzada_Suhail1","~Vrinda_Goel1","~Amit_Sethi2"]},"Track":{"value":"regular paper (up to 6 pages)"},"authors":{"value":["Pirzada Suhail","Vrinda Goel","Amit Sethi"]}},"tmdate":1746112261839,"pdate":1741240536086,"tcdate":1739211127633,"writers":["ICLR.cc/2025/Workshop/SCSL","ICLR.cc/2025/Workshop/SCSL/Submission20/Authors"],"signatures":["ICLR.cc/2025/Workshop/SCSL/Submission20/Authors"],"forum":"dvafjL2zXP","license":"CC BY 4.0","number":20,"cdate":1739211127633,"readers":["everyone"],"invitations":["ICLR.cc/2025/Workshop/SCSL/-/Submission","ICLR.cc/2025/Workshop/SCSL/-/Post_Submission","ICLR.cc/2025/Workshop/SCSL/-/Edit","ICLR.cc/2025/Workshop/SCSL/Submission20/-/Camera-Ready_Version"],"mdate":1746112261839,"odate":1741240536086,"domain":"ICLR.cc/2025/Workshop/SCSL","id":"dvafjL2zXP","version":2},{"content":{"summary":{"value":"The paper introduces a video instance segmentation framework focused on improving instance tracking in complex video scenarios by integrating contextual information for each instance. The authors claim two key contributions: (i) a Context-Aware Instance Tracker (CAIT) that explicitly incorporates contextual information for each instance, and (ii) a Prototypical Cross-frame Contrastive (PCC) loss to improve the consistency of enriched features across frames, further enhancing tracking robustness. Experiments are performed on the YouTube-VIS 2019, 2021, OVIS, and VIPSeg datasets."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"See the weaknesses mentioned above."},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"(i) Comprehensive experiments across several video segmentation datasets.\n(ii) The paper is easy to follow.\n(iii) The authors have shared the code."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Weaknesses:\n\n\n(i) The authors should more clearly explain the novel contributions of the proposed approach compared to closely related work, such as (Choudhuri et al., 2023). Currently, the authors mention that “Choudhuri et al., 2023 introduces a context-aware relative query which reflects global and local context by aggregating multi-level image features from consecutive frames” without explicitly explaining how the contextual information captured by the proposed method is superior. Additionally, the authors are expected to include a comparison with Choudhuri et al., 2023, and recent approaches in the state-of-the-art comparison table. Similarly, the novel contributions of PCC compared to closely related loss formulations also need to be clarified.  \nReference: Choudhuri et al., 2023: Context-Aware Relative Object Queries to Unify Video Instance and Panoptic Segmentation, CVPR 2023.\n\n(ii) It is unclear whether the performance gain is due to the additional number of duplicate queries or from capturing contextual information, as claimed. For example, some conventional object detection methods, such as GroupDETR (ICCV 2023) [https://arxiv.org/abs/2207.13085], demonstrate the benefits of using additional object queries.\n\n(iii) The method seems somewhat ad-hoc to me . For instance, it involves obtaining instance-level masks for each frame using Mask2Former, then using edge-filtered instance masks to capture surrounding features of instances, which are subsequently leveraged to improve mask prediction and tracking. Would performance suffer if the initial masks obtained from Mask2Former were inaccurate? In cases of heavy occlusion, I assume that temporal information from previous frames would be beneficial for accurate mask prediction as well. However, since the initial masks are predicted independently by Mask2Former, could this affect overall performance?\n\n(iv) It appears that the proposed method mainly targets improving tracking or association performance, rather than mask quality. If so, why isn’t the proposed approach evaluated on multi-object tracking datasets like MOTS2020 and KITTI-MOTS in addition to the VIS datasets, similar to Choudhuri et al.?"}},"nonreaders":[],"tmdate":1731427291486,"tcdate":1731087399495,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission637/Reviewer_rJtA"],"signatures":["ICLR.cc/2025/Conference/Submission637/Reviewer_rJtA"],"forum":"VhQelEo27A","number":4,"license":"CC BY 4.0","cdate":1731087399495,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission637/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427291486,"domain":"ICLR.cc/2025/Conference","replyto":"VhQelEo27A","id":"iP3jD1nDa7","forumContent":{"TLDR":{"value":"We propose a context-aware learning pipeline for occlusion handling in the VIS task."},"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video Instance Segmentation","Contrastive Learning","Instance Prototype","Context-aware Learning"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"In this paper, we introduce the Context-Aware Video Instance Segmentation (CAVIS), a novel framework designed to enhance instance association by integrating contextual information adjacent to each object. To efficiently extract and leverage this information, we propose the Context-Aware Instance Tracker (CAIT), which merges contextual data surrounding the instances with the core instance features to improve tracking accuracy. Additionally, we introduce the Prototypical Cross-frame Contrastive (PCC) loss, which ensures consistency in object-level features across frames, thereby significantly enhancing instance matching accuracy. CAVIS demonstrates superior performance over state-of-the-art methods on all benchmark datasets in video instance segmentation (VIS) and video panoptic segmentation (VPS). Notably, our method excels on the OVIS dataset, which is known for its particularly challenging videos."},"_bibtex":{"value":"@misc{\nlee2025contextaware,\ntitle={Context-Aware Video Instance Segmentation},\nauthor={Seunghun Lee and Jiwan Seo and Kiljoon Han and Minwoo Choi and Sunghoon Im},\nyear={2025},\nurl={https://openreview.net/forum?id=VhQelEo27A}\n}"},"title":{"value":"Context-Aware Video Instance Segmentation"},"pdf":{"value":"/pdf/11cd1c2e65f6b7f1e81d00edbbd7f162483f2af3.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"lee|contextaware_video_instance_segmentation"},"authorids":{"value":["~Seunghun_Lee3","~Jiwan_Seo2","~Kiljoon_Han1","~Minwoo_Choi1","~Sunghoon_Im1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Seunghun Lee","Jiwan Seo","Kiljoon Han","Minwoo Choi","Sunghoon Im"]}},"version":2},{"content":{"data_release":{"value":"We authorize the release of our submission and author names to the public in the event of acceptance."},"TLDR":{"value":"We create a synthetic dataset from BEDLAM animations and test several action recognition models for bias."},"venue":{"value":"PFATCV@ECCV26"},"email_sharing":{"value":"We authorize the sharing of all author emails with Program Chairs."},"keywords":{"value":["Human Action Recognition","Synthetic dataset","Bias analysis"]},"supplementary_material":{"value":"/attachment/0645c27720d4968b421d57984dcd8056df63a759.pdf"},"abstract":{"value":"Human Action Recognition (HAR) models may exploit human appearance when it correlates with action labels, yet such shortcuts are difficult to isolate in recorded video. We introduce a controlled synthetic-video audit in which the same motion, actor configuration, camera, and background are preserved while only the rendered skin texture changes. We evaluate models through three complementary protocols. First, we fine-tune six Kinetics-pretrained RGB backbones on binary action tasks where skin tone is made predictive during training, then test whether performance drops when the tone-to-action assignment is reversed. Second, we train linear probes on frozen representations to measure how readily the same shortcut can be extracted without adapting the encoder. Third, we compare the temporal self-similarity of per-frame features under skin-tone and action changes. Several models exhibit directional sensitivity, although the effects are small and vary substantially across architectures, tasks, and motion instances. Among the models tested, the language-supervised CLIP and SigLIP models show the largest linear-probe drops, while the temporal diagnostic shows a broadly similar pattern. Strong photometric augmentation reduces the induced shortcut for some architectures but does not remove it reliably. These results establish a controlled methodology for auditing appearance-related shortcuts in HAR and motivate careful dataset construction alongside post hoc mitigation."},"_bibtex":{"value":"@inproceedings{\nbaltaretu2026controlled,\ntitle={Controlled Auditing of Skin-Tone Shortcut Sensitivity in Human Action Recognition},\nauthor={Ana B{\\u{a}}lt{\\u{a}}rețu and Pascal Benschop and Justin Dauwels and Jan van Gemert},\nbooktitle={ECCV26 Workshop on Privacy, Fairness, Accountability and Transparency in Computer Vision},\nyear={2026},\nurl={https://openreview.net/forum?id=BLcczzfXIV}\n}"},"title":{"value":"Controlled Auditing of Skin-Tone Shortcut Sensitivity in Human Action Recognition"},"pdf":{"value":"/pdf/36183e4e0063de1d0899fdb3d01e0f6a906d8324.pdf"},"venueid":{"value":"thecvf.com/ECCV/2026/Workshop/PFATCV"},"paperhash":{"value":"bltreu|controlled_auditing_of_skintone_shortcut_sensitivity_in_human_action_recognition"},"authorids":{"value":["~Ana_Băltărețu1","~Pascal_Benschop1","~Justin_Dauwels1","~Jan_van_Gemert1"]},"authors":{"value":["Ana Băltărețu","Pascal Benschop","Justin Dauwels","Jan van Gemert"]}},"tmdate":1788028447723,"pdate":1788028446873,"tcdate":1783246380949,"writers":["thecvf.com/ECCV/2026/Workshop/PFATCV","thecvf.com/ECCV/2026/Workshop/PFATCV/Submission15/Authors"],"signatures":["thecvf.com/ECCV/2026/Workshop/PFATCV/Submission15/Authors"],"forum":"BLcczzfXIV","license":"CC BY 4.0","number":15,"cdate":1783246380949,"readers":["everyone"],"invitations":["thecvf.com/ECCV/2026/Workshop/PFATCV/-/Submission","thecvf.com/ECCV/2026/Workshop/PFATCV/-/Submission_Change_Before_Bidding","thecvf.com/ECCV/2026/Workshop/PFATCV/-/Submission_Change_Before_Reviewing","thecvf.com/ECCV/2026/Workshop/PFATCV/Submission15/-/Camera_Ready_Revision","thecvf.com/ECCV/2026/Workshop/PFATCV/-/Submission_Release"],"mdate":1788028447723,"odate":1788028446873,"domain":"thecvf.com/ECCV/2026/Workshop/PFATCV","id":"BLcczzfXIV","version":2},{"content":{"summary":{"value":"Motivation:  Popular video benchmarks, such as MSRVTT and TGIF, often fail to effectively evaluate AI models’ temporal reasoning abilities due to the lack of fine-grained temporal annotations.\nThey  introduce a new TemporalBench benchmark for fine-grained temporal event understanding in videos. TemporalBench, sourced from a diverse video datasets, consists of ∼10K pairs of video description questions, derived from ∼2K high-quality human-annotated video captions.\nTheir results show that state-of-the-art models like GPT-4o achieve only 38.0% multiple binary QA accuracy on TemporalBench, demonstrating a significant human-AI gap in temporal understanding."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"See above."},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. Their benchmark enriches several fundamental video understanding tasks due to its detailed captions.\n2. The experimental results showcase the specific performance of various video understanding models, accompanied by corresponding result analysis.\n3. There are some interesting findings regarding Multiple Binary QA."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. There is a lack of detailed comparison with the temporal understanding aspect in the following paper: MVBench: A Comprehensive Multi-modal Video Understanding Benchmark. Additionally, it lacks a specific definition for temporal understanding tasks compared to this benchmark.\n2. It is unclear what specific example questions are used for the \"binary QA accuracy (BA)\" and \"multiple binary QA settings\" in Table 2.\n3. This seems more like a hallucination-reduction captioning benchmark rather than a temporal benchmark.\n4. The approach description and analysis for generating negative captions are insufficient. The types of negatives will determine how different models perform on this benchmark, and the benchmark results also lack corresponding analysis."}},"nonreaders":[],"tmdate":1731427455816,"tcdate":1730806124326,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission1635/Reviewer_Eio5"],"signatures":["ICLR.cc/2025/Conference/Submission1635/Reviewer_Eio5"],"forum":"Wto5U7q6I2","number":5,"license":"CC BY 4.0","cdate":1730806124326,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission1635/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427455816,"domain":"ICLR.cc/2025/Conference","replyto":"Wto5U7q6I2","id":"hLNVtD8Qsy","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"TLDR":{"value":"We introduce a finegrained multimodal video understanding benchmark"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["video","benchmark","multimodel"]},"supplementary_material":{"value":"/attachment/1b00a92cc0682497e1612273bbcb5db6abfd1201.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Understanding fine-grained temporal dynamics is crucial for video understanding. Yet, popular video benchmarks, such as MSRVTT and TGIF, often fail to effectively evaluate AI models' temporal reasoning abilities due to the lack of fine-grained temporal annotations. \nAs a result, text-based models, leveraging strong language priors, often perform comparably to video models, and image-trained models have been reported to outperform their video-trained counterparts on MSRVTT and TGIF. This paper introduces a new TemporalBench benchmark for fine-grained temporal event understanding in videos. TemporalBench, sourced from a diverse video datasets, consists of $\\sim$10K pairs of video description questions, derived from $\\sim$2K high-quality human-annotated video captions.  Uniquely, our benchmark provides fine-grained temporal annotations to evaluate models' temporal reasoning abilities. Our results show that state-of-the-art models like GPT-4o achieve only 38.0\\% multiple binary QA accuracy on TemporalBench, demonstrating a significant human-AI gap in temporal understanding. We hope that TemporalBench is instrumental to fostering research on improving models' temporal reasoning capabilities."},"_bibtex":{"value":"@misc{\ncai2024temporalbench,\ntitle={TemporalBench: Towards Fine-grained Temporal Understanding for  Multimodal Video  Models},\nauthor={Mu Cai and Reuben Tan and Jianrui Zhang and Bocheng Zou and Kai Zhang and Feng Yao and Fangrui Zhu and Jing Gu and Yiwu Zhong and Yuzhang Shang and Yao Dou and Jaden Park and Jianfeng Gao and Yong Jae Lee and Jianwei Yang},\nyear={2024},\nurl={https://openreview.net/forum?id=Wto5U7q6I2}\n}"},"title":{"value":"TemporalBench: Towards Fine-grained Temporal Understanding for  Multimodal Video  Models"},"pdf":{"value":"/pdf/433bcaa84f8ff694d5c6245d586377638b770240.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"cai|temporalbench_towards_finegrained_temporal_understanding_for_multimodal_video_models"},"authorids":{"value":["~Mu_Cai1","~Reuben_Tan1","~Jianrui_Zhang1","~Bocheng_Zou1","~Kai_Zhang10","~Feng_Yao1","~Fangrui_Zhu1","~Jing_Gu2","~Yiwu_Zhong1","~Yuzhang_Shang1","~Yao_Dou1","~Jaden_Park1","~Jianfeng_Gao1","~Yong_Jae_Lee2","~Jianwei_Yang1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Mu Cai","Reuben Tan","Jianrui Zhang","Bocheng Zou","Kai Zhang","Feng Yao","Fangrui Zhu","Jing Gu","Yiwu Zhong","Yuzhang Shang","Yao Dou","Jaden Park","Jianfeng Gao","Yong Jae Lee","Jianwei Yang"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a geometry-aware 4D video generation model tailored for robot manipulation, which enforces multi-view 3D consistency during training via cross-view pointmap alignment and leverages pretrained video diffusion models for temporal coherence, enabling the generation of spatio-temporally aligned RGB-D sequences from novel viewpoints without camera pose inputs. It evaluates the model on both simulated (e.g., StoreCerealBoxUnderShelf) and real-world (e.g., TwistCapOffBottle) robotic tasks. Additionally, the predicted 4D videos can be used with an off-the-shelf 6DoF pose tracker (e.g., FoundationPose) to extract robot end-effector trajectories."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"see weaknesses"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. Enables geometry-consistent 4D video generation for multi-object dynamic robot manipulation scenes (a gap in prior 3D-aware methods limited to single objects and static backgrounds) by enforcing cross-view pointmap alignment during training, ensuring spatio-temporal consistency across novel camera viewpoints without relying on camera poses at inference.\n2. Integrates pretrained video diffusion models’ temporal priors with a dedicated geometry-consistent loss for pointmaps, achieving joint optimization of RGB video quality, depth accuracy, and multi-view 3D consistency.\n3. Bridges 4D video generation to practical robot manipulation: predicted RGB-D sequences can extract robot end-effector trajectories via an off-the-shelf 6DoF pose tracker (e.g., FoundationPose) and infer gripper open/close states, leading to higher task success rates than visuomotor policy baselines."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Relies on multi-view RGB-D datasets with varied camera viewpoints for training, which are challenging to collect in real-world settings due to hardware constraints and calibration requirements, and high-quality depth data acquisition in real scenarios remains difficult.\n2. The inference speed of the video generation model is relatively slow, making closed-loop planning for robot manipulation impractical compared to end-to-end behavior cloning policies.\n3. The baseline comparisons lack direct competition with state-of-the-art methods specifically designed for multi-object dynamic robot manipulation scenarios; selected baselines (e.g., 4D Gaussian, SVD) either have different task scopes or lack 3D consistency modeling, weakening the persuasiveness of performance superiority.\n4. Key implementation details for reproducibility are insufficiently disclosed, such as the specific camera sampling parameters (e.g., exact pitch/yaw ranges), training hyperparameter schedules (e.g., learning rate decay strategy), and threshold values for gripper state inference (e.g., distance threshold δ), which may hinder result validation.\n5. The multi-view cross-attention mechanism, a core component for 3D consistency, lacks unique design details; it is not clarified how it adapts to pointmap geometric features (e.g., whether attention weights correlate with 3D distances), making it indistinguishable from generic cross-attention modules."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762918653582,"tcdate":1761367382289,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission6362/Reviewer_RXK1"],"signatures":["ICLR.cc/2026/Conference/Submission6362/Reviewer_RXK1"],"forum":"18gC6pZVVc","number":1,"license":"CC BY 4.0","cdate":1761367382289,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission6362/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762918653582,"domain":"ICLR.cc/2026/Conference","replyto":"18gC6pZVVc","id":"lYVHdHz24a","forumContent":{"TLDR":{"value":"We propose a 4D video generation model that enforces geometric consistency across views to generate spatio-temporally aligned RGB-D sequences, enabling downstream applications of robot manipulation tasks via pose tracking."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Video Generation","Robot Manipulation","3D Perception"]},"primary_area":{"value":"applications to robotics, autonomy, planning"},"abstract":{"value":"Understanding and predicting dynamics of the physical world can enhance a robot's ability to plan and interact effectively in complex environments. While recent video generation models have shown strong potential in modeling dynamic scenes, generating videos that are both temporally coherent and geometrically consistent across camera views remains a significant challenge. To address this, we propose a 4D video generation model that enforces multi-view 3D consistency of generated videos by supervising the model with cross-view pointmap alignment during training. Through this geometric supervision, the model learns a shared 3D scene representation, enabling it to generate spatio-temporally aligned future video sequences from novel viewpoints given a single RGB-D image per view, and without relying on camera poses as input. Compared to existing baselines, our method produces more visually stable and spatially aligned predictions across multiple simulated and real-world robotic datasets. We further show that the predicted 4D videos can be used to recover robot end-effector trajectories using an off-the-shelf 6DoF pose tracker, yielding robot manipulation policies that generalize well to novel camera viewpoints."},"_bibtex":{"value":"@inproceedings{\nliu2026geometryaware,\ntitle={Geometry-aware 4D Video Generation for Robot Manipulation},\nauthor={Zeyi Liu and Shuang Li and Eric Cousineau and Siyuan Feng and Benjamin Burchfiel and Shuran Song},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=18gC6pZVVc}\n}"},"title":{"value":"Geometry-aware 4D Video Generation for Robot Manipulation"},"pdf":{"value":"/pdf/dc9cfece56983c0928896b5621f4d8dcfefdcd40.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"liu|geometryaware_4d_video_generation_for_robot_manipulation"},"authorids":{"value":["~Zeyi_Liu1","~Shuang_Li5","~Eric_Cousineau1","~Siyuan_Feng5","~Benjamin_Burchfiel1","~Shuran_Song3"]},"authors":{"value":["Zeyi Liu","Shuang Li","Eric Cousineau","Siyuan Feng","Benjamin Burchfiel","Shuran Song"]}},"version":2},{"content":{"venue":{"value":"CVPR 2026"},"abstract":{"value":"Text-video retrieval tasks have seen significant improvements due to the recent development of large-scale vision-language pre-trained models. Traditional methods primarily focus on video representations or cross-modal alignment, while recent works shift toward enriching text expressiveness to better match the rich semantics in videos. However, these methods use only interactions between text and frames/video, and ignore rich interactions among the internal frames within a video, so the final expanded text cannot capture frame contextual information, leading to disparities between text and video. In response, we introduce Energy-Aware Fine-Grained Relationship Learning Network (EagleNet) to generate accurate and context-aware enriched text embeddings. Specifically, the proposed Fine-Grained Relationship Learning mechanism (FRL) first constructs a text-frame graph by the generated text candidates and frames, then learns relationships among texts and frames, which are finally used to aggregate text candidates into an enriched text embedding that incorporates frame contextual information. To further improve fine-grained relationship learning in FRL, we design Energy-Aware Matching (EAM) to model the energy of text-frame interactions and thus accurately capture the distribution of real text-video pairs. Moreover, for more effective cross-modal alignment and stable training, we replace the conventional softmax-based contrastive loss with the sigmoid loss. Extensive experiments have demonstrated the superiority of EagleNet across MSRVTT, DiDeMo, MSVD, and VATEX."},"_bibtex":{"value":"@inproceedings{\nchen2026eaglenet,\ntitle={EagleNet: Energy-Aware Fine-Grained Relationship Learning Network for Text-Video Retrieval},\nauthor={Yuhan Chen and Pengwen Dai and Chuan Wang and Dayan Wu and Xiaochun Cao},\nbooktitle={Conference on Computer Vision and Pattern Recognition 2026},\nyear={2026},\nurl={https://openreview.net/forum?id=BlFtZtNEHM}\n}"},"title":{"value":"EagleNet: Energy-Aware Fine-Grained Relationship Learning Network for Text-Video Retrieval"},"pdf":{"value":"https://openaccess.thecvf.com/content/CVPR2026/papers/Chen_EagleNet_Energy-Aware_Fine-Grained_Relationship_Learning_Network_for_Text-Video_Retrieval_CVPR_2026_paper.pdf"},"venueid":{"value":"thecvf.com/CVPR/2026/Conference"},"paperhash":{"value":"chen|eaglenet_energyaware_finegrained_relationship_learning_network_for_textvideo_retrieval"},"authorids":{"value":["~Yuhan_Chen5","~Pengwen_Dai1","~Chuan_Wang1","~Dayan_Wu1","~Xiaochun_Cao3"]},"authors":{"value":["Yuhan Chen","Pengwen Dai","Chuan Wang","Dayan Wu","Xiaochun Cao"]}},"tmdate":1789656614170,"pdate":1789656438081,"tcdate":1765220460875,"writers":["thecvf.com/CVPR/2026/Conference","thecvf.com/CVPR/2026/Conference/Submission31869/Authors"],"signatures":["thecvf.com/CVPR/2026/Conference/Submission31869/Authors"],"forum":"BlFtZtNEHM","license":"CC BY 4.0","number":31869,"cdate":1765220460875,"readers":["everyone"],"invitations":["thecvf.com/CVPR/2026/Conference/-/Submission","thecvf.com/CVPR/2026/Conference/Submission31869/-/Full_Submission","thecvf.com/CVPR/2026/Conference/-/Post_Submission","thecvf.com/CVPR/2026/Conference/-/Edit","thecvf.com/CVPR/2026/Conference/Submission31869/-/Supplementary_Material","thecvf.com/CVPR/2026/Conference/-/Compute_Flag"],"mdate":1789656614170,"odate":1789656438081,"domain":"thecvf.com/CVPR/2026/Conference","id":"BlFtZtNEHM","version":2},{"content":{"summary":{"value":"This paper proposes the new task of detecting, tracking and captioning objects in a video. The authors propose evaluation metrics for this task, a baseline method and a training strategy involving disjoint supervision."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Please see Weaknesses."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"This work tackles an important problem that was missing from the literature. The paper is well-written and easy to follow. Extensive experiments have been performed."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"My only concern is that the end-to-end tracking algorithm (listed as a contribution) seems to be naive and not novel enough. There are other methods that perform identity association within the model (like MinVIS[1], CAROQ [2], trackformer [3]). The authors have explored some other ways of integrating temporal information/tracking in Table 2, but how about different ways of association within the model (e.g., query vector propagation in trackformer [3])?\n\n\n[1] MinVIS: A Minimal Video Instance Segmentation Framework without Video-based Training, Huang et al., NeurIPS 2022.\n\n[2] CAROQ: Context-Aware Relative Object Queries to Unify Video Instancd and Panoptic Segmentation, Choudhuri et al., CVPR 2023.\n\n[3] TrackFormer: Multi-Object Tracking with Transformers, Meinhardt et al., CVPR 2022"}},"nonreaders":[],"tmdate":1731428246390,"tcdate":1730586754579,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission4872/Reviewer_zAgu"],"signatures":["ICLR.cc/2025/Conference/Submission4872/Reviewer_zAgu"],"forum":"auZZ2gN0ZN","number":2,"license":"CC BY 4.0","cdate":1730586754579,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission4872/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428246390,"domain":"ICLR.cc/2025/Conference","replyto":"auZZ2gN0ZN","id":"6F4pIyVh52","forumContent":{"venue":{"value":"ICLR 2025 Spotlight"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["object captioning","video","tracking"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"We propose a new task and model for dense video object captioning -- detecting, tracking and captioning trajectories of objects in a video. This task unifies spatial and temporal localization in video, whilst also requiring fine-grained visual understanding that is best described by natural language. We propose a unified model, and demonstrate how our end-to-end approach is more accurate and temporally coherent than a multi-stage pipeline combining state-of-the-art detection, tracking, and captioning models. Moreover, we propose a training strategy based on a mixture of disjoint tasks, which allows us to leverage diverse, large-scale datasets which supervise different parts of our model. Although each pretraining task only provides weak supervision, they are complementary and, when combined, result in noteworthy zero-shot ability and serve as strong initialization for additional finetuning to further improve accuracy. We carefully design new metrics capturing all components of our task, and show how we can repurpose existing video grounding datasets (e.g. VidSTG and VLN) for our new task. We show that our model improves upon a number of strong baselines for this new task. Furthermore, we can apply our model to the task of spatial grounding, outperforming prior state-of-the-art on VidSTG and VLN, without explicitly training for it. Our code is available at https://github.com/google-research/scenic."},"_bibtex":{"value":"@inproceedings{\nzhou2025dense,\ntitle={Dense Video Object Captioning from Disjoint Supervision},\nauthor={Xingyi Zhou and Anurag Arnab and Chen Sun and Cordelia Schmid},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=auZZ2gN0ZN}\n}"},"title":{"value":"Dense Video Object Captioning from Disjoint Supervision"},"pdf":{"value":"/pdf/f62b7b1e28340d6f211f8cb661ed6c84c2199b9a.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"zhou|dense_video_object_captioning_from_disjoint_supervision"},"authorids":{"value":["~Xingyi_Zhou2","~Anurag_Arnab1","~Chen_Sun1","~Cordelia_Schmid1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Xingyi Zhou","Anurag Arnab","Chen Sun","Cordelia Schmid"]}},"version":2},{"content":{"summary":{"value":"This paper introduces EditVerse, a unified multimodal framework for text-guided image and video generation and editing. The model represents text, image, and video modalities as a single interleaved token sequence and applies self-attention for in-context learning and cross-modal reasoning. To address the lack of video editing data, the authors design a large-scale automated data generation pipeline. Extensive experiments demonstrate that EditVerse achieves strong performance and exhibits emergent editing capabilities that generalize beyond the training set."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"* How is the VLM used for data scoring and filtering? Was it used off-the-shelf or fine-tuned for editing relevance? What is its reliability or correlation with human judgment?\n* How does model performance and runtime scale with sequence length (e.g., multi-video input or long textual instructions)? Can efficiency be improved beyond full-sequence attention?\n* Have you compared the single-layer modality projectors against deeper or non-linear alternatives?\n* The reported improvements are modest; how sensitive are the evaluation metrics, and do the differences translate into perceptual gains?\n* Given that the data pipeline heavily depends on pretrained vision models, how do you ensure the resulting dataset and EditVerse itself do not inherit or amplify their biases?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"* Comprehensive experiments with solid quantitative and qualitative evaluations.\n* Clear presentation and well-structured writing.\n* The unified token representation for text, image, and video is conceptually elegant.\n* Interesting empirical insight: the model demonstrates emergent abilities on unseen editing tasks, suggesting strong cross-modal generalization."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"* While the unified sequence formulation is elegant, it primarily extends existing transformer-based multimodal modeling paradigms. The contribution lies more in engineering integration and data scaling than in introducing fundamentally new learning principles.\n* The model requires ~30 GB GPU memory and ~118 seconds to edit a single 360p video on an A100 80 GB GPU. This raises serious scalability concerns for long-duration or high-resolution videos, as well as practical deployment constraints.\n* The approach concatenates all modality tokens into one long sequence, resulting in $O(L^2)$ attention cost. It is unclear how EditVerse maintains efficiency with long inputs, multiple video clips, or multi-turn text–video interleaving. No efficiency analysis (FLOPs, latency, or memory growth curves) is provided.\n* The use of single-layer linear projectors for modality alignment may be oversimplified. There is no ablation on alternative projector depths or modality-specific encoders to validate whether this minimal design is sufficient for cross-modal alignment.\n* The video editing dataset is entirely generated via an automated pipeline using pretrained models (Grounded-SAM-2, VACE, ReCamMaster, etc.). This synthetic data may not reflect the diversity and imperfections of real-world edits, leading to overfitting on artificial editing patterns.\n* Each data generation stage depends on prior model outputs, so artifacts such as inaccurate masks, poor inpainting, or unrealistic motion can cascade. The paper provides no quantitative analysis or quality control to assess data noise accumulation. \n* The paper mentions using a VLM to assign quality scores for filtering generated data, but it is unclear whether the VLM was adapted for the data filtering, how accurate its scores are, or how thresholds were chosen."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762920054798,"tcdate":1762147304281,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission8067/Reviewer_U1BY"],"signatures":["ICLR.cc/2026/Conference/Submission8067/Reviewer_U1BY"],"forum":"blJXE07r7I","number":5,"license":"CC BY 4.0","cdate":1762147304281,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission8067/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762920054798,"domain":"ICLR.cc/2026/Conference","replyto":"blJXE07r7I","id":"NSkeiBjsa5","forumContent":{"venue":{"value":"ICLR 2026 Oral"},"keywords":{"value":["Video Editing","Content Generation","Artificial Intelligence"]},"supplementary_material":{"value":"/attachment/54f1a6cabd37b58af95d8fc7a1065d3ac273ac39.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Recent advances in foundation models highlight a clear trend toward unification and scaling, showing emergent capabilities across diverse domains. While image generation and editing have rapidly transitioned from task-specific to unified frameworks, video generation and editing remain fragmented due to architectural limitations and data scarcity. In this work, we introduce EditVerse, a unified framework for image and video generation and editing within a single model. By representing all modalities, i.e., text, image, and video, as a unified token sequence, EditVerse leverages self-attention to achieve robust in-context learning, natural cross-modal knowledge transfer, and flexible handling of inputs and outputs with arbitrary resolutions and durations. To address the lack of video editing training data, we design a scalable data pipeline that curates 232K video editing samples and combines them with large-scale image and video datasets for joint training. Furthermore, we present EditVerseBench, the first benchmark for instruction-based video editing covering diverse tasks and resolutions. Extensive experiments and user studies demonstrate that EditVerse achieves state-of-the-art performance, surpassing existing open-source and commercial models, while exhibiting emergent editing and generation abilities across modalities."},"_bibtex":{"value":"@inproceedings{\nju2026editverse,\ntitle={EditVerse: Unifying Image and Video Editing and Generation with In-Context Learning},\nauthor={Xuan Ju and Tianyu Wang and Yuqian Zhou and He Zhang and Qing Liu and Nanxuan Zhao and Zhifei Zhang and Yijun Li and Yuanhao Cai and Shaoteng Liu and Daniil Pakhomov and Zhe Lin and Soo Ye Kim and Qiang Xu},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=blJXE07r7I}\n}"},"title":{"value":"EditVerse: Unifying Image and Video Editing and Generation with In-Context Learning"},"pdf":{"value":"/pdf/b94811b83a0530da511e76edf48c05b2bbbd4725.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"ju|editverse_unifying_image_and_video_editing_and_generation_with_incontext_learning"},"authorids":{"value":["~Xuan_Ju1","~Tianyu_Wang3","~Yuqian_Zhou2","~He_Zhang16","~Qing_Liu1","~Nanxuan_Zhao1","~Zhifei_Zhang2","~Yijun_Li2","~Yuanhao_Cai1","~Shaoteng_Liu1","~Daniil_Pakhomov2","~Zhe_Lin1","~Soo_Ye_Kim1","~Qiang_Xu1"]},"authors":{"value":["Xuan Ju","Tianyu Wang","Yuqian Zhou","He Zhang","Qing Liu","Nanxuan Zhao","Zhifei Zhang","Yijun Li","Yuanhao Cai","Shaoteng Liu","Daniil Pakhomov","Zhe Lin","Soo Ye Kim","Qiang Xu"]}},"version":2},{"content":{"summary":{"value":"This paper proposes Video Mixture of Block Attention (VMoBA), a novel sparse attention mechanism tailored for Video Diffusion Models (VDMs). The authors introduce three key innovations: (i) a layer-wise recurrent 1D- 2D-3D block partitioning scheme, (ii) global block selection that selects the most salient query-key block interactions across all queries per head, and (iii) threshold-based dynamic block selection that adapts the number of attended blocks based on cumulative similarity. Experiments on long-sequence video generation show that VMoBA achieves faster training and fewer FLOPs than full attention, while matching or even surpassing its generation quality. VMoBA also performs competitively in training-free inference settings."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"- The hyperparameter \\tau controls the trade-off between speed and quality. If there exists a general value for all video types? If not, how to choose a proper \\tau for new videos or resolutions?\n- The paper notes that attention heads have different concentration levels. Have the authors analyzed whether VMoBA’s threshold-based selection leads to more specialized head behavior compared to full attention?\n- Have you tested VMoBA within a few-step distilled or consistency diffusion frameworks to verify compatibility with fast-sampling variants?\n- Could the cyclical 1-2-3D partitioning be learned end-to-end rather than fixed by layer index?"},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":3},"strengths":{"value":"- The three proposed modifications (1D-2D-3D partitioning, global selection, threshold-based sparsity) are well-motivated by the observed limitations of applying MoBA naively to video. Each component is clearly linked to a specific empirical observation.\n- The authors evaluate VMoBA in both training-based and training-free settings across multiple resolutions, using standard metrics (VBench, PSNR) and complete ablation studies. VMoBA achieves FLOPs reduction and training-time speedup with minimal loss in generation quality. Qualitative results further support the claims.\n- The writing is good and is easy to follow. Method illustration and implementation details are well documented."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The study omits some recent linear or hybrid video attentions that could serve as stronger baselines, such as STA[1] and RainFusion[2].\n- The paper should include more human evaluation. Human judgment on video quality and video consistency is crucial for assessing the performance.\n- In the global selection part, this module prioritizes key blocks with the highest overall significance, but may overlook certain keys that are locally relevant to queries yet have low global scores. The high-frequency details in generated videos may be affected.\n- Have you tried other fusion methods to fuse the tokens in a block? Does mean pooling harm the diversity within a block?\n\n[1] Fast Video Generation with Sliding Tile Attention\n\n[2] RainFusion: Adaptive Video Generation Acceleration via Multi-Dimensional Visual Redundancy"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762922333808,"tcdate":1761907586095,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission11178/Reviewer_9PZV"],"signatures":["ICLR.cc/2026/Conference/Submission11178/Reviewer_9PZV"],"forum":"oQaRElUdmh","number":3,"license":"CC BY 4.0","cdate":1761907586095,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission11178/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762922333808,"domain":"ICLR.cc/2026/Conference","replyto":"oQaRElUdmh","id":"LIEVBLfnsy","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Video Generation","Sparse Attention","Training Acceleration","MoBA"]},"primary_area":{"value":"generative models"},"abstract":{"value":"The quadratic complexity of full attention mechanisms poses a significant bottleneck for Video Diffusion Models (VDMs) aiming to generate long-duration, high-resolution videos. While various sparse attention methods have been proposed, many are designed as training-free inference accelerators or do not optimally capture the unique spatio-temporal characteristics inherent in video data when trained natively. This paper introduces Video Mixture of Block Attention (VMoBA), a novel sparse attention mechanism specifically adapted for VDMs. Motivated by an in-depth analysis of attention patterns within pre-trained video transformers, which revealed strong spatio-temporal locality, varying query importance, and head-specific concentration levels, VMoBA enhances the original MoBA framework with three key modifications: (1) a layer-wise recurrent block partition scheme (1D-2D-3D) to dynamically adapt to diverse spatio-temporal attention patterns and improve efficiency; (2) global block selection to prioritize the most salient query-key block interactions across an entire attention head; and (3) threshold-based block selection to dynamically determine the number of attended blocks based on their cumulative similarity. Extensive experiments demonstrate that VMoBA significantly accelerates the training of VDMs on longer sequences, achieving 2.92$\\times$ FLOPs and 1.48$\\times$ latency speedup, while attaining comparable or even superior generation quality to full attention. Furthermore, VMoBA exhibits competitive performance in training-free inference, offering 2.40$\\times$ FLOPs and 1.35$\\times$ latency speedup for high-res video generation."},"_bibtex":{"value":"@inproceedings{\nwu2026vmoba,\ntitle={{VM}o{BA}: Mixture-of-Block Attention for Video Diffusion Models},\nauthor={Jianzong Wu and Liang Hou and Haotian Yang and Ye Tian and Pengfei Wan and Di ZHANG and Yunhai Tong},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=oQaRElUdmh}\n}"},"title":{"value":"VMoBA: Mixture-of-Block Attention for Video Diffusion Models"},"pdf":{"value":"/pdf/668781fc4c0bc623c7ebd2a2f230f6e2515ad6fd.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"wu|vmoba_mixtureofblock_attention_for_video_diffusion_models"},"authorids":{"value":["~Jianzong_Wu1","~Liang_Hou1","~Haotian_Yang1","~Ye_Tian15","~Pengfei_Wan1","~Di_ZHANG3","~Yunhai_Tong1"]},"authors":{"value":["Jianzong Wu","Liang Hou","Haotian Yang","Ye Tian","Pengfei Wan","Di ZHANG","Yunhai Tong"]}},"version":2},{"content":{"TLDR":{"value":"SMORE is the first minimal-pair video benchmark to evaluate social, moral, and rational reasoning in MLLMs, revealing a substantial gap between current models and humans."},"venue":{"value":"ICLR 2027 Conference Submission"},"keywords":{"value":["multimodal large language models","video understanding","social reasoning","moral reasoning","rational reasoning","minimal pairs","social cognition"]},"primary_area":{"value":"applications to neuroscience & cognitive science"},"abstract":{"value":"Understanding human behavior involves interpreting actions in relation to agents’ beliefs, goals and social norms. Small differences in behavior can change these interpretations even when the surrounding context remains similar. To test whether multimodal large language models (MLLMs) reliably distinguish such cases, we introduce SMORE, a minimal-pair benchmark for evaluating Social, MOral, and Rational reasoning. It comprises 427 video pairs (854 videos) and 1,281 evaluation questions spanning belief-conditioned rationality, helping and hindering and judgements of intentionality, responsibility and norm violations. Videos within each pair share the same actors, environment, and overall interaction while differing in a single socially-relevant factor that changes the target judgement. This design probes whether model judgements track relevant behavioral differences under closely matched conditions. We evaluate models under two settings. First, we have paired-video evaluation, where both videos from a minimal pair are presented jointly to measure relative discrimination performance. Second, we have single-video evaluation, where each video is evaluated independently to measure absolute performance. From the second setting, we additionally compute Joint-Pair Accuracy, which gives credit only when both videos in a pair are individually answered correctly. On SMORE, humans reach 96% Joint-Pair Accuracy while the best model, Gemini-3.5-Flash, reaches only 63.2%. These results suggest that, despite strong performance on video action understanding, current MLLMs still fall substantially short of humans in reasoning about the subtle differences in human behavior that determine social, moral, and rational judgments."},"_bibtex":{"value":"@inproceedings{\nanonymous2026smore,\ntitle={{SMORE}: Minimal-Pair Videos for Social, Moral and Rational Reasoning in {MLLM}s},\nauthor={Anonymous},\nbooktitle={Submitted to The Fifteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=1XQKFnGKk4},\nnote={under review}\n}"},"title":{"value":"SMORE: Minimal-Pair Videos for Social, Moral and Rational Reasoning in MLLMs"},"pdf":{"value":"/pdf/00f9c8408511b777987375a082fda9b47d244cf2.pdf"},"venueid":{"value":"ICLR.cc/2027/Conference/Submission"}},"tmdate":1791233782223,"tcdate":1789749687449,"writers":["ICLR.cc/2027/Conference","ICLR.cc/2027/Conference/Submission46063/Authors"],"signatures":["ICLR.cc/2027/Conference/Submission46063/Authors"],"forum":"1XQKFnGKk4","license":"CC BY 4.0","number":46063,"cdate":1789749687449,"readers":["everyone"],"invitations":["ICLR.cc/2027/Conference/-/Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Bidding","ICLR.cc/2027/Conference/Submission46063/-/Full_Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Reviewing","ICLR.cc/2027/Conference/-/Edit"],"mdate":1791233782223,"odate":1791031328286,"domain":"ICLR.cc/2027/Conference","id":"1XQKFnGKk4","version":2},{"content":{"summary":{"value":"Vivid-VR is a generative video restoration method. To address the issue of distribution drift during fine-tuning, the authors propose a concept distillation strategy that uses the pre-trained T2V model to synthesize aligned text-video pairs, thereby preserving texture realism and temporal coherence. The control mechanism is further enhanced with a novel feature projector to filter degradation artifacts and a dual-branch connector for dynamic control feature retrieval. Extensive experiments demonstrate that Vivid-VR achieves superior performance in texture realism and temporal consistency compared to existing methods on various benchmarks."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. For the concept distillation process, what will happen if the generated textual descriptions have a major difference from the visual content? Then the pretrained T2V model may generate videos that are quite different from the original low HQ videos. Will it influence the semantic accuracy of the final high HQ videos?\n\n2. Why CogVideoX1.5-5B is selected as the pretrained T2V model? What will the model perform when selecting other alternatives? Besides, what will the video restoration perform when adopting pretrained T2V models in different sizes?\n\n3. What is the rationale of selecting DiT as the main architecture of the video restoration model, considering there are many alternatives, such as MMDiT?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. Leveraging the capabilities of pretrained T2V models to enhance the video restoration performance is an interesting approach. Using the textual description as a connector, the authors find an effective way to transfer the T2V model’s pretrained knowledge to the video restoration model, which I believe benefits the community.\n\n2. The proposed method outperforms several previous advances by a large margin in a wide range of benchmarks. The experiments are comprehensive, and the qualitative demos are also impressive.\n\n3. The ablation studies are also exhaustive. The effectiveness of each branch is clearly verified."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Though the idea of transferring the T2V model's capability to downstream tasks is interesting, the proposed method seems to be trivial and similar to other methods. Using a pretrained model to corrupt and reconstruct the visual content is a common way, especially in image enhancement and restoration tasks, and it is also widely adopted to add the textual description during the reconstruction. It is more likely to be a transition from the image restoration task to the video restoration task, which makes the technical contribution less competitive.\n\n2. How the concept distillation really works remains ambiguous. Why putting the textual description into the DiT block can ''transfer the T2V model’s conceptual knowledge to the video restoration model''? What concept is transferred to the video restoration model? What will happen if we use a video captioner to generate captions and add a text encoder to the DiT block instead of leveraging the pretrained T2V model?"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762916914739,"tcdate":1761793356793,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission3680/Reviewer_vgeK"],"signatures":["ICLR.cc/2026/Conference/Submission3680/Reviewer_vgeK"],"forum":"YV5Zgv8pdg","number":2,"license":"CC BY 4.0","cdate":1761793356793,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission3680/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762916914739,"domain":"ICLR.cc/2026/Conference","replyto":"YV5Zgv8pdg","id":"npXZOUIlOz","forumContent":{"TLDR":{"value":"We present Vivid-VR, a DiT-based generative video restoration method."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Video Restoration","Diffusion Transformer","Text-to-Video","ControlNet","Concept Distillation"]},"supplementary_material":{"value":"/attachment/1ac7838c26d644b9c27a33c04143e290587515fb.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"We present Vivid-VR, a DiT-based generative video restoration method built upon an advanced T2V foundation model, where ControlNet is leveraged to control the generation process, ensuring content consistency. However, conventional fine-tuning of such controllable pipelines frequently suffers from distribution drift due to limitations in imperfect multimodal alignment, resulting in compromised texture realism and temporal coherence. To tackle this challenge, we propose a concept distillation training strategy that utilizes the pretrained T2V model to synthesize training samples with embedded textual concepts, thereby distilling its conceptual understanding to preserve texture and temporal quality. To enhance generation controllability, we redesign the control architecture with two key components: 1) a control feature projector that filters degradation artifacts from input video latents to minimize their propagation through the generation pipeline, and 2) a new ControlNet connector employing a dual-branch design. This connector synergistically combines MLP-based feature mapping with cross-attention mechanism for dynamic control feature retrieval, enabling both content preservation and adaptive control signal modulation. Extensive experiments show that Vivid-VR performs favorably against existing approaches on both synthetic and real-world benchmarks, as well as AIGC videos, achieving impressive texture realism, visual vividness, and temporal consistency. The codes and checkpoints are publicly available at https://github.com/csbhr/Vivid-VR."},"_bibtex":{"value":"@inproceedings{\nbai2026vividvr,\ntitle={Vivid-{VR}: Distilling Concepts from Text-to-Video Diffusion Transformer for Photorealistic Video Restoration},\nauthor={Haoran Bai and Xiaoxu Chen and Canqian Yang and Zongyao He and Sibin Deng and Ying Chen},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=YV5Zgv8pdg}\n}"},"title":{"value":"Vivid-VR: Distilling Concepts from Text-to-Video Diffusion Transformer for Photorealistic Video Restoration"},"pdf":{"value":"/pdf/c0383acde3b4b19fa13cbf9ba5a770825b772ee9.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"bai|vividvr_distilling_concepts_from_texttovideo_diffusion_transformer_for_photorealistic_video_restoration"},"authorids":{"value":["~Haoran_Bai1","~Xiaoxu_Chen1","~Canqian_Yang1","~Zongyao_He1","~Sibin_Deng1","~Ying_Chen15"]},"authors":{"value":["Haoran Bai","Xiaoxu Chen","Canqian Yang","Zongyao He","Sibin Deng","Ying Chen"]}},"version":2},{"content":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"TLDR":{"value":"VLA models may easily learn shortcuts by exploiting spurious correlations. GateFlow uses transport distance to detect and suppress shortcuts while enhancing genuine understanding."},"keywords":{"value":["Vision-Language-Action","Flow Matching","Foundation Models","Embodied AI","Robotics"]},"primary_area":{"value":"applications to robotics, autonomy, planning"},"abstract":{"value":"Vision-Language-Action (VLA) models promise general-purpose robotic intelligence by leveraging pretrained vision-language representations. However, these models suffer from shortcut learning—exploiting spurious correlations between visual patterns and actions rather than developing semantic understanding. This occurs because VLA models optimize an Evidence Lower Bound (ELBO) proxy instead of the true likelihood, creating an optimization gap that enables memorized patterns to masquerade as genuine solutions. To mitigate this problem, we introduce GateFlow, a transport-guided gating mechanism that detects and suppresses shortcut learning by measuring the Wasserstein distance between observation and action representations. Low transport distance indicates semantic understanding and receives strong enhancement, while high distance reveals shortcuts and triggers suppression. This selective gating closes the ELBO-NLL gap by guiding optimization toward true likelihood minimization. We provide theoretical guarantees showing that GateFlow concentrates gradients on semantic features while eliminating spurious patterns. Empirically, GF-VLA achieves state-of-the-art performance on various tasks, with substantial improvements on long-range tasks or complex scenarios under non-stationary perturbations. GateFlow integrates seamlessly into existing VLA architectures with minimal computational overhead, offering a practical solution to more general robotic learning."},"_bibtex":{"value":"@misc{\nzhang2025gateflow,\ntitle={GateFlow: Mitigating Shortcut Learning in {VLA} Models via Gated Flow Matching},\nauthor={Wanpeng Zhang and Ye Wang and Hao Luo and Haoqi Yuan and Yicheng Feng and Sipeng Zheng and Qin Jin and Zongqing Lu},\nyear={2025},\nurl={https://openreview.net/forum?id=qOSy2PX4xS}\n}"},"title":{"value":"GateFlow: Mitigating Shortcut Learning in VLA Models via Gated Flow Matching"},"pdf":{"value":"/pdf/691cf8ebbc33f2ff0fab9ef100c422af853434ff.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"zhang|gateflow_mitigating_shortcut_learning_in_vla_models_via_gated_flow_matching"},"authorids":{"value":["~Wanpeng_Zhang1","~Ye_Wang12","~Hao_Luo6","~Haoqi_Yuan1","~Yicheng_Feng1","~Sipeng_Zheng1","~Qin_Jin1","~Zongqing_Lu2"]},"authors":{"value":["Wanpeng Zhang","Ye Wang","Hao Luo","Haoqi Yuan","Yicheng Feng","Sipeng Zheng","Qin Jin","Zongqing Lu"]}},"tmdate":1763121700289,"tcdate":1757953073129,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission6112/Authors"],"signatures":["ICLR.cc/2026/Conference/Submission6112/Authors"],"forum":"qOSy2PX4xS","license":"CC BY 4.0","number":6112,"cdate":1757953073129,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/-/Submission","ICLR.cc/2026/Conference/-/Post_Submission","ICLR.cc/2026/Conference/Submission6112/-/Full_Submission","ICLR.cc/2026/Conference/-/Withdrawn_Submission"],"mdate":1763121700289,"odate":1759896705795,"domain":"ICLR.cc/2026/Conference","id":"qOSy2PX4xS","version":2},{"content":{"summary":{"value":"This paper theoretecally and empirically showed that the inductive bias of default- ERM maximizing the margin causes shortcut learning in  a linear perception task.\nIt proposed uniform margins that leads to models that depend more on the stable than the shortcut feature and suggested loss functions encourage uniform-margin solutions."},"presentation":{"value":"3 good"},"contribution":{"value":"3 good"},"soundness":{"value":"3 good"},"strengths":{"value":"This paper analyzes shortcut learning theoretically in terms of margin maximization.\nI have not seen an analysis from this perspective before.\nIt also proposes the concept of uniform margin from theory and suggests a method to prevent shortcut learning."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The theory itself has limited applicability due to the linear model, and we do not know how well it can actually be explained in general terms in actual deep learning models."},"confidence":{"value":"3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked."},"questions":{"value":"It seems that this paper forcus on a theoretical analysis in a situation like background shortcuts, where shortcut features make up a large portion of the data as shown in term B. \nLearning  features with a large B preferentially is intuitively reasonable in such a simple situation; however, reading the proof in the Appendix, it seems to require a very tough proof to show that. \nI would like to know where the difficulty of the proof lies, what you devised in the Appendix proof ,and where you derived the key inequality.\n\nShortcut features do not always occupy as much space as stable features as background features. For example, in the following paper, the percentage of partial input is much smaller.\n>Overinterpretation reveals image classification model pathologies Brandon Carter, Siddhartha Jain, Jonas Mueller, David Gifford\nIsn't Bz>y important for the theoretical analysis?\nIf Bz<y, what part of the proof is more difficult? Or what other assumptions are necessary?\n\n\n\n"},"rating":{"value":"6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations."},"code_of_conduct":{"value":"Yes"},"limitations":{"value":"The theory itself has limited applicability due to the linear model, and we do not know how well it can actually be explained in general terms in actual deep learning models."}},"nonreaders":[],"tmdate":1702411520565,"tcdate":1688617232745,"writers":["NeurIPS.cc/2023/Conference","NeurIPS.cc/2023/Conference/Submission15594/Reviewer_9QKU"],"signatures":["NeurIPS.cc/2023/Conference/Submission15594/Reviewer_9QKU"],"forum":"zyZkaqNnpa","number":1,"license":"CC BY 4.0","cdate":1688617232745,"mdate":1702411520565,"readers":["everyone"],"invitations":["NeurIPS.cc/2023/Conference/Submission15594/-/Official_Review","NeurIPS.cc/2023/Conference/-/Edit"],"domain":"NeurIPS.cc/2023/Conference","replyto":"zyZkaqNnpa","id":"N1iZVpqdC6","forumContent":{"venue":{"value":"NeurIPS 2023 poster"},"keywords":{"value":["shortcut learning","spurious correlations","perfect stable feature","perception tasks","implicit bias in optimization","improving inductive biases"]},"_bibtex":{"value":"@inproceedings{\npuli2023dont,\ntitle={Don{\\textquoteright}t blame Dataset Shift! Shortcut Learning due to Gradients and Cross Entropy},\nauthor={Aahlad Manas Puli and Lily H Zhang and Yoav Wald and Rajesh Ranganath},\nbooktitle={Thirty-seventh Conference on Neural Information Processing Systems},\nyear={2023},\nurl={https://openreview.net/forum?id=zyZkaqNnpa}\n}"},"title":{"value":"Don’t blame Dataset Shift! Shortcut Learning due to Gradients and Cross Entropy"},"paperhash":{"value":"puli|dont_blame_dataset_shift_shortcut_learning_due_to_gradients_and_cross_entropy"},"TLDR":{"value":"Implicit biases toward maximizing margins induce shortcut learning in ERM even in tasks with perfect stable features, controlling margins mitigates shortcuts"},"abstract":{"value":"Common explanations for shortcut learning assume that the shortcut improves prediction only under the training distribution. Thus, models trained in the typical way by minimizing log-loss using gradient descent, which we call default-ERM, should utilize the shortcut. However, even when the stable feature determines the label in the training distribution and the shortcut does not provide any additional information, like in perception tasks, default-ERM exhibits shortcut learning. Why are such solutions preferred when the loss can be driven to zero when using the stable feature alone? By studying a linear perception task, we show that default-ERM’s preference for maximizing the margin, even without overparameterization, leads to models that depend more on the shortcut than the stable feature. This insight suggests that default-ERM’s implicit inductive bias towards max-margin may be unsuitable for perception tasks. Instead, we consider inductive biases toward uniform margins. We show that uniform margins guarantee sole dependence on the perfect stable feature in the linear perception task and suggest alternative loss functions, termed margin control (MARG-CTRL), that encourage uniform-margin solutions. MARG-CTRL techniques mitigate shortcut learning on a variety of vision and language tasks, showing that changing inductive biases can remove the need for complicated shortcut-mitigating methods in perception tasks."},"pdf":{"value":"/pdf/956a72600022df2e6738151ad2ff6a638cc32154.pdf"},"venueid":{"value":"NeurIPS.cc/2023/Conference"},"authorids":{"value":["~Aahlad_Manas_Puli1","~Lily_H_Zhang1","~Yoav_Wald1","~Rajesh_Ranganath2"]},"authors":{"value":["Aahlad Manas Puli","Lily H Zhang","Yoav Wald","Rajesh Ranganath"]}},"version":2},{"content":{"summary":{"value":"This paper mainly introduces a new dataset called GroundMoRe for motion-centric pixel-level language grounding in videos. It collects a large number of video clips with question-answer pairs and object masks and also proposes a new baseline method MoRA to handle the new task on the new dataset. Extensive statistics and experiments have shown the advantages of the dataset and the effectiveness of the proposed method."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"Please refer to the weaknesses."},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1.The paper is well-written and clearly presents the motivations of the proposed dataset as a new motion-centric video benchmark. And there are sufficient examples and figures given in the manuscript to illustrate the dataset characteristics.\n\n2.I think the new task Motion-Grounded Video Reasoning proposed in this work is a more practical setting to benchmark the video multimodal models, since it considers the capability of temporal localization which is commonly neglected by previous methods.\n\n3.The proposed baseline model is reasonable and shows effectiveness in addressing Motion-Grounded Video Reasoning on GroundMoRe."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1.Although one of the most significant contributions claimed by the authors is to introduce implicit reasoning into the video question-answering and segmentation models, this idea has been previously explored by several works like VISA [1] and ViLLa [2]. It seems that there is no obvious difference in terms of the introduction of implicit textual inputs compared to these existing works, and this makes the contribution and novelty of this work not that impressive.\n\n2.In terms of the Sequential question type in the proposed GroundMoRe dataset, I have some concerns on the effectiveness of it for truly reflecting the model's ability to reason about the temporal relations of motions. For example, as shown in Figure 1, the query input \"Who dribbled the ball before he accelerates passing the man in pink shorts?\" involves two consecutive motions, but there is only one motion of \"dribbling the ball\" in the video clip so the model can easily localize the first motion without looking at the second motion to find the correct answer, while this behavior cannot be regarded as temporal relation reasoning. I think similar problem will also occur in other examples mentioned in the manuscript like \"the woman opened the refrigerator before taking out the milk\". In my view, some samples should be explicitly constructed for this question type, where a same motion occurs twice but one of them is accompanied with another context motion, so the model has to rely on the temporal relation with the context motion to localize the answer.\n\n3.To my understanding, the most interesting contribution of this work is to take the temporal grounding into consideration during the benchmarking process. However, it seems that the evaluation and investigation of temporal grounding capabilities are quite simple. For example, I don't see the authors discuss any metrics regarding the temporal grounding capabilities, such as the typically adopted temporal Intersection over Union (tIoU). Actually I think there should be a more comprehensive metric considering the spatio-temporal localization ability in this work, which can be derived from a simple extension to the existing vIoU metric proposed in Spatio-Temporal Video Grounding [3, 4]. But the current metrics discussed in this work are still focusing on the spatial segmentation.\n\n4.According to the experimental results presented by Table 4, I feel kind of confused that for the Descriptive question type, the performance improvement is still remarkable when replacing the implicit questions with the referring expressions, i.e., the first row and second row for all methods in Table 4, and sometimes this improvement is even more obvious than other question types. Given that the Descriptive questions already convey a lot of description details, why could this phenomenon happen? Intuitively I think the difference between a descriptive question and a referring expression is quite minor, maybe there is just something like whether the concrete subject word is mentioned or not comparing these two kinds of textual inputs?\n\nReferences\n\n[1] VISA: Reasoning Video Object Segmentation via Large Language Models, Yan et al.\n\n[2] ViLLa: Video Reasoning Segmentation with Large Language Model, Zheng et al.\n\n[3] Where Does It Exist: Spatio-Temporal Video Grounding for Multi-Form Sentences, Zhang et al.\n\n[4] Human-Centric Spatio-Temporal Video Grounding With Visual Transformers, Tang et al."}},"nonreaders":[],"tmdate":1731427910480,"tcdate":1731063297636,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission7125/Reviewer_mqJ7"],"signatures":["ICLR.cc/2025/Conference/Submission7125/Reviewer_mqJ7"],"forum":"tEei1bolt3","number":4,"license":"CC BY 4.0","cdate":1731063297636,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission7125/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427910480,"domain":"ICLR.cc/2025/Conference","replyto":"tEei1bolt3","id":"D2aOWU1Hzc","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Motion","Video Grounding","Video Reasoning"]},"supplementary_material":{"value":"/attachment/1aad1cbcdd0b87f5eb4c884b6a886c85a596c3c5.zip"},"primary_area":{"value":"datasets and benchmarks"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"In this paper, we introduce Motion-Grounded Video Reasoning, a new motion understanding task that requires generating visual answers (video segmentation masks) according to the input question, and hence needs implicit spatiotemporal reasoning and grounding. This task extends existing spatiotemporal grounding work focusing on explicit action/motion grounding, to a more general format by enabling implicit reasoning via questions. To facilitate the development of the new task, we collect a large-scale dataset called GroundMoRe, which comprises 1,673 video clips, 243K object masks that are deliberately designed with 4 question types (Causal, Sequential, Counterfactual, and Descriptive) for benchmarking deep and comprehensive motion reasoning abilities. GroundMoRe uniquely requires models to generate visual answers, providing a more concrete and visually interpretable response than plain texts. It evaluates models on both spatiotemporal grounding and reasoning, fostering to address complex challenges in motion-related video reasoning, temporal perception, and pixel-level understanding. Furthermore, we introduce a novel baseline model named Motion-Grounded Video Reasoning Assistant (MoRA). MoRA incorporates the multimodal reasoning ability from the Multimodal LLM, the pixel-level perception capability from the grounding model (SAM), and the temporal perception ability from a lightweight localization head. MoRA achieves respectable performance on GroundMoRe outperforming the best existing visual grounding baseline model by an average of 28.8% relatively. We hope this novel and challenging task will pave the way for future advancements in robust and general motion understanding via video reasoning segmentation."},"_bibtex":{"value":"@misc{\ndeng2024motiongrounded,\ntitle={Motion-Grounded Video Reasoning: Understanding and Perceiving Motion at Pixel Level},\nauthor={Andong Deng and Tongjia Chen and Shoubin Yu and Taojiannan Yang and Lincoln Spencer and Yapeng Tian and Ajmal Saeed Mian and Mohit Bansal and Chen Chen},\nyear={2024},\nurl={https://openreview.net/forum?id=tEei1bolt3}\n}"},"title":{"value":"Motion-Grounded Video Reasoning: Understanding and Perceiving Motion at Pixel Level"},"pdf":{"value":"/pdf/2849202c7adcb654fa04efd8ea9660f6f344fb45.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"deng|motiongrounded_video_reasoning_understanding_and_perceiving_motion_at_pixel_level"},"authorids":{"value":["~Andong_Deng2","~Tongjia_Chen1","~Shoubin_Yu1","~Taojiannan_Yang1","~Lincoln_Spencer1","~Yapeng_Tian1","~Ajmal_Saeed_Mian1","~Mohit_Bansal2","~Chen_Chen18"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Andong Deng","Tongjia Chen","Shoubin Yu","Taojiannan Yang","Lincoln Spencer","Yapeng Tian","Ajmal Saeed Mian","Mohit Bansal","Chen Chen"]}},"version":2},{"content":{"venue":{"value":"IEEE Transactions on Circuits and Systems for Video Technology"},"pdf":{"value":"https://ieeexplore.ieee.org/iel7/76/9812753/09632538.pdf"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"xu|learning_an_occlusionaware_network_for_video_deblurring"},"html":{"value":"https://doi.org/10.1109/TCSVT.2021.3132102"},"abstract":{"value":"Video deblurring is a challenging task since the blur is caused by camera shake, object motions, etc. The success of the state-of-the-art methods stems mainly from exploiting the temporal information of neighboring frames through alignment. When there exists occlusion among the sequence, these approaches become less effective for inaccurate alignment. In this paper, we propose an effective occlusion-aware network to handle the occlusion for video deblurring. The proposed module first generates a coarse pixel-wise alignment filter to explore the temporal information and then learns an adaptive affine transformation to deal with the occluded areas. In addition, a self-attention mechanism is developed to better model the occluded pixels. To further improve the performance, we progress a multi-scale strategy and train the network in an end-to-end manner. Both quantitative and qualitative experimental results show that the proposed method achieves favorable performance against state-of-the-art methods on the benchmark datasets. The code and trained models are available at: https://github.com/XQLuck/code.git"},"title":{"value":"Learning an Occlusion-Aware Network for Video Deblurring"},"authors":{"value":[{"fullname":"Qian Xu"},{"fullname":"Jinshan Pan"},{"fullname":"Yuntao Qian","username":"~Yuntao_Qian1"}]}},"tmdate":1789093158000,"pdate":1656633600000,"externalIds":["doi:10.1109/tcsvt.2021.3132102"],"tcdate":1762322422221,"writers":["~","OpenReview.net/Public_Article/ORCID.org","OpenReview.net/Support"],"signatures":["~Yuntao_Qian1"],"forum":"2o9dI1x5GU","license":"CC BY-SA 4.0","number":5504,"cdate":1656705073568,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/ORCID.org/-/Record","OpenReview.net/Public_Article/-/Edit","OpenReview.net/-/Edit"],"mdate":1789093158000,"domain":"OpenReview.net/Public_Article","id":"2o9dI1x5GU","version":2},{"content":{"summary":{"value":"The paper investigates how a small, two-layer attention-only transformer learns an n-gram recall in-context learning (ICL) task, where the model must retrieve the correct output by matching the final \\(n=4\\) tokens of a cue seen earlier in the same sequence. Building on previously proposed attention-based solutions for n-gram recall (Varré et al.), the paper constructs a minimal model that implements those solutions. This model admits a fully analytic treatment under gradient flow and captures the learning behavior observed in practice. Experiments training the model are consistent with the analytic predictions."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. Ablations (A1--A4): How do ablations of A1, A2, A3, and A4 affect empirical training outcomes? Does the phenomenology seen in the full simple model persist under these ablations?\n\n2. Beyond n-gram ICL: How can the findings be extended to tasks outside n-gram recall? In particular, some classes of algorithmic ICL tasks (e.g., linear regression) have known transformer implementations (e.g., Lu et al., 2025). How can the paper’s findings be generalized to that class?\n\n3. Timing of transitions: Does the model predict the time points of performance ``jumps''? If so, do these predictions match experimental results?\n\n4. Initializations: Can you demonstrate how varying initialization affects the results in Figure 3?\n\nReferences\n[1] Varre, A., Yüce, G., Flammarion, N. “Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points.” 2025\n\n[2] Lu, Yue M., et al. \"Asymptotic theory of in-context learning by linear attention.\" Proceedings of the National Academy of Sciences 122.28 (2025): e2502599122."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. The motivation for introducing a minimal transformer model is clear. The proposed minimal model capture n-gram recall ICL solution.\n\n2. The minimal transformer model is fully analytic under gradient flow and yields closed-form training dynamics. Analyzing these dynamics reveals a conservation law, initialization dependence, and phase ordering"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The abstracted transformer architecture may be too minimal. While recent work aims to reduce complexity without losing fidelity, it remains unclear how much intuition from this abstraction transfers to real LM systems.\n\n2. The focus is limited to n-gram ICL, a setting already known to work as described; the incremental insights appear modest.\n\n3. Experiments are conducted in a specific setting without varying key factors such as \\(N\\) (the n-gram length) or architectural choices."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762942639558,"tcdate":1761501411231,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission23386/Reviewer_g8Xr"],"signatures":["ICLR.cc/2026/Conference/Submission23386/Reviewer_g8Xr"],"forum":"1pTzWVvwEd","number":1,"license":"CC BY 4.0","cdate":1761501411231,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission23386/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762942639558,"domain":"ICLR.cc/2026/Conference","replyto":"1pTzWVvwEd","id":"uZZs670nFO","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["incremental learning","stage-wise learning","training dynamics","incontext learning","associative recall"]},"primary_area":{"value":"learning theory"},"abstract":{"value":"Transformers acquire in-context learning abilities in abrupt phases during training, often unfolding over multiple stages, during which certain keys circuits like induction heads emerge. In this work, we characterize the dynamics behind the emergence of such circuits during these stages. We focus on a synthetic associative recall task, where sequences are drawn from random maps between a permutation group and a vocabulary range and the model is required to complete the mapping of a permutation by retrieving it from the context. On this task, we study the trajectories of gradient flow of a simplified two-layer, attention-only transformer. Leveraging symmetries in both the transformer architecture and the data, we derive conservation laws that guide the dynamics of the parameters. These conservation laws crucially reveal how initialization —both in shape and scale— determines the order of learning as well as the timescales over which such circuits emerge revealing the implicit curriculum. Finally, we provide empirical evidence across different architectural choices, validating  our simplifications and generalizing the insights from our analysis beyond the simple setting."},"_bibtex":{"value":"@misc{\nvarre2026incremental,\ntitle={Incremental Learning in Transformers for In-Context Associative Recall},\nauthor={Aditya Varre and Nicolas Flammarion},\nyear={2026},\nurl={https://openreview.net/forum?id=1pTzWVvwEd}\n}"},"title":{"value":"Incremental Learning in Transformers for In-Context Associative Recall"},"pdf":{"value":"/pdf/5b60533070ea5ec48d0e722b1381fdb9649baaa2.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"varre|incremental_learning_in_transformers_for_incontext_associative_recall"},"authorids":{"value":["~Aditya_Varre1","~Nicolas_Flammarion1"]},"authors":{"value":["Aditya Varre","Nicolas Flammarion"]}},"version":2},{"content":{"summary":{"value":"This papers explores the text-to-sounding-video generation task, in the context of jointly fine-tuning a pair of pretrained text-to-audio and  text-to-video generation models, with the following contributions:\n- Use different text prompts for each tower as opposed to a single prompt in previous works.\n- A Hierarchical Visual-Grounded Captioning (HVGC) framework meant to generate pairs of disentangled captions: one for video with emphasis on the visual aspect, one for audio with an emphasis on sound.\n- Incorporate a dual cross attention mechanism in the joint model architecture for modality interaction, named BridgeDiT."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"- What does the 1000 value correspond to between equations 5 and 6?\n- It is common to employ temporally-aware positional encodings when interacting between video and audio modalities (that usually operate at different frame rates). Does the proposed model employ such design?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"I believe that despite its apparent simplicity this paper is significant for the joint audio-video generation task.\n\nUsing different prompts for audio and video is clever and dual cross attention is a natural intermodal network design. This paper conducts rigorous ablations on both aspects to demonstrate their benefit for the task.\n\nMoreover the HVGC builds on the intuition that audio captioning is more accurate when being in context of the visual scene."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The major weakness of the paper is the lack of novelty.\n- Using different prompts for audio and video is clever but common sense.\n- The visually grounded audio captioning tasks lack a related works section. \n- Dual cross-attention architectures have been employed in the past for other multimodal tasks, yet this paper cites almost none of them (e.g. `Generative Spoken Dialogue Language Modeling` by Nguyen et al., `Dual-Stream Diffusion Net for Text-to-Video Generation` by Liu et al., `Domain Adaptation via Bidirectional Cross-Attention Transformer` by Wang et al., `Multimodal Transformer for Unaligned Multimodal Language Sequences`, `ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks`, `Pano-AVQA: Grounded Audio-Visual Question Answering on 360◦ Videos`). Moreover the two years old MM-diffusion also employs dual cross attention for the same task."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1764358686467,"tcdate":1762136304180,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission533/Reviewer_SK8e"],"signatures":["ICLR.cc/2026/Conference/Submission533/Reviewer_SK8e"],"forum":"jcVOMVkljY","number":4,"license":"CC BY 4.0","cdate":1762136304180,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission533/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1764358686467,"domain":"ICLR.cc/2026/Conference","replyto":"jcVOMVkljY","id":"nOlBACz8vq","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"TLDR":{"value":"This work presents a systematic approach to text-to-sounding-video generation task, tackling modal conditioning and interaction challenges through disentangled prompting and symmetric feature fusion to achieve superior synchronization."},"keywords":{"value":["Multi-modal learning","sounding video generation"]},"primary_area":{"value":"generative models"},"abstract":{"value":"This study focuses on a challenging yet promising task, Text-to-Sounding-Video (T2SV) generation, which aims to generate a video with synchronized audio from text conditions, meanwhile ensuring both modalities are aligned with text. \nDespite progress in joint audio-video training, two critical challenges still remain unaddressed: (1) a single, shared text caption where the text for video is equal to the text for audio $(T_V = T_A)$ often creates modal interference, confusing the pretrained backbones, and (2) the optimal mechanism for cross-modal feature interaction remains unclear.\nTo address these challenges, we first propose the Hierarchical Visual-Grounded Captioning (HVGC) framework that generates pairs of disentangled captions, a video caption ($T_V$), and an audio caption ($T_A$), eliminating interference at the conditioning stage.\nBased on HVGC, we further introduce BridgeDiT, a novel dual-tower diffusion transformer, which employs a Dual CrossAttention (DCA) mechanism that acts as a robust ``bridge\" to enable a symmetric, bidirectional exchange of information, achieving both semantic and temporal synchronization.\nExtensive experiments on three benchmark datasets, supported by human evaluations, demonstrate that our method achieves state-of-the-art results on most metrics. Comprehensive ablation studies further validate the effectiveness of our contributions, offering key insights for the future T2SV task. All the codes and checkpoints will be publicly released."},"_bibtex":{"value":"@misc{\nguan2026taming,\ntitle={Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction},\nauthor={Kaisi Guan and Xihua Wang and Zhengfeng Lai and Xin Cheng and Peng Zhang and Xiaojiang Liu and Ruihua Song and Meng Cao},\nyear={2026},\nurl={https://openreview.net/forum?id=jcVOMVkljY}\n}"},"title":{"value":"Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction"},"pdf":{"value":"/pdf/44c683f892c30063f615cc3f57a5aaf3664fbcc4.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"guan|taming_texttosounding_video_generation_via_advanced_modality_condition_and_interaction"},"authorids":{"value":["~Kaisi_Guan1","~Xihua_Wang2","~Zhengfeng_Lai1","~Xin_Cheng6","~Peng_Zhang3","~Xiaojiang_Liu2","~Ruihua_Song1","~Meng_Cao2"]},"authors":{"value":["Kaisi Guan","Xihua Wang","Zhengfeng Lai","Xin Cheng","Peng Zhang","Xiaojiang Liu","Ruihua Song","Meng Cao"]}},"version":2},{"content":{"venue":{"value":"CVPR 2026"},"abstract":{"value":"Diffusion Transformers (DiTs) achieve strong video generation performance but suffer from prohibitive computation cost due to dense spatiotemporal tokenization. Most existing works rely on uniform patchification, tokenizing non-overlapping spatiotemporal with a fixed patch size regardless of the underlying content. This content-agnostic tokenization results in substantial redundant computation, especially in visually simple or static areas. To address this inefficiency while preserving the video generation quality, we propose DynaPatch, a fine-grained dynamic patchification framework that adaptively selects patch sizes for each spatiotemporal region based on content complexity. A lightweight router predicts patch sizes directly from the latents encoded by 3D Variational Autoencoder (VAE), and is jointly optimized with the diffusion model through diffusion loss, an attention-guided saliency alignment loss, and a token-budget regularizer. Learnable patchify/unpatchify layers integrate seamlessly with standard DiT backbones, allowing flexible tokenization without architectural changes. Experiments demonstrate that DynaPatch can effectively reduce redundant computations while preserving fine details, achieving 1.3–1.8x acceleration with minimal quality degradation. On VBench, DynaPatch attains a Total Score of 83.42 at 30% token reduction, significantly outperforming prior patchification and token pruning approaches. These results indicate that content-aware patchification offers an effective direction for efficient and scalable video diffusion. Project page: https://shengli99.github.io/DynaPatch/."},"_bibtex":{"value":"@inproceedings{\nli2026contentaware,\ntitle={Content-Aware Dynamic Patchification for Efficient Video Diffusion},\nauthor={Sheng Li and Connelly Barnes and Mamshad Nayeem Rizve and Hongwu Peng and Zhengang Li and Ohi Dibua and Alireza Ganjdanesh and Xulong Tang and Yan Kang and Yifan Gong},\nbooktitle={Conference on Computer Vision and Pattern Recognition 2026},\nyear={2026},\nurl={https://openreview.net/forum?id=B5RU3MEWOb}\n}"},"title":{"value":"Content-Aware Dynamic Patchification for Efficient Video Diffusion"},"pdf":{"value":"https://openaccess.thecvf.com/content/CVPR2026/papers/Li_Content-Aware_Dynamic_Patchification_for_Efficient_Video_Diffusion_CVPR_2026_paper.pdf"},"venueid":{"value":"thecvf.com/CVPR/2026/Conference"},"paperhash":{"value":"li|contentaware_dynamic_patchification_for_efficient_video_diffusion"},"authorids":{"value":["~Sheng_Li16","~Connelly_Barnes2","~Mamshad_Nayeem_Rizve1","~Hongwu_Peng1","~Zhengang_Li2","~Ohi_Dibua1","~Alireza_Ganjdanesh1","~Xulong_Tang1","~Yan_Kang1","~Yifan_Gong2"]},"authors":{"value":["Sheng Li","Connelly Barnes","Mamshad Nayeem Rizve","Hongwu Peng","Zhengang Li","Ohi Dibua","Alireza Ganjdanesh","Xulong Tang","Yan Kang","Yifan Gong"]}},"tmdate":1789656693190,"pdate":1789656438702,"tcdate":1765223032676,"writers":["thecvf.com/CVPR/2026/Conference","thecvf.com/CVPR/2026/Conference/Submission42753/Authors"],"signatures":["thecvf.com/CVPR/2026/Conference/Submission42753/Authors"],"forum":"B5RU3MEWOb","license":"CC BY 4.0","number":42753,"cdate":1765223032676,"readers":["everyone"],"invitations":["thecvf.com/CVPR/2026/Conference/-/Submission","thecvf.com/CVPR/2026/Conference/Submission42753/-/Full_Submission","thecvf.com/CVPR/2026/Conference/-/Post_Submission","thecvf.com/CVPR/2026/Conference/Submission42753/-/Supplementary_Material","thecvf.com/CVPR/2026/Conference/-/Edit","thecvf.com/CVPR/2026/Conference/-/Compute_Flag"],"mdate":1789656693190,"odate":1789656438702,"domain":"thecvf.com/CVPR/2026/Conference","id":"B5RU3MEWOb","version":2},{"content":{"comment":{"value":"Dear Reviewer,\n\nThank you for your valuable and very detailed feedback. We will address each comment, and refer to respective changes to the paper or experiments. To make it easy to spot changes in the updated paper PDF, we highlight changes between red brackets “[[...]]”. \n\n### **Additional metrics:**\nThe reviewer suggested five additional metrics to our primary metric (Minimal Pair Score):\n\n1. **Single-answer-accuracy & pair agreement**: We add two additional tables (Table 7 and 8) where we report both single-answer accuracy and pair agreement, and add a discussion in the paper.\n2. **Uncertainty measures**: Most VideoLLMs have “do_sample=False” during text generation, so temperature is effectively 0; this makes sense in a multi-choice QA format (model only needs to produce a single letter response). However, we conduct an ablation with various prompting strategies and report these results in the paper; we observe little variance in downstream performance.\n3. **Confidence-accuracy calibration curves:** We show the relation between confidence (logprob of the answer letter A vs B) and accuracy in Figure 6 of the updated paper.\n4. **Pair difficulty distribution:** While MVP does not explicitly have a split such as “hard” vs “easy”, we add an analysis in the paper (p. 13) with three proxies for difficulty: video length on the vision side and word frequency/sentence difficulty on the text side. \n\n### **Quantity and proportion of noisy or ambiguous dataset cases:**\nTo quantify the level of noise in our data, we had already collected human accuracy on MVP (see Table 4). Additionally we now provide the explicit human accuracy for each of the 9 data sources in the appendix, and further analyze certain subsets in detail with manual inspection. This is added at the end of Section 3. Similarly, we expanded on the procedure for mining and constructing minimal pairs (e.g. Language Table) in the updated Appendix.\n\n### **Replacing multimodal-LLM-ensemble for single-frame-filtering with a more traditional method:**\nWhile it is possible to explore alternatives for identifying single-frame-solvable examples, MVP involves open-world concepts, and therefore such action classifiers would similarly need to perform open-world concept discrimination. One option could be to leverage a Video-CLIP model non-parametrically, but we would expect performance to be worse and more sensitive to wording.\n\n### **Expand the GPT-4o and Gemini evaluation to a representative subset, standardizing the frame rate, resolution, and prompt:**\nThe GPT-4o and Gemini evaluations are conducted on MVP-mini, which is a representative subset of the full dataset, we will clarify this in the paper; we standardize the prompts across models.\nRegarding frame rate and resolution, API models are often trained with different default parameters and often provide (or internally apply) their own custom preprocessing to the video. We adopt the default parameters for each (API) model, but have now also conducted a small test on Gemini: we run a subset of examples from each category split of MVP with different frame numbers (8,16,32) as well as 1fps (6 different combinations), and find no difference in prediction. Only when we change Gemini’s default resolution “low”, do we see a significant drop in performance from 37.5% to 12.5%.\n\n### **More detailed analysis of where model fails:**\nSpecifically for the intuitive physics subsets you mentioned as hard for models, we provide a short discussion of recent literature in the updated paper (p. 9).\nWe also expand our existing “Fine-grained failure analysis” section in Section 4.1 where we discuss various subsets, and show intuitive physics examples.\nFinally, we also note that we added additional shortcut baselines to the main results Table 4 on MVP.\n\n### **More extensive prompting/CoT analysis:**\nWith Gemini we observed that performance did not improve when allowing more time to reason before the answer.\nSo we tested if the same trend would hold for our strongest open-source VideoLLM (InternVL2.5-8B) for two prompts:\n1. The same as provided to Gemini, so allowing short reasoning of 1-3 sentences.\n2.  A CoT prompt to reason for longer “step-by-step” and listing all relevant objects, key events etc.\n\nWe find that CoT prompts only marginally helps, similar to Gemini (summarized in updated paper, Table 6).\nWhile few-shot prompting is very interesting, we did not explore it since the majority of models are never trained with more than 1 video as input, and requires extensive compute resources..\n\n\nOverall, we hope this addressed your main suggestions, and that our updated paper reflects this adequately. We are happy to discuss further in the following days and provide more evidence."},"title":{"value":"Addressing main suggestions such as providing additional metrics on top of Minimal Pair Score"}},"parentInvitations":"TMLR/-/Official_Comment","tmdate":1758579589878,"tcdate":1758579589878,"writers":["TMLR","TMLR/Paper5382/Authors"],"signatures":["TMLR/Paper5382/Authors"],"forum":"gvFgNJcSw1","number":8,"license":"CC BY 4.0","cdate":1758579589878,"readers":["everyone"],"invitations":["TMLR/Paper5382/-/Official_Comment"],"mdate":1758579589878,"domain":"TMLR","replyto":"7lQXy4NT1H","id":"ojdxzmJAoL","forumContent":{"submission_length":{"value":"Regular submission (no more than 12 pages of main content)"},"venue":{"value":"Accepted by TMLR"},"abstract":{"value":"Existing benchmarks for assessing the spatio-temporal understanding and reasoning abilities of video language models are susceptible to score inflation due to the presence of shortcut solutions based on superficial visual or textual cues. This paper mitigates the challenges in accurately assessing model performance by introducing the Minimal Video Pairs (MVP) benchmark, a simple shortcut-aware video QA benchmark for assessing the physical understanding of video language models. The benchmark is comprised of 55K high-quality multiple-choice video QA examples focusing on physical world understanding. Examples are curated from nine video data sources, spanning first-person egocentric and exocentric videos, robotic interaction data, and cognitive science intuitive physics benchmarks. To mitigate shortcut solutions that rely on superficial visual or textual cues and biases, each sample in MVP has a minimal-change pair — a visually similar video accompanied by an identical question but an opposing answer. To answer a question correctly, a model must provide correct answers for both examples in the minimal-change pair; as such, models that solely rely on visual or textual biases would achieve below random performance. Human performance on MVP is 92.9%, while the best open-source state-of-the- art video-language model achieves 40.2% compared to random performance at 25%."},"_bibtex":{"value":"@article{\nkrojer2025a,\ntitle={A Shortcut-aware Video-{QA} Benchmark for Physical Understanding via Minimal Video Pairs},\nauthor={Benno Krojer and Mojtaba Komeili and Candace Ross and Quentin Garrido and Koustuv Sinha and Nicolas Ballas and Mido Assran},\njournal={Transactions on Machine Learning Research},\nissn={2835-8856},\nyear={2025},\nurl={https://openreview.net/forum?id=gvFgNJcSw1},\nnote={}\n}"},"title":{"value":"A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs"},"pdf":{"value":"/pdf/f01dd06430bcd7b2aba068243d47011dcc96be6e.pdf"},"venueid":{"value":"TMLR"},"paperhash":{"value":"krojer|a_shortcutaware_videoqa_benchmark_for_physical_understanding_via_minimal_video_pairs"},"authorids":{"value":["~Benno_Krojer1","~Mojtaba_Komeili1","~Candace_Ross1","~Quentin_Garrido1","~Koustuv_Sinha1","~Nicolas_Ballas1","~Mido_Assran1"]},"assigned_action_editor":{"value":"~Tae-Hyun_Oh3"},"authors":{"value":["Benno Krojer","Mojtaba Komeili","Candace Ross","Quentin Garrido","Koustuv Sinha","Nicolas Ballas","Mido Assran"]}},"version":2},{"content":{"summary":{"value":"This paper presents LongLive, a method for fine-tuning pretrained video diffusion models for long video generation. First, it introduces KV-Recache, which regenerates the KV cache using the previously generated frames together with the new prompt when a prompt change occurs. This allows the model to generate videos that remain visually coherent with past frames while adapting better to the new prompt, rather than being stuck with the previous inputs. Moreover, the authors progressively extend the generation length by fine-tuning DMDs with an increasing number of video clips, where the gradient is computed only over the last video segment. The paper also proposes several additional techniques, such as applying an attention sink by maintaining the first frame as a global anchor. Leveraging all these techniques, LongLive demonstrates real-time long video generation (up to 240 seconds) on a single H100 GPU with minimal quality degradation."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"- What is the empirical maximum video length that the proposed model can generate without suffering from severe quality degradation? For instance, can a model generate videos even longer than 240s?\n-Why does the author choose DMD instead of other techniques, e.g., DMD2?  Is there a specific reason?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"- The paper is generally well-written and easy to follow. \n- Real-time long video generation using pretrained diffusion models is a very challenging but important problem, and I think this paper provides quite effective techniques to solve this.\n- For me, several techniques are simple (which is good) yet novel and effective, such as KV recache.\n- I really like the supplementary material that the authors attached."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- There are several missing works in the literature on long video generation, such as AAPT [1], TECO [2], MALT [3], and Rolling Diffusion [4]. I recommend that the authors include a discussion of these and other relevant works that have not been covered in the paper.\n- I also believe that having a long context is another important aspect in developing long video generation models. While being efficient, the proposed model inherently has a relatively short context length (mainly due to the attention sink that focuses on the first frame), which limits its effective context window. The method is indeed effective at generating visually plausible long videos in real time; however, it might struggle in scenarios that require longer temporal dependencies — for example, when a person disappears and reappears in the same scene, or when generating long videos with multiple scene transitions.\n\n[1] Lin et al., Autoregressive Adversarial Post-Training for Real-Time Interactive Video Generation, 2025.  \n[2] Yan et al., Temporally Consistent Transformers for Video Generation, ICML 2023.  \n[3] Yu et al., MALT Diffusion: Memory-Augmented Latent Transformers for Any-Length Video Generation, 2025.  \n[4] Ruhe et al., Rolling Diffusion Models, ICML 2024."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762916058258,"tcdate":1761862706624,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission2157/Reviewer_AEHT"],"signatures":["ICLR.cc/2026/Conference/Submission2157/Reviewer_AEHT"],"forum":"nCAODkpsPJ","number":2,"license":"CC BY 4.0","cdate":1761862706624,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission2157/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762916058258,"domain":"ICLR.cc/2026/Conference","replyto":"nCAODkpsPJ","id":"1cFaoFzreB","forumContent":{"TLDR":{"value":"We present LongLive, that is a real-time, interactive, and AR framework for long video generation."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Real-time","Interactive","Long Video Generation"]},"supplementary_material":{"value":"/attachment/6fe8f481a0c438177d3866f80a6f3b3989b24501.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"We present LongLive, a frame-level autoregressive (AR) framework for real-time and interactive long video generation. Long video generation presents challenges in both efficiency and quality. Diffusion and Diffusion-Forcing models can produce high-quality videos but suffer from low efficiency due to bidirectional attention. Causal attention AR models support KV caching for faster inference but often degrade in quality on long videos due to memory challenges during long-video training. In addition, beyond static prompt-based generation, interactive capabilities, such as streaming prompt inputs, are critical for dynamic content creation, enabling users to guide narratives in real time. This interactive requirement significantly increases the complexity, especially in ensuring visual consistency and semantic coherence during prompt transitions. To address these challenges, LongLive adopts a causal, frame-level AR design that integrates a KV-recache mechanism that refreshes cached states with the new prompt for smooth, adherent switches streaming long tuning to enable long video training and to align training and inference (train-long–test-long); and short window attention paired with a frame-level attention sink, preserving long-range consistency while enabling faster generation. With these key designs, LongLive fine-tunes a 1.3B-parameter short-clip model to minute-long generation in just 32 GPU-days. At inference, LongLive sustains 20.7 FPS on a single NVIDIA H100, achieves strong performance on VBench in both short- and long-video settings. LongLive supports up to 240-second videos on a single H100 GPU. With FP8 quantization, LongLive boosts inference to 24.8 FPS with marginal quality loss."},"_bibtex":{"value":"@inproceedings{\nyang2026longlive,\ntitle={LongLive: Real-time Interactive Long Video Generation},\nauthor={Shuai Yang and Wei Huang and Ruihang Chu and Yicheng Xiao and Yuyang Zhao and Xianbang Wang and Muyang Li and Enze Xie and Ying-Cong Chen and Yao Lu and Song Han and Yukang Chen},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=nCAODkpsPJ}\n}"},"title":{"value":"LongLive: Real-time Interactive Long Video Generation"},"pdf":{"value":"/pdf/fc6c8a70c050feda94f14f7bbd7e13df1dae002c.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"yang|longlive_realtime_interactive_long_video_generation"},"authorids":{"value":["~Shuai_Yang7","~Wei_Huang36","~Ruihang_Chu1","~Yicheng_Xiao1","~Yuyang_Zhao1","~Xianbang_Wang1","~Muyang_Li2","~Enze_Xie1","~Ying-Cong_Chen1","~Yao_Lu13","~Song_Han5","~Yukang_Chen1"]},"authors":{"value":["Shuai Yang","Wei Huang","Ruihang Chu","Yicheng Xiao","Yuyang Zhao","Xianbang Wang","Muyang Li","Enze Xie","Ying-Cong Chen","Yao Lu","Song Han","Yukang Chen"]}},"version":2},{"content":{"venue":{"value":"ICPR 2020"},"pdf":{"value":"https://ieeexplore.ieee.org/iel7/9411940/9411911/09413075.pdf"},"venueid":{"value":"dblp.org/conf/ICPR/2020"},"paperhash":{"value":"joo|slimming_resnet_by_slimming_shortcut"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Donggyu_Joo:","https://dblp.org/search/pid/api?q=author:Doyeon_Kim:","~Junmo_Kim1"]},"html":{"value":"https://doi.org/10.1109/ICPR48806.2021.9413075"},"_bibtex":{"value":"@inproceedings{DBLP:conf/icpr/JooKK20,\n  author={Donggyu Joo and Doyeon Kim and Junmo Kim},\n  title={Slimming ResNet by Slimming Shortcut},\n  year={2020},\n  cdate={1577836800000},\n  pages={7677-7683},\n  url={https://doi.org/10.1109/ICPR48806.2021.9413075},\n  booktitle={ICPR},\n  crossref={conf/icpr/2020}\n}\n"},"abstract":{"value":"Conventional network pruning methods on convolutional neural networks (CNNs) reduce the number of input or output channels of convolution layers. With these approaches, the channels in the plain network can be pruned without any restrictions. However, in the case of the ResNet based networks which have shortcuts (skip connections), the channel slimming of existing pruning methods is limited to the inside of each residual block. Since the number of Flops and parameters are also highly related to the number of channels in the shortcuts, more investigation on pruning channels in shortcuts is required. In this paper, we propose a novel pruning method, Slimming Shortcut Pruning (SSPruning), for pruning channels in shortcuts on ResNet based networks. First, we separate the long shortcut into individual regions that can be pruned independently without considering its long connections. Then, by applying our Importance Learning Gate (ILG) which learns the importance of channels globally regardless of channel type and location (i.e., in the shortcut or inside of the block), we can finally achieve an optimally pruned model. Through various experiments, we have confirmed that our method yields outstanding results when we prune the shortcuts and inside of the block together."},"title":{"value":"Slimming ResNet by Slimming Shortcut"},"authors":{"value":["Donggyu Joo","Doyeon Kim","Junmo Kim"]}},"tmdate":1746985916974,"pdate":1577836800000,"tcdate":1746985854559,"writers":["~"],"signatures":["~Junmo_Kim1"],"forum":"uLWPDPspEu","license":"CC BY-SA 4.0","number":416937,"cdate":1577836800000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1746985916974,"domain":"DBLP.org","id":"uLWPDPspEu","version":2},{"content":{"venue":{"value":"IEEE Trans. Circuits Syst. Video Technol. 2024"},"pdf":{"value":"https://ieeexplore.ieee.org/iel8/76/10768433/10555372.pdf"},"venueid":{"value":"dblp.org/journals/TCSV/2024"},"paperhash":{"value":"qi|collaborative_debias_strategy_for_temporal_sentence_grounding_in_video"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Zhaobo_Qi:","https://dblp.org/search/pid/api?q=author:Yibo_Yuan:","https://dblp.org/search/pid/api?q=author:Xiaowen_Ruan:","~Shuhui_Wang1","https://dblp.org/search/pid/api?q=author:Weigang_Zhang:","https://dblp.org/search/pid/api?q=author:Qingming_Huang:"]},"html":{"value":"https://doi.org/10.1109/TCSVT.2024.3413074"},"_bibtex":{"value":"@article{DBLP:journals/tcsv/QiYRWZH24,\n  author={Zhaobo Qi and Yibo Yuan and Xiaowen Ruan and Shuhui Wang and Weigang Zhang and Qingming Huang},\n  title={Collaborative Debias Strategy for Temporal Sentence Grounding in Video},\n  year={2024},\n  month={November},\n  cdate={1730419200000},\n  journal={IEEE Trans. Circuits Syst. Video Technol.},\n  volume={34},\n  number={11},\n  pages={10972-10986},\n  url={https://doi.org/10.1109/TCSVT.2024.3413074}\n}\n"},"abstract":{"value":"Temporal sentence grounding in video has witnessed significant advancements, but suffers from substantial dataset bias, which undermines its generalization ability. Existing debias approaches primarily concentrate on well-known distribution and linguistic biases, while overlooking the relationship among different biases, limiting their debias capability. In this work, we delve into the existence of visual bias and combinatorial bias in the widely used datasets, and introduce a collaborative debias structure that can be seamlessly integrated into present methods. It encompasses four low-capacity models, a re-label module, and a main model. Each biased model deliberately leverages bias as shortcut information to accurately perform grounding, achieved by customizing the appropriate model structure and input data format to align with the bias characteristics. During the training phase, the gradient descent direction for optimizing the main model should align with the negative gradient descent direction of the biased model that is optimized by utilizing ground truth labels. Subsequently, the re-label module introduces a gradient aggregation function, consolidating the gradient descent direction from these biased models and constructing new labels to compel the main model to effectively capture multi-modality alignment features instead of relying on shortcut contents for grounding. Finally, we design two debias structures, P-Debias and C-Debias, to exploit the independence and inclusion relationships between different types of biases. Extensive experiments on multiple span-based models over Charades-CD and ActivityNet-CD demonstrate the exceptional debias capability of our strategy (https://github.com/qzhb/CDS)."},"title":{"value":"Collaborative Debias Strategy for Temporal Sentence Grounding in Video"},"authors":{"value":["Zhaobo Qi","Yibo Yuan","Xiaowen Ruan","Shuhui Wang","Weigang Zhang","Qingming Huang"]}},"tmdate":1767802553039,"pdate":1704067200000,"externalIds":["dblp:journals/tcsv/QiYRWZH24"],"tcdate":1767802541308,"writers":["~"],"signatures":["~Shuhui_Wang1"],"forum":"5DRKTyLjS4","license":"CC BY-SA 4.0","number":730848,"cdate":1730419200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1767802553039,"domain":"DBLP.org","id":"5DRKTyLjS4","version":2},{"content":{"venue":{"value":"ACL (1) 2025"},"pdf":{"value":"https://aclanthology.org/2025.acl-long.910.pdf"},"venueid":{"value":"dblp.org/conf/ACL/2025"},"paperhash":{"value":"sterner|minimal_pairbased_evaluation_of_codeswitching"},"authorids":{"value":["~Igor_Sterner1","https://dblp.org/search/pid/api?q=author:Simone_Teufel:"]},"html":{"value":"https://aclanthology.org/2025.acl-long.910/"},"_bibtex":{"value":"@inproceedings{DBLP:conf/acl/SternerT25,\n  author={Igor Sterner and Simone Teufel},\n  title={Minimal Pair-Based Evaluation of Code-Switching},\n  year={2025},\n  cdate={1735689600000},\n  pages={18575-18598},\n  url={https://aclanthology.org/2025.acl-long.910/},\n  booktitle={ACL (1)},\n  crossref={conf/acl/2025-1}\n}\n"},"abstract":{"value":"There is a lack of an evaluation methodology that estimates the extent to which large language models (LLMs) use code-switching (CS) in the same way as bilinguals. Existing methods do not have wide language coverage, fail to account for the diverse range of CS phenomena, or do not scale. We propose an intervention based on minimal pairs of CS. Each minimal pair contains one naturally occurring CS sentence and one minimally manipulated variant. We collect up to 1,000 such pairs each for 11 language pairs. Our human experiments show that, for every language pair, bilinguals consistently prefer the naturally occurring CS sentence. Meanwhile our experiments with current LLMs show that the larger the model, the more consistently it assigns higher probability to the naturally occurring CS sentence than to the variant. In accordance with theoretical claims, the largest probability differences arise in those pairs where the manipulated material consisted of closed-class words."},"title":{"value":"Minimal Pair-Based Evaluation of Code-Switching"},"authors":{"value":["Igor Sterner","Simone Teufel"]}},"tmdate":1759790619887,"pdate":1735689600000,"externalIds":["dblp:conf/acl/SternerT25"],"tcdate":1759790618358,"writers":["~"],"signatures":["~Igor_Sterner1"],"forum":"TGqptHhS3d","license":"CC BY-SA 4.0","number":637832,"cdate":1735689600000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1759790619887,"domain":"DBLP.org","id":"TGqptHhS3d","version":2},{"content":{"summary":{"value":"This paper introduces OmniFace, a comprehensive framework designed to bridge the gap between image and video face swapping using a Diffusion Transformer (DiT) architecture. The authors propose a novel SyncID-Pipe data pipeline for constructing explicitly supervised paired data (bidirectional quadruplets) by integrating strengths from state-of-the-art image face swapping models and an Identity-Anchored Video Synthesizer (IVS). OmniFace employs a modality-aware conditioning mechanism to disentangle and inject multimodal information, a synthetic-to-real curriculum learning scheme, and an Identity-Coherence Reinforcement Learning (IRL) approach to improve robustness and consistency. The work further contributes a new benchmark, IDBench-V, to better evaluate video face swapping systems. Experimental results on IDBench-V and ablation studies demonstrate state-of-the-art performance, high versatility, and adaptability to extended human-centric swapping tasks."},"soundness":{"value":4},"confidence":{"value":4},"questions":{"value":"1.Could the authors provide a more detailed breakdown of common failure cases for OmniFace, including qualitative examples and quantitative error rates for scenarios such as extreme lighting changes, multi-subject videos, or minority demographics?\n2.Is overfitting to the synthetic domain a concern?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1.The SyncID-Pipe pipeline ingeniously leverages strong image face swapping performance and adapts it for video by crafting bidirectional ID quadruplets, ensuring explicit and effective supervision for video models. This pipeline is well illustrated in Figure 2 and is a core element in improving alignment between image and video domains.\n2.The Diffusion Transformer–based formulation is appropriately tailored for video, with a carefully crafted Modality-Aware Conditioning (MC) module for disentangling spatiotemporal, structural, and identity information (explained and visualized in Figure 3).\n3.The introduction of IDBench-V (as shown in Figure 8) fills an acknowledged gap by systematically evaluating face swapping methods in diverse real-world scenarios.\n4.All core equations (e.g., adaptive pose attention, Q-value definition, IRL loss, guidance purification) are specified with notation, and Appendix A.5 provides a proper theoretical justification for the IRL concept, relating it to reward-weighted likelihood maximization."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1.Some aspects of the training process are not entirely transparent.\n2.The baseline selection (as shown in Table 1) is mostly well done, but some important recent works in video face swapping, especially those involving subject-agnostic or reenactment models (e.g., direct comparison with FSGAN, which is not referenced), are not included. Even though these may not be directly SOTA, including them would clarify how much improvement is due to architectural advances versus data curation. More detailed explanations for the omission of certain non-diffusion approaches would help.\n3.The limitations section (Appendix A.8) briefly notes issues with speed and lighting preservation but lacks concrete error analysis on where (or why) the model produces visible artifacts, identity drift, or temporal inconsistency."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762916565286,"tcdate":1762177800725,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission3136/Reviewer_mwKb"],"signatures":["ICLR.cc/2026/Conference/Submission3136/Reviewer_mwKb"],"forum":"iQs4Ro6E6y","number":4,"license":"CC BY 4.0","cdate":1762177800725,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission3136/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762916565286,"domain":"ICLR.cc/2026/Conference","replyto":"iQs4Ro6E6y","id":"mfuKAO53DT","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Video Face Swapping","Diffusion Transformer"]},"supplementary_material":{"value":"/attachment/b208865193f2f9e3e9a3d1e78daaca095173dcec.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Video Face Swapping (VFS) requires seamlessly injecting a source identity into a target video while meticulously preserving the original pose, expression, lighting, background, and dynamic information. Existing methods struggle to maintain identity similarity and attribute preservation while preserving temporal consistency. To address the challenge, we propose a comprehensive framework to seamlessly transfer the superiority of Image Face Swapping (IFS) to the video domain. We first introduce a novel data pipeline SyncID-Pipe that pre-trains an Identity-Anchored Video Synthesizer and combines it with IFS models to construct bidirectional ID quadruplets for explicit supervision. Building upon paired data, we propose the first Diffusion Transformer-based framework OmniFace, employing a core Modality-Aware Conditioning module to discriminatively inject multi-model conditions. Meanwhile, we propose a Synthetic-to-Real Curriculum mechanism and an Identity-Coherence Reinforcement Learning strategy to enhance visual realism and identity consistency under challenging scenarios. To address the issue of limited benchmarks, we introduce IDBench-V, a comprehensive benchmark encompassing diverse scenes. Extensive experiments demonstrate OmniFace outperforms state-of-the-art methods and further exhibits exceptional versatility, which can be seamlessly adapted to various swap-related tasks."},"_bibtex":{"value":"@misc{\nguo2026omniface,\ntitle={OmniFace: Bridging the Image-to-Video Gap for High-Fidelity Face Swapping via Diffusion Transformer},\nauthor={Xu Guo and Fulong Ye and Xinghui Li and Pengqi Tu and Pengze Zhang and Songtao Zhao and Xiangwang Hou and Qian HE},\nyear={2026},\nurl={https://openreview.net/forum?id=iQs4Ro6E6y}\n}"},"title":{"value":"OmniFace: Bridging the Image-to-Video Gap for High-Fidelity Face Swapping via Diffusion Transformer"},"pdf":{"value":"/pdf/3dd8e448110c16b0066ce59347eaa89dd4aab362.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"guo|omniface_bridging_the_imagetovideo_gap_for_highfidelity_face_swapping_via_diffusion_transformer"},"authorids":{"value":["~Xu_Guo7","~Fulong_Ye1","~Xinghui_Li2","~Pengqi_Tu1","~Pengze_Zhang1","~Songtao_Zhao1","~Xiangwang_Hou1","~Qian_HE3"]},"authors":{"value":["Xu Guo","Fulong Ye","Xinghui Li","Pengqi Tu","Pengze Zhang","Songtao Zhao","Xiangwang Hou","Qian HE"]}},"version":2},{"content":{"summary":{"value":"In this work, the authors propose a principled pipeline for text to video generation. Specifically, three techniques are proposed: (1) a LLM based per-frame prompt generation, so that the motion/dynamics of each frame can be better specified. (2) a noise joint sampling schedule, and a step-aware attention shift is proposed to enhance temporal consistency. (3) an interpolation module to generate longer videos. Extensive experiments and results have demonstrated the effectiveness of the proposed work."},"soundness":{"value":"4 excellent"},"confidence":{"value":"4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work."},"questions":{"value":"elaborated in the above sections."},"rating":{"value":"7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations."},"code_of_conduct":{"value":"Yes"},"presentation":{"value":"4 excellent"},"contribution":{"value":"3 good"},"strengths":{"value":"1. The pipeline is technically sound and generally makes sense to me. It is straightforward to use the strong prior from pretrained LLM to give more temporal/motion/appearance information to each frame, to facilitate video generation. The design of noise scheme, sampling, and attention shift is technically sound.\n\n2. The exposition of this paper is very good and clear to me."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. In my honest opinion, directing comparing this work w/ many baselines is kind of unfair, since this involves strong prior in LLM, and this is the key factor of improvement in video generation. I strongly recommend the authors to compare w/ some similar text2video pipelines that also factorize the unified single prompt condition into per-frame motion-aware conditions. For example, first generate a sequence of optical flow maps, then takes them as conditions to generate the video.\n\n2. Since I'm not the expert in NLP domain, I'm concerned about the ability of LLM to generate very reasonable or temporally-coherent per-frame prompts. I'm also curious about the feature trajectory of such text features. Say, you generate the per-frame prompt of p1, p2, p3, ..., pn, and use the text encoder to get text features f1, f2, ..., fn for cross-attention. Do they form a smooth trajectory in the high-dimensional text feature space? Do the authors think that, a smooth feature transition among frames can guarantee a smooth video? If so, a naive idea for improvement of this work is to add some regularizations to constrain the smoothness of text features across frames."},"limitations":{"value":"N/A"}},"nonreaders":[],"tmdate":1702410780597,"tcdate":1688624655489,"writers":["NeurIPS.cc/2023/Conference","NeurIPS.cc/2023/Conference/Submission1312/Reviewer_mgrm"],"signatures":["NeurIPS.cc/2023/Conference/Submission1312/Reviewer_mgrm"],"forum":"paa2OU5jN8","number":3,"license":"CC BY 4.0","cdate":1688624655489,"mdate":1702410780597,"readers":["everyone"],"invitations":["NeurIPS.cc/2023/Conference/Submission1312/-/Official_Review","NeurIPS.cc/2023/Conference/-/Edit"],"domain":"NeurIPS.cc/2023/Conference","replyto":"paa2OU5jN8","id":"D6z0MjnDeJ","forumContent":{"venue":{"value":"NeurIPS 2023 poster"},"keywords":{"value":["Text-to-Video","Zero-Shot Generation","Large Language Model","Latent Diffusion Models"]},"supplementary_material":{"value":"/attachment/f2d0c5a77c34a820786ffe50cd192e361bcd6be7.zip"},"_bibtex":{"value":"@inproceedings{\nhuang2023freebloom,\ntitle={Free-Bloom: Zero-Shot Text-to-Video Generator with {LLM} Director and {LDM} Animator},\nauthor={Hanzhuo Huang and Yufan Feng and Cheng Shi and Lan Xu and Jingyi Yu and Sibei Yang},\nbooktitle={Thirty-seventh Conference on Neural Information Processing Systems},\nyear={2023},\nurl={https://openreview.net/forum?id=paa2OU5jN8}\n}"},"title":{"value":"Free-Bloom: Zero-Shot Text-to-Video Generator with LLM Director and LDM Animator"},"paperhash":{"value":"huang|freebloom_zeroshot_texttovideo_generator_with_llm_director_and_ldm_animator"},"TLDR":{"value":"This work introduces Free-Bloom, a zero-shot, training-free, semantic-coherent text-to-video generator that is capable of generating high-quality, temporally consistent, and semantically aligned videos."},"abstract":{"value":"Text-to-video is a rapidly growing research area that aims to generate a semantic, identical, and temporal coherence sequence of frames that accurately align with the input text prompt. This study focuses on zero-shot text-to-video generation considering the data- and cost-efficient. To generate a semantic-coherent video, exhibiting a rich portrayal of temporal semantics such as the whole process of flower blooming rather than a set of ``moving images'', we propose a novel Free-Bloom pipeline that harnesses large language models (LLMs) as the director to generate a semantic-coherence prompt sequence, while pre-trained latent diffusion models (LDMs) as the animator to generate the high fidelity frames. Furthermore, to ensure temporal and identical coherence while maintaining semantic coherence, we propose a series of annotative modifications to adapting LDMs in the reverse process, including joint noise sampling, step-aware attention shift, and dual-path interpolation. Without any video data and training requirements, Free-Bloom generates vivid and high-quality videos, awe-inspiring in generating complex scenes with semantic meaningful frame sequences.  In addition, Free-Bloom is naturally compatible with LDMs-based extensions."},"pdf":{"value":"/pdf/303a32c6185f4e35880879ec412f25eccd89a86c.pdf"},"venueid":{"value":"NeurIPS.cc/2023/Conference"},"authorids":{"value":["~Hanzhuo_Huang1","~Yufan_Feng1","~Cheng_Shi4","~Lan_Xu2","~Jingyi_Yu5","~Sibei_Yang1"]},"authors":{"value":["Hanzhuo Huang","Yufan Feng","Cheng Shi","Lan Xu","Jingyi Yu","Sibei Yang"]}},"version":2},{"content":{"summary":{"value":"This paper proposed a framework for detecting spurious correlations or shortcuts implied in training datasets. The main hypothesis of this paper is that the information between input and embedding would be low provided that there are shortcuts in a dataset. The author leveraged a neural tangent kernel to estimate mutual information between input and embedding representation and empirically represented that their hypothesis is valid on the synthetic (MNIST with shortcuts), benchmarks (Waterbirds, CelebA, and NICO), and real-world medical datasets."},"presentation":{"value":"3 good"},"contribution":{"value":"2 fair"},"soundness":{"value":"2 fair"},"strengths":{"value":"- The empirical results support their hypothesis that the $I(X;Z) of a dataset with a shortcut is lower than that without a shortcut.\n- The method is simple and easy to follow."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The proposed method is limited in the real-world application scenario. The proposed method requires a 'without shortcuts dataset' to detect whether there are shortcuts in the training dataset. I am not sure how many cases we can prepare the 'without shortcuts dataset' before we know whether the training dataset has a shortcut.\n- (Kirichenko et al., 2022) represented that the model trained on Waterbirds using ERM has the ability to classify the foreground-only and background-only datasets. It conflicts with the main hypothesis that the model trained with the shortcut dataset will have a low $I(X;Z).\n- The experiment graphs show that the mutual information is less discriminative than the losses, which diminishes the necessity of using mutual information to detect the existence of shortcuts in the training dataset.\n\nMinor corrections\n- Page 5 MNIST with synthetic shortcut) Figure 1 -> Figure 2\n- Figure 2) Please denote the used Saliency map.\n- Definition 2) Please state that the higher $\\Gamma$, the better the generalization."},"confidence":{"value":"4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work."},"questions":{"value":"- The proposed algorithm and the experiment setting are different. If I denote an original training and test dataset as $D_{tr}$ and $D_{te}$, respectively, and the shortcut added training dataset as $D_{tr}^{sc}$, then the Figure 3(c) plots '$I(X_{test};Z)$ trained on $D_{tr}$' and '$I(X_{test};Z)$ trained on $D_{tr}^{sc}$'. However, the algorithm seems to be denoted to compare '$I(X_{test};Z)$ trained on $D_{tr}$' and '$I(X_{tr};Z)$ trained on $D_{tr}$'. \n- Algorithm 1 step 1) Why $\\mathcal{F}$ is initially required?\n- Figure 3 with shortcut line vs Figure 4 100% line) I think they are the same experiment, but the graphs differ."},"rating":{"value":"3: reject, not good enough"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636989325,"tcdate":1699433517795,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission8022/Reviewer_cP5y"],"signatures":["ICLR.cc/2024/Conference/Submission8022/Reviewer_cP5y"],"forum":"hr4HTShC6l","number":5,"license":"CC BY 4.0","cdate":1699433517795,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission8022/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636989325,"domain":"ICLR.cc/2024/Conference","replyto":"hr4HTShC6l","id":"8Tu0Ia8E50","forumContent":{"TLDR":{"value":"Proposed a mutual-information based method to detect shortcuts/spurious correlations."},"venue":{"value":"Submitted to ICLR 2024"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["shortcuts","spurious correlation","mutual information","information theory","neural tangent kernel"]},"supplementary_material":{"value":"/attachment/51f14e81aa32b370b97ad25ffc7becbdabd22c25.pdf"},"primary_area":{"value":"societal considerations including fairness, safety, privacy"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"The failure of deep neural networks to generalize to out-of-distribution (OOD) data is a well-known problem that raises concerns about the deployment of trained networks in safety-critical domains such as healthcare and autonomous vehicles. We study a particular kind of distribution shift — shortcuts or spurious correlations in the training data. These correlations are not present in real-world test data, so there\nis a performance drop due to distribution shift, also referred to as shortcut learning. Shortcut learning is often only exposed when models are evaluated in carefully controlled experimental settings, posing a serious dilemma for AI practitioners to properly assess the effectiveness of a trained model for real-world applications. In this work, we try to understand shortcut learning using information-theoretic tools and propose to use the mutual information (MI) between the learned representation and the input space as a domain-agnostic metric for detecting shortcuts in the training datasets. For studying the training dynamics of shortcut learning, we develop a Neural Tangent Kernel (NTK) based framework, which can be used to detect shortcuts and spurious correlations in the training data without requiring class labels\nof the test data. We empirically demonstrate on multiple datasets, such as MNIST, CelebA, NICO, Waterbirds, and BenchMD, that MI can effectively detect shortcuts. We benchmark against multiple OOD detection baselines to show that OOD detectors cannot detect shortcuts, and our method can be used in complementary with OOD detectors to identify all types of distribution shifts in the datasets, including\nshortcuts."},"_bibtex":{"value":"@misc{\nadnan2024detecting,\ntitle={Detecting Shortcuts using Mutual Information},\nauthor={Mohammed Adnan and Yani Ioannou and Kenyon Tsai and Angus Galloway and Hamid Tizhoosh and Rahul G Krishnan and Graham W. Taylor},\nyear={2024},\nurl={https://openreview.net/forum?id=hr4HTShC6l}\n}"},"title":{"value":"Detecting Shortcuts using Mutual Information"},"pdf":{"value":"/pdf/ca51045eef9d7fcc21c7c83eaf9e2cd5f0cd86f2.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Rejected_Submission"},"paperhash":{"value":"adnan|detecting_shortcuts_using_mutual_information"},"authorids":{"value":["~Mohammed_Adnan1","~Yani_Ioannou1","~Kenyon_Tsai1","~Angus_Galloway1","~Hamid_Tizhoosh1","~Rahul_G_Krishnan1","~Graham_W._Taylor1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Mohammed Adnan","Yani Ioannou","Kenyon Tsai","Angus Galloway","Hamid Tizhoosh","Rahul G Krishnan","Graham W. Taylor"]}},"version":2},{"content":{"summary":{"value":"This paper provides a comprehensive exploration of noise-robust contrastive learning and introduces an innovative approach for addressing noisy data pairs. The authors propose the incorporation of learnable stochastic weights, informed by a Bayesian inference framework, to dynamically adjust the significance of each data pair. To facilitate the learning of these weights, the paper reformulates the problem within a probabilistic framework and devises a stochastic expectation maximization algorithm.The paper evaluates the method on several multi-modal contrastive learning benchmarks."},"presentation":{"value":"3 good"},"contribution":{"value":"3 good"},"soundness":{"value":"3 good"},"strengths":{"value":"This paper introduce a new method to combat with noisy data pairs through learnable stochastic weight from the Bayesian inference framework.\n\nExtensive experiments have been conducted to verify the effectivenss of the proposed method."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. One primary concern revolves around the claimed novelty of the research problem. While the authors assert that they are the first to address the issue of contrastive learning with noisy data pairs, existing studies have already explored similar territories. For instance, RINCE (CVPR 2022) focuses on mitigating the impact of noisy views (false positives) and presents a theoretically-grounded robust infoNCE loss function. More recent studies in noisy correspondence learning tackle issues like mismatched pairs and false negatives in diverse tasks, including partially view-aligned clustering, video-text retrieval, visible-infrared person re-identification, and image-text matching. Furthermore, [1] formally investigates noisy many-to-many correspondences and introduces an innovative training framework that outperforms its CLIP counterpart in multiple settings.I believe it is crucial to provide a more in-depth clarification of the distinctions between this paper and previous works. Additionally, it would be valuable to expand the related works section to offer a deeper understanding of these differences.\n   [1] Robust Cross-Modal Representation Learning with Progressive Self-Distillation, CVPR 2022\n\n2. The paper lacks sufficient detail and background information. Notably, explanations and prerequisites related to concepts such as the gamma distribution/prior and Bayesian inference are notably absent. Furthermore, some equations are inadequately elucidated and appear to be missing essential derivations.\n3. The paper's comparative analysis is limited to CLIP and DeCL, neglecting other crucial baselines, like [1], which specifically address the same issue of noisy data pairs.\n4. It is recommended to supplement the paper with additional experimental analyses that illuminate the inner workings of the proposed method. For instance, incorporating visualizations of the learned weights for positive and negative pairs, akin to Fig. 5 in [1], would enhance the reader's understanding of the method's underlying mechanisms."},"confidence":{"value":"5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully."},"questions":{"value":"Please see the weaknesses"},"rating":{"value":"5: marginally below the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636227097,"tcdate":1698751791742,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission2836/Reviewer_6kBB"],"signatures":["ICLR.cc/2024/Conference/Submission2836/Reviewer_6kBB"],"forum":"CvxcWCDX0h","number":3,"license":"CC BY 4.0","cdate":1698751791742,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission2836/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636227097,"domain":"ICLR.cc/2024/Conference","replyto":"CvxcWCDX0h","id":"NDuUgIipas","forumContent":{"venue":{"value":"ICLR 2024 Conference Withdrawn Submission"},"TLDR":{"value":"we propose a novel solution by reformulating the standard CL into a probability framework, and introducing learnable random weights to associate with data pairs, so as to allow automatic inference of the degree of noisiness for each data pair."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["multi-modal learning","contrastive learning","foundation models"]},"primary_area":{"value":"representation learning for computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Contrastive learning~(CL) represents one of the most successful paradigms for self-supervised representation learning, which has been applied to SOTA multi-modal learning applications. One overlooked limitation of standard contrastive learning, however, is that it is not designed for robust learning in the presence of noisy data pairs. For example, not all negative samples are truly negative, {\\it e.g.}, within a mini-batch there can be negative samples that are semantically as positive as the positive sample. This is common in most web-sourced multi-modal datasets such as CC3M and YFCC that are frequently used for CL, due to the noisy nature when crawling the datasets. Consequently, the noise in the datasets could significantly impair the power of CL. To remedy this issue, we propose a novel solution by reformulating the standard CL into a probability framework, and introducing learnable random weights to associate with data pairs, so as to allow automatic inference of the degree of noisiness for each data pair. Within our probability framework, posterior inference of the random weights can be done efficiently with Bayesian data augmentation. Consequently, the model can be effectively optimized by a novel learning algorithm based on stochastic expectation maximization. We demonstrate the effectiveness of our approach on several standard multi-modal contrastive learning benchmarks, which significantly outperforms standard contrastive learning."},"_bibtex":{"value":"@misc{\njiang2024learning,\ntitle={Learning Multi-Modal Representation Alignments from Noisy Data-Pairs},\nauthor={Qian Jiang and Jingjing Meng and Alireza Bagheri Garakani and Yang Jiao and Yetian Chen and Yikai Ni and Yan Gao and Yi Sun and Changyou Chen},\nyear={2024},\nurl={https://openreview.net/forum?id=CvxcWCDX0h}\n}"},"title":{"value":"Learning Multi-Modal Representation Alignments from Noisy Data-Pairs"},"pdf":{"value":"/pdf/5b5dc724edb76920a4b1bd79e791974ea61f81aa.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Withdrawn_Submission"},"paperhash":{"value":"jiang|learning_multimodal_representation_alignments_from_noisy_datapairs"},"authorids":{"value":["~Qian_Jiang1","~Jingjing_Meng1","~Alireza_Bagheri_Garakani1","~Yang_Jiao6","~Yetian_Chen1","yika@amazon.com","~Yan_Gao8","~Yi_Sun13","~Changyou_Chen1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Qian Jiang","Jingjing Meng","Alireza Bagheri Garakani","Yang Jiao","Yetian Chen","Yikai Ni","Yan Gao","Yi Sun","Changyou Chen"]}},"version":2},{"content":{"summary":{"value":"This paper presents LightCache, a training-free framework for accelerating diffusion-based video generation with improved memory efficiency. By introducing Asynchronous Cache Swapping, Feature Chunking, and VAE Slicing, the method mitigates memory surges in denoising and decoding stages, achieving faster inference and lower GPU usage than DeepCache with minimal quality loss."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1. It is unclear what specific model the TextToVideo baseline refers to in the paper, as there is no corresponding citation or reference provided."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"The paper provides a systematic analysis of memory bottlenecks in diffusion-based video generation and proposes simple yet practical strategies to alleviate them without retraining, offering a useful engineering reference for future work on efficient inference."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The writing quality needs improvement, as several parts are confusing. For example, the latter part of the Introduction does not clearly explain how the proposed slicing strategy works, but only describes why it is needed.\n\n2. The discussion and citation of related work are insufficient. \n\n    a) In particular, for the choice of base models, many more recent and memory-intensive DiT-based architectures, such as HunyuanVideo and WAN 2.2, have emerged. It remains unclear whether the proposed method, which is primarily validated on U-Net–based models, can generalize effectively to these DiT architectures or to models with significantly larger parameter scales.\n\n    b) Secondly, cache-based acceleration methods have been extensively explored in video generation. It is recommended that the authors adopt a more advanced method specifically designed for video diffusion, such as TeaCache, for experimental comparison. Moreover, the paper lacks adequate discussion of recent cache-based approaches in this domain.\n\n3. Many of the techniques mentioned in the paper, such as asynchronous CPU offloading, have already been explored in prior engineering practices, and therefore the proposed method lacks sufficient novelty.\n\n4.Typo: Line 201 — “cacheing” should be corrected to “caching.”"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762917996349,"tcdate":1761450361615,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission5302/Reviewer_mtf3"],"signatures":["ICLR.cc/2026/Conference/Submission5302/Reviewer_mtf3"],"forum":"NomB0oiwqI","number":1,"license":"CC BY 4.0","cdate":1761450361615,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission5302/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762917996349,"domain":"ICLR.cc/2026/Conference","replyto":"NomB0oiwqI","id":"29zvCJFQXA","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Cache","Training-free"]},"primary_area":{"value":"optimization"},"abstract":{"value":"Training-free acceleration has emerged as an advanced research area in video generation. The redundancy of latents in diffusion model inference provides a natural entry point for acceleration. We decompose the inference process into the encoding, denoising, and decoding stages, and observe that cache-based acceleration methods often lead to substantial memory surges in the latter two stages. To address this problem, we analyze the characteristics of inference across different stages and propose stage-specific strategies for reducing memory consumption: 1) Asynchronous Cache Swapping. 2) Feature chunk. 3) Slicing latents to decode. At the same time, we tune the hyper-parameters to ensure that the time overhead introduced by these three strategies remains lower than the acceleration gains themselves. Compared with the baseline, our approach achieves faster inference speed and lower memory usage, while maintaining quality degradation within an acceptable range. The Code is available at https://anonymous.4open.science/r/LightCache-CD86."},"_bibtex":{"value":"@misc{\nxiao2025lightcache,\ntitle={LightCache: Memory-Efficient, Training-Free Acceleration for Video Generation},\nauthor={Yang Xiao and Gen Li and Kaiyuan Deng and Yushu Wu and Zheng Zhan and Yanzhi Wang and Xiaolong Ma and Bo Hui},\nyear={2025},\nurl={https://openreview.net/forum?id=NomB0oiwqI}\n}"},"title":{"value":"LightCache: Memory-Efficient, Training-Free Acceleration for Video Generation"},"pdf":{"value":"/pdf/1da358dc6650d560096acf4798f395c406e8408a.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"xiao|lightcache_memoryefficient_trainingfree_acceleration_for_video_generation"},"authorids":{"value":["~Yang_Xiao18","~Gen_Li4","~Kaiyuan_Deng2","~Yushu_Wu1","~Zheng_Zhan3","~Yanzhi_Wang3","~Xiaolong_Ma2","~Bo_Hui1"]},"authors":{"value":["Yang Xiao","Gen Li","Kaiyuan Deng","Yushu Wu","Zheng Zhan","Yanzhi Wang","Xiaolong Ma","Bo Hui"]}},"version":2},{"content":{"venue":{"value":"CoRR 2025"},"pdf":{"value":"https://arxiv.org/pdf/2502.11594v2"},"venueid":{"value":"dblp.org/journals/CORR/2025"},"paperhash":{"value":"li|imove_instancemotionaware_video_understanding"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Jiaze_Li:","https://dblp.org/search/pid/api?q=author:Yaya_Shi:","https://dblp.org/search/pid/api?q=author:Zongyang_Ma:","~Haoran_Xu8","https://dblp.org/search/pid/api?q=author:Feng_Cheng:","https://dblp.org/search/pid/api?q=author:Huihui_Xiao:","https://dblp.org/search/pid/api?q=author:Ruiwen_Kang:","https://dblp.org/search/pid/api?q=author:Fan_Yang_0094:","https://dblp.org/search/pid/api?q=author:Tingting_Gao:","https://dblp.org/search/pid/api?q=author:Di_Zhang_0026:"]},"html":{"value":"https://doi.org/10.48550/arXiv.2502.11594"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2502-11594,\n  publtype={informal},\n  author={Jiaze Li and Yaya Shi and Zongyang Ma and Haoran Xu and Feng Cheng and Huihui Xiao and Ruiwen Kang and Fan Yang and Tingting Gao and Di Zhang},\n  title={iMOVE: Instance-Motion-Aware Video Understanding},\n  year={2025},\n  month={February},\n  cdate={1738368000000},\n  journal={CoRR},\n  volume={abs/2502.11594},\n  url={https://doi.org/10.48550/arXiv.2502.11594}\n}\n"},"abstract":{"value":"Enhancing the fine-grained instance spatiotemporal motion perception capabilities of Video Large Language Models is crucial for improving their temporal and general video understanding. However, current models struggle to perceive detailed and complex instance motions. To address these challenges, we have made improvements from both data and model perspectives. In terms of data, we have meticulously curated iMOVE-IT, the first large-scale instance-motion-aware video instruction-tuning dataset. This dataset is enriched with comprehensive instance motion annotations and spatiotemporal mutual-supervision tasks, providing extensive training for the model's instance-motion-awareness. Building on this foundation, we introduce iMOVE, an instance-motion-aware video foundation model that utilizes Event-aware Spatiotemporal Efficient Modeling to retain informative instance spatiotemporal motion details while maintaining computational efficiency. It also incorporates Relative Spatiotemporal Position Tokens to ensure awareness of instance spatiotemporal positions. Evaluations indicate that iMOVE excels not only in video temporal understanding and general video understanding but also demonstrates significant advantages in long-term video understanding."},"title":{"value":"iMOVE: Instance-Motion-Aware Video Understanding"},"authors":{"value":["Jiaze Li","Yaya Shi","Zongyang Ma","Haoran Xu","Feng Cheng","Huihui Xiao","Ruiwen Kang","Fan Yang","Tingting Gao","Di Zhang"]}},"tmdate":1769608360320,"pdate":1767139200000,"externalIds":["dblp:journals/corr/abs-2502-11594"],"tcdate":1769608354035,"writers":["~"],"signatures":["~Haoran_Xu8"],"forum":"Wg56ACzvQb","license":"CC BY-SA 4.0","number":812191,"cdate":1738368000000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1769608360320,"domain":"DBLP.org","id":"Wg56ACzvQb","version":2},{"content":{"summary":{"value":"This paper introduces the Video Understanding Domain Generalization (VUDG) dataset, a dataset designed to evaluate domain generalization of video understanding models. The authors key motivation is that existing video understanding benchmarks either do not measure domain generalization, or are ineffective at measuring it due to large semantic gap between domains. VUDG aims to address this by collecting videos from 11 distinct domains while keeping the underlying semantic content consistent through a shared set of daily human activities. the authors employ a progressive multi-expert annotation framework that combines multiple large vision-language models for automated question generation, verification, and filtering, followed by a human verification. This process yields a training dataset of 31k QA pairs, and a testing dataset of 4k QA pairs. The authors additionally benchmark existing video understanding models on VUDG."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"What evidence exists to motivate the community to evaluate the domain generalization ability of their models with VUDG, instead of a benchmark like Video-MME that also contains multiple video domains? Can this semantic shift be quantified and shown to hurt a benchmarks ability to assess domain generalization? Or can it be qualitatively shown?\n\nHow can the authors be certain that the MCQs in VUDG do not suffer from visual or language bias? Are there any failsafes in the data curation to mitigate this?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"The proposed benchmark is the first large-scale benchmark designed to measure domain generalization in VLMs, consisting of a dedicated training/testing set and clear protocols for evaluation\n\nThe choice of domains and corresponding videos is high quality: The domains selected in VUDG are broad and diverse and a large portion of the videos used for testing are newly collected by the authors, reducing the risk of data leakage with existing VLM training data"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The reviewer understands the primary motivation behind VUDG is to unify the semantic space across domains, and this intuitively makes sense to the reviewer. However, it is not scientifically shown why existing benchmarks fail to assess domain generalization due to the lack of cross-domain semantic alignment\n* What evidence exists to motivate the community to evaluate the domain generalization ability of their models with VUDG, instead of a benchmark like Video-MME that also contains multiple video domains?\n* The t-SNE in Figure 6 does show the semantic alignment between domains in VUDG, but it is not compared with other benchmarks\n\nIt is unclear if the testing dataset exhibits language or visual bias in the MCQ questions, which is a common pitfall of LLM generated MCQs (Zohar et al., Apollo: An Exploration of Video Understanding in Large Multimodal Models, CVPR 2025). The authors mention there is a human verification step, but do not explicitly state that it addresses this bias.\n* Additionally there are no qualitative examples of MCQs in the paper. As this is a dataset paper the reviewer expected to see many more examples of QA pairs in the dataset\n\nThe choice of VLMs used for evaluation in the full fine-tuning and single-domain generalization settings is relatively shallow, reducing the insight gained from these settings"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762925836974,"tcdate":1762053690216,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission15564/Reviewer_M1bM"],"signatures":["ICLR.cc/2026/Conference/Submission15564/Reviewer_M1bM"],"forum":"0mUiXz1TNq","number":3,"license":"CC BY 4.0","cdate":1762053690216,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission15564/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762925836974,"domain":"ICLR.cc/2026/Conference","replyto":"0mUiXz1TNq","id":"PjvBU07Pmz","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Video Understanding","Dataset","Domain Generalization"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Video understanding has made remarkable progress in recent years, largely driven by advances in deep models and the availability of large-scale annotated datasets.\nHowever, the robustness of these models to domain shifts encountered in real-world video applications remains a critical yet underexplored problem, limiting their practical reliability.\nTo address this problem, we introduce \\textbf{V}ideo \\textbf{U}nderstanding \\textbf{D}omain \\textbf{G}eneralization (\\textbf{VUDG}), the first dataset designed specifically for evaluating domain generalization in video understanding.\nVUDG contains videos from 11 distinct domains that cover three types of domain shifts, and maintains semantic consistency across different domains to ensure fair and meaningful evaluation. We propose a multi-expert progressive annotation framework to efficiently annotate videos with structured question-answer pairs designed for domain generalization.\nExtensive experiments on 9 representative Large Vision-Language Models (LVLMs) and several traditional video question answering methods show that most models (including state-of-the-art LVLMs) suffer performance degradation under domain shifts. \nThese results highlight the challenges posed by VUDG and the difference in the robustness of current models to data distribution shifts. We believe VUDG provides a critical resource to benefit future research in domain generalization for video understanding."},"_bibtex":{"value":"@inproceedings{\nwang2026vudg,\ntitle={{VUDG}: A Dataset for Video Understanding Domain Generalization},\nauthor={Ziyi Wang and Zhi Gao and Boxuan Yu and Zirui Dai and Peiyao Wang and Yuxiang Song and Qingyuan Lu and Jin Chen and Xinxiao Wu},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=0mUiXz1TNq}\n}"},"title":{"value":"VUDG: A Dataset for Video Understanding Domain Generalization"},"pdf":{"value":"/pdf/feec85898ca8b7da1de77de6c74b1c8ebabd4c20.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"wang|vudg_a_dataset_for_video_understanding_domain_generalization"},"authorids":{"value":["~Ziyi_Wang18","~Zhi_Gao5","~Boxuan_Yu1","~Zirui_Dai2","~Peiyao_Wang5","~Yuxiang_Song2","~Qingyuan_Lu2","~Jin_Chen2","~Xinxiao_Wu1"]},"authors":{"value":["Ziyi Wang","Zhi Gao","Boxuan Yu","Zirui Dai","Peiyao Wang","Yuxiang Song","Qingyuan Lu","Jin Chen","Xinxiao Wu"]}},"version":2},{"content":{"comment":{"value":"We thank the reviewer for the time taken to review the paper and provide feedback. We are pleased that they found our approach to be novel and our experiment design and analysis to be robust and well-explained. We appreciate the feedback the reviewer provides to improve and expand upon our analysis to improve the robustness of the paper. All changes to the text are highlighted in red.\n\n> **Figure 1 is only referenced after Figure 4 and should appear later in the paper**\n\nReferences to figures have been adjusted so that the order of the figures makes logical sense.\n\n> **I am wondering what is the difference between the PD distribution of test samples with shortcuts and without shortcuts for each class when classified by the shortcut model. [...]**\n\nWe thank the reviewer for highlighting this analysis, which is currently missing from our paper. We have now included in the appendix Figures 12 and 13, which shows the PD distribution for each of the shortcut models trained on ISIC and CheXpert. We distinguish between the samples that feature a shortcut and the samples that do not. This figure helps to visualise the peaks specific to the shortcut features. \n\n> **Is there any correlation between the prediction depth of a sample and if it is classified correctly or not?**\n\nWe find that PD positively correlates with performance metrics (AUC and accuracy). Our experiments find that models with deeper average test set PDs perform better than those with shallow test set PDs. We have added Figures 14 and 15 to the appendix that show the correlation matrices for two different models at different learning rates on the CheXpert dataset. We show that we tend to see a positive correlation between PD and AUC and a negative correlation between PD and divergence from a clean baseline."},"title":{"value":"Expanding on PD of shortcut vs no shortcut samples, and correlation between PD and performance"}},"tmdate":1712664383379,"tcdate":1710541065990,"writers":["MIDL.io/2024/Conference","MIDL.io/2024/Conference/Submission270/Authors"],"signatures":["MIDL.io/2024/Conference/Submission270/Authors"],"forum":"3UMzxqDcpY","number":9,"license":"CC BY 4.0","cdate":1710541065990,"readers":["everyone"],"invitations":["MIDL.io/2024/Conference/Submission270/-/Official_Comment","MIDL.io/2024/Conference/-/Edit"],"mdate":1712664383379,"domain":"MIDL.io/2024/Conference","replyto":"b4pFzZSxdX","id":"cS6B9MzLrL","forumContent":{"venue":{"value":"MIDL 2024 Poster"},"keywords":{"value":["shortcut learning","bias","prediction depth","model interpretation","clinical machine learning","spurious correlations","model robustness","generalization"]},"abstract":{"value":"Many studies have reported human-level accuracy (or better) for AI-powered algorithms performing a specific clinical task, such as detecting pathology. However, these results often fail to generalize to other scanners or populations. Several mechanisms have been identified that confound generalization. One such is shortcut learning, where a network erroneously learns to depend on a fragile spurious feature, such as a text label added to the image, rather than scrutinizing the genuinely useful regions of the image. In this way, systems can exhibit misleadingly high test-set results while the labels are present but fail badly elsewhere where the relationship between the label and the spurious feature breaks down. In this paper, we investigate whether it is possible to detect shortcut learning and locate where the shortcut is happening in a neural network. We propose a novel methodology utilizing the sample difficulty metric Prediction Depth (PD) and KL divergence to identify specific layers of a neural network model where the learned features of a shortcut manifest. We demonstrate that our approach can effectively isolate these layers across several shortcuts, model architectures, and datasets. Using this, we show a correlation between the visual complexity of a shortcut, the depth of its feature manifestation within the model, and the extent to which a model relies on it. Finally, we highlight the nuanced relationship between learning rate and shortcut learning."},"_bibtex":{"value":"@inproceedings{\nboland2024there,\ntitle={There Are No Shortcuts to Anywhere Worth Going: Identifying Shortcuts in Deep Learning Models for Medical Image Analysis},\nauthor={Christopher Boland and Keith A Goatman and Sotirios A. Tsaftaris and Sonia Dahdouh},\nbooktitle={Medical Imaging with Deep Learning},\nyear={2024},\nurl={https://openreview.net/forum?id=3UMzxqDcpY}\n}"},"title":{"value":"There Are No Shortcuts to Anywhere Worth Going: Identifying Shortcuts in Deep Learning Models for Medical Image Analysis"},"latex_code":{"value":"/attachment/6430a253dcfaddf41854115954ca18c9fc001a76.zip"},"pdf":{"value":"/pdf/3c6888b96f569bccf12925981f91ed4bc7d9a6d0.pdf"},"copyright_form":{"value":"/attachment/27380095359df7fa78ed39bf9f36960888ca219b.pdf"},"venueid":{"value":"MIDL.io/2024/Conference"},"paperhash":{"value":"boland|there_are_no_shortcuts_to_anywhere_worth_going_identifying_shortcuts_in_deep_learning_models_for_medical_image_analysis"},"authorids":{"value":["~Christopher_Boland1","~Keith_A_Goatman1","~Sotirios_A._Tsaftaris1","sonia.dahdouh@mre.medical.canon"]},"authors":{"value":["Christopher Boland","Keith A Goatman","Sotirios A. Tsaftaris","Sonia Dahdouh"]}},"version":2},{"content":{"summary":{"value":"This paper presents novel self-supervised learning methods that model the uncertainty in data pairs generated from natural generative processes, using regularized latent variables. For example the uncertainty between consecutive frames in a video. Two variants are presented, one based on variational inference on the other on enforcing sparsity on the latent variable. Experiments on artificial and reel data are conducted to demonstrate the effectiveness of the approach in identifying latent factors of variations and modelling uncertainty."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"- Line 160: Then the solution is just to project and do prediction in the same space ?\n\n- In AdaSSL-V, how could you make more explicit the mechanism that regularizes the latent variable regularized, basically what is L_reg ?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- The paper tackles a key problem in self-supervised learning: modelling conditional uncertainty, which arises in many other problems related to causal prediction. The potential applications are therefore numerous: video prediction, video generation, world modelling and latent action prediction, efficient self-supervised learning ect. The paper could actually do a better job at motivating these applications.\n\n- The ideas presented in the paper are interesting, described in depth, well-motivated, and seem to be good candidate solutions, at least on the toy problems explored in the experimental section."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The paper is complex to understand, with lots of formalism (the probabilist framework, Proposition 2.1), complex vocabulary (heteroscedasticity, modular editing, DCI) that is not necessarily introduced, or introduces lots of new vocabulary (CRL, DGP, SSL from structural invariance, Adaptive SSL, natural pairs), all for ideas that are actually fairly simple. I feel like this over–complixification hinders the reading flow and makes it harder to deliver the message it intends to deliver.\n\n- Also, the paper put a lot of emphasis on particular SSL losses such as contrastive vs non-contrastive, as well as several variants of InfoNCE, which does not seem very relevant for this study, and adds too many factors of variation that make the conclusions of the paper less clear and convincing. Section 2.2 and 2.3 are probably unnecessary, and the new variants H-InfoNCE, AnInfoNCE, ect are not well motivated.\n\n- Finally, all this formalism is derived by the JEPA framework and the authors mention that they only take inspiration from JEPA, whereas I see these contributions as instanciations of JEPAs, just with various ways of regularizing the latent variables. In Figure 1, b) and c) are JEPAs.\n\n- The experiments are conducted on toy problems which limits the credibility of the approach. How does it behave on more concrete problems ? I don’t think people care about the artificial numerical problems of section 4.2, these should be more of a tool for you to debug the approach. Section 4.2 is interesting but very artificial, and 4.3 is Moving MNIST which is good again for debugging but nowhere near close to the actual interesting problems.  Finally, the focus is on velocity decoding, which is good for debugging but not as interesting as the problem of prediction. It would be more interesting to show experiments where a model predicts the future trajectory in moving MNIST, and being able to sample several possible future trajectories by sampling from the latent.\n\n- Related to this, the claims made at the beginning of the paper need to be toned down, for example Line 20 “and we empirically show its superiority on identifiability, generalization, fine-grained image understanding, and world modeling on videos”. Superiority against which concurrent method ? And on benchmarks that are too toy.\n\n- The paper ignores the vast literature existing on uncertainty modelling and latent variables. All the work in generative modes, video generative models, video prediction, latent action models.\n\n- In conclusion, the paper is tackling an interesting problem and presents interesting ideas but it is hard to be convinced by the toy experiments. These points would make it much stronger:\n\n- Remove the studies on InfoNCE variants, along with sections 2.2 and 2.3, and focus on AdaSSL. Maybe rename using the JEPA terminology and just name the latent variable regularization methods.\n\n- Remove section 4.2 and focus more on real data experiments.\n\n- Add more motivations in terms of potential applications\n- Acknowledge other literature in uncertainty modelling and world modelling.\n\n- Focus the experiments more on video world modelling, and the prediction capability, rather than training probes to recover properties such as velocity."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762920093621,"tcdate":1761847095929,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission8115/Reviewer_gRec"],"signatures":["ICLR.cc/2026/Conference/Submission8115/Reviewer_gRec"],"forum":"r3JUDAYjIH","number":3,"license":"CC BY 4.0","cdate":1761847095929,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission8115/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762920093621,"domain":"ICLR.cc/2026/Conference","replyto":"r3JUDAYjIH","id":"VH8slijS1V","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Self-supervised learning","representation learning","disentanglement"]},"primary_area":{"value":"unsupervised, self-supervised, semi-supervised, and supervised representation learning"},"abstract":{"value":"Joint-embedding self-supervised learning (SSL), the key paradigm for unsupervised representation learning from visual data, learns from invariances between semantically-related data pairs.\nWe study the one-to-many mapping problem in SSL,\nwhere each datum may be mapped to multiple valid targets. \nThis arises when data pairs come from naturally occurring generative processes, e.g., successive video frames.\nWe show that existing methods struggle to flexibly capture this conditional uncertainty. As a remedy, we introduce \na latent variable to account for this uncertainty and derive a variational lower bound on the mutual information between paired embeddings.\nOur derivation yields a simple regularization term for standard SSL objectives.\nThe resulting method, which we call\nAdaSSL, applies to both contrastive and \ndistillation-based\nSSL objectives, and we empirically show its versatility in\ncausal representation learning,\nfine-grained image understanding, and world modeling on videos."},"_bibtex":{"value":"@inproceedings{\nzhang2026selfsupervised,\ntitle={Self-Supervised Learning from Structural Invariance},\nauthor={Yipeng Zhang and Hafez Ghaemi and Jungyoon Lee and Shahab Bakhtiari and Eilif B. Muller and Laurent Charlin},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=r3JUDAYjIH}\n}"},"title":{"value":"Self-Supervised Learning from Structural Invariance"},"pdf":{"value":"/pdf/525ee37007e6a0c0a8c8cc1ebcb1762a2a6eb67e.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"zhang|selfsupervised_learning_from_structural_invariance"},"authorids":{"value":["~Yipeng_Zhang5","~Hafez_Ghaemi1","~Jungyoon_Lee1","~Shahab_Bakhtiari1","~Eilif_B._Muller1","~Laurent_Charlin1"]},"authors":{"value":["Yipeng Zhang","Hafez Ghaemi","Jungyoon Lee","Shahab Bakhtiari","Eilif B. Muller","Laurent Charlin"]}},"version":2},{"content":{"venue":{"value":"CoRR 2023"},"pdf":{"value":"http://arxiv.org/pdf/2306.03571v2"},"venueid":{"value":"dblp.org/journals/CORR/2023"},"paperhash":{"value":"adriaens|minimizing_hitting_time_between_disparate_groups_with_shortcut_edges"},"authorids":{"value":["~Florian_Adriaens1","https://dblp.org/search/pid/api?q=author:Honglian_Wang:","~Aristides_Gionis1"]},"html":{"value":"https://doi.org/10.48550/arXiv.2306.03571"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2306-03571,\n  publtype={informal},\n  author={Florian Adriaens and Honglian Wang and Aristides Gionis},\n  title={Minimizing Hitting Time between Disparate Groups with Shortcut Edges},\n  year={2023},\n  cdate={1672531200000},\n  journal={CoRR},\n  volume={abs/2306.03571},\n  url={https://doi.org/10.48550/arXiv.2306.03571}\n}\n"},"abstract":{"value":"Structural bias or segregation of networks refers to situations where two or more disparate groups are present in the network, so that the groups are highly connected internally, but loosely connected to each other. In many cases it is of interest to increase the connectivity of disparate groups so as to, e.g., minimize social friction, or expose individuals to diverse viewpoints. A commonly-used mechanism for increasing the network connectivity is to add edge shortcuts between pairs of nodes. In many applications of interest, edge shortcuts typically translate to recommendations, e.g., what video to watch, or what news article to read next. The problem of reducing structural bias or segregation via edge shortcuts has recently been studied in the literature, and random walks have been an essential tool for modeling navigation and connectivity in the underlying networks. Existing methods, however, either do not offer approximation guarantees, or engineer the objective so that it satisfies certain desirable properties that simplify the optimization~task. In this paper we address the problem of adding a given number of shortcut edges in the network so as to directly minimize the average hitting time and the maximum hitting time between two disparate groups. Our algorithm for minimizing average hitting time is a greedy bicriteria that relies on supermodularity. In contrast, maximum hitting time is not supermodular. Despite, we develop an approximation algorithm for that objective as well, by leveraging connections with average hitting time and the asymmetric k-center problem."},"title":{"value":"Minimizing Hitting Time between Disparate Groups with Shortcut Edges"},"authors":{"value":["Florian Adriaens","Honglian Wang","Aristides Gionis"]}},"tmdate":1774426151929,"pdate":1672531200000,"tcdate":1738762573310,"writers":["~"],"signatures":["~Aristides_Gionis1"],"forum":"iv9RB1zeAI","license":"CC BY-SA 4.0","number":291767,"cdate":1672531200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1774426151929,"domain":"DBLP.org","id":"iv9RB1zeAI","version":2},{"content":{"venue":{"value":"ICML 2026 regular"},"keywords":{"value":["multimodal","vision","audio","diffusion"]},"_bibtex":{"value":"@inproceedings{\ndellali2026salsav,\ntitle={{SALSA}-V: Shortcut-Augmented Long-form Synchronized Audio from Videos},\nauthor={Amir Dellali and Luca A Lanzend{\\\"o}rfer and Florian Gr{\\\"o}tschla and Roger Wattenhofer},\nbooktitle={Forty-third International Conference on Machine Learning},\nyear={2026},\nurl={https://openreview.net/forum?id=akfapJ7Uuf}\n}"},"title":{"value":"SALSA-V: Shortcut-Augmented Long-form Synchronized Audio from Videos"},"paperhash":{"value":"dellali|salsav_shortcutaugmented_longform_synchronized_audio_from_videos"},"originally_submitted_PDF":{"value":"/pdf/f753b0d6e0744f91fcf380490af4eff956286e2c.pdf"},"primary_area":{"value":"deep_learning->generative_models_and_autoencoders"},"abstract":{"value":"We propose SALSA-V, a multimodal video-to-audio generation model capable of synthesizing highly synchronized, high-fidelity long-form audio from silent video content. Our approach introduces a masked diffusion objective, enabling audio-conditioned generation and the seamless synthesis of audio sequences of unconstrained length. Additionally, by integrating a shortcut loss into our training process, we achieve rapid generation of high-quality audio samples in as few as eight sampling steps, paving the way for near-real-time applications without requiring dedicated fine-tuning or retraining. We demonstrate that SALSA-V significantly outperforms existing state-of-the-art methods in both audiovisual alignment and synchronization with video content in quantiative evaluation and a human listening study. Furthermore, our use of random masking during training enables our model to match spectral characteristics of reference audio samples, broadening its applicability to professional audio synthesis tasks such as Foley generation and sound design."},"pdf":{"value":"/pdf/cbdf810841b46e5b36ee3a29a373b21d825704a1.pdf"},"lay_summary":{"value":"Adding sound to silent video is difficult because the audio must not only sound realistic, but also happen at exactly the right moment. We built SALSA-V, a deep learning model that can create synchronized audio for a given video, such as footsteps, impacts, voices, or environmental sounds, while keeping those sounds aligned with the visible events. Unlike many previous systems that work best on short clips, SALSA-V can extend audio over longer videos by using earlier generated sound as context. It can also use a short example sound to match a desired style or recording quality, giving users more control over the result. The model is designed to produce high-quality audio in only a few generation steps, making fast feedback possible. This could make sound design easier for films, games, user-generated videos, and accessibility tools."},"venueid":{"value":"ICML.cc/2026/Conference"},"authorids":{"value":["~Amir_Dellali1","~Luca_A_Lanzendörfer1","~Florian_Grötschla1","~Roger_Wattenhofer1"]},"authors":{"value":["Amir Dellali","Luca A Lanzendörfer","Florian Grötschla","Roger Wattenhofer"]}},"tmdate":1790067426553,"pdate":1777576831659,"tcdate":1769205271485,"writers":["ICML.cc/2026/Conference","ICML.cc/2026/Conference/Submission28585/Authors"],"signatures":["ICML.cc/2026/Conference/Submission28585/Authors"],"forum":"akfapJ7Uuf","license":"CC BY 4.0","number":28585,"cdate":1769205271485,"readers":["everyone"],"invitations":["ICML.cc/2026/Conference/-/Submission","ICML.cc/2026/Conference/-/Post_Submission","ICML.cc/2026/Conference/Submission28585/-/Full_Submission","ICML.cc/2026/Conference/-/Edit","ICML.cc/2026/Conference/Submission28585/-/Camera_Ready_Revision"],"mdate":1790067426553,"odate":1782342019633,"domain":"ICML.cc/2026/Conference","id":"akfapJ7Uuf","version":2},{"content":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["video-to-audio","audio generation","diffusion"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"We propose SALSA-V, a multimodal video-to-audio generation model capable of synthesizing highly synchronized, high-fidelity long-form audio from silent video content. Our approach introduces a masked diffusion objective, enabling audio-conditioned generation and the seamless synthesis of audio sequences of unconstrained length. Additionally, by integrating a shortcut loss into our training process, we achieve rapid generation of high-quality audio samples in as few as eight sampling steps, paving the way for near-real-time applications without requiring dedicated fine-tuning or retraining. We demonstrate that SALSA-V significantly outperforms existing state-of-the-art methods in both audiovisual alignment and synchronization with video content in quantiative evaluation and a human listening study. Furthermore, our use of random masking during training enables our model to match spectral characteristics of reference audio samples, broadening its applicability to professional audio synthesis tasks such as Foley generation and sound design."},"_bibtex":{"value":"@misc{\ndellali2026salsav,\ntitle={{SALSA}-V: Shortcut-Augmented Long-form Synchronized Audio from Videos},\nauthor={Amir Dellali and Luca A Lanzend{\\\"o}rfer and Florian Gr{\\\"o}tschla and Roger Wattenhofer},\nyear={2026},\nurl={https://openreview.net/forum?id=3FcjKYmNY5}\n}"},"title":{"value":"SALSA-V: Shortcut-Augmented Long-form Synchronized Audio from Videos"},"pdf":{"value":"/pdf/780501365dd6cdc6dc1c38131015c6ed7be54b52.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"dellali|salsav_shortcutaugmented_longform_synchronized_audio_from_videos"},"authorids":{"value":["~Amir_Dellali1","~Luca_A_Lanzendörfer1","~Florian_Grötschla1","~Roger_Wattenhofer1"]},"authors":{"value":["Amir Dellali","Luca A Lanzendörfer","Florian Grötschla","Roger Wattenhofer"]}},"tmdate":1770804965272,"tcdate":1758203356725,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission11727/Authors"],"signatures":["ICLR.cc/2026/Conference/Submission11727/Authors"],"forum":"3FcjKYmNY5","license":"CC BY 4.0","number":11727,"cdate":1758203356725,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/-/Submission","ICLR.cc/2026/Conference/-/Post_Submission","ICLR.cc/2026/Conference/Submission11727/-/Full_Submission","ICLR.cc/2026/Conference/Submission11727/-/Rebuttal_Revision","ICLR.cc/2026/Conference/-/Edit"],"mdate":1770804965272,"odate":1759896705795,"domain":"ICLR.cc/2026/Conference","id":"3FcjKYmNY5","version":2},{"content":{"venue":{"value":"CoRR 2024"},"pdf":{"value":"http://arxiv.org/pdf/2411.02018v2"},"venueid":{"value":"dblp.org/journals/CORR/2024"},"paperhash":{"value":"song|shortcut_learning_in_incontext_learning_a_survey"},"authorids":{"value":["~Rui_Song4","~Yingji_Li2","https://dblp.org/search/pid/api?q=author:Lida_Shi:","https://dblp.org/search/pid/api?q=author:Fausto_Giunchiglia:","https://dblp.org/search/pid/api?q=author:Hao_Xu_0012:"]},"html":{"value":"https://doi.org/10.48550/arXiv.2411.02018"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2411-02018,\n  publtype={informal},\n  author={Rui Song and Yingji Li and Lida Shi and Fausto Giunchiglia and Hao Xu},\n  title={Shortcut Learning in In-Context Learning: A Survey},\n  year={2024},\n  cdate={1704067200000},\n  journal={CoRR},\n  volume={abs/2411.02018},\n  url={https://doi.org/10.48550/arXiv.2411.02018}\n}\n"},"abstract":{"value":"Shortcut learning refers to the phenomenon where models employ simple, non-robust decision rules in practical tasks, which hinders their generalization and robustness. With the rapid development of large language models (LLMs) in recent years, an increasing number of studies have shown the impact of shortcut learning on LLMs. This paper provides a novel perspective to review relevant research on shortcut learning in In-Context Learning (ICL). It conducts a detailed exploration of the types of shortcuts in ICL tasks, their causes, available benchmarks, and strategies for mitigating shortcuts. Based on corresponding observations, it summarizes the unresolved issues in existing research and attempts to outline the future research landscape of shortcut learning."},"title":{"value":"Shortcut Learning in In-Context Learning: A Survey"},"authors":{"value":["Rui Song","Yingji Li","Lida Shi","Fausto Giunchiglia","Hao Xu"]}},"tmdate":1741332157259,"pdate":1704067200000,"tcdate":1739178728711,"writers":["~"],"signatures":["~yingji_Li1"],"forum":"iOMMaR5Rfr","license":"CC BY-SA 4.0","number":314961,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1741332157259,"domain":"DBLP.org","id":"iOMMaR5Rfr","version":2},{"content":{"summary":{"value":"This paper focuses on text-video retrieval for the challenging unsupervised domain adaptation setting. For this problem, the model is trained on a source only set of supervised video/text pairs and is adapted to a target domain which consists of videos and text labels with no pairing ground truth. The proposed method, named Uncertainty-aware Alignment Network (UAN), uses the assumption that many videos can correspond to a textual item and vice-versa, which allows for the multi-granularity modelling. They evaluate their method on three different pairs of datasets between MSR-VTT, MSVD, and TGIF. Their method is found to outperform all other unsupervised domain adaptation methods for video retrieval because they adapt classifier domain adaptation approaches."},"presentation":{"value":"4 excellent"},"contribution":{"value":"3 good"},"soundness":{"value":"3 good"},"strengths":{"value":"The experiments that have been run are convincing and thorough and even include evaluation of the method for image-text domain adaptation which is nice to see and proves that the proposed approach was created from the ground up for embedding spaces instead of borrowed from unsupervised domain adaptation approaches for classification. As a final note, the experiments were all run 5 times and averaged for robustness.\n\nThe proposed model works well for unsupervised domain adaptation video retrieval and makes sense to break from the one-to-one assumption during the training to better find potential positives."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"What are the values of a and b given in equation 7/8 and how are these chosen?\n\nEquations 10 and 11 could be better introduced/explained, additionally, it's easy to miss that T and K refer to the batch and for an individual example respectively.\n\nThe discussion on failure cases of the model within the paper is rather short. There is one given failure case within figure 6 which is very briefly explained (so much so that I'm not sure how the model failed on this case). I think these could be better highlighted as another future avenue for related work and better understanding of the method as a reader.\n\nThe method calls for many-to-one and one-to-many relationships between the two modalities during training, but for evaluation it is assumed that there is only a one-to-one relationship between modalities as with previous work. There has been some work that mentions the possibility of many-to-many relationships during training for images [a] and videos [b] and it might be worth discussing this point within the limitations which also mentions the one-to-many relationships at the early training stage.\n\n[a] Chun, Sanghyuk, et al. \"Probabilistic embeddings for cross-modal retrieval.\" Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2021.\n[b] Wray, Michael, Hazel Doughty, and Dima Damen. \"On semantic similarity in video retrieval.\" Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2021.\n\nOther Comments\nLine 6: missing space between comma and 'as'\n[24] and [25] are duplicated citations\nLine 142: inner production -> inner product"},"confidence":{"value":"5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully."},"questions":{"value":"1. Can the failure cases of the method be expanded and more information given regarding the failure case in Figure 6\n2. What are the values of a and b given in equation 7/8 when training the model and how are these chosen?\n3. With the domain adaptation requiring the one-to-one relationship between video and text during evaluation to be relaxed into a many-to-many relationship during training, do you think this has a method on overall method performance?"},"rating":{"value":"6: Weak Accept: Technically solid, moderate-to-high impact paper, with no major concerns with respect to evaluation, resources, reproducibility, ethical considerations."},"code_of_conduct":{"value":"Yes"},"limitations":{"value":"Limitations are discussed within the paper in the final paragraph, however, not much is said for the method beyond selecting pairs in the first few epochs likely causes issues. I think this could be expanded further to round out the paper."}},"nonreaders":[],"tmdate":1702410809547,"tcdate":1688412294009,"writers":["NeurIPS.cc/2023/Conference","NeurIPS.cc/2023/Conference/Submission1873/Reviewer_FYKf"],"signatures":["NeurIPS.cc/2023/Conference/Submission1873/Reviewer_FYKf"],"forum":"iQlK3VJxV7","number":1,"license":"CC BY 4.0","cdate":1688412294009,"mdate":1702410809547,"readers":["everyone"],"invitations":["NeurIPS.cc/2023/Conference/Submission1873/-/Official_Review","NeurIPS.cc/2023/Conference/-/Edit"],"domain":"NeurIPS.cc/2023/Conference","replyto":"iQlK3VJxV7","id":"fuyqHdKfjr","forumContent":{"venue":{"value":"NeurIPS 2023 poster"},"keywords":{"value":["video-text retrieval; cross-domain;Unsupervised Domain Adaptation Video-text Retrieval;"]},"_bibtex":{"value":"@inproceedings{\nhao2023uncertaintyaware,\ntitle={Uncertainty-Aware Alignment  Network  for Cross-Domain Video-Text Retrieval},\nauthor={Xiaoshuai Hao and Wanqian Zhang},\nbooktitle={Thirty-seventh Conference on Neural Information Processing Systems},\nyear={2023},\nurl={https://openreview.net/forum?id=iQlK3VJxV7}\n}"},"title":{"value":"Uncertainty-Aware Alignment  Network  for Cross-Domain Video-Text Retrieval"},"paperhash":{"value":"hao|uncertaintyaware_alignment_network_for_crossdomain_videotext_retrieval"},"TLDR":{"value":"In this paper, we address the challenge task of Unsupervised Domain Adaptation Video-text Retrieval (UDAVR), assuming that training (source) data and testing (target) data are from different domains."},"abstract":{"value":"Video-text retrieval is an important but challenging research task in the multimedia community.  In this paper, we address the challenge task of Unsupervised Domain Adaptation Video-text Retrieval (UDAVR), assuming that training (source) data and testing (target) data are from different domains. Previous approaches are mostly derived from classification based domain adaptation methods, which are neither multi-modal nor suitable for retrieval task.  In addition, as to the pairwise misalignment issue in target domain, i.e., no pairwise annotations between target videos and texts, the existing method assumes that a video corresponds to a text. Yet we empirically find that in the real scene, one text usually corresponds to multiple videos and vice versa. To tackle this one-to-many issue, we propose a novel method named Uncertainty-aware Alignment Network (UAN). Specifically, we first introduce the multimodal mutual information module to balance the minimization of domain shift in a smooth manner. To tackle the multimodal uncertainties pairwise misalignment in target domain, we propose the Uncertainty-aware Alignment Mechanism (UAM) to fully exploit the semantic information of both modalities in target domain. Extensive experiments in the context of domain-adaptive video-text retrieval demonstrate that our proposed method consistently outperforms multiple baselines, showing a superior generalization ability for target data."},"pdf":{"value":"/pdf/8e69ebbe33b1ba360a90af65d4977177eb771a6c.pdf"},"venueid":{"value":"NeurIPS.cc/2023/Conference"},"authorids":{"value":["~Xiaoshuai_Hao1","~Wanqian_Zhang1"]},"authors":{"value":["Xiaoshuai Hao","Wanqian Zhang"]}},"version":2},{"content":{"summary":{"value":"This paper introduces Artemis, a video-language model for video-based referential understanding. It can describe a target referred by a bounding box in a video. A referential video dataset named VideoRef45K is collected to train the model. Artemis is evaluated on HC-STVG benchmark and outperforms baselines adapted from image-based referring models. Besides referential understanding, Artemis can also perform general video question answering, and serve as a component in multi-round and long-form video understanding."},"soundness":{"value":2},"confidence":{"value":5},"questions":{"value":"1. Does the target-specific features in \\<region\\> tokens include the positional information of the bounding boxes?"},"rating":{"value":7},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":4},"strengths":{"value":"1. A video-based referential understanding dataset, VideoRef45K, is established. It facilitates the development of the area, providing referential pretraining data.\n2. The design of target-specific feature branch in the model architecture is well-motivated.\n3. Although no video-based referring model exists, the paper adapts image-based referring models to video as baselines. The adaption method is quite reasonable.\n4. Artemis outperforms the adapted image-based baselines significantly on HC-STVG.\n5. Artemis can still perform general video question answering and achieve better performance after training on the video referring task. This demonstrates that video referring can boost the reasoning capability of video-language models.\n6. Combined with existing video-language models, Artemis can perform multi-round video understanding with grounding and long-form video understanding."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The RoI selection step of Artemis clusters the object bounding boxes from different frames. However, the clustering algorithm only considers the bounding box coordinates but does not take the visual content in the bounding boxes into account. In some cases, the bounding box of an object may remain unchanged for a long time but the object state keeps changing, e.g., a person standing at a certain location performs a series of actions. Clustering these frames together can compressing them into one would lose valuable information.\n2. The proposed Artemis is built based on video-language models Video-LLaVA and Video-ChatGPT. However, these video-language models are not used as baselines in the experiments. Although they are not originally developed for video referring, there is a simple approach to adapt them for the task. As suggested by [1], directly drawing a circle or a rectangle on an image can help VLMs focus on the indicated object. Therefore, one can adapt the video-language models for video referring by drawing the object bounding boxes on the video frames and ask the model \"What is the target indicated by the red rectangle doing?\". Artemis should be compared with this simple baseline to demonstrate the effectiveness of the RoI feature branch in its model architecture.\n\n[1] Shtedritski et al. What does CLIP know about a red circle? Visual prompt engineering for VLMs. ICCV 2023."},"limitations":{"value":"Limitations are discussed in the paper."}},"nonreaders":[],"tmdate":1730878833373,"tcdate":1720111354786,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission3231/Reviewer_xzsp"],"signatures":["NeurIPS.cc/2024/Conference/Submission3231/Reviewer_xzsp"],"forum":"FaNhyXY6Y1","number":1,"license":"CC BY 4.0","cdate":1720111354786,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission3231/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730878833373,"domain":"NeurIPS.cc/2024/Conference","replyto":"FaNhyXY6Y1","id":"XuzFDXRBMz","forumContent":{"venue":{"value":"NeurIPS 2024 poster"},"TLDR":{"value":"This paper proposes a challenging setting for video-based referring and establishes an effective MLLM named Artemis."},"keywords":{"value":["Video Referring","Multimodal","RoI Selection Mechanism"]},"supplementary_material":{"value":"/attachment/2120e9d0c563ce94e32fb40dc3c9f329faaa2a59.zip"},"primary_area":{"value":"machine_vision"},"abstract":{"value":"Videos carry rich visual information including object description, action, interaction, etc., but the existing multimodal large language models (MLLMs) fell short in referential understanding scenarios such as video-based referring. In this paper, we present Artemis, an MLLM that pushes video-based referential understanding to a finer level. Given a video, Artemis receives a natural-language question with a bounding box in any video frame and describes the referred target in the entire video. The key to achieving this goal lies in extracting compact, target-specific video features, where we set a solid baseline by tracking and selecting spatiotemporal features from the video. We train Artemis on the newly established ViderRef45K dataset with 45K video-QA pairs and design a computationally efficient, three-stage training procedure. Results are promising both quantitatively and qualitatively. Additionally, we show that Artemis can be integrated with video grounding and text summarization tools to understand more complex scenarios. Code and data are available at https://github.com/NeurIPS24Artemis/Artemis."},"_bibtex":{"value":"@inproceedings{\nqiu2024artemis,\ntitle={Artemis:  Towards Referential Understanding in Complex Videos},\nauthor={Jihao Qiu and Yuan Zhang and Xi Tang and Lingxi Xie and Tianren Ma and Pengyu Yan and David Doermann and Qixiang Ye and Yunjie Tian},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=FaNhyXY6Y1}\n}"},"title":{"value":"Artemis:  Towards Referential Understanding in Complex Videos"},"pdf":{"value":"/pdf/6adf08928f9e63c0985b07403580c3d66bc2a7c8.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"qiu|artemis_towards_referential_understanding_in_complex_videos"},"authorids":{"value":["~Jihao_Qiu1","~Yuan_Zhang31","~Xi_Tang2","~Lingxi_Xie1","~Tianren_Ma1","~Pengyu_Yan1","~David_Doermann2","~Qixiang_Ye1","~Yunjie_Tian1"]},"authors":{"value":["Jihao Qiu","Yuan Zhang","Xi Tang","Lingxi Xie","Tianren Ma","Pengyu Yan","David Doermann","Qixiang Ye","Yunjie Tian"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a training-free framework that replaces isolated key frames with temporally coherent key clips to better capture motion and temporal context under a fixed token budget.\nBy introducing an adaptive trade-off between clip length and spatial resolution, F2C consistently improves Video-LMM performance on multiple long-form video benchmarks, demonstrating superior efficiency and temporal reasoning ability."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"See weakness"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1.\tThe paper introduces F2C that replaces isolated key frames with temporally coherent key clips, addressing the “needle in a haystack” problem in long-form video understanding.\n2.\tThe method is technically sound and well-analyzed, offering an adaptive trade-off between spatial resolution and clip length, with clear ablations verifying the impact of each design choice.\n3.\tF2C shows consistent improvements across multiple benchmarks while remaining computationally efficient and easy to integrate into existing Video-LMM pipelines."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. As introduced in the abstract, long-form video understanding tasks are constrained by the “needle in a haystack” problem. The authors should provide a clearer visualization or comparison to illustrate how their method alleviates this issue (e.g., similar to Figure 4 in [1], Figure 8 in [2], or visualization examples in [3]).\n\n2. Compared with frame-level token merging methods [4], it remains unclear what specific advantages this token selection strategy provides.\n\t\n3. The paper lacks some necessary references to related work, especially other training-free long-form video understanding models [5] [6] [7].\n\t\n4. It is also unclear whether the λr and λl choices are dependent on the benchmark used.\n\n[1] Zhang, Peiyuan, et al. \"Long context transfer from language to vision.\" arXiv preprint arXiv:2406.16852 (2024).\n\n[2] Zhao, Zijia, et al. \"Needle in a video haystack: A scalable synthetic evaluator for video mllms.\" arXiv preprint arXiv:2406.09367 (2024).\n\n[3] https://github.com/bigai-nlco/NeedleInAVideoHaystack\n\n[4] Song, Enxin, et al. \"Moviechat+: Question-aware sparse memory for long video question answering.\" IEEE Transactions on Pattern Analysis and Machine Intelligence (2025).\n\n[5] Santos, Saul, et al. \"$\\infty $-Video: A Training-Free Approach to Long Video Understanding via Continuous-Time Memory Consolidation.\" arXiv preprint arXiv:2501.19098 (2025).\n\n[6] Xu, Mingze, et al. \"Slowfast-llava: A strong training-free baseline for video large language models.\" arXiv preprint arXiv:2407.15841 (2024).\n\n[7] Zhang, Yiming, et al. \"Beyond training: Dynamic token merging for zero-shot video understanding.\" Proceedings of the IEEE/CVF International Conference on Computer Vision. 2025."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762918996762,"tcdate":1761705055440,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission6707/Reviewer_D9X5"],"signatures":["ICLR.cc/2026/Conference/Submission6707/Reviewer_D9X5"],"forum":"BAdePgN4uR","number":3,"license":"CC BY 4.0","cdate":1761705055440,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission6707/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762918996762,"domain":"ICLR.cc/2026/Conference","replyto":"BAdePgN4uR","id":"Si04NpUFsP","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"TLDR":{"value":"We propose F2C, which selects temporally coherent clips with adaptive resolution, outperforming frame-based sampling for long-form video understanding under fixed token budgets."},"keywords":{"value":["Video Large Language Model","Frame Selection"]},"primary_area":{"value":"foundation or frontier models, including LLMs"},"abstract":{"value":"Video Large Language Models (VLMs) have achieved remarkable results on a variety of vision language tasks, yet their practical use is limited by the \"needle in a haystack\" problem: the massive number of visual tokens produced from raw video frames exhausts the model’s context window. Existing solutions alleviate this issue by selecting a sparse set of frames, thereby reducing token count, but such frame-wise selection discards essential temporal dynamics, leading to suboptimal reasoning about motion and event continuity. In this work we systematically explore the impact of temporal information and demonstrate that extending selection from isolated key frames to key clips, which are short, temporally coherent segments, improves video understanding.\nTo maintain a fixed computational budget while accommodating the larger token footprint of clips, we propose an adaptive resolution strategy that dynamically balances spatial resolution and clip length, ensuring a constant token count per video. Experiments on three long-form video benchmarks demonstrate that our training-free approach, F2C, outperforms uniform sampling up to 8.1%, 5.6%, and 10.3% on Video-MME, LongVideoBench and MLVU benchmarks, respectively. These results highlight the importance of preserving temporal coherence in frame selection and provide a practical pathway for scaling Video LLMs to real world video understanding applications."},"_bibtex":{"value":"@misc{\nsun2025from,\ntitle={From Frames to Clips: Efficient Key Clip Selection for Long-Form Video Understanding},\nauthor={Guangyu Sun and Archit Singhal and Burak Uzkent and Mubarak Shah and Chen Chen and Garin N. Kessler},\nyear={2025},\nurl={https://openreview.net/forum?id=BAdePgN4uR}\n}"},"title":{"value":"From Frames to Clips: Efficient Key Clip Selection for Long-Form Video Understanding"},"pdf":{"value":"/pdf/8a673fbbfe961f36d3db6761131e622d16a88413.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"sun|from_frames_to_clips_efficient_key_clip_selection_for_longform_video_understanding"},"authorids":{"value":["~Guangyu_Sun3","~Archit_Singhal1","~Burak_Uzkent1","~Mubarak_Shah3","~Chen_Chen18","~Garin_N._Kessler1"]},"authors":{"value":["Guangyu Sun","Archit Singhal","Burak Uzkent","Mubarak Shah","Chen Chen","Garin N. Kessler"]}},"version":2},{"content":{"venue":{"value":"CoRR 2023"},"pdf":{"value":"http://arxiv.org/pdf/2312.14223v2"},"venueid":{"value":"dblp.org/journals/CORR/2023"},"paperhash":{"value":"weng|fast_diffusionbased_counterfactuals_for_shortcut_removal_and_generation"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Nina_Weng:","https://dblp.org/search/pid/api?q=author:Paraskevas_Pegios:","https://dblp.org/search/pid/api?q=author:Aasa_Feragen:","~Eike_Petersen1","https://dblp.org/search/pid/api?q=author:Siavash_Bigdeli:"]},"html":{"value":"https://doi.org/10.48550/arXiv.2312.14223"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2312-14223,\n  publtype={informal},\n  author={Nina Weng and Paraskevas Pegios and Aasa Feragen and Eike Petersen and Siavash Bigdeli},\n  title={Fast Diffusion-Based Counterfactuals for Shortcut Removal and Generation},\n  year={2023},\n  cdate={1672531200000},\n  journal={CoRR},\n  volume={abs/2312.14223},\n  url={https://doi.org/10.48550/arXiv.2312.14223}\n}\n"},"abstract":{"value":"Shortcut learning is when a model -- e.g. a cardiac disease classifier -- exploits correlations between the target label and a spurious shortcut feature, e.g. a pacemaker, to predict the target label based on the shortcut rather than real discriminative features. This is common in medical imaging, where treatment and clinical annotations correlate with disease labels, making them easy shortcuts to predict disease. We propose a novel detection and quantification of the impact of potential shortcut features via a fast diffusion-based counterfactual image generation that can synthetically remove or add shortcuts. Via a novel inpainting-based modification we spatially limit the changes made with no extra inference step, encouraging the removal of spatially constrained shortcut features while ensuring that the shortcut-free counterfactuals preserve their remaining image features to a high degree. Using these, we assess how shortcut features influence model predictions. This is enabled by our second contribution: An efficient diffusion-based counterfactual explanation method with significant inference speed-up at comparable image quality as state-of-the-art. We confirm this on two large chest X-ray datasets, a skin lesion dataset, and CelebA. Our code is publicly available at fastdime.compute.dtu.dk."},"title":{"value":"Fast Diffusion-Based Counterfactuals for Shortcut Removal and Generation"},"authors":{"value":["Nina Weng","Paraskevas Pegios","Aasa Feragen","Eike Petersen","Siavash Bigdeli"]}},"tmdate":1727524191499,"pdate":1672531200000,"tcdate":1727524183474,"writers":["~"],"signatures":["~Eike_Petersen1"],"forum":"bCTmkExKVZ","license":"CC BY-SA 4.0","number":101818,"cdate":1672531200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1727524191499,"domain":"DBLP.org","id":"bCTmkExKVZ","version":2},{"content":{"venue":{"value":"NeurIPS 2025 poster"},"keywords":{"value":["Panoramic Video Generation","Video Diffusion Models"]},"supplementary_material":{"value":"/attachment/09bfb028f1ad03dc96999ac8df322c7a5f20ef51.zip"},"primary_area":{"value":"deep_learning"},"abstract":{"value":"Panoramic video generation enables immersive 360$^\\circ$ content creation, valuable in applications that demand scene-consistent world exploration. However, existing panoramic video generation models struggle to leverage pre-trained generative priors from conventional text-to-video models for high-quality and diverse panoramic videos generation, due to limited dataset scale and the gap in spatial feature representations. In this paper, we introduce PanoWan to effectively lift pre-trained text-to-video models to the panoramic domain, equipped with minimal modules. PanoWan employs latitude-aware sampling to avoid latitudinal distortion, while its rotated semantic denoising and padded pixel-wise decoding ensure seamless transitions at longitude boundaries. To provide sufficient panoramic videos for learning these lifted representations, we contribute PanoVid, a high-quality panoramic video dataset with captions and diverse scenarios. Consequently, PanoWan achieves state-of-the-art performance in panoramic video generation and demonstrates robustness for zero-shot downstream tasks."},"_bibtex":{"value":"@inproceedings{\nxia2025panowan,\ntitle={PanoWan: Lifting Diffusion Video Generation Models to 360\\${\\textasciicircum}{\\textbackslash}circ\\$ with Latitude/Longitude-aware Mechanisms},\nauthor={Yifei Xia and Shuchen Weng and Siqi Yang and Jingqi Liu and Chengxuan Zhu and Minggui Teng and Zijian Jia and Han Jiang and Boxin Shi},\nbooktitle={The Thirty-ninth Annual Conference on Neural Information Processing Systems},\nyear={2025},\nurl={https://openreview.net/forum?id=7VLxvVEtHh}\n}"},"title":{"value":"PanoWan: Lifting Diffusion Video Generation Models to 360$^\\circ$ with Latitude/Longitude-aware Mechanisms"},"pdf":{"value":"/pdf/9623830127c1ad542bd54f0439a7e3fed1f6ceaa.pdf"},"venueid":{"value":"NeurIPS.cc/2025/Conference"},"paperhash":{"value":"xia|panowan_lifting_diffusion_video_generation_models_to_360^\\circ_with_latitudelongitudeaware_mechanisms"},"authorids":{"value":["~Yifei_Xia1","~Shuchen_Weng1","~Siqi_Yang2","~Jingqi_Liu2","~Chengxuan_Zhu1","~Minggui_Teng1","~Zijian_Jia1","~Han_Jiang11","~Boxin_Shi3"]},"authors":{"value":["Yifei Xia","Shuchen Weng","Siqi Yang","Jingqi Liu","Chengxuan Zhu","Minggui Teng","Zijian Jia","Han Jiang","Boxin Shi"]}},"tmdate":1783627451529,"pdate":1758216569978,"tcdate":1745141091492,"writers":["NeurIPS.cc/2025/Conference","NeurIPS.cc/2025/Conference/Submission3256/Authors"],"signatures":["NeurIPS.cc/2025/Conference/Submission3256/Authors"],"forum":"7VLxvVEtHh","license":"CC BY 4.0","number":3256,"cdate":1745141091492,"readers":["everyone"],"invitations":["NeurIPS.cc/2025/Conference/-/Submission","NeurIPS.cc/2025/Conference/-/Post_Submission","NeurIPS.cc/2025/Conference/Submission3256/-/Full_Submission","NeurIPS.cc/2025/Conference/Submission3256/-/Supplementary_Material","NeurIPS.cc/2025/Conference/-/Edit","NeurIPS.cc/2025/Conference/Submission3256/-/Camera_Ready_Revision"],"mdate":1783627451529,"odate":1761704751461,"domain":"NeurIPS.cc/2025/Conference","id":"7VLxvVEtHh","version":2},{"content":{"TLDR":{"value":"We propose IFCap, leveraging Image-like Retrieval and Frequency-based Entity Filtering to enhance zero-shot captioning by bridging the modality gap and improving caption quality significantly."},"venue":{"value":"Video-Langauge Models Poster"},"keywords":{"value":["Video-to-text generation","including video captioning and description"]},"supplementary_material":{"value":"/attachment/172348f0a45aa3fdff676ddcd1cb530cd4b528bc.zip"},"abstract":{"value":"Recent advancements in image captioning have explored text-only training methods to overcome the limitations of paired image-text data. However, existing text-only training methods often overlook the modality gap between using text data during training and employing images during inference. To address this issue, we propose a novel approach called Image-like Retrieval, which aligns text features with visually relevant features to mitigate the modality gap. Our method further enhances the accuracy of generated captions by designing a fusion module that integrates retrieved captions with input features. Additionally, we introduce a Frequency-based Entity Filtering technique that significantly improves caption quality. We integrate these methods into a unified framework, which we refer to as $\\text{IFCap}$ ($\\textbf{I}$mage-like Retrieval and $\\textbf{F}$requency-based Entity Filtering for Zero-shot $\\textbf{Cap}$tioning). Through extensive experimentation, our straightforward yet powerful approach has demonstrated its efficacy, outperforming the state-of-the-art methods by a significant margin in both image captioning and video captioning compared to zero-shot captioning based on text-only training."},"_bibtex":{"value":"@inproceedings{\nlee2025ifcap,\ntitle={{IFC}ap: Image-like Retrieval and Frequency-based Entity Filtering for Zero-shot Captioning},\nauthor={Soeun Lee and Si-Woo Kim and Taewhan Kim and Dong-Jin Kim},\nbooktitle={Workshop on Video-Language Models @ NeurIPS 2024},\nyear={2025},\nurl={https://openreview.net/forum?id=IbLmsNVlcI}\n}"},"title":{"value":"IFCap: Image-like Retrieval and Frequency-based Entity Filtering for Zero-shot Captioning"},"pdf":{"value":"/pdf/26fa2ac249dd141d9e5d7fde44ff317e5e93dcaa.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Workshop/Video-Langauge_Models"},"paperhash":{"value":"lee|ifcap_imagelike_retrieval_and_frequencybased_entity_filtering_for_zeroshot_captioning"},"authorids":{"value":["~Soeun_Lee1","~Si-Woo_Kim1","~Taewhan_Kim2","~Dong-Jin_Kim1"]},"track":{"value":"Long Paper Track (up to 9 pages)"},"authors":{"value":["Soeun Lee","Si-Woo Kim","Taewhan Kim","Dong-Jin Kim"]}},"tmdate":1736861081006,"pdate":1730081753443,"tcdate":1726315380758,"writers":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission52/Authors"],"signatures":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission52/Authors"],"forum":"IbLmsNVlcI","license":"CC BY 4.0","number":52,"cdate":1726315380758,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Post_Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission52/-/Full_Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Edit","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission52/-/Camera-Ready_Revision"],"mdate":1736861081006,"odate":1736861080991,"domain":"NeurIPS.cc/2024/Workshop/Video-Langauge_Models","id":"IbLmsNVlcI","version":2},{"content":{"venue":{"value":"Video-Langauge Models Poster"},"pdf":{"value":"/pdf/00b117289bd7d5e16d15a447180df3fef4f71d41.pdf"},"keywords":{"value":["video language model","knowledge distillation","action classification"]},"venueid":{"value":"NeurIPS.cc/2024/Workshop/Video-Langauge_Models"},"paperhash":{"value":"wei|dualmodel_distillation_for_efficient_action_classification_with_hybrid_edgecloud_solution"},"authorids":{"value":["~Timothy_Wei1","~Hsien_Xin_Peng1","elainexuyz@gmail.com","~Bryan_Zhao1","~Lei_Ding10","~Diji_Yang1"]},"abstract":{"value":"As Artificial Intelligence models, such as Large Video-Language models (VLMs), grow in size, their deployment in real-world applications becomes increasingly challenging due to hardware limitations and computational costs. To address this, we design a hybrid edge-cloud solution that leverages the efficiency of smaller models for local processing while deferring to larger, more accurate cloud-based models when necessary. Specifically, we propose a novel unsupervised data generation method, Dual-Model Distillation (DMD), to train a lightweight switcher model that can predict when the edge model’s output is uncertain and selectively offload inference to the large model in the cloud. Experimental results on the action classification task show that our framework not only requires less computational overhead, but also improves accuracy compared to using a large model alone. Our framework provides a scalable and adaptable solution for action classification in resource-constrained environments, with potential applications beyond healthcare. Noteworthy, while DMD-generated data is used for optimizing performance and resource usage in our pipeline, we expect the concept of DMD to further support future research on knowledge alignment across multiple models."},"_bibtex":{"value":"@inproceedings{\nwei2025dualmodel,\ntitle={Dual-Model Distillation for Efficient Action Classification with Hybrid Edge-Cloud Solution},\nauthor={Timothy Wei and Hsien Xin Peng and Elaine Xu and Bryan Zhao and Lei Ding and Diji Yang},\nbooktitle={Workshop on Video-Language Models @ NeurIPS 2024},\nyear={2025},\nurl={https://openreview.net/forum?id=vJoFQR15bw}\n}"},"title":{"value":"Dual-Model Distillation for Efficient Action Classification with Hybrid Edge-Cloud Solution"},"track":{"value":"Short Paper Track (up to 3 pages)"},"authors":{"value":["Timothy Wei","Hsien Xin Peng","Elaine Xu","Bryan Zhao","Lei Ding","Diji Yang"]}},"tmdate":1736861080759,"pdate":1730081753172,"tcdate":1726012794527,"writers":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission42/Authors"],"signatures":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission42/Authors"],"forum":"vJoFQR15bw","license":"CC BY 4.0","number":42,"cdate":1726012794527,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Post_Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission42/-/Full_Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Edit","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission42/-/Camera-Ready_Revision"],"mdate":1736861080759,"odate":1736861080746,"domain":"NeurIPS.cc/2024/Workshop/Video-Langauge_Models","id":"vJoFQR15bw","version":2},{"content":{"venue":{"value":"CoRR 2025"},"pdf":{"value":"https://arxiv.org/pdf/2502.13363v1"},"venueid":{"value":"dblp.org/journals/CORR/2025"},"paperhash":{"value":"zhang|pretrained_imagetext_models_are_secretly_video_captioners"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Chunhui_Zhang:","https://dblp.org/search/pid/api?q=author:Yiren_Jian:","https://dblp.org/search/pid/api?q=author:Zhongyu_Ouyang:","~Soroush_Vosoughi1"]},"html":{"value":"https://doi.org/10.48550/arXiv.2502.13363"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2502-13363,\n  publtype={informal},\n  author={Chunhui Zhang and Yiren Jian and Zhongyu Ouyang and Soroush Vosoughi},\n  title={Pretrained Image-Text Models are Secretly Video Captioners},\n  year={2025},\n  month={February},\n  cdate={1738368000000},\n  journal={CoRR},\n  volume={abs/2502.13363},\n  url={https://doi.org/10.48550/arXiv.2502.13363}\n}\n"},"abstract":{"value":"Developing video captioning models is computationally expensive. The dynamic nature of video also complicates the design of multimodal models that can effectively caption these sequences. However, we find that by using minimal computational resources and without complex modifications to address video dynamics, an image-based model can be repurposed to outperform several specialised video captioning systems. Our adapted model demonstrates top tier performance on major benchmarks, ranking 2nd on MSRVTT and MSVD, and 3rd on VATEX. We transform it into a competitive video captioner by post training a typical image captioning model BLIP2 with only 6,000 video text pairs and simply concatenating frames (significantly fewer data than other methods), which use 2.5 to 144 million pairs. From a resource optimization perspective, this video captioning study focuses on three fundamental factors: optimizing model scale, maximizing data efficiency, and incorporating reinforcement learning. This extensive study demonstrates that a lightweight, image based adaptation strategy can rival state-of-the-art video captioning systems, offering a practical solution for low-resource scenarios."},"title":{"value":"Pretrained Image-Text Models are Secretly Video Captioners"},"authors":{"value":["Chunhui Zhang","Yiren Jian","Zhongyu Ouyang","Soroush Vosoughi"]}},"tmdate":1768967195846,"pdate":1767139200000,"externalIds":["dblp:journals/corr/abs-2502-13363"],"tcdate":1768966359200,"writers":["~"],"signatures":["~Soroush_Vosoughi1"],"forum":"NcTkpOZQ0s","license":"CC BY-SA 4.0","number":774991,"cdate":1738368000000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1768967195846,"domain":"DBLP.org","id":"NcTkpOZQ0s","version":2},{"content":{"venue":{"value":"NAACL (Short Papers) 2025"},"pdf":{"value":"https://aclanthology.org/2025.naacl-short.26.pdf"},"venueid":{"value":"dblp.org/conf/NAACL/2025"},"paperhash":{"value":"zhang|pretrained_imagetext_models_are_secretly_video_captioners"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Chunhui_Zhang:","https://dblp.org/search/pid/api?q=author:Yiren_Jian:","https://dblp.org/search/pid/api?q=author:Zhongyu_Ouyang:","~Soroush_Vosoughi1"]},"html":{"value":"https://doi.org/10.18653/v1/2025.naacl-short.26"},"_bibtex":{"value":"@inproceedings{DBLP:conf/naacl/ZhangJOV25,\n  author={Chunhui Zhang and Yiren Jian and Zhongyu Ouyang and Soroush Vosoughi},\n  title={Pretrained Image-Text Models are Secretly Video Captioners},\n  year={2025},\n  cdate={1735689600000},\n  pages={292-305},\n  url={https://doi.org/10.18653/v1/2025.naacl-short.26},\n  booktitle={NAACL (Short Papers)},\n  crossref={conf/naacl/2025-2}\n}\n"},"abstract":{"value":"Developing video captioning models is computationally expensive. The dynamic nature of video also complicates the design of multimodal models that can effectively caption these sequences. However, we find that by using minimal computational resources and without complex modifications to address video dynamics, an image-based model can be repurposed to outperform several specialised video captioning systems. Our adapted model demonstrates top-tier performance on major benchmarks, ranking 2nd on MSR-VTT and MSVD, and 3rd on VATEX. We transform it into a competitive video captioner by post-training a typical image captioning model BLIP-2 with only 6,000 video-text pairs and simply concatenating frames—significantly fewer data than other methods, which use 2.5 to 144 million pairs. From a resource optimization perspective, this video captioning study focuses on three fundamental factors: optimizing model scale, maximizing data efficiency, and incorporating reinforcement learning. This extensive study demonstrates that a lightweight, image-based adaptation strategy can rival state-of-the-art video captioning systems, offering a practical solution for low-resource scenarios."},"title":{"value":"Pretrained Image-Text Models are Secretly Video Captioners"},"authors":{"value":["Chunhui Zhang","Yiren Jian","Zhongyu Ouyang","Soroush Vosoughi"]}},"tmdate":1768967191480,"pdate":1767139200000,"externalIds":["dblp:conf/naacl/ZhangJOV25"],"tcdate":1768966358755,"writers":["~"],"signatures":["~Soroush_Vosoughi1"],"forum":"jTZmd4Fmib","license":"CC BY-SA 4.0","number":774978,"cdate":1735689600000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1768967191480,"domain":"DBLP.org","id":"jTZmd4Fmib","version":2},{"content":{"summary":{"value":"This paper presents **OpenBiomedVid**, a large-scale dataset of **1,031 hours of biomedical educational videos** from YouTube, containing **22K clips** and **79K Q/A pairs**. The dataset is created through a **human-in-the-loop pipeline** that combines Whisper transcription, GPT-4o caption refinement, and expert verification. The authors also introduce two benchmarks, **SurgeryVideoQA** and **MIMICEchoQA**, for biomedical video-language evaluation. Fine-tuning **Qwen2-VL (2B/7B)** models on OpenBiomedVid yields strong performance gains across biomedical video and image benchmarks, showing that open educational videos can be an effective training signal."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. **Will the dataset and benchmarks be fully released?** If so, under what license and with what access restrictions?  \n2. How is **personal or sensitive information** (e.g., PHI, identifiable subjects) detected and filtered in the YouTube videos?  \n3. Does the fine-tuned model show **transferability** to unseen clinical or institutional datasets?  \n4. What is the ratio of educational diagrams/narration-only videos to true clinical imaging?  \n5. Are there **quantitative measures of data quality**, such as annotation agreement or verification accuracy?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. Addresses an underexplored area: **biomedical video-language modeling**.  \n2. The dataset is **large, diverse, and well-organized**, potentially a valuable community contribution.  \n3. The **data curation pipeline** is systematic, combining LLM-based processing with human verification.  \n4.  Empirical results are consistent across multiple benchmarks and architectures."},"flag_for_ethics_review":{"value":["Yes, Privacy, security and safety"]},"weaknesses":{"value":"1.  **Evaluation bias** — heavy dependence on GPT-4o for caption refinement, Q/A generation, and evaluation creates potential bias and reproducibility concerns.  \n2. **Lack of deep analysis** — minimal discussion of failure cases, temporal reasoning, or cross-domain generalization.  \n3. **Ethical and legal clarity** — the discussion of data licensing, PHI risk, and content ownership is insufficient.  \n4. **Incremental insight** — while scale is impressive, the core finding (“educational videos help”) feels somewhat intuitive and under-analyzed."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762942036656,"tcdate":1761942500332,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission22059/Reviewer_QzV3"],"signatures":["ICLR.cc/2026/Conference/Submission22059/Reviewer_QzV3"],"forum":"u4PmZOmtko","number":4,"license":"CC BY 4.0","cdate":1761942500332,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission22059/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762942036656,"domain":"ICLR.cc/2026/Conference","replyto":"u4PmZOmtko","id":"QafLh2cICm","forumContent":{"TLDR":{"value":"Instruction-tuning Qwen-2-VL on 1,031 hours of pedagogical biomedical videos dramatically boosts video and image understanding and includes new expert-curated benchmarks, with all data and code released."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["vision-language models","biomedicine","datasets","evaluations"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Publicly available biomedical videos, such as those on YouTube, serve as valuable educational resources for medical students. Unlike standard machine learning datasets, these videos are designed for human learners, often mixing medical imagery with narration, explanatory diagrams, and contextual framing. In this work, we investigate whether such pedagogically rich, yet non-standardized and heterogeneous videos can effectively teach general-domain vision-language models biomedical knowledge. To this end, we introduce OpenBiomedVid, a biomedical video instruction tuning dataset comprising 1031 hours of video-caption and Q/A pairs, curated through a multi-step human-in-the-loop pipeline. Diverse biomedical video datasets are rare, and OpenBiomedVid fills an important gap by providing instruction-style supervision grounded in real-world educational content. Surprisingly, despite the informal and heterogeneous nature of these videos, the fine-tuned Qwen-2-VL models exhibit substantial performance improvements across most benchmarks. The 2B model achieves gains of 98.7% on video tasks, 71.2% on image tasks, and 0.2% on text tasks. The 7B model shows improvements of 37.09% on video and 11.2% on image tasks, with a slight degradation of 2.7% on text tasks compared to their respective base models. To address the lack of standardized biomedical video evaluation datasets, we also introduce two new expert curated benchmarks, MIMICEchoQA and SurgeryVideoQA. On these benchmarks, the 2B model achieves gains of 99.1% and 98.1%, while the 7B model shows gains of 22.5% and 52.1%, respectively, demonstrating the models' ability to generalize and perform biomedical video understanding on cleaner and more standardized datasets than those seen during training. These results suggest that educational videos created for human learning offer a surprisingly effective training signal for biomedical VLMs. We release OpenBiomedVid, MIMICEchoQA, SurgeryVideoQA, the fine-tuned models, and the complete codebase to support future research."},"_bibtex":{"value":"@misc{\nthapa2026how,\ntitle={How Well Can General Vision-Language Models Learn Medicine By Watching Public Educational Videos?},\nauthor={Rahul Thapa and Andrew Li and Qingyang Wu and Bryan He and Yuki Sahashi and Christina Binder and Angela Zhang and Ben Athiwaratkun and Shuaiwen Leon Song and David Ouyang and James Zou},\nyear={2026},\nurl={https://openreview.net/forum?id=u4PmZOmtko}\n}"},"title":{"value":"How Well Can General Vision-Language Models Learn Medicine By Watching Public Educational Videos?"},"pdf":{"value":"/pdf/fcf643dd9df9746824583d9bb992dfdc21d7b0d9.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"thapa|how_well_can_general_visionlanguage_models_learn_medicine_by_watching_public_educational_videos"},"authorids":{"value":["~Rahul_Thapa1","~Andrew_Li4","~Qingyang_Wu1","~Bryan_He1","~Yuki_Sahashi1","~Christina_Binder1","~Angela_Zhang1","~Ben_Athiwaratkun1","~Shuaiwen_Leon_Song1","~David_Ouyang1","~James_Zou1"]},"authors":{"value":["Rahul Thapa","Andrew Li","Qingyang Wu","Bryan He","Yuki Sahashi","Christina Binder","Angela Zhang","Ben Athiwaratkun","Shuaiwen Leon Song","David Ouyang","James Zou"]}},"version":2},{"content":{"venue":{"value":"CoRR 2022"},"pdf":{"value":"https://arxiv.org/pdf/2211.16220v1"},"venueid":{"value":"dblp.org/journals/CORR/2022"},"paperhash":{"value":"shinoda|which_shortcut_solution_do_question_answering_models_prefer_to_learn"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Kazutoshi_Shinoda:","https://dblp.org/search/pid/api?q=author:Saku_Sugawara:","~Akiko_Aizawa1"]},"html":{"value":"https://doi.org/10.48550/arXiv.2211.16220"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2211-16220,\n  publtype={informal},\n  author={Kazutoshi Shinoda and Saku Sugawara and Akiko Aizawa},\n  title={Which Shortcut Solution Do Question Answering Models Prefer to Learn?},\n  year={2022},\n  cdate={1640995200000},\n  journal={CoRR},\n  volume={abs/2211.16220},\n  url={https://doi.org/10.48550/arXiv.2211.16220}\n}\n"},"abstract":{"value":"Question answering (QA) models for reading comprehension tend to learn shortcut solutions rather than the solutions intended by QA datasets. QA models that have learned shortcut solutions can achieve human-level performance in shortcut examples where shortcuts are valid, but these same behaviors degrade generalization potential on anti-shortcut examples where shortcuts are invalid. Various methods have been proposed to mitigate this problem, but they do not fully take the characteristics of shortcuts themselves into account. We assume that the learnability of shortcuts, i.e., how easy it is to learn a shortcut, is useful to mitigate the problem. Thus, we first examine the learnability of the representative shortcuts on extractive and multiple-choice QA datasets. Behavioral tests using biased training sets reveal that shortcuts that exploit answer positions and word-label correlations are preferentially learned for extractive and multiple-choice QA, respectively. We find that the more learnable a shortcut is, the flatter and deeper the loss landscape is around the shortcut solution in the parameter space. We also find that the availability of the preferred shortcuts tends to make the task easier to perform from an information-theoretic viewpoint. Lastly, we experimentally show that the learnability of shortcuts can be utilized to construct an effective QA training set; the more learnable a shortcut is, the smaller the proportion of anti-shortcut examples required to achieve comparable performance on shortcut and anti-shortcut examples. We claim that the learnability of shortcuts should be considered when designing mitigation methods."},"title":{"value":"Which Shortcut Solution Do Question Answering Models Prefer to Learn?"},"authors":{"value":["Kazutoshi Shinoda","Saku Sugawara","Akiko Aizawa"]}},"tmdate":1768998995713,"pdate":1672444800000,"externalIds":["dblp:journals/corr/abs-2211-16220"],"tcdate":1768998978613,"writers":["~"],"signatures":["~Akiko_Aizawa1"],"forum":"aBsOjdUaDA","license":"CC BY-SA 4.0","number":783924,"cdate":1640995200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1768998995713,"domain":"DBLP.org","id":"aBsOjdUaDA","version":2},{"content":{"venue":{"value":"AAAI 2023"},"pdf":{"value":"https://ojs.aaai.org/index.php/AAAI/article/download/26590/26362"},"venueid":{"value":"dblp.org/conf/AAAI/2023"},"paperhash":{"value":"shinoda|which_shortcut_solution_do_question_answering_models_prefer_to_learn"},"authorids":{"value":["~Kazutoshi_Shinoda1","~Saku_Sugawara1","~Akiko_Aizawa1"]},"html":{"value":"https://doi.org/10.1609/aaai.v37i11.26590"},"_bibtex":{"value":"@inproceedings{DBLP:conf/aaai/ShinodaSA23,\n  author={Kazutoshi Shinoda and Saku Sugawara and Akiko Aizawa},\n  title={Which Shortcut Solution Do Question Answering Models Prefer to Learn?},\n  year={2023},\n  cdate={1672531200000},\n  pages={13564-13572},\n  url={https://doi.org/10.1609/aaai.v37i11.26590},\n  booktitle={AAAI},\n  crossref={conf/aaai/2023}\n}\n"},"abstract":{"value":"Question answering (QA) models for reading comprehension tend to exploit spurious correlations in training sets and thus learn shortcut solutions rather than the solutions intended by QA datasets. QA models that have learned shortcut solutions can achieve human-level performance in shortcut examples where shortcuts are valid, but these same behaviors degrade generalization potential on anti-shortcut examples where shortcuts are invalid. Various methods have been proposed to mitigate this problem, but they do not fully take the characteristics of shortcuts themselves into account. We assume that the learnability of shortcuts, i.e., how easy it is to learn a shortcut, is useful to mitigate the problem. Thus, we first examine the learnability of the representative shortcuts on extractive and multiple-choice QA datasets. Behavioral tests using biased training sets reveal that shortcuts that exploit answer positions and word-label correlations are preferentially learned for extractive and multiple-choice QA, respectively. We find that the more learnable a shortcut is, the flatter and deeper the loss landscape is around the shortcut solution in the parameter space. We also find that the availability of the preferred shortcuts tends to make the task easier to perform from an information-theoretic viewpoint. Lastly, we experimentally show that the learnability of shortcuts can be utilized to construct an effective QA training set; the more learnable a shortcut is, the smaller the proportion of anti-shortcut examples required to achieve comparable performance on shortcut and anti-shortcut examples. We claim that the learnability of shortcuts should be considered when designing mitigation methods."},"title":{"value":"Which Shortcut Solution Do Question Answering Models Prefer to Learn?"},"authors":{"value":["Kazutoshi Shinoda","Saku Sugawara","Akiko Aizawa"]}},"tmdate":1768998980224,"pdate":1672531200000,"tcdate":1737938619579,"writers":["~"],"signatures":["~Kazutoshi_Shinoda1"],"forum":"qOqk8UYR8c","license":"CC BY-SA 4.0","number":282130,"cdate":1672531200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1768998980224,"domain":"DBLP.org","id":"qOqk8UYR8c","version":2},{"content":{"summary":{"value":"The paper proposes UNIC (UNified In-Context Video Editing)， a framework for unifying diverse video editing tasks (e.g., ID insertion/deletion/swap, stylization, re-camera control, propagation) within a single diffusion transformer model. \nInstead of relying on task-specific adapters or DDIM inversion, UNIC concatenates noisy video, reference video, and multi-modal condition tokens into a single token sequence, enabling unified learning through native transformer attention. Experiments on a unified benchmark of six tasks show that UNIC achieves competitive or superior results to baselines like VACE, AnyV2V, and ReCamMaster, while offering emergent task composition and parameter efficiency."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"please see weaknesses section."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The idea of in-context unification across multiple video editing tasks via token concatenation is elegant and clear\n\n2. The introduction of Condition Bias and Task-Aware RoPE is well-motivated and clearly implemented.\n\n3. The benchmark spans six distinct tasks, covering both local (ID editing) and global (stylization, propagation) settings. it is comprehensive .\n4. The demonstration of task composition is interesting, highlights that the model generalizes beyond discrete training tasks"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The paper doesn’t compare against or reference newer state-of-the-art T2V and editing models like Wan series\n\n2. While the in-context formulation is elegant, much of the architecture directly borrows from existing full-attention DiT and OmniGen paradigms. The novelty primarily lies in combining these ideas for video, rather than a fundamentally new mechanism.\n\n3. Missing some important related work discussion like VEGGIE [1], which is also unified video editing framework + abilities in in-context video editing \n\n4. Some metrics (e.g., CLIP, DINO, ArtFID) are weak proxies for perceptual quality. The paper doesn’t include human evaluation or temporal consistency metrics (like VBench perceptual coherence, even though it is for video generation, it still captures the generation quality of a video editing model).\n\n[1] VEGGIE: Instructional Editing and Reasoning Video Concepts with Grounded Generation. ICCV25."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762922903578,"tcdate":1762136189974,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission11891/Reviewer_njmg"],"signatures":["ICLR.cc/2026/Conference/Submission11891/Reviewer_njmg"],"forum":"Vb4nE3WWf5","number":4,"license":"CC BY 4.0","cdate":1762136189974,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission11891/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762922903578,"domain":"ICLR.cc/2026/Conference","replyto":"Vb4nE3WWf5","id":"WMEvQMYn1g","forumContent":{"TLDR":{"value":"a parameter-efficient and unified framework for video editing tasks"},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["video editing; video generation; diffusion models"]},"primary_area":{"value":"generative models"},"abstract":{"value":"Recent advances in text-to-video generation have sparked interest in generative video editing tasks. Previous methods often rely on task-specific architectures (e.g., additional adapter modules) or dedicated customizations (e.g., DDIM inversion), which limit the integration of versatile editing conditions and the unification of various editing tasks. In this paper, we introduce UNified In-Context Video Editing (UNIC), a simple yet effective framework that unifies diverse video editing tasks within a single model in an in-context manner. To achieve this unification, we represent the inputs of various video editing tasks as three types of tokens: the source video tokens, the noisy video latent, and the multi-modal conditioning tokens that vary according to the specific editing task. Based on this formulation, our key insight is to integrate these three types into a single consecutive token sequence and jointly model them using the native attention operations of DiT, thereby eliminating the need for task-specific adapter designs. Nevertheless, direct task unification under this framework is challenging, leading to severe token collisions and task confusion due to the varying video lengths and diverse condition modalities across tasks. To address these, we introduce task-aware RoPE to facilitate consistent temporal positional encoding, and condition bias that enables the model to clearly differentiate different editing tasks. This allows our approach to adaptively perform different video editing tasks by referring the source video and varying condition tokens \"in context\", and support flexible task composition. To validate our method, we construct a unified video editing benchmark containing six representative video editing tasks. Results demonstrate that our unified approach achieves comparable performance with task specialists and exhibits emergent task composition abilities."},"_bibtex":{"value":"@inproceedings{\nye2026unified,\ntitle={Unified In-Context Video Editing},\nauthor={Zixuan Ye and Xuanhua He and Quande Liu and Qiulin Wang and Xintao Wang and Pengfei Wan and Di ZHANG and Kun Gai and Qifeng Chen and Wenhan Luo},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=Vb4nE3WWf5}\n}"},"title":{"value":"Unified In-Context Video Editing"},"pdf":{"value":"/pdf/b037cc220716c285a2c27112cbc052e5b1e7ef62.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"ye|unified_incontext_video_editing"},"authorids":{"value":["~Zixuan_Ye1","~Xuanhua_He1","~Quande_Liu1","~Qiulin_Wang1","~Xintao_Wang1","~Pengfei_Wan1","~Di_ZHANG3","~Kun_Gai1","~Qifeng_Chen1","~Wenhan_Luo1"]},"authors":{"value":["Zixuan Ye","Xuanhua He","Quande Liu","Qiulin Wang","Xintao Wang","Pengfei Wan","Di ZHANG","Kun Gai","Qifeng Chen","Wenhan Luo"]}},"version":2},{"content":{"summary":{"value":"This paper introduces SemCache, a training-free semantic-aware caching framework that accelerates video diffusion inference by adapting caching strategies to prompt semantics and motion dynamics. It combines Prompt Semantic-Aware (PSA) caching, which estimates scene complexity via cross-attention variations, and Temporal Motion Metric (TMM), which measures latent motion intensity to guide computation allocation."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"See above"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. This paper  uses prompt semantics and motion perception to guide caching in diffusion transformers, bridging high-level text understanding and low-level feature reuse.\n\n2. In the exeperiments part, it demonstrates  speedup and high-quality preservation across two large-scale video diffusion backbones and multiple metrics."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1.In the experimental section, several important details are missing, and the experiments do not adequately highlight or validate the main motivation of the paper (see detailed comments 1–3).\n\n2.The connection between cross-attention variation and semantic complexity, while intuitive, lacks a rigorous analytical foundation or sensitivity analysis.\n\n\nDetailed comments:\n\n1. In the experimental section, several details are unclear. For instance, for the FLOPs metric, it is not specified whether the reported values are in terabytes, gigabytes, or another unit. Additionally, key implementation details such as the generated video length and the prompt size are not provided.\n\n2. More importantly, the paper is motivated by the observation that prompts with different semantic complexities can influence caching behavior. However, in the experimental results, the authors do not explicitly evaluate performance(not only limit to latency) under prompts with diverse semantic dynamics (e.g., static vs. dynamic scenes). As a result, it is difficult to verify whether the proposed semantic-aware mechanism truly adapts to varying prompt semantics as claimed.\n\n\n3. Table 1 is not placed correctly in the paper.\n\n4. In the early example where an LLM is used to analyze prompt semantic scoring, it is unclear how the method determines which cache entries to retain, since all prompts are static. Would the model then keep all cached results at every step?"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762922365363,"tcdate":1761745719834,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission11218/Reviewer_6Jrv"],"signatures":["ICLR.cc/2026/Conference/Submission11218/Reviewer_6Jrv"],"forum":"WyfmWX2ncn","number":1,"license":"CC BY 4.0","cdate":1761745719834,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission11218/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762922365363,"domain":"ICLR.cc/2026/Conference","replyto":"WyfmWX2ncn","id":"amyQ1Ssz4p","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Video diffusion model acceleration ;Video generation ;Diffusion transformers"]},"primary_area":{"value":"generative models"},"abstract":{"value":"Diffusion models have achieved significant progress in video generation tasks, but slow inference speed remains a major challenge. Existing cache-based acceleration methods for video diffusion have demonstrated considerable improvements in inference speed. An existing efficient caching strategy involves reusing model outputs by estimating and leveraging the fluctuating differences among model outputs across timesteps. However, the strategy relies on extensive calibration sets and neglects the fact that prompt semantic variations affect the variational differences in model outputs. This phenomenon is illustrated by comparing videos generated from different prompts: while \"a horse running on the grassland\" produces highly dynamic content, \"a person reading in a coffee shop\" results in relatively static scenes.  Building on this observation, we propose a novel training-free SemCache method that can adaptively adjust caching strategies by perceiving prompt semantics changes. Key innovations include a Prompt Semantic-Aware (PSA) caching that evaluates prompt semantics and then dynamically decides a caching strategy tailored to the current timestep based on semantic information. We further introduce a Temporal Motion Metric (TMM) scheme to guide the compute allocation along the temporal dimension based on motion information, which not only ensures motion consistency in videos but also further reduces inference time. Experimental results demonstrate that SemCache achieves 2.45× and 2.66× speedups on HunyuanVideo and Wan2.1 respectively, while maintaining high video quality. Our code will be made publicly available."},"_bibtex":{"value":"@misc{\nwang2026semcacheadaptive,\ntitle={SemCache:Adaptive Semantic-Aware Caching for Efficient Video Diffusion},\nauthor={Zhenxue Wang and Yufan Liu and Fengjuan Wang and Congyan Lang and Wenyang Luo and Jinming Lou and Yuming Li},\nyear={2026},\nurl={https://openreview.net/forum?id=WyfmWX2ncn}\n}"},"title":{"value":"SemCache:Adaptive Semantic-Aware Caching for Efficient Video Diffusion"},"pdf":{"value":"/pdf/464449e979d3ddcded13b279ed4cfa9c140bac55.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"wang|semcacheadaptive_semanticaware_caching_for_efficient_video_diffusion"},"authorids":{"value":["~Zhenxue_Wang2","~Yufan_Liu1","~Fengjuan_Wang2","~Congyan_Lang2","~Wenyang_Luo1","~Jinming_Lou2","~Yuming_Li5"]},"authors":{"value":["Zhenxue Wang","Yufan Liu","Fengjuan Wang","Congyan Lang","Wenyang Luo","Jinming Lou","Yuming Li"]}},"version":2},{"content":{"summary":{"value":"The paper presents FOCUS (Frame-Optimistic Confidence Upper-bound Selection), a training-free and model-agnostic keyframe selection approach designed to improve long-video understanding in multimodal large language models (MLLMs). The authors formulate keyframe selection as a combinatorial pure-exploration problem within a multi-armed bandit framework, leveraging empirical means and Bernstein-style confidence bounds to balance exploration and exploitation. A two-stage coarse-to-fine procedure is further proposed to enable efficient and parallel computation. Experiments on LongVideoBench and Video-MME demonstrate consistent accuracy gains across multiple MLLMs. Overall, the paper provides an efficient and theoretically grounded solution to long-video understanding under tight token budget constraints."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"As mentioned in the Weaknesses section, I believe additional ablation studies would greatly help clarify the method’s behavior and limitations. In particular, it would be valuable to see how different choices of M (number of clips) or other hyperparameters affect both performance and efficiency.\nAdditionally, it would be interesting to explore how the proposed method performs on spatially focused video understanding tasks, where temporal redundancy is less dominant. Such experiments could provide insights into the generality and potential boundaries of the FOCUS framework."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The paper is original in formulating keyframe selection for long-video understanding as a combinatorial pure-exploration multi-armed bandit problem. This is a novel and reasonable perspective that provides new theoretical and algorithmic insights for researchers working on efficient video representation and token budgeting.\n2. The proposed two-stage coarse-to-fine procedure effectively addresses the non-parallelizable nature of sequential arm-pulling and updating, providing a practical solution that improves efficiency with minimal performance loss. This design is both elegant and empirically effective.\n3. The paper also shows strong theoretical grounding and empirical validation. The theoretical analysis is clear and complete, and the experiments are comprehensive, covering multiple benchmarks and models. The results align well with the theoretical claims, reinforcing the soundness and significance of the contribution."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. In Section 2.2, the paper assumes that “frame-level utility within the same arm share the same distribution.” It is unclear how this assumption is ensured in practice, especially regarding how the M non-overlapping fixed-length clips are partitioned. For instance, when the video contains shot changes or scene transitions, it is not clear how these are handled or whether the authors explored alternative segmentation strategies.\n2. The experiments are limited to LongVideoBench and Video-MME, both of which focus on long-form video QA. Evaluating the method on datasets with different characteristics—such as spatial reasoning benchmarks (e.g., VSI-Bench) or shorter video datasets—would provide a better understanding of the method’s generalizability and potential limitations.\n3. I am curious about how the number of clips (M) affects performance and efficiency. Since the method’s core formulation relies on partitioning videos into M fixed-length clips, an ablation study on M (and possibly related hyperparameters) would make the analysis more complete.\n4. There are a few minor typos (e.g., a missing period around line 025). The authors are encouraged to carefully proofread the paper to minimize such small errors."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762927849005,"tcdate":1762024612780,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission18064/Reviewer_Ccxf"],"signatures":["ICLR.cc/2026/Conference/Submission18064/Reviewer_Ccxf"],"forum":"1OQKqLFcbB","number":4,"license":"CC BY 4.0","cdate":1762024612780,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission18064/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762927849005,"domain":"ICLR.cc/2026/Conference","replyto":"1OQKqLFcbB","id":"hV6E2QZ0Gv","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Keyframe Selection","Multimodal large language models","Long Video Understanding","Combinatorial Pure-exploration"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Multimodal large language models (MLLMs) represent images and video frames as visual tokens. Scaling from single images to hour-long videos, however, inflates the token budget far beyond practical limits. Popular pipelines therefore either uniformly subsample or apply keyframe selection with retrieval-style scoring using smaller vision-language models. However, these keyframe selection methods still rely on pre-filtering before selection to reduce the inference cost and can miss the most informative moments.\n\nWe propose FOCUS, Frame-Optimistic Confidence Upper-bound Selection, a training-free, model-agnostic keyframe selection module that selects query-relevant frames under a strict token budget. FOCUS formulates keyframe selection as a combinatorial pure-exploration (CPE) problem in multi-armed bandits: it treats short temporal clips as arms, and uses empirical means and Bernstein confidence radius to identify informative regions while preserving exploration of uncertain areas. The resulting two-stage exploration-exploitation procedure reduces from a sequential policy with theoretical guarantees, first identifying high-value temporal regions, then selecting top-scoring frames within each region. Extensive experiments across four long-video question-answering benchmarks and four popular MLLMs demonstrate that FOCUS delivers substantial accuracy improvements while processing less than 2% of video frames. For videos longer than 20 minutes, it achieves an 11.9% gain in accuracy on LongVideoBench, demonstrating its effectiveness as a keyframe selection method and providing a simple and general solution for scalable long-video understanding with MLLMs."},"_bibtex":{"value":"@inproceedings{\nzhu2026focus,\ntitle={{FOCUS}: Efficient Keyframe Selection for Long Video Understanding},\nauthor={Zirui Zhu and Hailun Xu and Yang Luo and Yong Liu and Kanchan Sarkar and Zhenheng Yang and Yang You},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=1OQKqLFcbB}\n}"},"title":{"value":"FOCUS: Efficient Keyframe Selection for Long Video Understanding"},"pdf":{"value":"/pdf/1a4f41e9567a6de0b3b4cef826420a76cec302fb.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"zhu|focus_efficient_keyframe_selection_for_long_video_understanding"},"authorids":{"value":["~Zirui_Zhu2","~Hailun_Xu1","~Yang_Luo4","~Yong_Liu13","~Kanchan_Sarkar1","~Zhenheng_Yang3","~Yang_You1"]},"authors":{"value":["Zirui Zhu","Hailun Xu","Yang Luo","Yong Liu","Kanchan Sarkar","Zhenheng Yang","Yang You"]}},"version":2},{"content":{"venue":{"value":"CoRR 2024"},"pdf":{"value":"http://arxiv.org/pdf/2406.18944v3"},"venueid":{"value":"dblp.org/journals/CORR/2024"},"paperhash":{"value":"liu|investigating_and_defending_shortcut_learning_in_personalized_diffusion_models"},"authorids":{"value":["~Yixin_Liu4","~Ruoxi_Chen1","https://dblp.org/search/pid/api?q=author:Lichao_Sun_0001:"]},"html":{"value":"https://doi.org/10.48550/arXiv.2406.18944"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2406-18944,\n  publtype={informal},\n  author={Yixin Liu and Ruoxi Chen and Lichao Sun},\n  title={Investigating and Defending Shortcut Learning in Personalized Diffusion Models},\n  year={2024},\n  cdate={1704067200000},\n  journal={CoRR},\n  volume={abs/2406.18944},\n  url={https://doi.org/10.48550/arXiv.2406.18944}\n}\n"},"abstract":{"value":"Personalized diffusion models have gained popularity for adapting pre-trained text-to-image models to generate images of specific topics with minimal training data. However, these models are vulnerable to minor adversarial perturbations, leading to degraded performance on corrupted datasets. Such vulnerabilities are further exploited to craft protective perturbations on sensitive images like portraits that prevent unauthorized generation. In response, diffusion-based purification methods have been proposed to remove these perturbations and retain generation performance. However, existing works turn to over-purifying the images, which causes information loss. In this paper, we take a closer look at the fine-tuning process of personalized diffusion models through the lens of shortcut learning. And we propose a hypothesis explaining the manipulation mechanisms of existing perturbation methods, demonstrating that perturbed images significantly deviate from their original prompts in the CLIP-based latent space. This misalignment during fine-tuning causes models to associate noisy patterns with identifiers, resulting in performance degradation. Based on these insights, we introduce a systematic approach to maintain training performance through purification. Our method first purifies the images to realign them with their original semantic meanings in latent space. Then, we introduce contrastive learning with negative tokens to decouple the learning of clean identities from noisy patterns, which shows a strong potential capacity against adaptive perturbation. Our study uncovers shortcut learning vulnerabilities in personalized diffusion models and provides a firm evaluation framework for future protective perturbation research. Code is available at https://github.com/liuyixin-louis/DiffShortcut."},"title":{"value":"Investigating and Defending Shortcut Learning in Personalized Diffusion Models"},"authors":{"value":["Yixin Liu","Ruoxi Chen","Lichao Sun"]}},"tmdate":1739143885220,"pdate":1704067200000,"tcdate":1727491047653,"writers":["~"],"signatures":["~Ruoxi_Chen1"],"forum":"4XV3OSc3Sp","license":"CC BY-SA 4.0","number":97988,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1739143885220,"domain":"DBLP.org","id":"4XV3OSc3Sp","version":2},{"content":{"summary":{"value":"This paper proposed a new vision-language pre-training framework, called COSA. In particular, COSA augmented original image-text pairs by concatenating multiple examples as pseudo video-text pairs. Extensive experiments were conducted covering both video-language and image-language tasks, and demonstrated the effectiveness of the proposed method."},"presentation":{"value":"3 good"},"contribution":{"value":"2 fair"},"soundness":{"value":"2 fair"},"strengths":{"value":"- The paper is well written and easy to follow. In addition, the proposed method was supported by comprehensive experiments together with ablation studies, which made the paper a complete work.\n- The method COSA itself was simple yet effective to improve the learned representations for downstream tasks, and at the same time, it did not introduce extra computational costs."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The method was more like a trick of data augmentation instead of a significant technical contribution, as it just simply concatenated images and their corresponding captions and it was not very surprising to observe performance improvements.\n- As it was mentioned in the paper that apart from modified objectives, COSA also included original objectives for pre-training on image-text pairs. It was a complicated design to have so many training objectives and it was unclear how they were weighted (seemed to be equally weighted). Even though there was an ablation study of training objectives in Table 7, it still did not explain well the contributions of each item.\n- The method leveraged the average pooled [CLS] token for each image as the final representation for the pseudo video. In this way, there was actually no temporal information considered. And the selected downstream tasks were less dependent on temporal information in the meanwhile. It would be better if tasks such as temporal action localization were included to show whether COSA can improve those tasks. In addition, since temporal information did not play any role in current method, I am afraid that using augmentations like mixup for videos/images might lead to similar performance gain, as shown in [1].\n- Previous works showed that using CLIP initialization could lead to better performance. Among compared baseline methods, some of them such as MILES [2] actually used ViT trained on ImageNet for image classification and it was not a fair comparison to COSA with CLIP initialization.\n\n[1] Hao, Xiaoshuai, et al. \"Mixgen: A new multi-modal data augmentation.\" Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. 2023.\n\n[2] Ge, Yuying, et al. \"Miles: Visual bert pre-training with injected language semantics for video-text retrieval.\" European Conference on Computer Vision. Cham: Springer Nature Switzerland, 2022."},"confidence":{"value":"3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked."},"questions":{"value":"- Is it possible for the authors to include tasks which rely much on temporal information like temporal action localization? This would provide better understanding of the proposed method.\n- It would be better if results with different initializations can be presented to remove my concern about better CLIP initialization.\n- It was worth trying data augmentations like mixup and it might lead to similar performance gain as demonstrated in the paper."},"rating":{"value":"6: marginally above the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1700678123487,"tcdate":1698552787035,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission5044/Reviewer_gWLj"],"signatures":["ICLR.cc/2024/Conference/Submission5044/Reviewer_gWLj"],"forum":"bDkisS75zy","number":2,"license":"CC BY 4.0","cdate":1698552787035,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission5044/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1700678123487,"domain":"ICLR.cc/2024/Conference","replyto":"bDkisS75zy","id":"jyfSb5qPmA","forumContent":{"venue":{"value":"ICLR 2024 poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Vision-Language Foundation Model","Video-Language Pretraining"]},"primary_area":{"value":"representation learning for computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Due to the limited scale and quality of video-text training corpus, most  vision-language  foundation  models employ  image-text datasets for pretraining and primarily focus on modeling visually semantic representations while disregarding temporal semantic representations and correlations. To address this issue, we propose COSA, a COncatenated SAmple pretrained vision-language foundation model. COSA can jointly model visual contents and event-level temporal cues using only image-text corpora.  We achieve this by sequentially concatenating multiple image-text pairs as inputs for pretraining. This transformation effectively converts existing image-text corpora into a pseudo  video-paragraph corpus, enabling richer scene transformations and explicit event-description correspondence. Extensive experiments demonstrate that COSA consistently improves performance across a broad range of semantic vision-language downstream tasks, including paragraph-to-video retrieval, text-to-video/image retrieval, video/image captioning and video QA. Notably, COSA achieves state-of-the-art results on various competitive benchmarks. Code and model are released at https://github.com/TXH-mercury/COSA."},"_bibtex":{"value":"@inproceedings{\nchen2024cosa,\ntitle={{COSA}: Concatenated Sample Pretrained Vision-Language Foundation Model},\nauthor={Sihan Chen and Xingjian He and Handong Li and Xiaojie Jin and Jiashi Feng and Jing Liu},\nbooktitle={The Twelfth International Conference on Learning Representations},\nyear={2024},\nurl={https://openreview.net/forum?id=bDkisS75zy}\n}"},"title":{"value":"COSA: Concatenated Sample Pretrained Vision-Language Foundation Model"},"pdf":{"value":"/pdf/3fbcf031f5ace68ec8d5ec3b8e59609d4f43d942.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference"},"paperhash":{"value":"chen|cosa_concatenated_sample_pretrained_visionlanguage_foundation_model"},"authorids":{"value":["~Sihan_Chen3","~Xingjian_He1","~Handong_Li1","~Xiaojie_Jin1","~Jiashi_Feng1","~Jing_Liu1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Sihan Chen","Xingjian He","Handong Li","Xiaojie Jin","Jiashi Feng","Jing Liu"]}},"version":2},{"content":{"summary":{"value":"The paper introduces **KVComm**, a novel communication protocol for multi-agent systems based on Large Language Models (LLMs). Instead of relying on natural language or hidden states, KVComm selectively shares key-value (KV) pairs across models using a strategy guided by attention importance scores and a Gaussian prior. The method significantly reduces communication and computation overhead while maintaining or surpassing task performance on diverse benchmarks."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. **How transferable is the KVComm strategy across model families?** The paper evaluates KVComm mostly on pairs of identical or fine-tuned models. How would this framework perform if the sender and receiver are *structurally different models*, e.g., GPT-style vs. Mistral-style? Would attention alignment or KV dimensions cause integration issues?\n\n2. **Can the selection of layers be made context-adaptive instead of fixed?**\n   The Gaussian prior is fixed during calibration. Could performance improve if KV layer selection is dynamically adjusted based on the current query or context (e.g., using entropy or other uncertainty measures)?\n\n3. **Why use a Gaussian prior for layer selection?**\n   While intuitive and empirically justified, the Gaussian prior assumes a unimodal distribution over informative layers. Have other priors or non-parametric approaches (e.g., entropy-weighted or data-driven learned selection) been considered?\n\n4. **Does the integration of KV pairs affect positional encoding or cause mismatch issues?**\n   Since KV pairs include positional biases (especially in rotary or absolute encodings), how does the concatenation affect the receiver’s decoding stability and accuracy?\n\n5. **What is the impact on latency and memory footprint during runtime?**\n   The paper discusses FLOPs and communication savings, but does the selective KV sharing induce memory overhead (e.g., extra caching or tensor reshaping), particularly for large batch sizes or streaming inference?"},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"S1: Using KV pairs rather than natural language or hidden states is a compelling, less explored direction that avoids known bottlenecks like information loss and concentration bias.\n\nS2: The use of attention importance with a Gaussian prior to guide layer selection is well-motivated and supported by empirical validation.\n\nS3: The authors benchmark across multiple datasets and LLM pairs, including ablations and comparisons with strong baselines like Skyline and AC.\n\nS4: The approach offers substantial computation and communication savings (2.5x–6x), making it relevant for scalable deployment scenarios.\n\nS5: The paper provides detailed setup information, dataset samples, and makes its code and synthetic datasets available."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"To be honest, I carefully read this paper but I have not found some serious weakness. Below are just some suggestions that may lead to a more solid work.\n\n\n\nSuggest 1: The method is primarily validated on same or fine-tuned model pairs. Scalability to truly heterogeneous models remains unexplored.\n\nSuggest 2: The selection strategy is static post-calibration. Dynamic or context-aware layer selection could further enhance flexibility.\n\nSuggest 3: Several datasets are synthetically created or downsampled, which might not fully reflect real-world task complexity or communication needs."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762919847450,"tcdate":1761977924263,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission7805/Reviewer_1XmS"],"signatures":["ICLR.cc/2026/Conference/Submission7805/Reviewer_1XmS"],"forum":"F7rUng23nw","number":3,"license":"CC BY 4.0","cdate":1761977924263,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission7805/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762919847450,"domain":"ICLR.cc/2026/Conference","replyto":"F7rUng23nw","id":"6aesngmAnL","forumContent":{"TLDR":{"value":"We propose KVComm, a communication framework that enables efficient inter-LLM collaboration by selectively sharing key-value pairs, achieving near upper-bound performance with significantly reduced communication cost."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Large Language Models","Multi-Agent Systems","Inter-LLM Communication","Multi-agent Debate"]},"supplementary_material":{"value":"/attachment/ad7387bbb5fb018dbb30197b19d6cfa84a168443.zip"},"primary_area":{"value":"foundation or frontier models, including LLMs"},"abstract":{"value":"Large Language Models (LLMs) are increasingly deployed in multi-agent systems, where effective inter-model communication is crucial. Existing communication protocols either rely on natural language, incurring high inference costs and information loss, or on hidden states, which suffer from information concentration bias and inefficiency. To address these limitations, we propose KVComm, a novel communication framework that enables efficient communication between LLMs through selective sharing of KV pairs. KVComm leverages the rich information encoded in the KV pairs while avoiding the pitfalls of hidden states. We introduce a KV layer-wise selection strategy based on attention importance scores with a Gaussian prior to identify the most informative KV pairs for communication. Extensive experiments across diverse tasks and model pairs demonstrate that KVComm achieves comparable performance to the upper-bound method, which directly merges inputs to one model without any communication, while transmitting as few as 30\\% of layers' KV pairs. Our study highlights the potential of KV pairs as an effective medium for inter-LLM communication, paving the way for scalable and efficient multi-agent systems."},"_bibtex":{"value":"@inproceedings{\nshi2026kvcomm,\ntitle={{KVC}omm: Enabling Efficient {LLM} Communication through Selective {KV} Sharing},\nauthor={Xiangyu Shi and Marco Chiesa and Gerald Q. Maguire Jr. and Dejan Kostic},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=F7rUng23nw}\n}"},"title":{"value":"KVComm: Enabling Efficient LLM Communication through Selective KV Sharing"},"pdf":{"value":"/pdf/fb0426d7900d0d9d2c2252532541eb292b23765a.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"shi|kvcomm_enabling_efficient_llm_communication_through_selective_kv_sharing"},"authorids":{"value":["~Xiangyu_Shi2","~Marco_Chiesa1","~Gerald_Q._Maguire_Jr.1","~Dejan_Kostic1"]},"authors":{"value":["Xiangyu Shi","Marco Chiesa","Gerald Q. Maguire Jr.","Dejan Kostic"]}},"version":2},{"content":{"venue":{"value":"ACL (Findings) 2025"},"pdf":{"value":"https://aclanthology.org/2025.findings-acl.1228.pdf"},"venueid":{"value":"dblp.org/conf/ACL/2025"},"paperhash":{"value":"li|imove_instancemotionaware_video_understanding"},"authorids":{"value":["~Jiaze_Li3","https://dblp.org/search/pid/api?q=author:Yaya_Shi:","~Zongyang_Ma3","~Haoran_Xu8","https://dblp.org/search/pid/api?q=author:Yandong_Bai:","https://dblp.org/search/pid/api?q=author:Huihui_Xiao:","https://dblp.org/search/pid/api?q=author:Ruiwen_Kang:","https://dblp.org/search/pid/api?q=author:Fan_Yang_0094:","https://dblp.org/search/pid/api?q=author:Tingting_Gao:","https://dblp.org/search/pid/api?q=author:Di_Zhang_0026:"]},"html":{"value":"https://aclanthology.org/2025.findings-acl.1228/"},"_bibtex":{"value":"@inproceedings{DBLP:conf/acl/LiSMXBXKYGZ25,\n  author={Jiaze Li and Yaya Shi and Zongyang Ma and Haoran Xu and Yandong Bai and Huihui Xiao and Ruiwen Kang and Fan Yang and Tingting Gao and Di Zhang},\n  title={iMOVE : Instance-Motion-Aware Video Understanding},\n  year={2025},\n  cdate={1735689600000},\n  pages={23959-23975},\n  url={https://aclanthology.org/2025.findings-acl.1228/},\n  booktitle={ACL (Findings)},\n  crossref={conf/acl/2025f}\n}\n"},"abstract":{"value":"Enhancing the fine-grained instance spatiotemporal motion perception capabilities of Video Large Language Models is crucial for improving their temporal and general video understanding. However, current models struggle to perceive detailed and complex instance motions. To address these challenges, we have made improvements from both data and model perspectives. In terms of data, we have meticulously curated iMOVE-IT, the first large-scale instance-motion-aware video instruction-tuning dataset. This dataset is enriched with comprehensive instance motion annotations and spatiotemporal mutual-supervision tasks, providing extensive training for the model’s instance-motion-awareness. Building on this foundation, we introduce iMOVE, an instance-motion-aware video foundation model that utilizes Event-aware Spatiotemporal Efficient Modeling to retain informative instance spatiotemporal motion details while maintaining computational efficiency. It also incorporates Relative Spatiotemporal Position Tokens to ensure awareness of instance spatiotemporal positions. Evaluations indicate that iMOVE excels not only in video temporal understanding and general video understanding but also demonstrates significant advantages in long-term video understanding. We will release the data, code, and model weights after acceptance."},"title":{"value":"iMOVE : Instance-Motion-Aware Video Understanding"},"authors":{"value":["Jiaze Li","Yaya Shi","Zongyang Ma","Haoran Xu","Yandong Bai","Huihui Xiao","Ruiwen Kang","Fan Yang","Tingting Gao","Di Zhang"]}},"tmdate":1775647307136,"pdate":1735689600000,"externalIds":["dblp:conf/acl/LiSMXBXKYGZ25"],"tcdate":1762668955524,"writers":["~"],"signatures":["~Zongyang_Ma3"],"forum":"LLmRsQWVto","license":"CC BY-SA 4.0","number":680078,"cdate":1735689600000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1775647307136,"domain":"DBLP.org","id":"LLmRsQWVto","version":2},{"content":{"summary":{"value":"This paper introduces Chimera, a benchmark for diagnosing shortcut learning in vision-language models on diagram understanding tasks. Using minimal pairs, it reveals models' overreliance on visual or structural cues over semantic understanding and proposes prompting strategies to mitigate such behavior."},"ethical_considerations":{"value":"No, there are no or only very minor ethics concerns"},"dataset_code_accessibility":{"value":"Yes"},"responsible_reviewing_acknowledgement":{"value":"Yes"},"code_of_conduct_acknowledgement":{"value":"Yes"},"confidence":{"value":4},"rating":{"value":5},"final_justification":{"value":"The authors have addressed my concern and promised to include related works in their revised version of the main paper."},"limitations_weaknesses":{"value":"While the supplementary material includes a list of related work, the main paper lacks a dedicated Related Work section. This makes it difficult for readers to contextualize the Chimera benchmark within prior research on shortcut behaviors and LLM evaluation. Including a concise but informative discussion of related benchmarks and prior shortcut analyses in the main text would significantly enhance clarity and strengthen the paper’s positioning in the literature."},"strengths_contributions":{"value":"- The paper is, to the best of my knowledge, the first to systematically investigate shortcut learning in diagram-based reasoning tasks.\n- While shortcut behaviors have been widely studied in language-only settings (e.g., NLI, QA), this work expands the scope to multimodal reasoning.\n- The benchmark design isolates subtle differences in diagram inputs to reveal when models overfit to non-semantic cues. The analysis shows that even strong models (e.g., GPT-4V, Claude) can fail to generalize in structured visual contexts."}},"parentInvitations":"NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/-/Official_Review","nonreaders":[],"tmdate":1761794806882,"tcdate":1751492890894,"writers":["NeurIPS.cc/2025/Datasets_and_Benchmarks_Track","NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/Submission1798/Reviewer_JmgV"],"signatures":["NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/Submission1798/Reviewer_JmgV"],"forum":"dYjSDFqwUT","number":2,"license":"CC BY 4.0","cdate":1751492890894,"readers":["everyone"],"invitations":["NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/Submission1798/-/Official_Review","NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/-/Edit","NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/Submission1798/Official_Review2/-/Review_Revision"],"mdate":1761794806882,"domain":"NeurIPS.cc/2025/Datasets_and_Benchmarks_Track","replyto":"dYjSDFqwUT","id":"KisOLHBkB9","forumContent":{"venue":{"value":"Submitted to NeurIPS 2025 Datasets and Benchmarks Track"},"keywords":{"value":["Multimodal Benchmark","Diagram Comprehension","Visual Reasoning","Evaluation","Vision-Language Models"]},"supplementary_material":{"value":"/attachment/49d450962128197465311b481c38179c1973ee41.zip"},"primary_area":{"value":"datasets_&_benchmarks_for_language"},"abstract":{"value":"Diagrams convey symbolic information in a visual format rather than a linear stream of words, making them especially challenging for AI models to process. While recent evaluations suggest that vision-language models (VLMs) perform well on diagram-related benchmarks, their reliance on knowledge, reasoning, or modality shortcuts raises concerns about whether they genuinely understand and reason over diagrams. To address this gap, we introduce Chimera, a comprehensive benchmark comprising 7,500 high-quality diagrams sourced from Wikipedia; each diagram is annotated with its symbolic content represented by semantic triples along with multi-level questions designed to assess four fundamental aspects of diagram comprehension: entity recognition, relation understanding, knowledge grounding, and visual reasoning. We use Chimera to measure the presence of three types of shortcuts in visual question answering: \n(1) the visual-memorization shortcut, where VLMs rely on memorized visual patterns;\n(2) the knowledge-recall shortcut, where models leverage memorized factual knowledge instead of interpreting the diagram; and\n(3) the Clever-Hans shortcut, where models exploit superficial language patterns or priors without true comprehension.\nWe evaluate 15 open-source VLMs from 7 model families on Chimera and find that their seemingly strong performance largely stems from shortcut behaviors -- visual-memorization shortcuts have slight impact, knowledge-recall shortcuts play a moderate role, and Clever-Hans shortcuts contribute significantly. These findings expose critical limitations in current VLMs and underscore the need for more robust evaluation protocols that benchmark genuine comprehension of complex visual inputs (e.g., diagrams) rather than question-answer shortcuts."},"_bibtex":{"value":"@misc{\nchi2026chimera,\ntitle={Chimera: Revisiting Shortcut Learning by Vision-Language Models for Diagram Understanding},\nauthor={Ziheng Chi and Yifan Hou and Chenxi Pang and Shaobo Cui and Mubashara Akhtar and Mrinmaya Sachan},\nyear={2026},\nurl={https://openreview.net/forum?id=dYjSDFqwUT}\n}"},"title":{"value":"Chimera: Revisiting Shortcut Learning by Vision-Language Models for Diagram Understanding"},"pdf":{"value":"/pdf/ea4f2cd540be7aecb6dbf4d21f66997205f9358f.pdf"},"croissant_file":{"value":"/attachment/e89b2f6fb03ead7f57be96ccf170c35e416e4c9e.json"},"code_URL":{"value":"https://anonymous.4open.science/r/CHIMERA-Anonymous-4D73/README.md"},"venueid":{"value":"NeurIPS.cc/2025/Datasets_and_Benchmarks_Track/Rejected_Submission"},"paperhash":{"value":"chi|chimera_revisiting_shortcut_learning_by_visionlanguage_models_for_diagram_understanding"},"authorids":{"value":["~Ziheng_Chi1","~Yifan_Hou1","~Chenxi_Pang1","~Shaobo_Cui1","~Mubashara_Akhtar1","~Mrinmaya_Sachan3"]},"dataset_URL":{"value":"https://huggingface.co/datasets/hfprivatehf/CHIMERA-Anonymous"},"authors":{"value":["Ziheng Chi","Yifan Hou","Chenxi Pang","Shaobo Cui","Mubashara Akhtar","Mrinmaya Sachan"]}},"version":2},{"content":{"summary":{"value":"The authors propose a method for adding controllable Bokeh in videos. They fine tune a video diffusion model to accept a video and a corresponding explicit scene geometry conditioning and output a depth-aware Bokeh video in a single diffusion step. The authors propose a multi-stage training strategy that facilitates robustness to noisy input geometry and high temporal bokeh consistency. The authors provide extensive evaluations showing state-of-the-art results for controllable video bokeh, with control of the focus plane and bokeh intensity."},"soundness":{"value":4},"confidence":{"value":3},"questions":{"value":"Please see weaknesses section."},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":3},"strengths":{"value":"1.\tThe authors leverage the prior of a video diffusion model in a novel way, to perform temporally consistent, controllable video bokeh. \n2.\tThe authors showcase good bokeh results. They also provide a supplementary video with their results and comparisons to other methods, which is very important for the qualitative assessment of their claims for temporally consistent bokeh addition.\n3.\tThe authors show extensive quantitative evaluations, emphasizing their lead over other competing methods.\n4.\tThe authors provide several ablations for their training strategy choices."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Major:\n1.\tThe authors do not provide limitations for their method. Are there any scenarios where the model fails to generate a good bokeh video? Maybe in videos with fast motion, such as a car race.\nMinor:\n1.\tFigure 1 is not referenced.\n2.\tSM figures 10,11 – red border not corresponding to zoom-in area."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762918099615,"tcdate":1761667403918,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission5506/Reviewer_7dEg"],"signatures":["ICLR.cc/2026/Conference/Submission5506/Reviewer_7dEg"],"forum":"h05AulYT7g","number":2,"license":"CC BY 4.0","cdate":1761667403918,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission5506/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762918099615,"domain":"ICLR.cc/2026/Conference","replyto":"h05AulYT7g","id":"SudpaIq5c9","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Computational photography"]},"supplementary_material":{"value":"/attachment/c7336d4080c97fd4256939cb9c3cefb00df1817d.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Diffusion models have recently emerged as powerful tools for camera simulation, enabling both geometric transformations and realistic optical effects. Among these, image-based bokeh rendering has shown promising results, but diffusion for video bokeh remains unexplored. Existing image-based methods are plagued by temporal flickering and inconsistent blur transitions, while current video editing methods lack explicit control over the focus plane and bokeh intensity. These issues limit their applicability for controllable video bokeh. In this work, we propose a one-step diffusion framework for generating temporally coherent, depth-aware video bokeh rendering. The framework employs a multi-plane image (MPI) representation adapted to the focal plane to condition the video diffusion model, thereby enabling it to exploit strong 3D priors from pretrained backbones. To further enhance temporal stability, depth robustness, and detail preservation, we introduce a progressive training strategy. Experiments on synthetic and real-world benchmarks demonstrate superior temporal coherence, spatial accuracy, and controllability, outperforming prior baselines. This work represents the first dedicated diffusion framework for video bokeh generation, establishing a new baseline for temporally coherent and controllable depth-of-field effects. Project page is available at this website https://vivocameraresearch.github.io/any2bokeh/."},"_bibtex":{"value":"@inproceedings{\nyang2026anytobokeh,\ntitle={Any-to-Bokeh: Arbitrary-Subject Video Refocusing with Video Diffusion Model},\nauthor={Yang Yang and Siming Zheng and Qirui Yang and Jinwei Chen and Boxi Wu and Xiaofei He and Deng Cai and Bo Li and Peng-Tao Jiang},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=h05AulYT7g}\n}"},"title":{"value":"Any-to-Bokeh: Arbitrary-Subject Video Refocusing with Video Diffusion Model"},"pdf":{"value":"/pdf/b8270462538ddaf5a4463edac51557409651493c.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"yang|anytobokeh_arbitrarysubject_video_refocusing_with_video_diffusion_model"},"authorids":{"value":["~Yang_Yang41","~Siming_Zheng3","~Qirui_Yang1","~Jinwei_Chen3","~Boxi_Wu1","~Xiaofei_He2","~Deng_Cai4","~Bo_Li20","~Peng-Tao_Jiang1"]},"authors":{"value":["Yang Yang","Siming Zheng","Qirui Yang","Jinwei Chen","Boxi Wu","Xiaofei He","Deng Cai","Bo Li","Peng-Tao Jiang"]}},"version":2},{"content":{"venue":{"value":"ICML 2024 Poster"},"abstract":{"value":"Graph Neural Networks (GNNs) have demonstrated remarkable performance in graph classification tasks. However, ensuring the explainability of their predictions remains a challenge. To address this, graph rationalization methods have been introduced to generate concise subsets of the original graph, known as rationales, which serve to explain the predictions made by GNNs. Existing rationalizations often rely on shortcuts in data for prediction and rationale composition. In response, de-shortcut rationalization methods have been proposed, which commonly leverage counterfactual augmentation to enhance data diversity for mitigating the shortcut problem. Nevertheless, these methods have predominantly focused on centralized datasets and have not been extensively explored in the Federated Learning (FL) scenarios. To this end, in this paper, we propose a Federated Graph Rationalization (FedGR) with anti-shortcut augmentations to achieve self-explaining GNNs, which involves two data augmenters. These augmenters are employed to produce client-specific shortcut conflicted samples at each client, which contributes to mitigating the shortcut problem under the FL scenarios. Experiments on real-world benchmarks and synthetic datasets validate the effectiveness of FedGR under the FL scenarios."},"_bibtex":{"value":"@inproceedings{\nyue2024federated,\ntitle={Federated Self-Explaining {GNN}s with Anti-shortcut Augmentations},\nauthor={Linan Yue and Qi Liu and Weibo Gao and Ye Liu and Kai Zhang and Yichao Du and Li Wang and Fangzhou Yao},\nbooktitle={Forty-first International Conference on Machine Learning},\nyear={2024},\nurl={https://openreview.net/forum?id=ZxDqSBgFSM}\n}"},"title":{"value":"Federated Self-Explaining GNNs with Anti-shortcut Augmentations"},"pdf":{"value":"/pdf/04585205e977b2e131373b686eda50f707f2586b.pdf"},"venueid":{"value":"ICML.cc/2024/Conference"},"paperhash":{"value":"yue|federated_selfexplaining_gnns_with_antishortcut_augmentations"},"authorids":{"value":["~Linan_Yue1","~Qi_Liu3","~Weibo_Gao1","~Ye_Liu10","~Kai_Zhang12","~Yichao_Du1","~Li_Wang18","~Fangzhou_Yao1"]},"authors":{"value":["Linan Yue","Qi Liu","Weibo Gao","Ye Liu","Kai Zhang","Yichao Du","Li Wang","Fangzhou Yao"]}},"tmdate":1719287219732,"pdate":1714610266970,"tcdate":1705303446505,"writers":["ICML.cc/2024/Conference","ICML.cc/2024/Conference/Submission433/Authors"],"signatures":["ICML.cc/2024/Conference/Submission433/Authors"],"forum":"ZxDqSBgFSM","license":"CC BY 4.0","number":433,"cdate":1705303446505,"readers":["everyone"],"invitations":["ICML.cc/2024/Conference/-/Submission","ICML.cc/2024/Conference/-/Post_Submission","ICML.cc/2024/Conference/-/Edit","ICML.cc/2024/Conference/Submission433/-/Camera_Ready_Revision"],"mdate":1719287219732,"odate":1717693027359,"domain":"ICML.cc/2024/Conference","id":"ZxDqSBgFSM","version":2},{"content":{"venue":{"value":"ICDIM 2008"},"pdf":{"value":"https://ieeexplore.ieee.org/iel5/4733742/4746691/04746772.pdf"},"venueid":{"value":"dblp.org/conf/ICDIM/2008"},"paperhash":{"value":"choi|mire_a_minimal_rule_engine_for_contextaware_mobile_devices"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Changbai_Choi:","https://dblp.org/search/pid/api?q=author:Insuk_Park:","https://dblp.org/search/pid/api?q=author:Soon_J._Hyun:","~Dongman_Lee1","https://dblp.org/search/pid/api?q=author:David_Hyun_Sim:"]},"html":{"value":"https://doi.org/10.1109/ICDIM.2008.4746772"},"_bibtex":{"value":"@inproceedings{DBLP:conf/icdim/ChoiPHLS08,\n  author={Changbai Choi and Insuk Park and Soon J. Hyun and Dongman Lee and David Hyun Sim},\n  title={MiRE: A Minimal Rule Engine for context-aware mobile devices},\n  year={2008},\n  cdate={1199145600000},\n  pages={172-177},\n  url={https://doi.org/10.1109/ICDIM.2008.4746772},\n  booktitle={ICDIM},\n  crossref={conf/icdim/2008}\n}\n"},"abstract":{"value":"Context-aware services have been introduced into mobile devices, such as cellular phones. Context management and rule processing are the key research issues for the resource-limited devices to have context-aware capability. In this paper, we introduce a context-aware rule processing engine, called MiRE (minimal rule engine). The engine implements the minimum cores of the conventional rule processing facilities and some special policies to achieve resource-saving and light-weight inference engine suitable for resource-limited mobile devices. MiRE, as a part of context-aware middleware system, provides a flexible architecture to adopt various context-aware applications while rules and contexts are dynamically registered at run-time. It maintains the amount of context (i.e., facts) and rules to be optimal in order to save computing resources and processing time. Its design strategy allows MiRE to be conveniently equipped into resource-limited context-aware mobile devices. The performance evaluation carried out on a context-aware cellular phone shows the intended efficiency has been well achieved."},"title":{"value":"MiRE: A Minimal Rule Engine for context-aware mobile devices"},"authors":{"value":["Changbai Choi","Insuk Park","Soon J. Hyun","Dongman Lee","David Hyun Sim"]}},"tmdate":1768385223810,"pdate":1199145600000,"externalIds":["dblp:conf/icdim/ChoiPHLS08"],"tcdate":1768385170749,"writers":["~"],"signatures":["~Dongman_Lee1"],"forum":"muy6uCMBbo","license":"CC BY-SA 4.0","number":751104,"cdate":1199145600000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1768385223810,"domain":"DBLP.org","id":"muy6uCMBbo","version":2},{"content":{"summary":{"value":"The paper introduces AccidentBench, a large-scale video QA benchmark for multimodal understanding and reasoning, specifically in safety‑critical settings. The benchmark focuses on vehicle accidents (83%) and extends also to airspace (10.2%) and waterway (6.8%) scenarios, with more than 2,000 real-world videos and 19,000 human‑annotated multiple-choice Q&A pairs spanning temporal, spatial, and intent/goal reasoning. The dataset is split into three difficulties category ( easy/medium/hard), enabling systematic evaluation with controlled increases in precision requirements and task complexity. The paper evaluates a broad range of SOTA models, presents detailed performance breakdowns by difficulty, reasoning type, and video length, and includes qualitative error analyses."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. The author should provide the proportion or breakdown of QAs that explicitly address accidents, violations, or hazard-related reasoning versus generic spatiotemporal questions.\n\n\n2. Could the authors report random-guess or chance-normalized baselines for each difficulty level?\nSince the number of answer options varies across settings, a chance baseline would help interpret how far each model is from random performance, and make difficulty comparisons fairer."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":3},"strengths":{"value":"1. Unique focus on Safety-critical Scenarios: AccidentBench distinctively centers on accident and safety-critical environments, integrating land, air, and water domains into a unified benchmark. This cross-domain safety emphasis is novel relative to prior driving-centric (e.g., DriveLM, DriveBench) or general video QA benchmarks (e.g., MVBench, LongVideoBench) that lack such a unified safety context. The emphasis on vehicle accidents addresses a particularly important real-world application domain for autonomous systems.\n\n2. Strong Dataset Scale and Diagnostic Design: The benchmark is large-scale and comprehensive (~2 k videos and ~19 k QA pairs) with diverse weather, viewpoints, and dynamic contexts across domains. Its controlled difficulty levels and detailed breakdowns (by domain, difficulty, reasoning type, and video length) enable fine-grained diagnostic evaluation of model reasoning weaknesses. The annotation quality is ensured through human annotation by highly educated annotators. The qualitative failure analyses further highlight concrete reasoning gaps across spatial, temporal, and intent understanding\n\n3. Accessibility: The dataset, code, and evaluation scripts are publicly released via the project site, ensuring immediate accessibility and fostering reproducibility and future comparison.\n\n4. Comprehensive Model Evaluation: The paper evaluates an extensive range of both proprietary models (GPT-5, GPT-4o, Gemini 2.5 Pro) and open-source models (InternVL, LLaVA, Qwen), providing a thorough empirical assessment."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Safety-Critical Claim Not Well Quantified: While the benchmark is framed as the first safety-critical multimodal benchmark, many example questions (e.g., in Fig. 1) assess generic spatial or temporal reasoning (e.g., \"How many boats are observed in this video?\" or directional positioning queries) rather than explicit accident causality, hazard identification, or safety violations. The paper does not quantify what proportion of QAs explicitly involve accident-related reasoning, causal safety analysis, or hazard assessment versus general spatiotemporal understanding. A clearer breakdown distinguishing safety-specific reasoning questions from general perception and reasoning tasks would better substantiate the benchmark's core contribution and differentiate it from existing video understanding benchmarks.\n\n2. Potential Sampling Bias in Evaluation (Table 7): In Table 7, the sampled subset consistently achieves higher accuracy than the full dataset. While the paper states this validates the sampling strategy, no statistical significance tests, confidence intervals, or variance analyses are provided to explain why a supposedly random sample would systematically perform better. This raises concerns about potential sampling bias, task distribution imbalances, or selection effects.\n\n\n3. Limited Clarity on Annotation Protocol and Reliability: Although annotators are described as “highly educated,” the paper provides no details on annotation guidelines, inter-annotator agreement rate, or validation steps for complex intent/goal questions. These details are crucial for ensuring label consistency in a benchmark that emphasizes reasoning quality"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762924066810,"tcdate":1761966226333,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission13442/Reviewer_ie21"],"signatures":["ICLR.cc/2026/Conference/Submission13442/Reviewer_ie21"],"forum":"f5lIozG83H","number":4,"license":"CC BY 4.0","cdate":1761966226333,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission13442/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762924066810,"domain":"ICLR.cc/2026/Conference","replyto":"f5lIozG83H","id":"zlwr6p74OF","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Multimodal Understanding and Reasoning","Large-Scale Dataset","Traffic Accident","Land Space","Airplane Navigation","Ship Motion"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Rapid advances in multimodal models demand benchmarks that rigorously evaluate understanding and reasoning in safety-critical, dynamic real-world settings. We present AccidentBench, a large-scale benchmark that combines vehicle accident scenarios with Beyond domains, safety-critical settings in air and water that emphasize spatial and temporal reasoning (e.g., navigation, orientation, multi-vehicle motion). The benchmark contains approximately 2000 videos and over 19000 human-annotated question--answer pairs spanning multiple video lengths (short/medium/long) and difficulty levels (easy/medium/hard). Tasks systematically probe core capabilities: temporal, spatial, and intent understanding and reasoning.  By unifying accident-centric traffic scenes with broader safety-critical scenarios in air and water, AccidentBench offers a comprehensive, physically grounded testbed for evaluating models under real-world variability. Evaluations of state-of-the-art models (e.g., Gemini-2.5 Pro and GPT-5) show that even the strongest models achieve only about 18% accuracy on the hardest tasks and longest videos, revealing substantial gaps in real-world temporal, spatial, and intent reasoning. AccidentBench is designed to expose these critical gaps and drive the development of multimodal models that are safer, more robust, and better aligned with real-world safety-critical challenges. The code and dataset are available at: http://accident-bench.site"},"_bibtex":{"value":"@misc{\ngu2026accidentbench,\ntitle={AccidentBench: Benchmarking Multimodal Understanding and Reasoning in Vehicle Accidents and Beyond},\nauthor={Shangding Gu and Xiaohan Wang and Donghao Ying and Haoyu Zhao and Runing Yang and Ming Jin and Boyi Li and Marco Pavone and Serena Yeung-Levy and Jun Wang and Dawn Song and Costas Spanos},\nyear={2026},\nurl={https://openreview.net/forum?id=f5lIozG83H}\n}"},"title":{"value":"AccidentBench: Benchmarking Multimodal Understanding and Reasoning in Vehicle Accidents and Beyond"},"pdf":{"value":"/pdf/f92d19d32838720da5d3cb35e4ad61301871e9ba.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"gu|accidentbench_benchmarking_multimodal_understanding_and_reasoning_in_vehicle_accidents_and_beyond"},"authorids":{"value":["~Shangding_Gu1","~Xiaohan_Wang2","~Donghao_Ying1","~Haoyu_Zhao3","~Runing_Yang1","~Ming_Jin2","~Boyi_Li1","~Marco_Pavone1","~Serena_Yeung-Levy1","~Jun_Wang2","~Dawn_Song1","~Costas_Spanos1"]},"authors":{"value":["Shangding Gu","Xiaohan Wang","Donghao Ying","Haoyu Zhao","Runing Yang","Ming Jin","Boyi Li","Marco Pavone","Serena Yeung-Levy","Jun Wang","Dawn Song","Costas Spanos"]}},"version":2},{"content":{"venue":{"value":"CoRR 2024"},"pdf":{"value":"http://arxiv.org/pdf/2407.15566v1"},"venueid":{"value":"dblp.org/journals/CORR/2024"},"paperhash":{"value":"liu|not_all_pairs_are_equal_hierarchical_learning_for_averageprecisionoriented_video_retrieval"},"authorids":{"value":["~Yang_Liu121","https://dblp.org/search/pid/api?q=author:Qianqian_Xu:","~Peisong_Wen1","https://dblp.org/search/pid/api?q=author:Siran_Dai:","~Qingming_Huang1"]},"html":{"value":"https://doi.org/10.48550/arXiv.2407.15566"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2407-15566,\n  publtype={informal},\n  author={Yang Liu and Qianqian Xu and Peisong Wen and Siran Dai and Qingming Huang},\n  title={Not All Pairs are Equal: Hierarchical Learning for Average-Precision-Oriented Video Retrieval},\n  year={2024},\n  cdate={1704067200000},\n  journal={CoRR},\n  volume={abs/2407.15566},\n  url={https://doi.org/10.48550/arXiv.2407.15566}\n}\n"},"abstract":{"value":"The rapid growth of online video resources has significantly promoted the development of video retrieval methods. As a standard evaluation metric for video retrieval, Average Precision (AP) assesses the overall rankings of relevant videos at the top list, making the predicted scores a reliable reference for users. However, recent video retrieval methods utilize pair-wise losses that treat all sample pairs equally, leading to an evident gap between the training objective and evaluation metric. To effectively bridge this gap, in this work, we aim to address two primary challenges: a) The current similarity measure and AP-based loss are suboptimal for video retrieval; b) The noticeable noise from frame-to-frame matching introduces ambiguity in estimating the AP loss. In response to these challenges, we propose the Hierarchical learning framework for Average-Precision-oriented Video Retrieval (HAP-VR). For the former challenge, we develop the TopK-Chamfer Similarity and QuadLinear-AP loss to measure and optimize video-level similarities in terms of AP. For the latter challenge, we suggest constraining the frame-level similarities to achieve an accurate AP loss estimation. Experimental results present that HAP-VR outperforms existing methods on several benchmark datasets, providing a feasible solution for video retrieval tasks and thus offering potential benefits for the multi-media application."},"title":{"value":"Not All Pairs are Equal: Hierarchical Learning for Average-Precision-Oriented Video Retrieval"},"authors":{"value":["Yang Liu","Qianqian Xu","Peisong Wen","Siran Dai","Qingming Huang"]}},"tmdate":1777731582044,"pdate":1704067200000,"tcdate":1731503869401,"writers":["~"],"signatures":["~Qingming_Huang2"],"forum":"ysPRN8qjo4","license":"CC BY-SA 4.0","number":227570,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1777731582044,"domain":"DBLP.org","id":"ysPRN8qjo4","version":2},{"content":{"venue":{"value":"ACM Multimedia 2024"},"venueid":{"value":"dblp.org/conf/MM/2024"},"paperhash":{"value":"liu|not_all_pairs_are_equal_hierarchical_learning_for_averageprecisionoriented_video_retrieval"},"authorids":{"value":["~Yang_Liu121","~Qianqian_Xu2","~Peisong_Wen1","~Siran_Dai1","~Qingming_Huang1"]},"html":{"value":"https://doi.org/10.1145/3664647.3681110"},"_bibtex":{"value":"@inproceedings{DBLP:conf/mm/LiuXWDH24,\n  author={Yang Liu and Qianqian Xu and Peisong Wen and Siran Dai and Qingming Huang},\n  title={Not All Pairs are Equal: Hierarchical Learning for Average-Precision-Oriented Video Retrieval},\n  year={2024},\n  cdate={1704067200000},\n  pages={3828-3837},\n  url={https://doi.org/10.1145/3664647.3681110},\n  booktitle={ACM Multimedia},\n  crossref={conf/mm/2024}\n}\n"},"abstract":{"value":"The rapid growth of online video resources has significantly promoted the development of video retrieval methods. As a standard evaluation metric for video retrieval, Average Precision (AP) assesses the overall rankings of relevant videos at the top list, making the predicted scores a reliable reference for the users. However, recent video retrieval methods utilize pair-wise losses that treat all sample pairs equally, leading to an evident gap between the training objective and evaluation metric. To effectively bridge this gap, in this work, we aim to address two primary challenges: a) The current similarity measure and AP-based loss are suboptimal for video retrieval; b) The noticeable noise from frame-to-frame matching introduces ambiguity in estimating the AP loss. In response to these challenges, we propose the Hierarchical learning framework for Average-Precision-oriented Video Retrieval (HAP-VR). For the former challenge, we develop the TopK-Chamfer Similarity and QuadLinear-AP loss to measure and optimize video-level similarities in terms of AP. For the latter challenge, we suggest constraining the frame-level similarities to achieve an accurate AP loss estimation. Experimental results present that HAP-VR outperforms existing methods on several benchmark datasets, providing a feasible solution for video retrieval tasks and thus offering potential benefits for the multi-media application."},"title":{"value":"Not All Pairs are Equal: Hierarchical Learning for Average-Precision-Oriented Video Retrieval"},"authors":{"value":["Yang Liu","Qianqian Xu","Peisong Wen","Siran Dai","Qingming Huang"]}},"tmdate":1758531948463,"pdate":1704067200000,"tcdate":1731503867548,"writers":["~"],"signatures":["~Qingming_Huang2"],"forum":"E7OnkSUq9O","license":"CC BY-SA 4.0","number":227504,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1758531948463,"domain":"DBLP.org","id":"E7OnkSUq9O","version":2},{"content":{"venue":{"value":"MM2024 Oral"},"supplementary_material":{"value":"/attachment/e913a0313ea0a516b66d5623e8748c241b785fb0.zip"},"abstract":{"value":"The rapid growth of online video resources has significantly promoted the development of video retrieval methods. As a standard evaluation metric for video retrieval, Average Precision (AP) assesses the overall rankings of relevant videos at the top list, making the predicted scores a reliable reference for users. However, recent video retrieval methods utilize pair-wise losses that treat all sample pairs equally, leading to an evident gap between the training objective and evaluation metric. To effectively bridge this gap, in this work, we aim to address two primary challenges: a) The current similarity measure and AP-based loss are suboptimal for video retrieval; b) The noticeable noise from frame-to-frame matching introduces ambiguity in estimating the AP loss. In response to these challenges, we propose the Hierarchical learning framework for Average-Precision-oriented Video Retrieval (HAP-VR). For the former challenge, we develop the TopK-Chamfer Similarity and QuadLinear-AP loss to measure and optimize video-level similarities in terms of AP. For the latter challenge, we suggest constraining the frame-level similarities to achieve an accurate AP loss estimation. Experimental results present that HAP-VR outperforms existing methods on several benchmark datasets, providing a feasible solution for video retrieval tasks and thus offering potential benefits for the multi-media application."},"relevance_to_conference":{"value":"The rapid extension of online video resources has posed tough challenges for efficient video search and analysis, highlighting the need for advanced content-based video retrieval methods, which serve as a crucial component for various multi-media applications such as recommendation systems, video processing, and online education. In this work, we propose a hierarchical learning framework for video retrieval based on Average Precision (AP) optimization, to fill the gap between training objectives and evaluation metrics that the previous works have overlooked. Specifically, we design a video-oriented similarity measure and a surrogate AP loss with proper gradients, constraining both the video-level and frame-level similarities to achieve an accurate AP loss estimation. Experimental results reveal that our framework frequently surpasses existing methods on several downstream tasks, which provides a feasible and promising solution for large-scale video understanding and management. We hope our work could bring potential benefits for broader applications and support subsequent research to further contribute to the multi-media community."},"_bibtex":{"value":"@inproceedings{\nliu2024not,\ntitle={Not All Pairs are Equal: Hierarchical Learning for Average-Precision-Oriented Video Retrieval},\nauthor={Yang Liu and Qianqian Xu and Peisong Wen and Siran Dai and Qingming Huang},\nbooktitle={ACM Multimedia 2024},\nyear={2024},\nurl={https://openreview.net/forum?id=6vcZHJTlPj}\n}"},"title":{"value":"Not All Pairs are Equal: Hierarchical Learning for Average-Precision-Oriented Video Retrieval"},"secondary_subject_area":{"value":["[Systems] Data Systems Management and Indexing"]},"pdf":{"value":"/pdf/1462452dfc7ada2da57c5118c6d236d738aca975.pdf"},"venueid":{"value":"acmmm.org/ACMMM/2024/Conference"},"paperhash":{"value":"liu|not_all_pairs_are_equal_hierarchical_learning_for_averageprecisionoriented_video_retrieval"},"primary_subject_area":{"value":"[Content] Media Interpretation"},"authorids":{"value":["~Yang_Liu121","~Qianqian_Xu2","~Peisong_Wen1","~Siran_Dai1","~Qingming_Huang2"]},"authors":{"value":["Yang Liu","Qianqian Xu","Peisong Wen","Siran Dai","Qingming Huang"]}},"tmdate":1721528154092,"pdate":1721513516068,"tcdate":1712395892387,"writers":["acmmm.org/ACMMM/2024/Conference","acmmm.org/ACMMM/2024/Conference/Submission2607/Authors"],"signatures":["acmmm.org/ACMMM/2024/Conference/Submission2607/Authors"],"forum":"6vcZHJTlPj","license":"CC BY 4.0","number":2607,"cdate":1712395892387,"readers":["everyone"],"invitations":["acmmm.org/ACMMM/2024/Conference/-/Submission","acmmm.org/ACMMM/2024/Conference/-/Post_Submission","acmmm.org/ACMMM/2024/Conference/Submission2607/-/Revision","acmmm.org/ACMMM/2024/Conference/Submission2607/-/Supplementary_Material","acmmm.org/ACMMM/2024/Conference/-/Edit"],"mdate":1721528154092,"odate":1721513516068,"domain":"acmmm.org/ACMMM/2024/Conference","id":"6vcZHJTlPj","version":2},{"content":{"summary":{"value":"This paper introduces CamPilot, a novel framework aimed at enhancing camera controllability in video diffusion models through reward feedback learning. The authors identify a persistent challenge in aligning generated video content with specified camera trajectories, which undermines 3D consistency in downstream tasks such as scene reconstruction. To address this, they propose a camera-aware 3D decoder that projects video latent and camera poses into 3D Gaussians (3DGS), enabling efficient rendering and reward computation. The framework is evaluated on RealEstate10K and WorldScore, demonstrating improved camera alignment and visual fidelity."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. How does the method handle dynamic scenes involving object motion or non-rigid transformations? It is recommended that the authors include evaluation results evaluated on open-sourced dynamic RealCam-VID[1].\n2. It is recommended that more comparison with other recent methods focused on camera control be added, such as RealCam-I2V[2]. We recommend including quantitative and qualitative results on the open-sourced dynamic RealCam-VID[1] dataset, alongside an analysis detailing the advantages of the proposed CamPilot framework in final paper.\n3. How well does the model perform under conditions of rapid camera motion or large trajectory shifts (e.g., simulated drone flights or rapid rotations)? The authors are encouraged to provide quantitative or qualitative results in such scenarios.\n4. How is camera scale handled during training and inference? The authors are recommended to clarify whether the model rely on absolute or relative scale inputs, and how does it ensure scale consistency across scenes?\n\n[1] RealCam-Vid: High-resolution Video Dataset with Dynamic Scenes and Metric-scale Camera Movements\n\n[2] RealCam-I2V: Real-World Image-to-Video Generation with Interactive Complex Camera Control"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. Introduces a feed-forward 3D Gaussian-based decoder that efficiently evaluates camera-video alignment without reliance on computationally intensive post-processing tools like COLMAP.\n2. Applies reward feedback learning (ReFL) to optimize camera adherence, which represents a previously underexplored direction in video diffusion.\n3. Enables high-quality 3D scene reconstruction directly from video latents and camera poses, bypassing computationally expensive per-scene optimization."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The evaluation is limited to static scene datasets, this potentially limits its applicability for real-world video generation tasks.\n2. The method assumes precise extrinsic and intrinsic camera parameters are available, which may not be a valid assumption in real-world applications."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762916481837,"tcdate":1761902066872,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission2989/Reviewer_62yC"],"signatures":["ICLR.cc/2026/Conference/Submission2989/Reviewer_62yC"],"forum":"qqij8fCGDl","number":4,"license":"CC BY 4.0","cdate":1761902066872,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission2989/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762916481837,"domain":"ICLR.cc/2026/Conference","replyto":"qqij8fCGDl","id":"wzsVNPIcQf","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Video generation; 3D scene exploration; Reward feedback learning"]},"supplementary_material":{"value":"/attachment/190335a20c10687ec00345e0bee18e333335cb61.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Recent advancements in camera-controlled video diffusion models have significantly improved video-camera alignment and enabled more accurate 3D scene generation, driven by potential downstream applications such as virtual reality.\nHowever, we reveal that existing approaches often struggle to precisely adhere to the given camera conditions, leading to inconsistencies in the 3D geometry.\nInspired by Reward Feedback Learning in diffusion models, which has demonstrated strong potential in aligning model outputs with task-specific objectives, we build upon this paradigm and aim to further improve camera controllability.\nDirectly borrowing existing ReFL approaches faces several challenges. First, current reward models lack the capacity to assess video-camera alignment. Second, decoding latent into RGB videos for reward computation introduces substantial computational overhead. Third, 3D geometric information is typically neglected during video decoding.\nTo address these limitations, we introduce a camera-aware 3D decoder that efficiently decodes video latent into 3D representations for reward computation. Specifically, we project the video latent and camera pose into 3D Gaussians, which supports efficient rendering from arbitrary views. \nIn this process, the camera pose not only acts as an input variable but also serves as a projection parameter for determining the mean of each Gaussian.\nIf the generated video does not match the camera conditions, the 3D structure becomes geometrically inconsistent, leading to blurry rendered images.\nBased on this property, we explicitly optimizing pixel-level consistency between rendered novel views and ground-truth ones as reward feedback.\nTo accommodate the stochastic nature, we further introduce a visibility term that selectively supervises only deterministic regions derived via geometric warping.\nExtensive experiments conducted on the RealEstate10K and WorldScore benchmarks demonstrate the effectiveness of our proposed method in enhancing both camera controllability and generation quality."},"_bibtex":{"value":"@misc{\nge2025campilot,\ntitle={CamPilot: Improving Camera Control in Video Diffusion Model with Efficient Camera Reward Feedback},\nauthor={Wenhang Ge and Guibao Shen and Jiawei Feng and Luozhou Wang and Hao LU and Xingye Tian and Xin Tao and Pengfei Wan and Ying-Cong Chen},\nyear={2025},\nurl={https://openreview.net/forum?id=qqij8fCGDl}\n}"},"title":{"value":"CamPilot: Improving Camera Control in Video Diffusion Model with Efficient Camera Reward Feedback"},"pdf":{"value":"/pdf/3919cc68511ac2752505bc0f7b4e7f570a2cd725.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"ge|campilot_improving_camera_control_in_video_diffusion_model_with_efficient_camera_reward_feedback"},"authorids":{"value":["~Wenhang_Ge1","~Guibao_Shen1","~Jiawei_Feng1","~Luozhou_Wang2","~Hao_LU8","~Xingye_Tian2","~Xin_Tao3","~Pengfei_Wan1","~Ying-Cong_Chen1"]},"authors":{"value":["Wenhang Ge","Guibao Shen","Jiawei Feng","Luozhou Wang","Hao LU","Xingye Tian","Xin Tao","Pengfei Wan","Ying-Cong Chen"]}},"version":2},{"content":{"summary":{"value":"The paper tackles the inefficiency of DiTs used in video diffusion model. The speedup of the presented method comes from two sources: 1) pruning the large full 3D attention of VDM DiTs and 2) distilling the model into a multi-step consistency model.\nThe authors identify a repetitive tile-like pattern, termed \"Attention Tile,\" in the 3D attention maps of video data. Leveraging this pattern, they propose a new family of sparse 3D attention mechanisms that reduce the computational complexity from quadratic to linear with respect to the number of video frames.\nTo further accelerate the inference process, the paper introduces a multi-step consistency distillation (MCD) technique. By dividing the sampling trajectory into segments and performing consistency distillation within each, the number of sampling steps required for video generation is significantly reduced.\nResults show that the method achieves good speedup without suffer much performance, using limited training data."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"N/A"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The paper makes a significant contribution by discovering the \"Attention Tile\" phenomenon in 3D full attention Diffusion Transformers (DiTs) for video data. This insight into the redundancy and repetitive patterns within attention maps is a valuable addition to the understanding of how attention mechanisms function in video generation models.\n2. Building on the Attention Tile observation, the authors propose a new family of sparse 3D attention mechanisms that reduce computational complexity from quadratic to linear concerning the number of video frames. This is a substantial improvement that directly addresses the inefficiency issues in existing models.\n3. The introduction of the EFFICIENT-VDIT framework is a well-thought-out approach that combines multi-step consistency distillation, layer-wise sparse attention mask searching, and knowledge distillation. This pipeline effectively accelerates inference while maintaining high metrics.\n4. Achieving these results using only 0.1% of the pretraining data is notable. It indicates that the method is not only computationally efficient but also data-efficient, which is advantageous when large datasets are not readily available."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The paper could benefit from a more in-depth discussion of the trade-offs involved, such as the balance between sparsity level and video quality or the impact on different types of video content (e.g., fast-moving vs. static scenes). For instance, why don't you directly use the demo videos on OpenSORA's websites and compare the qualitative results? They provided both static scenes with only relative camera poses and more dynamic scenes, e.g. filming of an explosion scene.\n2. The method relies on the observation that the Attention Tile pattern is data-independent. If this assumption does not hold for certain types of video data (e.g., highly dynamic scenes), the efficiency gains might not translate, potentially limiting the method's applicability.\n3. The use of only 0.1% of the pretraining data raises concerns about the generalization capabilities of the accelerated model. While performance loss is minimal on tested datasets, the model may underperform on unseen data or less common video scenarios.\n4. While the paper uses VBench and FVD for evaluation, these metrics may not capture all aspects of video quality, such as temporal coherence in more complex scenes or perceptual quality under different conditions. Including additional metrics or user studies could provide a more comprehensive assessment. This is especially concerning combined with weakness #2, since FVD is commonly known as a weak metric that focuses strongly on independent frames rather than overall video coherence. Overall, the evaluation seems to favor more static videos rather than highly dynamic videos, and I suspect the attention pruning would encourage such results too. A metric that takes motion into account is Content-Debiased FVD [1], but ideally, this is more suitable via a user study (even though I do not think this is necessary for the rebuttal stage, but better prepare it for another iteration of the paper).\n5. Inherit my point in #2 and #4, the paper does not provide any video data, making it challenging to assess the actual quality of the generated contents. From my point of view, a VDM paper should always be accompanied with as many videos as possible within the supplemental material size limit. Again, a good set would be the demo videos on OpenSORA's websites. They provided a wide range of descriptions and all the corresponding text prompts --- supposedly those prompts would work well on OpenSORA.\n\n[1] Ge et al., On the Content Bias in Fréchet Video Distance, in CVPR 2024."}},"nonreaders":[],"tmdate":1731427665240,"tcdate":1730557155551,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission2845/Reviewer_AbNB"],"signatures":["ICLR.cc/2025/Conference/Submission2845/Reviewer_AbNB"],"forum":"2ezRxhlAxJ","number":3,"license":"CC BY 4.0","cdate":1730557155551,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission2845/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427665240,"domain":"ICLR.cc/2025/Conference","replyto":"2ezRxhlAxJ","id":"rhOh3OIBBY","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Efficient inference","video generation","diffusion","Transformer"]},"primary_area":{"value":"infrastructure, software libraries, hardware, systems, etc."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Despite the promise of synthesizing high-fidelity videos, Diffusion Transformers (DiTs) with 3D full attention suffer from expensive inference due to the complexity of attention computation and numerous sampling steps. For example, the popular Open-Sora-Plan model consumes more than 9 minutes for generating a single video of 29 frames. This paper addresses the inefficiency issue from two aspects: 1) Prune the 3D full attention based on the redundancy within video data; We identify a prevalent tile-style repetitive pattern in the 3D attention maps for video data, and advocate a new family of sparse 3D attention that holds a linear complexity w.r.t. the number of video frames. 2) Shorten the sampling process based on multi-step consistency distillation; We split the entire sampling trajectory into several segments and perform consistency distillation within each one to activate few-step generation capacities. We further devise a three-stage training pipeline to conjoin the low-complexity attention and few-step generation capacities. Notably, with 0.1% pretraining data, we turn the Open-Sora-Plan-1.2 model into an efficient one that is 7.4x −7.8x faster for 29 and 93 frames 720p video generation with a marginal performance trade-off in VBench. In addition, we demonstrate that our approach is amenable to distributed inference, achieving an additional 3.91x speedup when running on 4 GPUs with sequence parallelism."},"_bibtex":{"value":"@misc{\nding2025efficientvdit,\ntitle={Efficient-vDiT: Efficient Video Diffusion Transformers With Attention Tile},\nauthor={Hangliang Ding and Dacheng Li and Runlong Su and Zhijie Deng and Ion Stoica and Hao Zhang},\nyear={2025},\nurl={https://openreview.net/forum?id=2ezRxhlAxJ}\n}"},"title":{"value":"Efficient-vDiT: Efficient Video Diffusion Transformers With Attention Tile"},"pdf":{"value":"/pdf/1009bb8176c4e1e6924213856a051436e5042cb8.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"ding|efficientvdit_efficient_video_diffusion_transformers_with_attention_tile"},"authorids":{"value":["~Hangliang_Ding1","~Dacheng_Li1","~Runlong_Su1","~Zhijie_Deng1","~Ion_Stoica1","~Hao_Zhang2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Hangliang Ding","Dacheng Li","Runlong Su","Zhijie Deng","Ion Stoica","Hao Zhang"]}},"version":2},{"content":{"summary":{"value":"This paper introduces MoCa, a framework for camera-controlled video generation that improves object stability by focusing on maintaining consistency in view, appearance, and motion. The model uses a dual-branch architecture with semantic guidance to preserve object identity and a disentanglement mechanism to separate object dynamics from camera movement."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"- The semantic guidance strategy (Section 3.2) uses a ReferenceNet to maintain object identity. The paper says it uses \"reference video frames\" (line 228), but it's not clear what this means in practice. Is the input just the first frame of the video, or a specific set of keyframes? This detail is crucial for understanding how the model gets its \"identity guidance.\"\n   - The high-frequency object-aware mask is a key part of the motion disentanglement. Could you provide more detail on how this mask is actually used in the \"Hybrid Condition Fusion\" step (line 263)? For example, is it used as a soft attention map to guide the DenoisingNet, or is it combined with other features in a different way? A clearer explanation of this mechanism would be very helpful."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The video examples shown in the paper supp are strong. The generated videos look much more stable and realistic than the comparison methods.\n2.  The main idea of framing the problem around \"object consistency\" (view, appearance, and motion) is a smart way to tackle the challenge. It breaks down a complex 3D problem into more manageable 2D properties that we can observe in the final video. \n3.  The method for separating object motion from camera motion is clever. Using a 2D Discrete Wavelet Transform (2D-DWT) to create a \"high-frequency object-aware mask\" is an interesting technical contribution."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The paper relies on Object Consistency (OC) and Background Consistency (BC) scores from VBench to prove its main contribution. However, as the VBench paper itself explains, these metrics just measure feature similarity (using DINO and CLIP) across frames. This means they mainly check if an object is consistently present, not if its motion is natural or if it's free from distortion. A video with a \"frozen\" object sliding unnaturally across the screen could still get a high OC score, which doesn't really support the claim of improved motion consistency.\n2. The method is built by fine-tuning CogVideoX, a very large (~5B) foundation model. While this helps achieve impressive results, it makes it hard to judge how much of the performance comes from the new MoCa architecture versus the power of the base model, especially considering the nature of randomness in the generation.\n3. The paper could use another round of proofreading. There are several places where the citation formatting is incorrect (e.g., in Section 4.1, should be a proper \\citep command). These small errors, along with some sections that are a bit dense, can make the paper harder to read and feel less polished, such the following part:\n\n   - The semantic guidance strategy (Section 3.2) uses a ReferenceNet to maintain object identity. The paper says it uses \"reference video frames\" (line 228), but it's not clear what this means in practice. Is the input just the first frame of the video, or a specific set of keyframes? This detail is crucial for understanding how the model gets its \"identity guidance.\"\n   - The high-frequency object-aware mask is a key part of the motion disentanglement. Could you provide more detail on how this mask is actually used in the \"Hybrid Condition Fusion\" step (line 263)? For example, is it used as a soft attention map to guide the DenoisingNet, or is it combined with other features in a different way? A clearer explanation of this mechanism would be very helpful."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762922127787,"tcdate":1761957653464,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission10931/Reviewer_QNWA"],"signatures":["ICLR.cc/2026/Conference/Submission10931/Reviewer_QNWA"],"forum":"DZcpnudp7f","number":3,"license":"CC BY 4.0","cdate":1761957653464,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission10931/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762922127787,"domain":"ICLR.cc/2026/Conference","replyto":"DZcpnudp7f","id":"aegM7SqlfU","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"TLDR":{"value":"We propose MoCa, a framework that enables precise camera control in text-to-video generation by modeling object view, appearance, and motion consistency to bridge 2D pixels and 3D scenes without explicit 3D supervision."},"keywords":{"value":["Text-To-Video","Camera-Control","Video Generation","Generative Model"]},"supplementary_material":{"value":"/attachment/d213752a1813617c7d8dd8ee3061c97279409eff.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Camera control is important in text-to-video generation for achieving realistic scene navigation and view synthesis. \nThis control is defined by parameters that describe movement through 3D space, thereby introducing 3D consistency into the generation process.\nA core challenge for existing methods is achieving 3D consistency within the 2D pixel domain.  Strategies that directly integrate camera conditions into text-to-video models often produce artifacts, while those relying on explicit 3D supervision face challenges with generalization. \nBoth limitations originate from the gap between the 2D pixel space and the underlying 3D world.\nThe key insight is that the projection of a smooth 3D camera movement produces consistency in object view, appearance, and motion across 2D frames. Inspired by this insight, we propose MoCa, a dual-branch framework that bridges this gap by modeling object consistency to implicitly learn 3D relationships between the camera and the scene.\nTo ensure view consistency, we design a Spatial-Temporal Camera Encoder with Plücker embedding, which encodes camera trajectories into a geometrically grounded latent representation. For appearance consistency, we introduce a semantic guidance strategy that leverages persistent vision-language features to maintain object identity and texture across frames. To address motion consistency, we propose an object-aware motion disentanglement mechanism that separates object dynamics from global camera movement, ensuring precise camera control and natural object motion.\nExperiments show that MoCa achieves accurate camera control while preserving video quality, offering a practical and effective solution for camera-controllable video generation."},"_bibtex":{"value":"@inproceedings{\nchengzhijing2026moca,\ntitle={MoCa: Modeling Object Consistency for 3D Camera Control in Video Generation},\nauthor={Zhijing Cheng and Xuancheng Zhang and Donglin Di and Chen Wei and Hao Li and Xun Yang},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=DZcpnudp7f}\n}"},"title":{"value":"MoCa: Modeling Object Consistency for 3D Camera Control in Video Generation"},"pdf":{"value":"/pdf/882a5b31f2ea04b8ae24c7c484f35c94706563c1.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"cheng|moca_modeling_object_consistency_for_3d_camera_control_in_video_generation"},"authorids":{"value":["~Zhijing_Cheng3","~Xuancheng_Zhang1","~Donglin_Di1","~Chen_Wei10","~Hao_Li81","~Xun_Yang1"]},"authors":{"value":["Zhijing Cheng","Xuancheng Zhang","Donglin Di","Chen Wei","Hao Li","Xun Yang"]}},"version":2},{"content":{"summary":{"value":"The paper proposes MVSplat360, a method for wide-sweepign or 360-degree novel view synthesis on general scenes from sparse input views. It extends an existing state-of-the-art approach MVSplat to render 3D feature Gaussians as conditioning for a refinement network in form of a pre-trained video diffusion model, which is fine-tuned jointly. The approach is evaluated on RealEstate10k, an existing benchmark of home walkthrough videos, and on a newly proposed benchmark for 360-degree novel view synthesis, leveraging an existing dataset of diverse scenes. Quantitative and qualitative results show improvements over regression-based and generative, splatting-based baselines and plausable completions in case of incomplete observations."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"- Are the structural parameters of the Gaussians purely trained via the SVD loss through the features?\n  - Does the for features extended 3D Gaussians rasterizer compute gradients through the features to the structural parameters of the Gaussians?\n  - If so, what is the intuition why this apparently works, while baselines (latentSplat and ReconFusion) use auxiliary losses?\n- Are novel views refined jointly by the video diffusion model?\n  - If yes, how do you deal with long videos?\n  - Is temporal ordering leveraged or is the model agnostic to temporal permutations of target camera poses?\n  - What are implications regarding 3D consistency?\n- How is the training done on RealEstate10k?\n  - Do you train two different models (also for baselines) for inter- and extrapolation?"},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"- The paper tackles a challenging and interesting task of generalizable 360-degree novel view synthesis from sparse views on diverse real-world scenes.\n- The proposed method combines the strengths of splatting-based generalizable 3D reconstruction and large-scale pre-trained video diffusion models as a genrative refinement network:\n  - pixelSplat [1] and MVSplat [2] have shown strong performance in view interpolation.\n  - latentSplat [3] has shown the advantages of using a decoder generative network.\n  - Video diffusion models like Stable Video Diffusion [4] have been trained on massive data and learn a strong prior useful for plausible view extrapolation and scene completion.\n- The problem definition and main approach explained well.\n- The evaluation validates the effectiveness of MVSplat360\n  - The proposed benchmark on the DL3DV dataset [5] is challenging and supports the claims of wide-sweeping or 360-degree novel view synthesis on large real-world scenes.\n  - The approach outperforms all baselines in almost all metrics quantitatively on DL3DV and RealEstate10k.\n  - Qualitative results show high-quality novel views with less artifacts than baselines for both datasets with plausible completions for extrapolation on RealEstate10k.\n  - The supplement provides more convincing qualitative results including a video.\n- The paper includes an ablation study validating all the proposed design choices and showing desired performance scaling w.r.t. increasing numbers of input views.\n\n1. pixelSplat: 3D Gaussian Splats from Image Pairs for Scalable Generalizable 3D Reconstruction, CVPR 2024\n2. MVSplat: Efficient 3D Gaussian Splatting from Sparse Multi-View Images, ECCV 2024\n3. latentSplat: Autoencoding Variational Gaussians for Fast Generalizable 3D Reconstruction, ECCV 2024\n4. Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets, arxiv 2023\n5. DL3DV-10K: A Large-Scale Scene Dataset for Deep Learning-based 3D Vision, CVPR 2024"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Confusing title: Given the contributions, the title should rather focus on the proposed method.\n  - The title of the paper suggests a focus on benchmarking existing methods.\n  - The paper still proposes a new approach MVSplat360.\n  - Evaluation is done on one existing benchmark (RealEstate10k) and one newly proposed one leveraging the already existing DL3DV dataset.\n  - The problem of generalizable novel view synthesis given sparse views is not novel (see baselines [1], [2], [3]) such that the benchmark creation only consists of the definition of input and target views.\n\n- Insufficient contextualization relative to prior work:\n  - Inaccurate use of the term 'concurrent work':\n    - In line 60, ReconFusion [4] and latentSplat [3] (again in line 244) are referred to as concurrent work.\n    - This is not really the case anymore, because\n      - ReconFusion [4] appeared on arxiv in the beginning of December 2023 and was presented at CVPR 2024.\n      - latentSplat [3] appeared on arxiv end of March 2024, 3 days after MVSplat [2], which is one of the building blocks of the proposed method.\n  - Missing references to related work regarding conditioning via feature rendering and CLIP embeddings:\n    - In lines 163-168 and 196-201, the authors propose to render latent features as spatial conditions for a generative model (Stable Video Diffusion), which is later ablated in lines 296-299.\n       - This idea is not novel and the paper is missing important references in this context:\n         - GeNVS [5] proposed to use rendered pixelNeRF features as conditoning for a 2D diffusion model.\n         - ReconFusion [4] adopts this to condition a text-to-image latent diffusion model.\n         - latentSplat[3] also renders feature maps that are decoded to a novel view in a GAN setup.\n    - In lines 193, the authors describe the use of CLIP image embeddings of the input views as global conditioning.\n      - ReconFusion [4] desribes a very similar procedure. \n\n- Lack of clarity regarding the proposed method:\n  - From the paper and supplement, it is not clear, for which loss terms the architecture is optimized during training.\n    - Since the paragraph about \"Multi-frame diffusion model\" starting in line 174 mainly describes SVD [6], it is not clear, whether the loss in equation (2) is also the only loss for MVSplat360.\n  - In line 167f., the authors claim that joint training of MVSplat and SVD can further enhance geometry through the feature conditioning.\n    - Is that the only source of gradient signals to the structural parameters of the 3D Gaussians (location, scale, rotation...)?\n    - ReconFusion [4] and latentSplat [3] both additionally optimize direct RGB renderings to regress the ground truth for better geometry gradients.\n  - The role of the video diffusion model compared to a single-view diffusion model is not clear.\n    - The paper misses to explain, whether and if so how novel views are refined jointly (see questions).\n  \n- The proposed approach is incremental.\n  - Compared to ReconFusion, MVSplat360 uses MVSplat instead of pixelNeRF and a video instead of an image diffusion model.\n  - Compared to latentSplat, it replaces the pixelSplat encoder with the improved MVSplat and a small CNN-based decoder with SVD [6].\n\n- Unclear conclusion of benchmarking:\n  - latentSplat builds upon pixelSplat, which is shown to be worse in depth estimation than MVSplat.\n    - Regarding this baseline, it is difficult to conclude, how much the reconstruction and the refinement modules contribute to the performance.\n  - If the goal of the paper is benchmarking, a component-wise evaluation (backbone feature extractors, depth estimation, refinement networks)  would be more insightful.\n\n1. pixelSplat: 3D Gaussian Splats from Image Pairs for Scalable Generalizable 3D Reconstruction, CVPR 2024\n2. MVSplat: Efficient 3D Gaussian Splatting from Sparse Multi-View Images, ECCV 2024\n3. latentSplat: Autoencoding Variational Gaussians for Fast Generalizable 3D Reconstruction, ECCV 2024\n4. ReconFusion: 3D Reconstruction with Diffusion Priors, CVPR 2024\n5. GeNVS: Generative Novel View Synthesis with 3D-Aware Diffusion Models, ICCV 2023\n6. Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets, arxiv 2023"},"limitations":{"value":"The authors addressed limitations and societal impacts in the appendix."}},"nonreaders":[],"tmdate":1730878769063,"tcdate":1722035468222,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission2348/Reviewer_4RiL"],"signatures":["NeurIPS.cc/2024/Conference/Submission2348/Reviewer_4RiL"],"forum":"B0OWOkMwhz","number":4,"license":"CC BY 4.0","cdate":1722035468222,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission2348/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730878769063,"domain":"NeurIPS.cc/2024/Conference","replyto":"B0OWOkMwhz","id":"da0YnomrSc","forumContent":{"venue":{"value":"NeurIPS 2024 poster"},"TLDR":{"value":"MVSplat360 is a feed-forward approach for 360° novel view synthesis of diverse real-world scenes, using only sparse observations."},"keywords":{"value":["novel view synthesis","feed-forward 3DGS","3D gaussians splatting","latent video diffusion model"]},"primary_area":{"value":"generative_models"},"abstract":{"value":"We introduce MVSplat360, a feed-forward approach for 360° novel view synthesis (NVS) of diverse real-world scenes, using only sparse observations. This setting is inherently ill-posed due to minimal overlap among input views and insufficient visual information provided, making it challenging for conventional methods to achieve high-quality results. Our MVSplat360 addresses this by effectively combining geometry-aware 3D reconstruction with temporally consistent video generation. Specifically, it refactors a feed-forward 3D Gaussian Splatting (3DGS) model to render features directly into the latent space of a pre-trained Stable Video Diffusion (SVD) model, where these features then act as pose and visual cues to guide the denoising process and produce photorealistic 3D-consistent views. Our model is end-to-end trainable and supports rendering arbitrary views with as few as 5 sparse input views. To evaluate MVSplat360's performance, we introduce a new benchmark using the challenging DL3DV-10K dataset, where MVSplat360 achieves superior visual quality compared to state-of-the-art methods on wide-sweeping or even 360° NVS tasks. Experiments on the existing benchmark RealEstate10K also confirm the effectiveness of our model. Readers are highly recommended to view the video results at [donydchen.github.io/mvsplat360](https://donydchen.github.io/mvsplat360)."},"_bibtex":{"value":"@inproceedings{\nchen2024mvsplat,\ntitle={{MVS}plat360: Feed-Forward 360 Scene Synthesis from Sparse Views},\nauthor={Yuedong Chen and Chuanxia Zheng and Haofei Xu and Bohan Zhuang and Andrea Vedaldi and Tat-Jen Cham and Jianfei Cai},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=B0OWOkMwhz}\n}"},"title":{"value":"MVSplat360: Feed-Forward 360 Scene Synthesis from Sparse Views"},"pdf":{"value":"/pdf/8eab4c80c2d13fe008de8ba51c5e01546d382d2f.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"chen|mvsplat360_feedforward_360_scene_synthesis_from_sparse_views"},"authorids":{"value":["~Yuedong_Chen1","~Chuanxia_Zheng1","~Haofei_Xu1","~Bohan_Zhuang1","~Andrea_Vedaldi1","~Tat-Jen_Cham1","~Jianfei_Cai1"]},"authors":{"value":["Yuedong Chen","Chuanxia Zheng","Haofei Xu","Bohan Zhuang","Andrea Vedaldi","Tat-Jen Cham","Jianfei Cai"]}},"version":2},{"content":{"venue":{"value":"CoRR 2020"},"pdf":{"value":"http://arxiv.org/pdf/2008.06941v2"},"venueid":{"value":"dblp.org/journals/CORR/2020"},"paperhash":{"value":"zhang|objectaware_multibranch_relation_networks_for_spatiotemporal_video_grounding"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Zhu_Zhang:","~Zhou_Zhao2","~Zhijie_Lin1","https://dblp.org/search/pid/api?q=author:Baoxing_Huai:","~Nicholas_Jing_Yuan2"]},"html":{"value":"https://arxiv.org/abs/2008.06941"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2008-06941,\n  publtype={informal},\n  author={Zhu Zhang and Zhou Zhao and Zhijie Lin and Baoxing Huai and Nicholas Jing Yuan},\n  title={Object-Aware Multi-Branch Relation Networks for Spatio-Temporal Video Grounding},\n  year={2020},\n  cdate={1577836800000},\n  journal={CoRR},\n  volume={abs/2008.06941},\n  url={https://arxiv.org/abs/2008.06941}\n}\n"},"abstract":{"value":"Spatio-temporal video grounding aims to retrieve the spatio-temporal tube of a queried object according to the given sentence. Currently, most existing grounding methods are restricted to well-aligned segment-sentence pairs. In this paper, we explore spatio-temporal video grounding on unaligned data and multi-form sentences. This challenging task requires to capture critical object relations to identify the queried target. However, existing approaches cannot distinguish notable objects and remain in ineffective relation modeling between unnecessary objects. Thus, we propose a novel object-aware multi-branch relation network for object-aware relation discovery. Concretely, we first devise multiple branches to develop object-aware region modeling, where each branch focuses on a crucial object mentioned in the sentence. We then propose multi-branch relation reasoning to capture critical object relationships between the main branch and auxiliary branches. Moreover, we apply a diversity loss to make each branch only pay attention to its corresponding object and boost multi-branch learning. The extensive experiments show the effectiveness of our proposed method."},"title":{"value":"Object-Aware Multi-Branch Relation Networks for Spatio-Temporal Video Grounding"},"authors":{"value":["Zhu Zhang","Zhou Zhao","Zhijie Lin","Baoxing Huai","Nicholas Jing Yuan"]}},"tmdate":1768535346509,"pdate":1577836800000,"tcdate":1731479062167,"writers":["~"],"signatures":["~Zhijie_Lin1"],"forum":"TLKe06uInv","license":"CC BY-SA 4.0","number":206909,"cdate":1577836800000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1768535346509,"domain":"DBLP.org","id":"TLKe06uInv","version":2},{"content":{"venue":{"value":"IJCAI 2020"},"pdf":{"value":"https://www.ijcai.org/proceedings/2020/0149.pdf"},"venueid":{"value":"dblp.org/conf/IJCAI/2020"},"paperhash":{"value":"zhang|objectaware_multibranch_relation_networks_for_spatiotemporal_video_grounding"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Zhu_Zhang:","~Zhou_Zhao2","~Zhijie_Lin1","https://dblp.org/search/pid/api?q=author:Baoxing_Huai:","~Jing_Yuan12"]},"html":{"value":"https://doi.org/10.24963/ijcai.2020/149"},"_bibtex":{"value":"@inproceedings{DBLP:conf/ijcai/ZhangZLHY20,\n  author={Zhu Zhang and Zhou Zhao and Zhijie Lin and Baoxing Huai and Jing Yuan},\n  title={Object-Aware Multi-Branch Relation Networks for Spatio-Temporal Video Grounding},\n  year={2020},\n  cdate={1577836800000},\n  pages={1069-1075},\n  url={https://doi.org/10.24963/ijcai.2020/149},\n  booktitle={IJCAI},\n  crossref={conf/ijcai/2020}\n}\n"},"abstract":{"value":"Spatio-temporal video grounding aims to retrieve the spatio-temporal tube of a queried object according to the given sentence. Currently, most existing grounding methods are restricted to well-aligned segment-sentence pairs. In this paper, we explore spatio-temporal video grounding on unaligned data and multi-form sentences. This challenging task requires to capture critical object relations to identify the queried target. However, existing approaches cannot distinguish notable objects and remain in ineffective relation modeling between unnecessary objects. Thus, we propose a novel object-aware multi-branch relation network for object-aware relation discovery. Concretely, we first devise multiple branches to develop object-aware region modeling, where each branch focuses on a crucial object mentioned in the sentence. We then propose multi-branch relation reasoning to capture critical object relationships between the main branch and auxiliary branches. Moreover, we apply a diversity loss to make each branch only pay attention to its corresponding object and boost multi-branch learning. The extensive experiments show the effectiveness of our proposed method."},"title":{"value":"Object-Aware Multi-Branch Relation Networks for Spatio-Temporal Video Grounding"},"authors":{"value":["Zhu Zhang","Zhou Zhao","Zhijie Lin","Baoxing Huai","Jing Yuan"]}},"tmdate":1768452483168,"pdate":1577836800000,"tcdate":1731479061806,"writers":["~"],"signatures":["~Zhijie_Lin1"],"forum":"iGZnL78l9u","license":"CC BY-SA 4.0","number":206907,"cdate":1577836800000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1768452483168,"domain":"DBLP.org","id":"iGZnL78l9u","version":2},{"content":{"venue":{"value":"IEEE Trans. Image Process. 2014"},"pdf":{"value":"https://ieeexplore.ieee.org/iel7/83/6657713/06605606.pdf"},"venueid":{"value":"dblp.org/journals/TIP/2014"},"paperhash":{"value":"hadizadeh|saliencyaware_video_compression"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Hadi_Hadizadeh:","~Ivan_V._Bajic1"]},"html":{"value":"https://doi.org/10.1109/TIP.2013.2282897"},"_bibtex":{"value":"@article{DBLP:journals/tip/HadizadehB14,\n  author={Hadi Hadizadeh and Ivan V. Bajic},\n  title={Saliency-Aware Video Compression},\n  year={2014},\n  cdate={1388534400000},\n  journal={IEEE Trans. Image Process.},\n  volume={23},\n  number={1},\n  pages={19-33},\n  url={https://doi.org/10.1109/TIP.2013.2282897}\n}\n"},"abstract":{"value":"In region-of-interest (ROI)-based video coding, ROI parts of the frame are encoded with higher quality than non-ROI parts. At low bit rates, such encoding may produce attention-grabbing coding artifacts, which may draw viewer's attention away from ROI, thereby degrading visual quality. In this paper, we present a saliency-aware video compression method for ROI-based video coding. The proposed method aims at reducing salient coding artifacts in non-ROI parts of the frame in order to keep user's attention on ROI. Further, the method allows saliency to increase in high quality parts of the frame, and allows saliency to reduce in non-ROI parts. Experimental results indicate that the proposed method is able to improve visual quality of encoded video relative to conventional rate distortion optimized video coding, as well as two state-of-the art perceptual video coding methods."},"title":{"value":"Saliency-Aware Video Compression"},"authors":{"value":["Hadi Hadizadeh","Ivan V. Bajic"]}},"tmdate":1747248625834,"pdate":1388534400000,"tcdate":1747248537036,"writers":["~"],"signatures":["~Ivan_V._Bajic1"],"forum":"cL4pkmK96w","license":"CC BY-SA 4.0","number":465152,"cdate":1388534400000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1747248625834,"domain":"DBLP.org","id":"cL4pkmK96w","version":2},{"content":{"summary":{"value":"The paper introduces Failure-Aware Inverse Reinforcement Learning (FA-IRL) to recover the latent reward behind RLHF by explicitly identifying and up-weighting “failures”—preference pairs that are ambiguous or misclassified—during reward learning. Concretely, it uses a dual-path reward model over frozen texts embeddings; failures are detected via margin uncertainty or disagreement with supervised labels. Evaluated on detoxification, with preference pairs built from RealToxicityPrompts (base vs. aligned “expert”) and a held-out Jigsaw set, FA-IRL improves F1/AUC and reduces STARC error versus standard IRL, and when used for re-RLHF it lowers toxicity to ~6% (vs. ~9% with standard IRL; ~4% with ground-truth reward)."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"N/A"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"* The paper brings a simple and effect idea to treating uncertain or wrong pairs as “failures” and giving them extra weight with a two-path reward, turning common near-ties and label noise into useful signal. \n* The method is clearly described and easy to re-produce, with a small theory piece that explains why focusing on failures reduces ambiguity and with careful experiments and ablations to back it up. \n* The gains show up in both standard metrics and downstream re-alignment (lower toxicity) across several model sizes, suggesting practical impact beyond this one detox task."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"* Preference pairs are synthetic (base vs. expert with the expert always preferred), which can make learning partly about distinguishing policies rather than general human preferences; include some human-curated pairs, allow non-expert-preferred pairs, and check if gains hold.\n* The dataset is small and targeted to detox (≈20k pairs per family, plus Jigsaw), so generality is unclear, including non-toxicity axis (e.g., factuality or refusal correctness) and report cross-task transfer could help.\n* The “ground-truth” reward that trains the expert policy is a public toxicity classifier, not human ratings; this risks learning the classifier’s biases rather than human preference, so add a small human eval could be useful and probe for bias (group-wise error, calibration) and agreement with humans. \n* Results use only a few small model families (≈135M–410M), so it’s unclear if gains hold for larger instruction-tuned models or other tasks. Including at least one 7B-class model would be needed for generality."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762927403123,"tcdate":1761972727751,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission17527/Reviewer_586b"],"signatures":["ICLR.cc/2026/Conference/Submission17527/Reviewer_586b"],"forum":"GQpkHR1H7t","number":3,"license":"CC BY 4.0","cdate":1761972727751,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission17527/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762927403123,"domain":"ICLR.cc/2026/Conference","replyto":"GQpkHR1H7t","id":"3vxIe8kA48","forumContent":{"TLDR":{"value":"We develop failure aware IRL to understand LLM alignment"},"venue":{"value":"ICLR 2026 Conference Desk Rejected Submission"},"keywords":{"value":["Inverse Reinforcement Learning","Failures; Alignment","LLMs","RLHF"]},"primary_area":{"value":"reinforcement learning"},"abstract":{"value":"Reinforcement Learning from Human Feedback (RLHF) aligns Large Language Models (LLMs) with human preferences, yet the underlying reward signals they internalize remain hidden, posing a critical challenge for interpretability and safety. Existing approaches attempt to extract these latent incentives using Inverse Reinforcement Learning (IRL), but treat all preference pairs equally, often overlooking the most informative signals: those examples the extracted reward model misclassifies or assigns nearly equal scores, which we term *failures*. We introduce a novel *failure-aware* IRL algorithm that focuses on misclassified or difficult examples to recover the latent rewards defining model behaviors. By learning from these failures, our failure-aware IRL extracts reward functions that better reflect the true objectives behind RLHF. We demonstrate that failure-aware IRL outperforms existing IRL baselines across multiple metrics when applied to LLM detoxification, without requiring external classifiers or supervision. Crucially, failure-aware IRL yields rewards that better capture the true incentives learned during RLHF, enabling more effective re-RLHF training than standard IRL. This establishes failure-aware IRL as a robust, scalable method for auditing model alignment and reducing ambiguity in the IRL process."},"_bibtex":{"value":"@misc{\nanonymous2026learning,\ntitle={Learning from Failures: Understanding {LLM} Alignment through Failure-Aware Inverse {RL}},\nauthor={Anonymous},\nyear={2026},\nurl={https://openreview.net/forum?id=GQpkHR1H7t}\n}"},"title":{"value":"Learning from Failures: Understanding LLM Alignment through Failure-Aware Inverse RL"},"pdf":{"value":"/pdf/4cb236295bec17d0977e5160c6623dfe93c3bcdd.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Desk_Rejected_Submission"},"paperhash":{"value":"patel|learning_from_failures_understanding_llm_alignment_through_failureaware_inverse_rl"},"authorids":{"value":["~Nyal_Patel1","~Matthieu_Bou1","~Arjun_Jagota1","~Satyapriya_Krishna2","~Sonali_Parbhoo2"]},"authors":{"value":["Nyal Patel","Matthieu Bou","Arjun Jagota","Satyapriya Krishna","Sonali Parbhoo"]}},"version":2},{"content":{"summary":{"value":"This paper presents VidGuard-R1, a multi-modal large language model (MLLM)–based system for AI-generated video detection and explanation. The work addresses growing societal risks from realistic video generation by developing a reasoning-capable authenticity detector that produces both accurate judgments and interpretable explanations. The model fine-tunes Qwen-VL using a two-stage framework: supervised chain-of-thought (CoT) initialization followed by reinforcement learning with Group Relative Policy Optimization (GRPO) and two specialized reward models—one emphasizing temporal artifacts (GRPO-TA) and another focusing on generation complexity (GRPO-Q). The authors construct a large dataset of real and synthetic videos generated by recent diffusion-based models and demonstrate that VidGuard-R1 achieves strong zero-shot and fine-tuned performance on several benchmarks, accompanied by qualitative examples illustrating its reasoning process."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"- The statement that “directly assigning a reward of 1 to real videos and 0 to fake ones presents challenges” is not justified theoretically or empirically; a short ablation or analysis on this design choice would clarify its necessity.\n\n- Will the dataset and code be publicly released for community use?"},"rating":{"value":6},"details_of_ethics_concerns":{"value":"None"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":3},"strengths":{"value":"- The paper tackles an important societal problem by adapting reasoning MLLMs for deepfake and synthetic video detection.\n- The combination of supervised CoT initialization and RL-based GRPO fine-tuning is conceptually sound and clearly explained, with reward models thoughtfully designed to encourage temporal and quality-aware reasoning.\n- The curated dataset is large, diverse, and standardized to minimize shortcut cues such as duration or resolution differences, making it well suited for robust learning. The dataset can be of good value to the community as well.\n- Experimental results show consistent and often state-of-the-art accuracy across multiple benchmarks, including GenVideo and GenVidBench, with competitive zero-shot generalization.\n- The inclusion of reasoning traces and qualitative explanations enhances transparency and potential user trust, distinguishing this work from prior black-box detectors.\n\n- The paper is well written and easy to follow, with clear explanations of both the methodology and its motivation. Figures and examples  effectively convey the reasoning process and highlight the interpretability of the model’s predictions."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Although interpretability is a stated goal, the paper does not evaluate the quality of explanations beyond qualitative examples. It remains unclear whether the rationales are causally consistent with correct predictions or merely plausible-sounding. Established metrics (e.g., LLM-as-judge, human evaluation, or coherence scores) should be adopted to substantiate this claim.\n\n- Several key ideas have been explored in prior work on image and video authenticity detection (e.g., SafeWatch, DeMamba-XCLIP, or recent surveys on AI-generated media detection). The paper could more clearly articulate what is fundamentally new in its contribution beyond integrating existing elements.\n\n- The related work section omits relevant prior literature on video safety and explainable detection, such as SafeWatch (ICLR 2025) and other text-to-video safety models, which share similar objectives and techniques. A more explicit comparison—methodologically and empirically—would strengthen positioning.\n\n- The training dataset, while large, is largely synthesized from one or two generative sources (e.g., HunyuanVideo-I2V, CogVideoX). This raises concerns about generalization to unseen generation models. Evaluating across broader generative distributions or conducting cross-model tests would better support claims of robustness."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762919113945,"tcdate":1761753870094,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission6859/Reviewer_1fnv"],"signatures":["ICLR.cc/2026/Conference/Submission6859/Reviewer_1fnv"],"forum":"gXjOsBcXIR","number":2,"license":"CC BY 4.0","cdate":1761753870094,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission6859/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762919113945,"domain":"ICLR.cc/2026/Conference","replyto":"gXjOsBcXIR","id":"ib9TQCAV3J","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"TLDR":{"value":"We present VidGuard-R1, the first multimodal LLM fine-tuned with reinforcement learning to detect and explain AI-generated videos, achieving state-of-the-art accuracy with interpretable reasoning."},"keywords":{"value":["Discriminator","MLLM"]},"supplementary_material":{"value":"/attachment/b252cebc0ec8f5d381df44826c485f20435a1093.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"The rapid proliferation of AI-generated video necessitates robust detection tools that offer both high accuracy and human-interpretable explanations. While existing MLLM-based detectors rely on supervised fine-tuning (SFT) or direct preference optimization (DPO), these methods are often bottlenecked by static, pre-labeled datasets that fail to capture the evolving, multi-step physical inconsistencies of modern generative models. To bridge this gap, we introduce VidGuard-R1, the first video authenticity detector to utilize group relative policy optimization (GRPO). Moving beyond passive preference matching, VidGuard-R1 employs a reinforcement learning framework that encourages the model to explore and rank multiple reasoning paths. By introducing specialized reward models for temporal stability and diffusion-aware complexity, we incentivize the model to discover 'physics-grounded' artifacts. Our contributions include: (1) a curated dataset of 140,000 challenging real/fake video pairs; (2) a GRPO-based training paradigm that achieves state-of-the-art zero-shot performance; and (3) a reasoning-first architecture that provides precise, verifiable rationales for its forensic judgments. Project website: https://vidguard-r1.github.io"},"_bibtex":{"value":"@inproceedings{\npark2026vidguardr,\ntitle={VidGuard-R1: {AI}-Generated Video Detection and Explanation via Reasoning {MLLM}s and {RL}},\nauthor={Kyoungjun Park and Yifan Yang and Juheon Yi and Shicheng Zheng and Muhammad Muaz and Yifei Shen and Dongqi Han and Caihua Shan and Lili Qiu},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=gXjOsBcXIR}\n}"},"title":{"value":"VidGuard-R1: AI-Generated Video Detection and Explanation via Reasoning MLLMs and RL"},"pdf":{"value":"/pdf/da179971e84b546eeed9cddcc1689f1de0dbf192.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"park|vidguardr1_aigenerated_video_detection_and_explanation_via_reasoning_mllms_and_rl"},"authorids":{"value":["~Kyoungjun_Park1","~Yifan_Yang9","~Juheon_Yi1","~Shicheng_Zheng4","~Muhammad_Muaz1","~Yifei_Shen1","~Dongqi_Han1","~Caihua_Shan1","~Lili_Qiu1"]},"authors":{"value":["Kyoungjun Park","Yifan Yang","Juheon Yi","Shicheng Zheng","Muhammad Muaz","Yifei Shen","Dongqi Han","Caihua Shan","Lili Qiu"]}},"version":2},{"content":{"summary":{"value":"This paper proposes PEVLM, a fine-tuning-free parallel encoding method to address the inefficiency of Vision-Language Models (VLMs) in long video understanding due to quadratic attention complexity. By partitioning input videos into context blocks and aligning attention scores with Full-Attention, PEVLM reduces complexity from O((TN)^2) to O(TN) with minimal accuracy loss. Experiments show up to 7.47x speedup, 40% latency reduction, and even improved accuracy in some cases, making PEVLM a promising solution for low-latency, long-video reasoning tasks."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"I noticed that all the baselines used in the paper are based on 7B models. How would the proposed method perform on larger or smaller models, and what impact would model size have on the approach?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The paper is well-written, with the methodology and experimental results presented in a clear and systematic manner.\n2. The research topic is highly practical, as introducing a training-free method that effectively reduces inference memory usage and latency provides significant value for real-world model deployment.\n3. The experiments in the paper highlight the effectiveness of PEVLM, achieving a 40% reduction in latency while maintaining accuracy on several long-video benchmarks."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The experimental evaluation is incomplete. While PEVLM is designed for long-video understanding, all the benchmarks focus solely on QA tasks. Token sparsification typically has limited impact on QA tasks; however, for tasks that rely on fine-grained visual details, such as video captioning or video OCR, it could introduce significant drawbacks. The authors should include results on such benchmarks to better demonstrate the method’s versatility across different task types.\n2. The paper briefly mentions a 40% latency reduction but lacks detailed analyses of latency and memory usage in the experiments. For instance, the impact of different hyper-parameter configurations on memory, latency, and accuracy across various models is not thoroughly explored. Token count alone does not provide sufficient insight into these metrics, and additional detailed analysis would strengthen the results."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762920658143,"tcdate":1761209020919,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission8905/Reviewer_89Qk"],"signatures":["ICLR.cc/2026/Conference/Submission8905/Reviewer_89Qk"],"forum":"oGbtnFrcLH","number":2,"license":"CC BY 4.0","cdate":1761209020919,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission8905/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762920658143,"domain":"ICLR.cc/2026/Conference","replyto":"oGbtnFrcLH","id":"CfvqM4zcrz","forumContent":{"TLDR":{"value":"PEVLM is a fast, fine-tuning-free method for long video understanding in VLMs, achieving up to 7.47× speedup and better accuracy by using parallel encoding with sequential position embeddings."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Vision-Language Models","Parallel Encoding","Large multimodal models","Latency-Constrained Inference","Long Context","Efficient Inference","Local and Global Attention"]},"primary_area":{"value":"generative models"},"abstract":{"value":"Vision-Language Models (VLMs) have demonstrated strong capabilities in multimodal understanding and generation tasks. However, their application to long video understanding remains hindered by the quadratic complexity of standard attention mechanisms. In this work, we introduce \\textbf{PEVLM}, a fine-tuning-free parallel encoding method designed to enhance the prefilling efficiency of VLMs in long video scenarios. To the best of our knowledge, this is the first work to adapt parallel encoding to VLMs. PEVLM partitions the input video into context blocks with a shared sink block, while preserving sequential position embeddings to align the attention score distribution with that of Full-Attention. This design reduces the complexity of attention from $O((T \\times N)^2)$ to $O(T \\times N)$ where $T$ is the number of frames and $N$ the number of tokens per frame, with minimal loss in accuracy.\nExtensive experiments across multiple state-of-the-art models and benchmarks demonstrate that PEVLM consistently outperforms existing parallel encoding approaches, achieving up to \\textbf{7.47x} speedup in attention computation and reducing end-to-end latency by \\textbf{44\\%} to \\textbf{50\\%}. Remarkably, PEVLM not only maintains high accuracy, but in some settings even surpasses Full-Attention performance. Under strict latency constraints, it achieves substantial gains, improving accuracy from \\textbf{23.26\\%} to \\textbf{61.03\\%}. These results underscore the effectiveness of PEVLM for low-latency, long-context video understanding, making it a promising solution for real-world applications."},"_bibtex":{"value":"@misc{\nkang2026pevlm,\ntitle={{PEVLM}: Parallel Encoding for Vision-Language Models},\nauthor={Letian Kang and Shixian Luo and Yiqiang Li and Shenxuan Zhou and Yuxin Yin and XiaoyangYu and Jin Yang and Yong Wu},\nyear={2026},\nurl={https://openreview.net/forum?id=oGbtnFrcLH}\n}"},"title":{"value":"PEVLM: Parallel Encoding for Vision-Language Models"},"pdf":{"value":"/pdf/de74e555aef35a8ebf1d0c54282446e2469a02a1.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"kang|pevlm_parallel_encoding_for_visionlanguage_models"},"authorids":{"value":["~Letian_Kang1","~Shixian_Luo2","~Yiqiang_Li4","~Shenxuan_Zhou1","~Yuxin_Yin1","~XiaoyangYu2","~Jin_Yang11","~Yong_Wu13"]},"authors":{"value":["Letian Kang","Shixian Luo","Yiqiang Li","Shenxuan Zhou","Yuxin Yin","XiaoyangYu","Jin Yang","Yong Wu"]}},"version":2},{"content":{"summary":{"value":"The paper introduces Collaborative Inference with Token-level Routing a framework designed to enhance the efficiency of large language model (LLM) inference while maintaining output quality. By implementing a token-level router that predicts the importance of individual tokens, CITER enables smaller language models (SLMs) to handle less critical tokens, reserving LLMs for essential ones. This approach formulates a reinforcement learning (RL) problem to minimize inference costs and introduces a shortcut for reward estimation, significantly accelerating training. Experiments on four benchmark datasets show that CITER can reduce LLM calls by up to 30% while preserving high accuracy or improve accuracy by 25% with the same call ratio. Additionally, ablation studies reveal that token-level routing is more flexible and effective than query-level routing, highlighting the benefits of considering long-term routing impacts."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"Refer to weaknesses."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- The paper is well written and easy to follow.\n- The RL-based router training method is novel, and a shortcut to the reward function is proposed to make training easier.\n- Experimental results show that the proposed method can achieve better performance under the same call to LLM."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- While the framework introduces a shortcut for estimating the reward function, the initial training of the token-level router still requires significant computational resources due to the need for reinforcement learning, which can be a barrier for practical implementation.\n- The effectiveness of CITER heavily relies on the accuracy of token importance predictions. If the router fails to accurately assess which tokens are critical, it could lead to suboptimal routing decisions, potentially compromising the quality of the generated outputs. More analytical experiments should be conducted to prove this point.\n- The main experiment in Figure 2 involves a few baselines. If a comparison with the LLM Inference Acceleration method can be added, the effectiveness of the proposed method will be more prominent.\n- Most experimental results show \"% Call to LLM\". I hope to show more intuitive metrics in the experiment, such as the amount of computation (FLOPs) or inference time/speed, which will help readers to have an intuitive feeling."}},"nonreaders":[],"tmdate":1731429162955,"tcdate":1730640245897,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission8785/Reviewer_9kzc"],"signatures":["ICLR.cc/2025/Conference/Submission8785/Reviewer_9kzc"],"forum":"J2FyEVg8HR","number":2,"license":"CC BY 4.0","cdate":1730640245897,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission8785/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731429162955,"domain":"ICLR.cc/2025/Conference","replyto":"J2FyEVg8HR","id":"3oMmEKxX5P","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["collaborative inference","efficient inference","token-level routing","large language model"]},"primary_area":{"value":"foundation or frontier models, including LLMs"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Large language models (LLMs) have achieved remarkable success in natural language processing tasks but suffer from high computational costs during inference, limiting their deployment in latency-constrained applications. To address this issue, we propose a novel \\textbf{C}ollaborative \\textbf{I}nference with \\textbf{T}oken-l\\textbf{E}vel \\textbf{R}outing (CITER) framework that introduces a token-level routing mechanism, enabling efficient collaboration between small and large language models (SLMs \\& LLMs). Specifically, CITER enables routing non-critical tokens to an SLM to reduce computational overhead, while critical tokens are processed by an LLM to maintain generation quality. We formulate the training of the router as a reinforcement learning task, where the router receives rewards based on both the quality of predictions and the inference cost of generation. This allows the router to learn to predict token-level routing scores and make routing decisions based on both the current token and the future impact of its decisions. To further accelerate reward evaluation process, we introduce a shortcut for reward function estimation, significantly reducing the cost of the reward estimation and improving the practicality of our approach. Extensive experiments across four benchmark datasets demonstrate that CITER reduces inference cost while preserving high-quality generation, offering a promising solution for real-time and resource-constrained applications."},"_bibtex":{"value":"@misc{\nzheng2025citer,\ntitle={{CITER}: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing},\nauthor={Wenhao Zheng and Yixiao Chen and Weitong Zhang and Souvik Kundu and Yun Li and Zhengzhong Liu and Eric P. Xing and Hongyi Wang and Huaxiu Yao},\nyear={2025},\nurl={https://openreview.net/forum?id=J2FyEVg8HR}\n}"},"title":{"value":"CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing"},"pdf":{"value":"/pdf/6164f6446b7d19716e02c5b91e4a924bdabe9f04.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"zheng|citer_collaborative_inference_for_efficient_large_language_model_decoding_with_tokenlevel_routing"},"authorids":{"value":["~Wenhao_Zheng4","~Yixiao_Chen2","~Weitong_Zhang2","~Souvik_Kundu2","~Yun_Li7","~Zhengzhong_Liu1","~Eric_Xing1","~Hongyi_Wang1","~Huaxiu_Yao1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Wenhao Zheng","Yixiao Chen","Weitong Zhang","Souvik Kundu","Yun Li","Zhengzhong Liu","Eric P. Xing","Hongyi Wang","Huaxiu Yao"]}},"version":2},{"content":{"summary":{"value":"This paper addresses the problem of target-unaware video generation, where existing models fail to ensure an actor interacts with a specific target object designated by the user. The authors propose a method that feeds a target object's mask as an additional condition and introduces a novel cross-attention loss. During training, this loss forces the attention map of a special text token ([TGT]) to align with the input target mask. This allows the model to connect a text command to a specific spatial location in the scene, thereby generating target-aware interaction videos."},"soundness":{"value":4},"confidence":{"value":5},"questions":{"value":"In the examples shown in the paper, the target object occupies a much smaller area relative to the actor. I am curious about how the model would perform when the target object takes up a significant portion of the frame (e.g., a car very close to the camera as the target object)."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":3},"strengths":{"value":"1. The paper is well-written, and the illustrative view clearly presents the entire pipeline.\n2. The paper defines a clear and practical problem (target-aware video generation), It has excellent application value in multiple fields.\n3. The ablation studies and quantitative analysis are thorough. The results clearly demonstrate the effectiveness of each module in the method and provides a reasonable explanation for the selection of hyperparameters."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Cross-attention loss cannot be considered as one of the innovations. There have already been numerous works in image generation and video generation that design losses at the attention level for fine-tuning models or optimizing generation results.\n2. The amount of data used for training is insufficient. With only 1K+ video clips, the highly adaptable DiT architecture is prone to overfitting. Additionally, the data primarily consists of single-person indoor scenes, so the effectiveness of this method in outdoor or more complex scenarios requires further validation.\n3. The examples demonstrating control over both the source actor and the target object are somewhat limited. It would be helpful to include more instances of interactions between objects (not just robotic arms) to showcase the generalization capability on the source actor."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762917135326,"tcdate":1762002225290,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission4015/Reviewer_AKUd"],"signatures":["ICLR.cc/2026/Conference/Submission4015/Reviewer_AKUd"],"forum":"311AxWM8FU","number":2,"license":"CC BY 4.0","cdate":1762002225290,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission4015/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762917135326,"domain":"ICLR.cc/2026/Conference","replyto":"311AxWM8FU","id":"LiBrqtENgS","forumContent":{"TLDR":{"value":"Our target-aware model generates a video in which an actor accurately interacts with the target, specified with its segmentation mask."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Controllable video diffusion models","Human-scene interaction","Robotics planning"]},"supplementary_material":{"value":"/attachment/7a4d8a24e01dd34a1dcf692bea8162b9b33d526c.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"We present a target-aware video diffusion model that generates videos from an input image, in which an actor interacts with a specified target while performing a desired action. The target is defined by a segmentation mask, and the action is described through a text prompt. Our key motivation is to incorporate target awareness into video generation, enabling actors to perform directed actions on designated objects. This enables video diffusion models to act as motion planners, producing plausible predictions of human-object interactions by leveraging the priors of large-scale video generative models. We build our target-aware model by extending a baseline model to incorporate the target mask as an additional input. To enforce target awareness, we introduce a special token that encodes the target's spatial information within the text prompt. We then fine-tune the model with our curated dataset using an additional cross-attention loss that aligns the cross-attention maps associated with this token with the input target mask. To further improve performance, we selectively apply this loss to the most semantically relevant attention regions and transformer blocks. Experimental results show that our target-aware model outperforms existing solutions in generating videos where actors interact accurately with the specified targets. We further demonstrate its efficacy in two downstream applications: zero-shot 3D HOI motion synthesis with physical plausibility and long-term video content creation."},"_bibtex":{"value":"@inproceedings{\nkim2026targetaware,\ntitle={Target-Aware Video Diffusion Models},\nauthor={Taeksoo Kim and Hanbyul Joo},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=311AxWM8FU}\n}"},"title":{"value":"Target-Aware Video Diffusion Models"},"pdf":{"value":"/pdf/c98ab4dc22c46001be3678bca8a6a1a4b97e01a3.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"kim|targetaware_video_diffusion_models"},"authorids":{"value":["~Taeksoo_Kim2","~Hanbyul_Joo2"]},"authors":{"value":["Taeksoo Kim","Hanbyul Joo"]}},"version":2},{"content":{"venue":{"value":"ICASSP 2022"},"pdf":{"value":"https://ieeexplore.ieee.org/iel7/9745891/9746004/09746537.pdf"},"venueid":{"value":"dblp.org/conf/ICASSP/2022"},"paperhash":{"value":"xu|adjacency_pairsaware_hierarchical_attention_networks_for_dialogue_intent_classification"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Jiabao_Xu:","~Peijie_Huang1","https://dblp.org/search/pid/api?q=author:Youming_Peng:","https://dblp.org/search/pid/api?q=author:Jiande_Ding:","https://dblp.org/search/pid/api?q=author:Boxi_Huang:","https://dblp.org/search/pid/api?q=author:Simin_Huang:"]},"html":{"value":"https://doi.org/10.1109/ICASSP43922.2022.9746537"},"_bibtex":{"value":"@inproceedings{DBLP:conf/icassp/XuHPDHH22,\n  author={Jiabao Xu and Peijie Huang and Youming Peng and Jiande Ding and Boxi Huang and Simin Huang},\n  title={Adjacency Pairs-Aware Hierarchical Attention Networks for Dialogue Intent Classification},\n  year={2022},\n  cdate={1640995200000},\n  pages={7622-7626},\n  url={https://doi.org/10.1109/ICASSP43922.2022.9746537},\n  booktitle={ICASSP},\n  crossref={conf/icassp/2022}\n}\n"},"abstract":{"value":"Dialogue intent classification is a fundamental and essential task in dialogue systems. Although sentence-level and document-level text classification have made dramatic progress in recent years with the help of deep learning technology, dialogue-level classification remains challenging. Dialogue has unique characteristics that distinguish it from other types of text. Dialogue is interactive, with feedback between speakers, and turn-taking. These unique features suggest that model architecture should take dialogue structure into account to learn a better representation. In this paper we propose an Adjacency Pairs-Aware Hierarchical Attention Network (AP-HAN) for dialogue intent classification. A dialogue reconstruction strategy is designed to match the question and answer utterances properly and then make the dialogue to be presented as a sequence of adjacent pairs. Then, the adjacency pairs features are incorporated into the hierarchical attention network. Experimental results on public CCL2018-Task1 corpus show the better performance of the proposed model."},"title":{"value":"Adjacency Pairs-Aware Hierarchical Attention Networks for Dialogue Intent Classification"},"authors":{"value":["Jiabao Xu","Peijie Huang","Youming Peng","Jiande Ding","Boxi Huang","Simin Huang"]}},"tmdate":1718677826156,"pdate":1640995200000,"tcdate":1718677824246,"writers":["~"],"signatures":["~Peijie_Huang1"],"forum":"srjHzZmWfE","license":"CC BY-SA 4.0","number":32135,"cdate":1640995200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1718677826156,"domain":"DBLP.org","id":"srjHzZmWfE","version":2},{"content":{"venue":{"value":"ICLR 2027 Conference Submission"},"keywords":{"value":["Interpretability","Data attribution"]},"supplementary_material":{"value":"/attachment/9774b49e683ec3cc08fe56c77b6c644bab76f38a.zip"},"primary_area":{"value":"other topics in machine learning (i.e., none of the above)"},"abstract":{"value":"A common response to spurious correlations is to weaken them, for instance by balancing the training data so that the shortcut feature no longer predicts the label. We find that in small transformers this can make the robust rule harder to learn. In our tasks, a shortcut feature agrees with the label on a fraction $r$ of training examples and contradicts it on a held-out adversarial split. When the label is the sum parity of an integer sequence and the shortcut is the parity of its largest element, the fraction of two-layer seeds that reach $100\\%$ adversarial accuracy grows from 0\\% at $r{=}0.5$ to 53\\% at $r{=}0.9$, whereas for one-layer models it drops from 33\\% to 0\\%. The benefit depends on how far $r$ is from chance rather than on the direction in which it deviates. On a ternary task, two-layer models generalize in 80--93\\% of seeds both when the shortcut is correlated and when it is anti-correlated with the label, and in no seeds at chance. In that setting they memorize the training set, a failure that disappears with 16$\\times$ more data. All models learn the shortcut-only predictor first, and in the runs that succeed the robust rule emerges while attention to the shortcut token remains high. Finally, models trained briefly with an informative shortcut and then moved to balanced data still generalize in most seeds, which suggests that the shortcut serves as an optimization scaffold rather than working by amplifying the gradient of the misclassified minority."},"_bibtex":{"value":"@inproceedings{\nanonymous2026when,\ntitle={When Data Imbalance Helps: Robust Generalization Through Shortcut Saturation},\nauthor={Anonymous},\nbooktitle={Submitted to The Fifteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=0f0qnvTA2T},\nnote={under review}\n}"},"title":{"value":"When Data Imbalance Helps: Robust Generalization Through Shortcut Saturation"},"pdf":{"value":"/pdf/0c6bece54e090c78632b40072dc98eb20c8ed7ba.pdf"},"venueid":{"value":"ICLR.cc/2027/Conference/Submission"}},"tmdate":1791231571415,"tcdate":1789658418367,"writers":["ICLR.cc/2027/Conference","ICLR.cc/2027/Conference/Submission32588/Authors"],"signatures":["ICLR.cc/2027/Conference/Submission32588/Authors"],"forum":"0f0qnvTA2T","license":"CC BY 4.0","number":32588,"cdate":1789658418367,"readers":["everyone"],"invitations":["ICLR.cc/2027/Conference/-/Submission","ICLR.cc/2027/Conference/-/Edit","ICLR.cc/2027/Conference/-/Submission_Change_Before_Bidding","ICLR.cc/2027/Conference/Submission32588/-/Full_Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Reviewing"],"mdate":1791231571415,"odate":1791031328286,"domain":"ICLR.cc/2027/Conference","id":"0f0qnvTA2T","version":2},{"content":{"summary":{"value":"The authors propose MrDPO to enhance the model's video captioning capability, and leverage this capability to generate high-quality video caption data, which is then used to train a general-purpose model, thereby improving its general video question-answering ability."},"soundness":{"value":2},"confidence":{"value":5},"questions":{"value":"1. What is the design rationale behind the Video Caption Benchmark? Why did you develop a benchmark that evaluates video caption quality based on the extraction of various types of events? Is GPT-3.5 truly capable of reliably extracting diverse events from video captions?\n2. Why does improving video captioning capability lead to enhanced video question-answering performance?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. video-SALMON 2 demonstrates strong video captioning capability, surpassing many well-known models such as Gemini 1.5, VideoLLaMA3, and Qwen2.5-VL.\n2. video-SALMON 2 provides a self-constructed caption benchmark, which helps advance the video captioning capabilities of other models.\n3. video-SALMON 2 exhibits powerful audio-visual question-answering performance, achieving strong results on VideoMME."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The paper’s first main contribution MrDPO (a DPO variant that updates the reference model during training to improve stability) is not novel; a similar approach was already proposed in TR-DPO [1]. Moreover, the paper’s primary focus is on enhancing video captioning performance, yet the design of MrDPO is unrelated to video captioning and appears to be a generic DPO method. The motivation behind MrDPO and the mechanism by which it improves video captioning capability remain unclear.\n\n2. The second main contribution (using video captioning to enhance video question-answering performance) is also problematic. On one hand, this idea is already well established in the field and has been thoroughly validated in prior work. On the other hand, the authors fail to explain why or how improved video captioning capability leads to better video QA performance, leaving the underlying rationale unaddressed.\n\n[1] Learn Your Reference Model for Real Good Alignment"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762926023082,"tcdate":1761832763997,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission15791/Reviewer_N9PM"],"signatures":["ICLR.cc/2026/Conference/Submission15791/Reviewer_N9PM"],"forum":"O2tca5tExP","number":3,"license":"CC BY 4.0","cdate":1761832763997,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission15791/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762926023082,"domain":"ICLR.cc/2026/Conference","replyto":"O2tca5tExP","id":"LWWnVe9Dq9","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["video captioning","audio-visual LLM","multi-round DPO","video QA"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"We present video-SALMONN 2, a family of audio-visual large language models that set new state-of-the-art (SOTA) results in video description and question answering (QA). Our core contribution is multi-round direct preference optimisation (MrDPO), paired with a caption-quality objective that jointly rewards completeness and factual accuracy. Unlike standard DPO with a fixed reference policy, MrDPO periodically refreshes the reference by bootstrapping from a newly re-initialised lightweight adapter trained on the latest preferences, avoiding reference staleness and enabling continual improvement. This strategy produces captions that are consistently more detailed and accurate than those from proprietary systems such as GPT-4o and Gemini-1.5 Pro. We further distil these gains by using our model to generate a high-quality video-caption corpus for supervised fine-tuning of new models, transferring benefits beyond captioning to strong performance on complex video-QA tasks. Across widely used audio-visual and visual-only understanding benchmarks (including Video-MME, WorldSense, AVUT, Video-Holmes, DailyOmni, MLVU, and LVBench), our 3B and 7B models achieve SOTA results at comparable scales, while the 72B model surpasses all other open-source systems. Our source code, models, and data will be released."},"_bibtex":{"value":"@misc{\ntang2026videosalmonn,\ntitle={video-{SALMONN} 2: Caption-Enhanced Audio-Visual Large Language Models},\nauthor={Changli Tang and Yixuan Li and Yudong Yang and Jimin Zhuang and Guangzhi Sun and Wei Li and Zejun MA and Chao Zhang},\nyear={2026},\nurl={https://openreview.net/forum?id=O2tca5tExP}\n}"},"title":{"value":"video-SALMONN 2: Caption-Enhanced Audio-Visual Large Language Models"},"pdf":{"value":"/pdf/eedab7f0de7a1b39f06aade78ee3aa8b05999dc2.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"tang|videosalmonn_2_captionenhanced_audiovisual_large_language_models"},"authorids":{"value":["~Changli_Tang1","~Yixuan_Li13","~Yudong_Yang1","~Jimin_Zhuang1","~Guangzhi_Sun1","~Wei_Li78","~Zejun_MA1","~Chao_Zhang20"]},"authors":{"value":["Changli Tang","Yixuan Li","Yudong Yang","Jimin Zhuang","Guangzhi Sun","Wei Li","Zejun MA","Chao Zhang"]}},"version":2},{"content":{"summary":{"value":"This paper introduces ActAnywhere, a video diffusion model designed to generate video backgrounds that adapt to the foreground subject's motion. By utilizing a sequence of foreground subject segmentation and a background image, the model produces realistic videos with coherent foreground-background interactions. Experiments on a large-scale dataset demonstrate the model's effectiveness, outperforming existing methods in generating realistic and dynamic backgrounds."},"soundness":{"value":4},"confidence":{"value":4},"questions":{"value":"Q1: I wonder about the impact of the quality of the foreground segmentation. It seems that the model relies heavily on the quality of the foreground segmentation.\n\nQ2:  How scalable is your model for generating longer video sequences? Have you tested its performance in generating videos of varying lengths, and what are the results?"},"rating":{"value":7},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":3},"strengths":{"value":"S1: The paper introduces a novel problem of automated subject-aware video background generation.\n\nS2: The methodology shows improvements in generating coherent videos with realistic subject-background interactions.\n\nS4: The contributions are significant, particularly for applications in the movie industry and visual effects.\n\nS4: The paper is comprehensive and well-written."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"W1: The paper lacks sufficient comparison with a broader range of existing methods, particularly those leveraging recent advancements in video generation and editing (though they are not for background generation, they can do).\n\nW2: It seems that the model relies heavily on the quality of the foreground video segmentation masks."},"limitations":{"value":"I think there is no potential negative societal impact and the authors have addressed the limitations."}},"nonreaders":[],"tmdate":1730880161009,"tcdate":1720679800853,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission21087/Reviewer_gAcX"],"signatures":["NeurIPS.cc/2024/Conference/Submission21087/Reviewer_gAcX"],"forum":"ntlFREw59A","number":2,"license":"CC BY 4.0","cdate":1720679800853,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission21087/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730880161009,"domain":"NeurIPS.cc/2024/Conference","replyto":"ntlFREw59A","id":"LCYWLPCI5k","forumContent":{"venue":{"value":"NeurIPS 2024 poster"},"keywords":{"value":["Video Background Generation","Video Generation","Video Synthesis","Video Editing"]},"supplementary_material":{"value":"/attachment/25ef7faf3ede09b16187f03b68bf377e404cb2fd.zip"},"primary_area":{"value":"machine_vision"},"abstract":{"value":"We study a novel problem to automatically generate video background that tailors to foreground subject motion. It is an important problem for the movie industry and visual effects community, which traditionally requires tedious manual efforts to solve. To this end, we propose ActAnywhere, a video diffusion model that takes as input a sequence of foreground subject segmentation and an image of a novel background and generates a video of the subject interacting in this background. We train our model on a large-scale dataset of 2.4M videos of human-scene interactions. Through extensive evaluation, we show that our model produces videos with realistic foreground-background interaction while strictly following the guidance of the condition image. Our model generalizes to diverse scenarios including non-human subjects, gaming and animation clips, as well as videos with multiple moving subjects. Both quantitative and qualitative comparisons demonstrate that our model significantly outperforms existing methods, which fail to accomplish the studied task. Please visit our project webpage at https://actanywhere.github.io."},"_bibtex":{"value":"@inproceedings{\npan2024actanywhere,\ntitle={ActAnywhere: Subject-Aware Video Background Generation},\nauthor={Boxiao Pan and Zhan Xu and Chun-Hao Paul Huang and Krishna Kumar Singh and Yang Zhou and Leonidas Guibas and Jimei Yang},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=ntlFREw59A}\n}"},"title":{"value":"ActAnywhere: Subject-Aware Video Background Generation"},"pdf":{"value":"/pdf/5bb1cb449e096181fbbdf1dc3cac36beee1a8cbe.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"pan|actanywhere_subjectaware_video_background_generation"},"authorids":{"value":["~Boxiao_Pan1","~Zhan_Xu1","~Chun-Hao_Paul_Huang2","~Krishna_Kumar_Singh4","~Yang_Zhou10","~Leonidas_Guibas1","~Jimei_Yang1"]},"authors":{"value":["Boxiao Pan","Zhan Xu","Chun-Hao Paul Huang","Krishna Kumar Singh","Yang Zhou","Leonidas Guibas","Jimei Yang"]}},"version":2},{"content":{"summary":{"value":"The paper proposes FOCUS (Frame-Optimistic Confidence Upper-bound Selection), a training-free, model-agnostic keyframe selection method for long video understanding with multimodal LLMs (MLLMs). To address the prohibitive token cost of processing all frames, FOCUS formulates keyframe selection as a combinatorial pure-exploration (CPE) problem in a multi-armed bandit setting, where short temporal clips are treated as arms. It employs a two-stage exploration–exploitation strategy: first using optimistic confidence upper bounds (UCB) to identify promising clips, then selecting top-scoring frames within them. Evaluated on LongVideoBench and Video-MME, FOCUS processes <2% of frames yet achieves consistent gains over uniform sampling and SOTA retrieval-based methods, e.g., +11.9% accuracy on videos >20 minutes. The method is simple, theoretically grounded, and plug-and-play compatible with existing MLLMs"},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"- FOCUS first partition the timeline into $M$ non-overlapping fixed-length clips. Does this destroy the spatiotemporal consistency of the video? Which makes it difficult to capture continuous segments with high information density?\n\n- As discussed in the Limitation section, FOCUS assumes that each frame of the video is independent, which is completely contrary to the nature of the video. Does this limit FOCUS’s applicability to short videos?\n\n- Table 1 lacks the experimental results of other frame selection methods, such as AKS and Q-Frame, based on the same baseline.\n\n- The paper provides a visualization of FOCUS's superiority over uniform sampling in Figure 3. It is meaningful to include the negative aspects of FOCUS, which helps readers better understand its limitations."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"The paper introduces FOCUS, a training-free keyframe selection method that formulates the task as a combinatorial pure-exploration bandit problem. It partitions the video into clips (arms), uses optimistic confidence upper bounds to identify informative regions with minimal sampling, and then selects top frames within those clips. This two-stage exploration-exploitation strategy enables MLLMs to achieve strong performance on long-video QA benchmarks while processing fewer than 2% of frames. It significantly outperforms uniform sampling and existing retrieval-based methods, especially on videos over 20 minutes."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The concept of ACF and the calculation of $r_t$ in Figure 1 require further explanation, which will help readers better understand the motivation of the paper. \n\n- This pre-filtering process before keyframe selection undermines the goal of identifying the most informative keyframes from all frames. The findings are interesting, but it is not certain whether the two-stage ARM selection proposed in the article will fall into the same limitations.\n\n- Lack of experimental results on MLVU, a commonly used long video understanding benchmark.\n\n- Minor Weaknesses\n  - Line 36: multimodal LLMs (MLLMs) -> multimodal large language models (MLLMs)\n  - Line 107: multimodal LLMs -> MLLMs\n\n[1] Zhou J, Shu Y, Zhao B, et al. Mlvu: A comprehensive benchmark for multi-task long video understanding[J]. arXiv e-prints, 2024: arXiv: 2406.04264."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762927851200,"tcdate":1761132933540,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission18064/Reviewer_aTVY"],"signatures":["ICLR.cc/2026/Conference/Submission18064/Reviewer_aTVY"],"forum":"1OQKqLFcbB","number":1,"license":"CC BY 4.0","cdate":1761132933540,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission18064/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762927851200,"domain":"ICLR.cc/2026/Conference","replyto":"1OQKqLFcbB","id":"BqRvEKngel","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Keyframe Selection","Multimodal large language models","Long Video Understanding","Combinatorial Pure-exploration"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Multimodal large language models (MLLMs) represent images and video frames as visual tokens. Scaling from single images to hour-long videos, however, inflates the token budget far beyond practical limits. Popular pipelines therefore either uniformly subsample or apply keyframe selection with retrieval-style scoring using smaller vision-language models. However, these keyframe selection methods still rely on pre-filtering before selection to reduce the inference cost and can miss the most informative moments.\n\nWe propose FOCUS, Frame-Optimistic Confidence Upper-bound Selection, a training-free, model-agnostic keyframe selection module that selects query-relevant frames under a strict token budget. FOCUS formulates keyframe selection as a combinatorial pure-exploration (CPE) problem in multi-armed bandits: it treats short temporal clips as arms, and uses empirical means and Bernstein confidence radius to identify informative regions while preserving exploration of uncertain areas. The resulting two-stage exploration-exploitation procedure reduces from a sequential policy with theoretical guarantees, first identifying high-value temporal regions, then selecting top-scoring frames within each region. Extensive experiments across four long-video question-answering benchmarks and four popular MLLMs demonstrate that FOCUS delivers substantial accuracy improvements while processing less than 2% of video frames. For videos longer than 20 minutes, it achieves an 11.9% gain in accuracy on LongVideoBench, demonstrating its effectiveness as a keyframe selection method and providing a simple and general solution for scalable long-video understanding with MLLMs."},"_bibtex":{"value":"@inproceedings{\nzhu2026focus,\ntitle={{FOCUS}: Efficient Keyframe Selection for Long Video Understanding},\nauthor={Zirui Zhu and Hailun Xu and Yang Luo and Yong Liu and Kanchan Sarkar and Zhenheng Yang and Yang You},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=1OQKqLFcbB}\n}"},"title":{"value":"FOCUS: Efficient Keyframe Selection for Long Video Understanding"},"pdf":{"value":"/pdf/1a4f41e9567a6de0b3b4cef826420a76cec302fb.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"zhu|focus_efficient_keyframe_selection_for_long_video_understanding"},"authorids":{"value":["~Zirui_Zhu2","~Hailun_Xu1","~Yang_Luo4","~Yong_Liu13","~Kanchan_Sarkar1","~Zhenheng_Yang3","~Yang_You1"]},"authors":{"value":["Zirui Zhu","Hailun Xu","Yang Luo","Yong Liu","Kanchan Sarkar","Zhenheng Yang","Yang You"]}},"version":2},{"content":{"TLDR":{"value":"We synthesize shortcut-resistant deep-search tasks that improve search-agent training."},"venue":{"value":"ICLR 2027 Conference Submission"},"keywords":{"value":["deep-search agents","question synthesis","shortcut mitigation","agentic training data","information seeking"]},"primary_area":{"value":"foundation or frontier models, including LLMs"},"abstract":{"value":"Training deep-search agents requires verifiable questions that demand\nsubstantial evidence acquisition, yet structurally complex tasks may still\nadmit cheap identifying routes when clues are overly selective, co-covered\nby the same evidence, directly searchable, or resolved from solver priors.\nWe formalize this gap with a shortcut-aware difficulty framework that\ndistinguishes intended construction complexity from identifying routes\navailable through retrieval. Guided by this framework, we introduce FORT, a Framework of Shortcut-Resistant Training-Data Synthesis,\nwhich targets four shortcut risks across entity selection, evidence graph\nconstruction, question formulation, and adversarial refinement. We evaluate\nrealized search behavior with solver-conditioned trajectory diagnostics rather\nthan graph complexity or search length alone. Compared with existing\nopen-source deep-search datasets, FORT shows later answer exposure,\nhigher retrieval effort, lower single-clue selectivity, greater evidence\ndispersion, and higher realized pre-answer dependency costs. Using the\nresulting trajectories, we train FORT-Searcher with supervised fine-tuning\nalone. Across model scales, FORT-Searcher consistently improves over its backbones, achieving the strongest overall performance among\nsimilarly sized agents."},"_bibtex":{"value":"@inproceedings{\nanonymous2026fortsearcher,\ntitle={{FORT}-Searcher: Synthesizing Shortcut-Resistant Search Tasks for Training Deep Search Agents},\nauthor={Anonymous},\nbooktitle={Submitted to The Fifteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=HiSOa9i7C2},\nnote={under review}\n}"},"title":{"value":"FORT-Searcher: Synthesizing Shortcut-Resistant Search Tasks for Training Deep Search Agents"},"pdf":{"value":"/pdf/c77e76417d8a8080d69817aaefffa76c003798d0.pdf"},"venueid":{"value":"ICLR.cc/2027/Conference/Submission"}},"tmdate":1791228804385,"tcdate":1789400860618,"writers":["ICLR.cc/2027/Conference","ICLR.cc/2027/Conference/Submission18621/Authors"],"signatures":["ICLR.cc/2027/Conference/Submission18621/Authors"],"forum":"HiSOa9i7C2","license":"CC BY 4.0","number":18621,"cdate":1789400860618,"readers":["everyone"],"invitations":["ICLR.cc/2027/Conference/-/Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Bidding","ICLR.cc/2027/Conference/Submission18621/-/Full_Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Reviewing","ICLR.cc/2027/Conference/-/Edit"],"mdate":1791228804385,"odate":1791031328286,"domain":"ICLR.cc/2027/Conference","id":"HiSOa9i7C2","version":2},{"content":{"summary":{"value":"This paper proposes **VoCap**, a unified framework for video object understanding that can generate spatio-temporal masks and natural language descriptions from any prompt.\nThe writing is clear and well-structured, with a coherent and novel task formulation supported by a logically developed methodology.\nExperimental results demonstrate state-of-the-art performance, confirming the effectiveness and potential of the proposed approach for video segmentation and captioning."},"soundness":{"value":2},"confidence":{"value":5},"questions":{"value":"Missing comparisons with DAM, a relevant and competitive method in multimodal video understanding, undermines the fairness and completeness of the evaluation.\n\nLacks visual comparisons with other methods."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"The paper explores an interesting direction by attempting to unify video object segmentation and captioning within a single framework.\n\nThe idea of leveraging different input modalities (text, box, mask) is conceptually appealing and potentially useful for future multimodal understanding tasks.\n\nThe paper is clearly written and easy to follow, with well-organized structure and visual illustrations."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The proposed VoCap framework mainly stacks existing techniques (SAM2 for segmentation and BLIP2-style text decoding) with minimal methodological innovation.\n\nThe model design lacks substantial novelty or clear insight into how segmentation and captioning are effectively integrated beyond simple module combination.\n\nThe experimental validation is insufficient and somewhat superficial; it primarily reports improvements on internal benchmarks without solid comparisons to recent or stronger baselines, like DAM[1]\n\nThe work does not provide a convincing analysis or explanation to justify the claimed synergy between segmentation and captioning modules.\n\n[1] Lian L, Ding Y, Ge Y, et al. Describe anything: Detailed localized image and video captioning[J]. arXiv preprint arXiv:2504.16072, 2025."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762926862803,"tcdate":1762017103299,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission16843/Reviewer_MS3b"],"signatures":["ICLR.cc/2026/Conference/Submission16843/Reviewer_MS3b"],"forum":"uLwuPbVegY","number":5,"license":"CC BY 4.0","cdate":1762017103299,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission16843/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762926862803,"domain":"ICLR.cc/2026/Conference","replyto":"uLwuPbVegY","id":"XgRhChWQ9O","forumContent":{"TLDR":{"value":"Video segmentation with location"},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["video segmentation; VOS; referring expression segmentation"]},"supplementary_material":{"value":"/attachment/0f72cf865474d633a757c245846dfdb44e1cdedb.pdf"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Understanding objects in videos in terms of fine-grained localization masks and\ndetailed semantic properties is a fundamental task in video understanding. In this\npaper, we propose VoCap, a flexible video model that consumes a video and a\nprompt of various modalities (text, box or mask), and produces a spatio-temporal\nmasklet with a corresponding object-centric caption. As such our model addresses\nsimultaneously the tasks of promptable video object segmentation, referring expression segmentation, and object captioning. Since obtaining data for this task is tedious and expensive, we propose to annotate an existing large-scale segmentation\ndataset (SAV) with pseudo object captions. We do so by preprocessing videos with\ntheir ground-truth masks to highlight the object of interest and feed this to a large\nVision Language Model (VLM). For an unbiased evaluation, we collect manual\nannotations on the validation set. We call the resulting dataset SAV-Caption. We\ntrain our VoCap model at scale on a SAV-Caption together with a mix of other\nimage and video datasets. Our model yields state-of-the-art results on referring\nexpression video object segmentation, is competitive on semi-supervised video\nobject segmentation, and establishes a benchmark for video object captioning. Our\ndataset will be made available."},"_bibtex":{"value":"@misc{\nuijlings2026vocap,\ntitle={VoCap: video object captioning and segmentation from any prompt},\nauthor={Jasper Uijlings and Xingyi Zhou and Xiuye Gu and Arsha Nagrani and Anurag Arnab and Alireza Fathi and David A Ross and Cordelia Schmid},\nyear={2026},\nurl={https://openreview.net/forum?id=uLwuPbVegY}\n}"},"title":{"value":"VoCap: video object captioning and segmentation from any prompt"},"pdf":{"value":"/pdf/fe82b3884d3a8d29461e043146cb4a7060c2e2c9.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"uijlings|vocap_video_object_captioning_and_segmentation_from_any_prompt"},"authorids":{"value":["~Jasper_Uijlings3","~Xingyi_Zhou2","~Xiuye_Gu1","~Arsha_Nagrani2","~Anurag_Arnab1","~Alireza_Fathi1","~David_A_Ross1","~Cordelia_Schmid1"]},"authors":{"value":["Jasper Uijlings","Xingyi Zhou","Xiuye Gu","Arsha Nagrani","Anurag Arnab","Alireza Fathi","David A Ross","Cordelia Schmid"]}},"version":2},{"content":{"venue":{"value":"NeurIPS 2025 poster"},"TLDR":{"value":"This paper tackle the five core issues that held shortcut models back."},"keywords":{"value":["generative models"]},"supplementary_material":{"value":"/attachment/e810cabee3368bd799ec22e4ad3f9237fab030c6.zip"},"primary_area":{"value":"applications"},"abstract":{"value":"Shortcut models represent a promising, non-adversarial paradigm for generative modeling, uniquely supporting one-step, few-step, and multi-step sampling from a single trained network. However, their widespread adoption has been stymied by critical performance bottlenecks. This paper tackles the five core issues that held shortcut models back: (1) the hidden flaw of compounding guidance, which we are the first to formalize, causing severe image artifacts; (2) inflexible fixed guidance that restricts inference-time control; (3) a pervasive frequency bias driven by a reliance on low-level distances in the direct domain, which biases reconstructions toward low frequencies; (4) divergent self-consistency arising from a conflict with EMA training; and (5) curvy flow trajectories that impede convergence. To address these challenges, we introduce iSM, a unified training framework that systematically resolves each limitation. Our framework is built on four key improvements: Intrinsic Guidance provides explicit, dynamic control over guidance strength, resolving both compounding guidance and inflexibility. A Multi-Level Wavelet Loss mitigates frequency bias to restore high-frequency details. Scaling Optimal Transport (sOT) reduces training variance and learns straighter, more stable generative paths. Finally, a Twin EMA strategy reconciles training stability with self-consistency. Extensive experiments on ImageNet 256x256 demonstrate that our approach yields substantial FID improvements over baseline shortcut models across one-step, few-step, and multi-step generation, making shortcut models a viable and competitive class of generative models."},"_bibtex":{"value":"@inproceedings{\nnguyen2025improved,\ntitle={Improved Training Technique for Shortcut Models},\nauthor={Anh Nguyen and Viet Van Nguyen and Duc Vu and Trung Tuan Dao and Chi Tran and Toan Tran and Anh Tuan Tran},\nbooktitle={The Thirty-ninth Annual Conference on Neural Information Processing Systems},\nyear={2025},\nurl={https://openreview.net/forum?id=ovUuNzZZbK}\n}"},"title":{"value":"Improved Training Technique for Shortcut Models"},"pdf":{"value":"/pdf/03787a580b05e99ae3d231935ae265d3b7b1eea0.pdf"},"venueid":{"value":"NeurIPS.cc/2025/Conference"},"paperhash":{"value":"nguyen|improved_training_technique_for_shortcut_models"},"authorids":{"value":["~Anh_Nguyen9","~Viet_Van_Nguyen1","~Duc_Vu1","~Trung_Tuan_Dao1","~Chi_Tran1","~Toan_Tran1","~Anh_Tuan_Tran2"]},"authors":{"value":["Anh Nguyen","Viet Van Nguyen","Duc Vu","Trung Tuan Dao","Chi Tran","Toan Tran","Anh Tuan Tran"]}},"tmdate":1783627358652,"pdate":1758217072854,"tcdate":1746877289463,"writers":["NeurIPS.cc/2025/Conference","NeurIPS.cc/2025/Conference/Submission15954/Authors"],"signatures":["NeurIPS.cc/2025/Conference/Submission15954/Authors"],"forum":"ovUuNzZZbK","license":"CC BY 4.0","number":15954,"cdate":1746877289463,"readers":["everyone"],"invitations":["NeurIPS.cc/2025/Conference/-/Submission","NeurIPS.cc/2025/Conference/-/Post_Submission","NeurIPS.cc/2025/Conference/Submission15954/-/Full_Submission","NeurIPS.cc/2025/Conference/Submission15954/-/Supplementary_Material","NeurIPS.cc/2025/Conference/-/Edit","NeurIPS.cc/2025/Conference/Submission15954/-/Camera_Ready_Revision"],"mdate":1783627358652,"odate":1761704925453,"domain":"NeurIPS.cc/2025/Conference","id":"ovUuNzZZbK","version":2},{"content":{"summary":{"value":"The paper ZoomV, a query-aware temporal zoom-in framework designed for efficient and accurate long video understanding. It retrieves relevant events and their associated temporal windows as candidates, and select higher-confidence temporal windows as the LVLM's final input to provide the answer. It conducts experiments on temporal grounding benchmarks as well as long video understanding benchmarks to demonstrate the effectiveness of the proposed method."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"1. The improvement on VideoMME is very limited, for example only 0.1 on Qwen2.5-VL and no improvement on MVBench, while achieves 11.3 on LVBench. It seems that the method is not generalized to all video benchmarks. Could you explain why it achieves improvement by large margin on LVBench, but not effective to VideoMME.\n2. Both VideoMME and LVBench are long-video understanding benchmarks that contain thousands of frames. However, the proposed method only samples 64 frames, which results in a substantial loss of visual details throughout the video and may prevent accurate grounding on evidence frames. Have the authors experimented with increasing the number of sampled frames? This could better demonstrate the effectiveness of the proposed approach.\n3. The evaluation involves recursively exploring video frames, meaning that the total number of processed frames exceeds 64. How many frames are explored on average?\n4. Considering the increased number of processed frames and the computational overhead, is it entirely fair to compare the results with the base model under a 64-frame input setting? A fairer comparison would be against the model’s officially reported best performance, for example, Qwen2.5-VL achieves 70.2 on MLVU and 65.1 on VideoMME.\n5. How about the inference efficiency on long video benchmark like VideoMME compared with base models?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The paper is clearly written and easy to follow.\n2. The proposed approach is reasonable and methodologically sound.\n3. Experiments are conducted on both temporal grounding and long video understanding benchmarks to demonstrate the effectiveness of the method."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. One of the main contributions claimed by the paper is the confidence-based temporal grounding approach. However, this concept has already been introduced in TimeSearch [1]. Therefore, it cannot be regarded as a novel contribution of this work. Moreover, the authors have not properly cited TimeSearch to acknowledge prior work.\n2. The technical novelty appears limited, as the main modification involves adding textual timestamps to each frame embedding, which was already employed in models such as Eagle2.5 [2] ([Eagle2.5 implementation](https://github.com/NVlabs/Eagle/blob/047e51070e8976978376cb828f7af92323c0f8ef/Eagle2_5/deployment/inference.py#L85))\n3. The method seems not consistently effective to all video benchmarks, and the improvement is very trivial in several benchmarks such as MVBench and VideoMME.\n4. Since the paper positions its approach as an agent-style method, it should also include comparisons with recent video agent frameworks such as Video-RAG [3].\n5. Given that the model is fine-tuned on a recent backbone (Qwen2.5-VL), which already exhibits strong temporal grounding capabilities, it would be more convincing to compare against recent models fine-tuned on the same base, such as VideoChat-R1 [4] and Time-R1 [5].\n6. The hierachical search is not a novel idea, which has already explored in TimeSearch [1], UniTime [6] and VideoChat-R1.5 [7]. They should be discussed in the related work and experiments.\n7. Since the proposed methods are fundamentally based on temporal grounding, the paper should include a discussion of the temporal grounding task in the related work section.\n8. The paper emphasizes efficiency in its title; however, it does not provide a comprehensive analysis of efficiency compared to the base models.\n\n[1] TimeSearch: Hierarchical Video Search with Spotlight and Reflection for Human-like Long Video Understanding, arXiv:2504.01407.\n\n[2] Eagle 2.5: Boosting Long-Context Post-Training for Frontier Vision-Language Models, arXiv:2504.15271.\n\n[3] Video-RAG: Visually-aligned Retrieval-Augmented Long Video Comprehension, NeurIPS 2025.\n\n[4] VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning, arXiv:2504.06958.\n\n[5] Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding, NeurIPS 2025.\n\n[6] Universal Video Temporal Grounding with Generative Multi-modal Large Language Models, arXiv:2506.18883.\n\n[7] VideoChat-R1.5: Visual Test-Time Scaling to Reinforce Multimodal Reasoning by Iterative Perception, arXiv:2509.21100."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762916694162,"tcdate":1761656195818,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission3376/Reviewer_F8Az"],"signatures":["ICLR.cc/2026/Conference/Submission3376/Reviewer_F8Az"],"forum":"Spg6FCsmyc","number":3,"license":"CC BY 4.0","cdate":1761656195818,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission3376/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762916694162,"domain":"ICLR.cc/2026/Conference","replyto":"Spg6FCsmyc","id":"mCy2KdKNLK","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Long video understanding","Zoom-in search","Large video language model"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Long video understanding poses a fundamental challenge for large video-language models (LVLMs) due to the overwhelming number of frames and the risk of losing essential context through naive downsampling.\nInspired by the way humans watch videos on mobile phones, constantly zooming in on frames of interest, we propose $\\textbf{ZoomV}$, a query-aware temporal zoom-in framework designed for efficient and accurate long video understanding. \nSpecifically, ZoomV operates in three stages: (1) Temporal interests grounding: guided by the query, ZoomV retrieves relevant events and their associated temporal windows as candidates. (2) Event interests spotlighting: within pools of candidate windows, each window is scored through the model itself reflection and filtered accordingly, where higher-confidence windows are more representative. (3) Compact representation: the selected events are encoded and temporally downsampled to preserve critical semantics while significantly reducing redundancy.\nExtensive experiments demonstrate that ZoomV substantially outperforms prior $\\textbf{video-agent–style}$ approaches. On temporal grounding, ZoomV unlocks the latent capability of LVLMs, achieving an $\\textbf{11.8\\\\%}$ $\\textbf{mIoU}$ gain on Charades-STA. Remarkably, ZoomV further boosts accuracy on LVBench by $\\textbf{9.7\\\\%}$, underscoring its effectiveness on long-video benchmarks."},"_bibtex":{"value":"@misc{\npan2026zoomv,\ntitle={ZoomV: Temporal Zoom-in for Efficient Long Video Understanding},\nauthor={Junwen Pan and Yuan Zhang and Rui Zhang and Xin Wan and Qizhe Zhang and Ming Lu and Shanghang Zhang and Qi She},\nyear={2026},\nurl={https://openreview.net/forum?id=Spg6FCsmyc}\n}"},"title":{"value":"ZoomV: Temporal Zoom-in for Efficient Long Video Understanding"},"pdf":{"value":"/pdf/2087940d526be508892893c64a3a7f278789b691.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"pan|zoomv_temporal_zoomin_for_efficient_long_video_understanding"},"authorids":{"value":["~Junwen_Pan1","~Yuan_Zhang20","~Rui_Zhang51","~Xin_Wan1","~Qizhe_Zhang1","~Ming_Lu2","~Shanghang_Zhang4","~Qi_She1"]},"authors":{"value":["Junwen Pan","Yuan Zhang","Rui Zhang","Xin Wan","Qizhe Zhang","Ming Lu","Shanghang Zhang","Qi She"]}},"version":2},{"content":{"summary":{"value":"The authors develop a method, Modality Composition Awareness (MCA), for mitigating the modality shortcut problem often incurred when primarily relying on contrastive learning based methods for late-fusion and MLLM-adaptation approaches to universal/cross-modal multimodal retrieval. Specifically, to encourage having the joint model learn a robust composed representation (and avoid having a single dominant modality), the authors propose a preference loss that: (1) enforces the embeddings of a multimodal composition to be more discriminative that any of its unimodal counterparts (lines 83-84 -- \"Modality Composition Preference\" (MCP)); and (2) a composition regularization objective rewards consistency between the composite embedding from the unified encoder and a composite prototype constructed from the unimodal embeddings (\"Modality Composition Regularization\" (MCR)). Methodologically, MCP is performed by adding a loss function similar to 'standard' contrastive loss that pulls the composed embedding closer to the paired item than the unimodal similarities (Equation 4) and MCR is performed by another contrastive loss objective where the composed embedding is pulled closer to its mixed prototype (gated fusion works better than mean pooling and multimodal factorized bilinear pooling (MFB) in Table 2) than other in-batch negatives. These three losses (i.e., 'standard' CL, MCP, MCR) are added together as a linear function. Experiments are conducted on the {cross-modal, composed} task parts of MMEB, focusing on the distinction between in-domain and out-of-domain (OOD) performance differences where the OVEN, FashionIQ, EDIS are the OOD datasets and additional datasets {MSCOCO-Grounding, Visual7W-Pointing, RefCOCO, RefCOCO-Matching} are used for zero-shot grounding settings. The core empirical conclusion is that has minimal effect on the in-domain cases, but demonstrates improvements on OOD and ZS-Grounding settings (Table 4) trained used a Qwen-VL-2B-Instruct model as a backbone and reporting results as accuracy@1."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"Associated with the weakness, questions include:\n- How does this performance compare to other MLLM adaptation methods for cross-modal retrieval? Would you be able to hypothesize/explain any differences -- in particular within the context of modality shortcuts?\n- How does MCA perform on other modality pairs? \n- Any discussion regarding hyperparameters would be useful (the discussion in A.3 is limited in my opinion)."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"Strengths of this work include:\n- The core idea is conceptually appealing for mitigating the modality shortcut problem in multimodal retrieval.\n- The empirical performance on the performed experiments consistently demonstrate that there is minimal loss for in-domain settings and improvements in OOD and ZS-Grounding settings across multiple datasets.\n- The 'imbalanced' modality quality experiments and associated observations are interesting from the perspective of modality shortcut issues.\n- The motivation/introduction section (Section 1 + Figure 1, Figure 2) and methods section (Section 3) is well-written and easy to understand."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Weaknesses of this work include:\n- There is far too much of the experimental details in the Appendices; the paper basically cannot be evaluated without reading the appendices. While I understand there are space constraints, having the main results in the Appendices also make the analysis and interpretation difficult to follow and validate the main claims. If I need to choose (even though I actually think the main paper can make room for both), section 4.1.1 and section 4.3 are less important than Appendix A5 and parts of A2, A3 & A4.\n- In Table 4, the only baseline models provided are dual-encoder approaches and 'standard' contrastive learning. However, MCA is a MLLM-adaptation method (lines 140-143). Thus, I think comparisons are needed to other MLLM-adaptation approaches (unless there is a good reason that I am missing). For cases where other systems perform better/worse, there should be some discussion to validate that MCA is better than all of the other loss functions that researchers augment 'vanilla CL' with.\n- It is claimed that this method applies beyond image-text, which is true mathematically, but not validated when it isn't that difficult to do so with MMEB-2 (and more parts of MMEB). I suppose the OOD/ZS issues would have to be considered, but this would be valuable.\n- Some guidance regarding setting the hyperparameters would be useful as there is performance variance across the loss function scaling factors in Table 4."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762922964618,"tcdate":1762714280967,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission11968/Reviewer_YQUt"],"signatures":["ICLR.cc/2026/Conference/Submission11968/Reviewer_YQUt"],"forum":"EzBRI0Llxk","number":4,"license":"CC BY 4.0","cdate":1762714280967,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission11968/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762922964618,"domain":"ICLR.cc/2026/Conference","replyto":"EzBRI0Llxk","id":"gD4lFdVLcZ","forumContent":{"TLDR":{"value":"Mitigating the modality shortcut problem in MLLM-based multimodal retrieval with modal composition awareness."},"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Multimodal Retrieval","Modality Shortcut","MLLMs","Modality Composition"]},"primary_area":{"value":"unsupervised, self-supervised, semi-supervised, and supervised representation learning"},"abstract":{"value":"Multimodal retrieval, which seeks to retrieve relevant content across modalities such as text or image, supports applications from AI search to contents production. Despite the success of separate-encoder approaches like CLIP align modality-specific embeddings with contrastive learning, recent multimodal large language models (MLLMs) enable a unified encoder that directly processes composed inputs. While flexible and advanced, we identify that unified encoders trained with conventional contrastive learning are prone to learn modality shortcut, leading to poor robustness under distribution shifts. We propose a modality composition awareness framework to mitigate this issue. Concretely, a preference loss enforces multimodal embeddings to outperform their unimodal counterparts, while a composition regularization objective aligns multimodal embeddings with prototypes composed from its unimodal parts. These objectives explicitly model structural relationships between the composed representation and its unimodal counterparts. Experiments on various benchmarks show gains in out-of-distribution retrieval, highlighting modality composition awareness as a effective principle for robust composed multimodal retrieval when utilizing MLLMs as the unified encoder."},"_bibtex":{"value":"@misc{\nwu2025mca,\ntitle={{MCA}: Modality Composition Awareness for Robust Composed Multimodal Retrieval},\nauthor={Qiyu Wu and Shuyang Cui and Satoshi Hayakawa and Wei-Yao Wang and Hiromi Wakaki and Yuki Mitsufuji},\nyear={2025},\nurl={https://openreview.net/forum?id=EzBRI0Llxk}\n}"},"title":{"value":"MCA: Modality Composition Awareness for Robust Composed Multimodal Retrieval"},"pdf":{"value":"/pdf/2c99bf21eae5f0db226f555f11aba423303addf3.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"wu|mca_modality_composition_awareness_for_robust_composed_multimodal_retrieval"},"authorids":{"value":["~Qiyu_Wu2","~Shuyang_Cui1","~Satoshi_Hayakawa1","~Wei-Yao_Wang1","~Hiromi_Wakaki1","~Yuki_Mitsufuji1"]},"authors":{"value":["Qiyu Wu","Shuyang Cui","Satoshi Hayakawa","Wei-Yao Wang","Hiromi Wakaki","Yuki Mitsufuji"]}},"version":2},{"content":{"venue":{"value":"IEEE Transactions on Circuits and Systems for Video Technology"},"pdf":{"value":"https://ieeexplore.ieee.org/iel7/76/10492617/10234439.pdf"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"liu|semanticaware_contrastive_learning_with_proposal_suppression_for_video_semantic_role_grounding"},"html":{"value":"https://doi.org/10.1109/TCSVT.2023.3310296"},"abstract":{"value":"Video semantic role grounding has gained substantial interest from both the academic and industrial communities. While existing methods have demonstrated considerable performance improvements, the influence of noisy and intra-object proposals, referring to proposals with the same object label, has yet to be explored in video semantic role grounding. In this study, we propose a semantic-aware contrastive learning network with proposal suppression to enhance the accuracy of grounding referenced objects. To fully exploit the semantic information in each semantic role, we introduce a novel semantic role encoding module that allows for precise representations of each semantic role. We also design a semantic-aware proposal suppression network to reduce the impact of noisy proposals on object representation learning. Additionally, we propose a proposal contrastive loss to improve cross-modal alignment and reduce the effect of irrelevant intra-object proposals. Extensive experiments on four datasets demonstrate that our model achieves significant improvements over state-of-the-art methods."},"title":{"value":"Semantic-Aware Contrastive Learning With Proposal Suppression for Video Semantic Role Grounding"},"authors":{"value":[{"fullname":"Meng Liu"},{"fullname":"Di Zhou"},{"fullname":"Jie Guo"},{"fullname":"Xin Luo"},{"fullname":"Zan Gao","username":"~Zan_Gao1"},{"fullname":"Liqiang Nie"}]}},"tmdate":1789091676272,"pdate":1711929600000,"externalIds":["doi:10.1109/tcsvt.2023.3310296"],"tcdate":1768547110035,"writers":["~","OpenReview.net/Public_Article/ORCID.org","OpenReview.net/Support"],"signatures":["~Zan_Gao1"],"forum":"R7PvHuPDCO","license":"CC BY-SA 4.0","number":30721,"cdate":1693426863716,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/ORCID.org/-/Record","OpenReview.net/Public_Article/-/Edit","OpenReview.net/-/Edit"],"mdate":1789091676272,"domain":"OpenReview.net/Public_Article","id":"R7PvHuPDCO","version":2},{"content":{"summary":{"value":"This paper proposes a video-conditioned text representation refinement method called VICTER for text-to-video retrieval. The authors aim to enhance the textual representation by using video information to highlight relevant words, thus improving retrieval accuracy. Specifically, the approach involves two key components: a video abstraction module that summarizes representative features from video frames and a video-conditioned text enhancement module that refines text embeddings based on these video features. The method shows state-of-the-art performance improvements on several benchmark datasets (MSR-VTT, DiDeMo, LSMDC)."},"soundness":{"value":1},"confidence":{"value":5},"questions":{"value":"1. How does the method handle the computational burden during inference when video-conditioned text enhancement is applied to potentially thousands of candidate videos? Is there any strategy to reduce the complexity in such cases?\n\n2.  How does the model avoid data leakage during inference, given that the proposed method conditions text representation on the video? Are there measures in place to ensure that the video-specific text enhancement does not unfairly bias the retrieval process, namely, during the inference code, the ground truth is provided first for metric measure purpose?"},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The paper addresses an important aspect of text-to-video retrieval, focusing on enhancing text representation using video information, which is often overlooked compared to video feature improvement.\n\n2. Empirical results demonstrate the effectiveness of the proposed method, achieving notable improvements over baseline models across multiple benchmark datasets.\n\n3. The VICTER module can be integrated with various existing frameworks, showing its versatility in different retrieval tasks, including both image-text and video-text models."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The proposed video-conditioned text representation method appears impractical for real-world text-video retrieval systems. In actual inference scenarios, it is unrealistic to assume prior knowledge of the video-text pairing, making it computationally infeasible to conduct video-conditioned text generation for each candidate video. This suggests either a data leakage risk (if not independently enhancing text for each candidate) or an unacceptably high computational burden (if done independently for each candidate).\n\n2. The video-conditioned text enhancement module, if applied independently for each candidate video, incurs significant computational overhead. This renders the approach infeasible for real-time or large-scale retrieval tasks where thousands of videos may need to be evaluated against a single text query.\n\n3.  The paper lacks a clear explanation of how to avoid data leakage during inference. The text enhancement process involves generating text conditioned on the video, which could lead to unfair retrieval results if the same process is used during inference without independently processing each candidate video. This issue raises concerns about the validity of the reported improvements."}},"nonreaders":[],"tmdate":1731427887858,"tcdate":1730580302995,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission3584/Reviewer_4R5H"],"signatures":["ICLR.cc/2025/Conference/Submission3584/Reviewer_4R5H"],"forum":"jOVJhKzc3Y","number":4,"license":"CC BY 4.0","cdate":1730580302995,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission3584/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427887858,"domain":"ICLR.cc/2025/Conference","replyto":"jOVJhKzc3Y","id":"GVGp4XiBOP","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Text-to-Video Retrieval","Video-conditioned Text Representation Enhancement"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Pre-trained vision-language models (VLMs), such as CLIP, have shown remarkable success in the text-video retrieval task due to their strong vision-language\nrepresentations learned from large-scale paired image-text samples. However,\ncompared to videos, text is often short and concise, making it difficult to fully\ncapture the rich and redundant semantics present in a video with thousands of\nframes. Recent advances have focused on utilizing text features to extract key information from these redundant video frames. However, text representation generated without considering video information can suffer from bias and lack the\nexpressiveness needed to capture key words that could enhance retrieval performance. In this study, we first conduct preliminary experiments to demonstrate\nthe importance of enhancing text representations. These experiments reveal that\ntext representation only generated from text input often misinterpret critical information. To address this, we propose a simple yet efficient method, VICTER, i.e.,\nvideo-conditioned text representation refinement, to enrich text representation using a versatile module. Specifically, we introduce a video abstraction module that\nextracts representative features from multiple video frames. This is followed by\na video-conditioned text enhancement module that refines the original text features by reassessing individual word features and extracting key words using the\ngenerated video features. Empirical evidence shows that VICTER not only effectively captures relevant key words from the input text but also complements\nvarious existing frameworks. Our experimental results demonstrate a significant\nimprovement of VICTER over several baseline frameworks (with 0.4% ∼ 1.0%\nimprovements on R@1). Furthermore, VICTER achieves state-of-the-art performance on three benchmark datasets, including MSRVTT, DiDeMo, and LSMDC. Code will be made available."},"_bibtex":{"value":"@misc{\nguo2024the,\ntitle={The Devil is in the Word: Video-Conditioned Text Representation Refinement for Text-to-Video Retrieval},\nauthor={Jianyuan Guo and Fangyun Wei and Jianbo Ma and Chang Xu},\nyear={2024},\nurl={https://openreview.net/forum?id=jOVJhKzc3Y}\n}"},"title":{"value":"The Devil is in the Word: Video-Conditioned Text Representation Refinement for Text-to-Video Retrieval"},"pdf":{"value":"/pdf/c75eef1bd7836c613c8a7835d66857ca5cf63641.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"guo|the_devil_is_in_the_word_videoconditioned_text_representation_refinement_for_texttovideo_retrieval"},"authorids":{"value":["~Jianyuan_Guo1","~Fangyun_Wei1","~Jianbo_Ma3","~Chang_Xu4"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Jianyuan Guo","Fangyun Wei","Jianbo Ma","Chang Xu"]}},"version":2},{"content":{"research_area_keywords":{"value":"vision question answering, video processing, multimodality"},"keywords":{"value":["Video Understanding Benchmark","Spatio-Temporal Compositionality"]},"B_use_or_create_scientific_artifacts":{"value":"Yes"},"languages_studied":{"value":"English"},"D4_ethics_review_board_approval":{"value":"N/A"},"C4_parameters_for_packages":{"value":"Yes"},"B6_elaboration":{"value":"3.3 Data Construction Pipeline"},"B1_cite_creators_of_artifacts":{"value":"Yes"},"B2_discuss_the_license_for_artifacts":{"value":"Yes"},"B4_elaboration":{"value":"A.1 Data Collection Processes"},"D_human_subjects_including_annotators":{"value":"Yes"},"B5_elaboration":{"value":"A.1 Data Collection Processes"},"A2_elaboration":{"value":"A.8 Impact Statement"},"B2_elaboration":{"value":"A.9 License and Intended Use"},"B3_elaboration":{"value":"A.9 License and Intended Use"},"C2_elaboration":{"value":"4 Experimental Setup"},"C4_elaboration":{"value":"4 Experimental Setup"},"D3_elaboration":{"value":"A.1 Data Collection Processes"},"E_ai_assistants_in_research_or_writing":{"value":"Yes"},"D1_elaboration":{"value":"A.1 Data Collection Processes"},"C1_elaboration":{"value":"4 Experimental Setup"},"C2_experimental_setup_and_hyperparameters":{"value":"Yes"},"B1_elaboration":{"value":"References"},"C1_model_size_and_budget":{"value":"Yes"},"venue":{"value":"ACL ARR 2026 May Submission"},"D3_data_consent":{"value":"Yes"},"_bibtex":{"value":"@inproceedings{\nanonymous2026timeblind,\ntitle={TimeBlind: A Spatio-Temporal Compositionality Benchmark for Video {LLM}s},\nauthor={Anonymous},\nbooktitle={Submitted to ACL Rolling Review - May 2026},\nyear={2026},\nurl={https://openreview.net/forum?id=kKMcY4A4fR},\nnote={under review}\n}"},"title":{"value":"TimeBlind: A Spatio-Temporal Compositionality Benchmark for Video LLMs"},"C3_descriptive_statistics":{"value":"N/A"},"contribution_types":{"value":["Data resources","Data analysis"]},"A1_limitations_section":{"value":"This paper has a limitations section."},"B6_statistics_for_data":{"value":"Yes"},"B4_data_contains_personally_identifying_info_or_offensive_content":{"value":"Yes"},"B3_artifact_use_consistent_with_intended_use":{"value":"Yes"},"A2_potential_risks":{"value":"Yes"},"E1_information_about_use_of_ai_assistants":{"value":"N/A"},"abstract":{"value":"Fine-grained spatio-temporal understanding is essential for video reasoning and embodied AI. Yet, while Multimodal Large Language Models (MLLMs) master static semantics, their grasp of temporal dynamics remains brittle. We present TimeBlind, a diagnostic benchmark for compositional spatio-temporal understanding. Inspired by cognitive science, TimeBlind categorizes fine-grained temporal understanding into three levels: recognizing atomic events, characterizing event properties, and reasoning about event interdependencies. Unlike benchmarks that conflate recognition with temporal reasoning, TimeBlind leverages a minimal-pairs paradigm: video pairs share identical static visual content but differ solely in temporal structure, utilizing complementary questions to neutralize language priors. Evaluating over 20 state-of-the-art MLLMs (e.g., GPT-5, Gemini 3 Pro) on 600 curated instances (2400 video-question pairs), reveals that the Instance Accuracy (correctly distinguishing both videos in a pair) of the best performing MLLM is only 48.2\\%, far below the human performance (98.2\\%). These results demonstrate that even frontier models exhibit significant deficiencies in temporal reasoning, positioning TimeBlind as a vital diagnostic tool for next-generation video understanding."},"D1_instructions_given_to_participants":{"value":"Yes"},"paper_type":{"value":"Long"},"C_computational_experiments":{"value":"Yes"},"pdf":{"value":"/pdf/c60598a56293d0efd7ac6cb933a1dfda762f1bf1.pdf"},"research_area":{"value":"Multimodality and Language Grounding to Vision, Robotics and Beyond"},"B5_documentation_of_artifacts":{"value":"Yes"},"EMNLP_2026_AI_Reviewing_Experiment":{"value":"no"},"D2_recruitment_and_payment":{"value":"N/A"},"venueid":{"value":"aclweb.org/ACL/ARR/2026/May/Submission"}},"tmdate":1790696925285,"tcdate":1779475041654,"writers":["aclweb.org/ACL/ARR/2026/May","aclweb.org/ACL/ARR/2026/May/Submission3608/Authors"],"signatures":["aclweb.org/ACL/ARR/2026/May/Submission3608/Authors"],"forum":"kKMcY4A4fR","license":"CC BY 4.0","number":3608,"cdate":1779475041654,"readers":["everyone"],"invitations":["aclweb.org/ACL/ARR/2026/May/-/Submission","aclweb.org/ACL/ARR/2026/May/-/Edit","aclweb.org/ACL/ARR/2026/May/-/Post_Submission","aclweb.org/ACL/ARR/2026/May/-/Preprint_Post_Submission","aclweb.org/ACL/ARR/2026/May/Submission3608/-/Blind_Submission_License_Agreement"],"mdate":1790696925285,"odate":1780385598983,"domain":"aclweb.org/ACL/ARR/2026/May","id":"kKMcY4A4fR","version":2},{"content":{"data_release":{"value":"We authorize the release of our submission and author names to the public in the event of acceptance."},"TLDR":{"value":"A Spatio-Temporal Compositionality Benchmark for Video LLMs"},"venue":{"value":"VidLLMs 2026 Poster"},"email_sharing":{"value":"We authorize the sharing of all author emails with Program Chairs."},"pdf":{"value":"/pdf/73021566f805d2d404cd118ec73ea369241aa203.pdf"},"keywords":{"value":["Video Understanding; Benchmark; Evaluation; VLM"]},"venueid":{"value":"thecvf.com/CVPR/2026/Workshop/VidLLMs"},"paperhash":{"value":"li|timeblind_a_spatiotemporal_compositionality_benchmark_for_video_llms"},"authorids":{"value":["~Baiqi_Li2","~Kangyi_Zhao1","~Ce_Zhang7","~Chancharik_Mitra1","~Jean_de_Dieu_Nyandwi1","~Gedas_Bertasius1"]},"abstract":{"value":"Fine-grained spatio-temporal understanding is essential for video reasoning and embodied AI. Yet, while Multimodal Large Language Models (MLLMs) master static semantics, their grasp of temporal dynamics remains brittle. We present TimeBlind, a diagnostic benchmark for compositional spatio-temporal understanding. Inspired by cognitive science, TimeBlind categorizes fine-grained temporal understanding into three levels: recognizing atomic events, characterizing event properties, and reasoning about event interdependencies. Unlike benchmarks that conflate recognition with temporal reasoning, TimeBlind leverages a minimal-pairs paradigm: video pairs share identical static visual content but differ solely in temporal structure, utilizing complementary questions to neutralize language priors. Evaluating over 20 state-of-the-art MLLMs (e.g., GPT-5, Gemini 3 Pro) on 600 curated instances (2400 video-question pairs), reveals that the Instance Accuracy (correctly distinguishing both videos in a pair) of the best performing MLLM is only 48.2\\%, far below the human performance (98.2\\%). These results demonstrate that even frontier models lack temporal reasoning, positioning TimeBlind as a vital diagnostic tool for next-generation video understanding. We will release the data and code."},"title":{"value":"TimeBlind: A Spatio-Temporal Compositionality Benchmark for Video LLMs"},"authors":{"value":["Baiqi Li","Kangyi Zhao","Ce Zhang","Chancharik Mitra","Jean de Dieu Nyandwi","Gedas Bertasius"]}},"tmdate":1780041772134,"pdate":1780041769715,"tcdate":1776132413552,"writers":["thecvf.com/CVPR/2026/Workshop/VidLLMs","thecvf.com/CVPR/2026/Workshop/VidLLMs/Submission16/Authors"],"signatures":["thecvf.com/CVPR/2026/Workshop/VidLLMs/Submission16/Authors"],"forum":"bQweUa7Ug3","license":"CC BY 4.0","number":16,"cdate":1776132413552,"readers":["everyone"],"invitations":["thecvf.com/CVPR/2026/Workshop/VidLLMs/-/Submission","thecvf.com/CVPR/2026/Workshop/VidLLMs/-/Submission_Change_Before_Reviewing","thecvf.com/CVPR/2026/Workshop/VidLLMs/-/Submission_Change_Before_Bidding","thecvf.com/CVPR/2026/Workshop/VidLLMs/Submission16/-/Camera_Ready_Revision","thecvf.com/CVPR/2026/Workshop/VidLLMs/-/Submission_Release"],"mdate":1780041772134,"odate":1780041769715,"domain":"thecvf.com/CVPR/2026/Workshop/VidLLMs","id":"bQweUa7Ug3","version":2},{"content":{"summary":{"value":"This paper presents a non-asymptotic theoretical lower bound that quantitatively links the minimal number of in-context demonstrations to the stability of In-Context Learning (ICL). Under high-dimensional sub-Gaussian feature assumptions, the authors derive a spectral condition ensuring stability, and propose a two-stage observable estimator to compute the required prompt length. They validate the theory through experiments."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"Could the authors explain the rationale behind their Goal (lines 111–115)? (See Weaknesses 1)"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"This paper provides an interesting theoretical framework to quantify the minimal number of in-context demonstrations required for stable in-context learning (ICL)."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The paper lacks a clear explanation of the Goal (lines 111–115). The authors only provide a brief and vague justification in lines 310–315, without any experimental evidence to support its validity. Moreover, the Goal itself raises several conceptual concerns. For example, the minimal nonzero eigenvalue of a matrix captures very limited information about the matrix’s overall structure—so is it reasonable to consider only this quantity? How should the parameter $\\delta$ be chosen? Shouldn’t the overall magnitude or spectrum of the matrix also be taken into account?\nFor instance, consider two matrices where one has a smallest eigenvalue of 0 and the other has 0.00001, while all other eigenvalues are identical. These two matrices would likely behave almost identically, with such a tiny difference possibly arising from numerical noise. However, under the authors’ theoretical framework, their minimal nonzero eigenvalues could differ significantly, which seems questionable.\n\n2. The presentation of this paper also has room for improvement. For example, before diving into the mathematical derivations, the authors should first clearly explain the intuition — what the goal is, why this goal is reasonable, and what motivates the chosen formulation."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762917511173,"tcdate":1761956777347,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission4683/Reviewer_XKBd"],"signatures":["ICLR.cc/2026/Conference/Submission4683/Reviewer_XKBd"],"forum":"ar8EnfwITb","number":2,"license":"CC BY 4.0","cdate":1761956777347,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission4683/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762917511173,"domain":"ICLR.cc/2026/Conference","replyto":"ar8EnfwITb","id":"8VjK5AhtiG","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["In-Context Learning; probabilistic methods; XAI; learning theory."]},"supplementary_material":{"value":"/attachment/4e2fce303302c5b11391a04316a78be3c43e5b37.zip"},"primary_area":{"value":"learning theory"},"abstract":{"value":"In-context learning (ICL) is flexible but its reliability is highly sensitive to prompt length. This paper establishes a non-asymptotic lower bound that links the minimal number of demonstrations to ICL stability under fixed high-dimensional sub-Gaussian representations. The bound gives explicit sufficient conditions in terms of spectral properties of the covariance, providing a computable criterion for practice. Building on this analysis, we propose a two-stage observable estimator with a one-shot calibration that produces practitioner-ready prompt-length estimates without distributional priors. Experiments across diverse datasets, encoders, and generators show close alignment between the predicted thresholds and empirical knee-points, with the theory acting as a conservative but reliable upper bound; the calibrated variant further tightens this gap. These results connect spectral coverage to stable ICL, bridge theory and deployment, and improve the interpretability and reliability of large-scale prompting in realistic finite-sample regimes."},"_bibtex":{"value":"@misc{\nwang2025theoretical,\ntitle={Theoretical Bounds for Stable In-Context Learning},\nauthor={Tongxi Wang and Zhuoyang Xia},\nyear={2025},\nurl={https://openreview.net/forum?id=ar8EnfwITb}\n}"},"title":{"value":"Theoretical Bounds for Stable In-Context Learning"},"pdf":{"value":"/pdf/48c68e4e75957d0b8a21ab323867649c3aa22520.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"wang|theoretical_bounds_for_stable_incontext_learning"},"authorids":{"value":["~Tongxi_Wang1","~Zhuoyang_Xia1"]},"authors":{"value":["Tongxi Wang","Zhuoyang Xia"]}},"version":2},{"content":{"TLDR":{"value":"Our target-aware model generates a video in which an actor accurately interacts with the target, specified with its segmentation mask."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Controllable video diffusion models","Human-scene interaction","Robotics planning"]},"supplementary_material":{"value":"/attachment/7a4d8a24e01dd34a1dcf692bea8162b9b33d526c.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"We present a target-aware video diffusion model that generates videos from an input image, in which an actor interacts with a specified target while performing a desired action. The target is defined by a segmentation mask, and the action is described through a text prompt. Our key motivation is to incorporate target awareness into video generation, enabling actors to perform directed actions on designated objects. This enables video diffusion models to act as motion planners, producing plausible predictions of human-object interactions by leveraging the priors of large-scale video generative models. We build our target-aware model by extending a baseline model to incorporate the target mask as an additional input. To enforce target awareness, we introduce a special token that encodes the target's spatial information within the text prompt. We then fine-tune the model with our curated dataset using an additional cross-attention loss that aligns the cross-attention maps associated with this token with the input target mask. To further improve performance, we selectively apply this loss to the most semantically relevant attention regions and transformer blocks. Experimental results show that our target-aware model outperforms existing solutions in generating videos where actors interact accurately with the specified targets. We further demonstrate its efficacy in two downstream applications: zero-shot 3D HOI motion synthesis with physical plausibility and long-term video content creation."},"_bibtex":{"value":"@inproceedings{\nkim2026targetaware,\ntitle={Target-Aware Video Diffusion Models},\nauthor={Taeksoo Kim and Hanbyul Joo},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=311AxWM8FU}\n}"},"title":{"value":"Target-Aware Video Diffusion Models"},"pdf":{"value":"/pdf/c98ab4dc22c46001be3678bca8a6a1a4b97e01a3.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"kim|targetaware_video_diffusion_models"},"authorids":{"value":["~Taeksoo_Kim2","~Hanbyul_Joo2"]},"authors":{"value":["Taeksoo Kim","Hanbyul Joo"]}},"tmdate":1775876934687,"pdate":1769435753125,"tcdate":1757584517661,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission4015/Authors"],"signatures":["ICLR.cc/2026/Conference/Submission4015/Authors"],"forum":"311AxWM8FU","license":"CC BY 4.0","number":4015,"cdate":1757584517661,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/-/Submission","ICLR.cc/2026/Conference/-/Post_Submission","ICLR.cc/2026/Conference/Submission4015/-/Full_Submission","ICLR.cc/2026/Conference/Submission4015/-/Rebuttal_Revision","ICLR.cc/2026/Conference/-/Edit","ICLR.cc/2026/Conference/Submission4015/-/Camera_Ready_Revision"],"mdate":1775876934687,"odate":1759896705795,"domain":"ICLR.cc/2026/Conference","id":"311AxWM8FU","version":2},{"content":{"summary":{"value":"This paper proposes CamPilot, a technique to improve camera control in video diffusion models. Concretely, this work uses Reward Feedback Learning (ReFL) to improve the camera control. For this, a 3DGS decoder on top of the video latents is used to directly generate 3D Gaussians instead of RGB videos. Then, the renderings of the 3D Gaussians are used for reward feedback. Using the depth to obtain a visibility mask, the generated output is masked out to only supervise on input content."},"soundness":{"value":2},"confidence":{"value":5},"questions":{"value":"I generally like the idea of using 3DGS output to improve 3D consistency/camera control precision of video models. However, the majority of the paper overclaims the architectural contributions since they just follow Wonderland. The main formulation of the reward feedback is simple, which is good, but it is not analyzed that deeply. Moreover, the supplementary material is not organized well and is difficult to digest. I highly recommend preparing a supplementary website with side-by-side comparisons.\n\nI would like authors to address the following question:\n\n- What are the architectural contributions described in Sec. 3.1-3.3? All the parts mentioned are just reusing other works.\n\nI am currently negative but happy to see what the authors say about the actual contributions of the work."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"- Novelty of using reward feedback for camera control: I have not seen a work that uses 3DGS to improve the underlying video model and its 3D consistency/camera control precision. The direction seems promising."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Camera-aware 3DGS decoder already proposed: The first point of the contribution list claims the camera-aware 3DGS decoder to be a contribution. However, the approach just follows Wonderland [1] without changes. The claim of the whole pipeline seems to be a big overclaim and the authors should have acknowledged Wonderland a lot more during the paper. Wonderland is mentioned and referenced in the paper, but there are no differences in the approaches. The main point of the paper is the reward-based feedback learning, not the 3DGS decoder.\n- Supplementary video presentation: The supplementary video presentation is one of the key components of a video generation submission. But there is no website but only separate videos and it is not clear what is happening. Moreover, there are no side-by-side comparisons with previous works. You should also show multiple generated scenes for one camera trajectory to show that they are consistent and the camera control works across scenes.\n- Unclear handling of dynamic objects: Currently, there is no guarantee that the scene remains static. While most results are designed for scenes without objects in them, that could move, the approach seems to not handle dynamic objects which would get animated by the video model in normal cases. Those animated objects would then lead to blurry 3DGS outputs and bad reward feedback. Hence, currently the model is restricted to static scenes.\n\n[1] Liang et al., Wonderland: Navigating 3D Scenes from a Single Image, CVPR 2025"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762916483495,"tcdate":1760566591712,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission2989/Reviewer_qxx8"],"signatures":["ICLR.cc/2026/Conference/Submission2989/Reviewer_qxx8"],"forum":"qqij8fCGDl","number":1,"license":"CC BY 4.0","cdate":1760566591712,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission2989/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762916483495,"domain":"ICLR.cc/2026/Conference","replyto":"qqij8fCGDl","id":"VFJbYDDDgV","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Video generation; 3D scene exploration; Reward feedback learning"]},"supplementary_material":{"value":"/attachment/190335a20c10687ec00345e0bee18e333335cb61.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Recent advancements in camera-controlled video diffusion models have significantly improved video-camera alignment and enabled more accurate 3D scene generation, driven by potential downstream applications such as virtual reality.\nHowever, we reveal that existing approaches often struggle to precisely adhere to the given camera conditions, leading to inconsistencies in the 3D geometry.\nInspired by Reward Feedback Learning in diffusion models, which has demonstrated strong potential in aligning model outputs with task-specific objectives, we build upon this paradigm and aim to further improve camera controllability.\nDirectly borrowing existing ReFL approaches faces several challenges. First, current reward models lack the capacity to assess video-camera alignment. Second, decoding latent into RGB videos for reward computation introduces substantial computational overhead. Third, 3D geometric information is typically neglected during video decoding.\nTo address these limitations, we introduce a camera-aware 3D decoder that efficiently decodes video latent into 3D representations for reward computation. Specifically, we project the video latent and camera pose into 3D Gaussians, which supports efficient rendering from arbitrary views. \nIn this process, the camera pose not only acts as an input variable but also serves as a projection parameter for determining the mean of each Gaussian.\nIf the generated video does not match the camera conditions, the 3D structure becomes geometrically inconsistent, leading to blurry rendered images.\nBased on this property, we explicitly optimizing pixel-level consistency between rendered novel views and ground-truth ones as reward feedback.\nTo accommodate the stochastic nature, we further introduce a visibility term that selectively supervises only deterministic regions derived via geometric warping.\nExtensive experiments conducted on the RealEstate10K and WorldScore benchmarks demonstrate the effectiveness of our proposed method in enhancing both camera controllability and generation quality."},"_bibtex":{"value":"@misc{\nge2025campilot,\ntitle={CamPilot: Improving Camera Control in Video Diffusion Model with Efficient Camera Reward Feedback},\nauthor={Wenhang Ge and Guibao Shen and Jiawei Feng and Luozhou Wang and Hao LU and Xingye Tian and Xin Tao and Pengfei Wan and Ying-Cong Chen},\nyear={2025},\nurl={https://openreview.net/forum?id=qqij8fCGDl}\n}"},"title":{"value":"CamPilot: Improving Camera Control in Video Diffusion Model with Efficient Camera Reward Feedback"},"pdf":{"value":"/pdf/3919cc68511ac2752505bc0f7b4e7f570a2cd725.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"ge|campilot_improving_camera_control_in_video_diffusion_model_with_efficient_camera_reward_feedback"},"authorids":{"value":["~Wenhang_Ge1","~Guibao_Shen1","~Jiawei_Feng1","~Luozhou_Wang2","~Hao_LU8","~Xingye_Tian2","~Xin_Tao3","~Pengfei_Wan1","~Ying-Cong_Chen1"]},"authors":{"value":["Wenhang Ge","Guibao Shen","Jiawei Feng","Luozhou Wang","Hao LU","Xingye Tian","Xin Tao","Pengfei Wan","Ying-Cong Chen"]}},"version":2},{"content":{"summary":{"value":"The method conditions a video diffusion backbone on both the last few temporal frames for motion continuity and projected views from a static‑only 3D scene memory for spatial consistency, where the memory is created by dynamic SLAM plus a masking pipeline that removes moving objects using optical‑vs‑warped flow differencing, backward point tracking, and SAM2 propagation to yield clean static point clouds. Static‑only point clouds are projected from spatially adjacent viewpoints to the target camera poses and concatenated with temporal latents through a frozen 3D VAE channel, enabling camera‑controllable generation without modifying the DiT backbone."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"I have highlighted the major questions in the weaknesses already, and I am summing them up below(I have summarised them shortly so that it is easier for the reviewer to quote the exact question they are answering. Please refer to the Weakness for the detailed problem and question asked.\n1. Why is the spatial structure not preserved in the first supplementary video where two chairs were supposed to be side-by-side, but instead a white pillar or cupboard appears?\n2. Is there any error in camera pose estimation that, when back-projected to image space, causes inaccurate projections and leads to inconsistency?\n3. Are the comparisons fair, given that your method takes both the initial few frames and the last few frames as input, while the other methods only use the last frame as prior?\n4. Are the better qualitative results mainly due to conditioning on more input frames rather than the proposed architectural changes?\n5. Why didn’t the authors train a baseline like CameraCtrl or another strong model with minimal modifications to take more input frames and verify whether improvements come from more inputs or architecture changes?\n6. What are the limitations of the proposed work?\n7. What are the potential future directions or improvements for this work?\n8. Why are there so few video results in the supplementary, only a handful instead of 15–20 videos as expected for a video generation paper?\n9. Are the generated videos coherent with each other when given the same initial video but different trajectories, especially in the static regions?\n10. Why hasn’t the paper compared its method with ReCamMaster, given that its evaluation code has been available since before the ICLR deadline?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":4},"strengths":{"value":"1. Proposed new problem statement which is relevant to Video Generation which the previous camera controllable methods which conditions only on the last few frames\n2. Dual spatio‑temporal conditioning addresses both motion continuity and scene consistency, overcoming the short‑window limitation of prior future‑video models.\n3. A static‑only 3D scene memory cleanly separates persistent geometry from dynamics, avoiding the leakage of outdated moving objects into future frames.\n4. Three‑stage dynamic masking (optical‑vs‑warped flow differencing, backward point tracking, SAM2 propagation) yields clean static point clouds without ghosting."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. In the first video in the supplementary video, I see that the spatial structure is not preserved. In the initial video, there were supposed to be 2 chairs side-by-side but I don’t see them in the generated video even though the 3D scene memory captures it. Instead a white pillor / cupboard seems to appear which is wrong. Why is it so? Is there any error in camera pose estimation that when back projected on image space, they are not projected accurately and may lead to inconsistency?\n2. I feel like the comparisons are not fair. The video comparisons are all against methods which only do camera control video generation with only the last frame as prior. But your method takes in the few initial videos frames along with the last few frames for video generation and so automatically your method would be more consistent since it is taking more frames as input. This is an unfair comparison in my opinion. The good qualitative results maybe just because you are conditioning your model with more priors and hence more consistent results or is it because of your proposed new architectural changes that enabled better consistent results. I feel the paper didn’t differentiate these two effectively. I agree that it is a novel problem statement and hence no baselines exist but the authors could have trained CameraCtrl or any of the best model to take in more inputs and trained the model with very minimal changes to effectively clarify if the difference in clarity is from more data inputs or actual model architectural changes. \n3. There are no limitations mentioned in the main paper or the appendix. I request the authors to propose the limitations of the work and the future work. It is important for the reviewers and the community to know the limitations so that there could be more work/discussions in that direction to resolve it. \n4. Minimal video results. For a paper proposing video generation, I would expect a minimum of 15-20 videos to be shown in the supplementary. It is difficult for the reviewer to judge if the results are cherry picked or does it actually generalise well on many videos and different camera trajectory. \n5. There are no results shown in the case where given an initial video and 2-3 different trajectories, are all 2-3 videos coherent to the initial video atleast in the region of static part. Are the generated videos coherent with each other in the common regions?\n6. What about a comparison against the newly released ReCamMaster? It is an ICCV paper I assume and hence the evaluation code has been released long before ICLR deadline. I would request the authors to compare this work with their method and show few comparison."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762928211442,"tcdate":1761512190951,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission18512/Reviewer_xUs1"],"signatures":["ICLR.cc/2026/Conference/Submission18512/Reviewer_xUs1"],"forum":"3XxoBwMusJ","number":1,"license":"CC BY 4.0","cdate":1761512190951,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission18512/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762928211442,"domain":"ICLR.cc/2026/Conference","replyto":"3XxoBwMusJ","id":"fpZshnfsEO","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Scene-Consistent Video Generation; Camera-Controllable Video Generation; Video Diffusion Models;"]},"supplementary_material":{"value":"/attachment/f2e2d8532cc22a8c11a6904f4d996453eae9a722.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"We present 3DScenePrompt, a framework for camera-controllable video generation that maintains scene consistency when extending arbitrary-length input videos along user-specified trajectories. Unlike existing video generative methods limited to conditioning on a single image or just a few frames, we introduce a dual spatio-temporal conditioning strategy that fundamentally rethinks how video models should reference prior content. Our approach conditions on both temporally adjacent frames for motion continuity and spatially adjacent content for scene consistency. However, when generating beyond temporal boundaries, directly using spatially adjacent frames would incorrectly preserve dynamic elements from the past. We address this through introducing a 3D scene memory that represents exclusively the static geometry extracted from the entire input video. To construct this memory, we leverage dynamic SLAM with our newly introduced dynamic masking strategy that explicitly separates static scene geometry from moving elements. The static scene representation can then be projected to any target viewpoint, providing geometrically-consistent warped views that serve as strong spatial prompts while allowing dynamic regions to evolve naturally from temporal context. This enables our model to maintain long-range spatial coherence and precise camera control without sacrificing computational efficiency or motion realism. Extensive experiments demonstrate that our framework significantly outperforms existing methods in scene consistency, camera controllability, and generation quality."},"_bibtex":{"value":"@inproceedings{\nlee2026d,\ntitle={3D Scene Prompting for Scene-Consistent Camera-Controllable Video Generation},\nauthor={JoungBin Lee and Jaewoo Jung and Jisang Han and Takuya Narihira and Kazumi Fukuda and Junyoung Seo and Sunghwan Hong and Yuki Mitsufuji and Seungryong Kim},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=3XxoBwMusJ}\n}"},"title":{"value":"3D Scene Prompting for Scene-Consistent Camera-Controllable Video Generation"},"pdf":{"value":"/pdf/91fe1cc2278440c1779e1bfb687d9ccee37326b3.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"lee|3d_scene_prompting_for_sceneconsistent_cameracontrollable_video_generation"},"authorids":{"value":["~JoungBin_Lee1","~Jaewoo_Jung2","~Jisang_Han1","~Takuya_Narihira2","~Kazumi_Fukuda1","~Junyoung_Seo1","~Sunghwan_Hong2","~Yuki_Mitsufuji1","~Seungryong_Kim1"]},"authors":{"value":["JoungBin Lee","Jaewoo Jung","Jisang Han","Takuya Narihira","Kazumi Fukuda","Junyoung Seo","Sunghwan Hong","Yuki Mitsufuji","Seungryong Kim"]}},"version":2},{"content":{"summary":{"value":"This paper introduces a frame selection method formulated as active perception, to efficiently handle long-form video understanding based on a video diffusion prior and a VLM. It is motivated to identify out-of-distribution (or, surprising) frames in a given video, using the diffusion prior as world knowledge (i.e., expected behavior). The authors first uniformly sample a small number of frames, and feed it through a diffusion model to interpolate and generate intermediate frame latents, also conditioned on question/answer text. Next, they compare the generated frames and the original frames in the diffusion latent space, identifying the most different frames as keyframes, subsequently using these for long-form video QA. The proposed method is validated on multiple benchmarks (eg: EgoSchema, NExT-QA, ActivityNet-QA, CLEVRER) based on various VLMs (GPT-4o, Gemini 1.5 Pro, and LLaVA-OV)."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"- How best can the authors justify efficiency measurements only w.r.t. the number of frames, and avoiding other metrics such as latency, computations or memory of the overall pipeline?\n- Can the authors clarify missing details about how they exactly use the diffusion prior (eg: which frames are generated: intermediate/subsequent, if any masking is used, exact nature of conditioning, number of denoising steps, model scale, what are the transformer layers in L146: are they already pretrained?)?\n- Can the authors visualize generated frames in pixel-space, and compare with real frames?\n- How does the proposed method address questions about generic scenarios, as it always try to sample frames that are very different from the norm?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- The proposed method is very interesting and refreshing, as it relies on a video-diffusion prior to encode world-knowledge in a multi-modal LLM setup.\n- The idea of keyframe sampling is extremely useful in long-form video understanding to mitigate associated costs.\n- The authors have conducted extensive experimentation on mutliple long-video VQA benchmarks, making use of both open-sorce and proprietary LLMs.\n- The paper is generally well-written and easy-to-follow."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- This paper mainly advertises the efficiency of the proposed pipeline as it relies of fewer frames. However, it also relies on many other components to sample such keyframes (eg: diffusion prior). It would be interesting to see whether it would translate to an efficiency gain that matters (eg: runtime, flops/memory of overall pipeline). Is there another way to justify efficiency measurement only in-terms of the number of frames (eg: significant reduction in prompting cost)?\n- I have concerns about the generated latent frames. (1) A lot of details about the diffusion based generation and interpolation is missing (eg: which frames are generated: intermediate/subsequent, if any masking is used, exact nature of conditioning, number of denoising steps, model scale, what are the transformer layers in L146: are they already pretrained?). (2) I doubt the fidelity/usefulness of the generated frames. Can the model generate high-fidelity missing frames with reasonable motion that we may see in a 3min video? This can be easily evaluated by conveting the latent back to the pixel space and visualizing. It would be interesting to see how different they are from actual frames.\n- In this method, initial frames are unoformly-sampled and the highest-different frames from these are fed to the VLM for video VQA. However, for questions that require normal/vanilla dynamics (eg: asking about the generic/prominent activity in a given video), should not have a better acuracy by sampling deviating (or, surprising) frames from the norm, right? This should be explained further, as otherwise the motivation of this work does not generalize.\n- The performance gain seems to be minimal in most of the comparisons in Table 1. Also the comparison with other frame selection methods is not entirely fair (eg: others using GPT-4 and the proposed method using GPT-4o). Can the authors make a fair comparison, by running a few experiments with GPT-4?\n\n**[Update]: Since the authors have addressed most of my concerns, I am raising my original rating from 3 to 6. There is still a confusion about efficiency (*e.g.* wall-clock time), yet I don't believe it warrents a reject rating--- as this work can be useful to the community. I suggest the authors encorporate all the discussions in the final version of the paper for improved clarity and better foundation of claims.**"}},"nonreaders":[],"tmdate":1733230055920,"tcdate":1730695152406,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission1117/Reviewer_sf66"],"signatures":["ICLR.cc/2025/Conference/Submission1117/Reviewer_sf66"],"forum":"KtqZrNjvjd","number":3,"license":"CC BY 4.0","cdate":1730695152406,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission1117/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1733230055920,"domain":"ICLR.cc/2025/Conference","replyto":"KtqZrNjvjd","id":"VRvi3V7c8M","forumContent":{"TLDR":{"value":"The paper proposes an \"active perception\" method that selects key frames using generative models to improve efficiency and performance of vision-language models in long-form video question answering."},"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video question answering","Vision Language Model"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Large vision-language models (VLMs) have advanced multimodal tasks such as video question answering (QA), yet they struggle with long-form videos due to the computational burden of processing excessive tokens. Inspired by active perception theory, which posits that models gain information by acquiring data that differ from their expectations, we introduce Video Active Perception (VAP), a training-free method to enhance long-form video QA using VLMs. Our approach treats key frame selection as data acquisition in active perception and leverages a lightweight text-conditioned video generation model to represent prior world knowledge. Empirically, VAP achieves state-of-the-art zero-shot results on long-form video QA datasets such as EgoSchema, NExT-QA, ActivityNet-QA and CLEVRER, achieving an increase of up to 5.6 X efficiency by frames per question over standard GPT-4o, Gemini 1.5 Pro, and LLaVA-OV. Moreover, VAP shows stronger reasoning abilities than previous methods and effectively selects key frames relevant to questions. These findings highlight the potential of leveraging active perception to improve efficiency and effectiveness of long-form video QA."},"_bibtex":{"value":"@misc{\nma2025video,\ntitle={Video Active Perception: Efficient Inference-Time Long-Form Video Understanding with Vision-Language Models},\nauthor={Martin Q. Ma and Willis Guo and Aditya Agrawal and Ankit Gupta and Paul Pu Liang and Russ Salakhutdinov and Louis-Philippe Morency},\nyear={2025},\nurl={https://openreview.net/forum?id=KtqZrNjvjd}\n}"},"title":{"value":"Video Active Perception: Efficient Inference-Time Long-Form Video Understanding with Vision-Language Models"},"pdf":{"value":"/pdf/0c35130c0a6315516545bebd5ff63f0df45726d6.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"ma|video_active_perception_efficient_inferencetime_longform_video_understanding_with_visionlanguage_models"},"authorids":{"value":["~Martin_Q._Ma1","~Willis_Guo1","~Aditya_Agrawal3","~Ankit_Gupta7","~Paul_Pu_Liang1","~Russ_Salakhutdinov1","~Louis-Philippe_Morency1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Martin Q. Ma","Willis Guo","Aditya Agrawal","Ankit Gupta","Paul Pu Liang","Russ Salakhutdinov","Louis-Philippe Morency"]}},"version":2},{"content":{"summary":{"value":"This paper proposes VideoGuard to protect videos from unauthorized editing. VideoGuard consists of two-stage pipelines. In the first stage, the authors use DDIM inversion process to get an initial noise latent from the original video. Then, the method optimizes the latent so that the denoising from the latent will result in a disrupted video, and therefore the video editing will be unsuccessful.  In the second stage, the method tries to find a perturbation that after adding it on top of the source video it can make the inversion of the video close to the optimized latent from stage 1. In addition, the added perturbation is required to be imperceptible. Therefore, VideoGuard can produce a similar source video that is free of unauthorized editing."},"soundness":{"value":2},"confidence":{"value":2},"questions":{"value":"N/A"},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. The proposed approach is interesting and reasonable to achieve the goal of protecting videos from unauthorized editing.\n2. VideoGuard is to protect the video in the video pixel space, and the enhanced video is similar to the source video with extra shield.\n3. The method treats the video as a whole without needing to process each frame individually."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"My major concern is the results of VideoGuard. Although VideoGuard is a reasonable approach, seems the performance is not very effective. \n\nFirst of all, the video editing used seems not good. The edited videos are not very realistic from the examples shown in the paper. If the video editing method is not strong enough, it might be easy to protect the source video from effective editing.\n\nSecond, the protection result of VideoGuard is not good. As shown in the second row of Fig.3, VideoGuard fails to protect the source video. The immunized video is successfully changed by the prompt. Also, as shown in Fig. 5, the edited video protected by VideoGuard changed a lot, deviating a lot from the source video.\n\nFor the quantitative result, compared to the baseline without protection, the number did not make a large difference. I suppose the metric number will be very difference when comparing with and without protection."}},"nonreaders":[],"tmdate":1731428557301,"tcdate":1730925912911,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission6023/Reviewer_C7GU"],"signatures":["ICLR.cc/2025/Conference/Submission6023/Reviewer_C7GU"],"forum":"VRTCXYvPxc","number":3,"license":"CC BY 4.0","cdate":1730925912911,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission6023/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428557301,"domain":"ICLR.cc/2025/Conference","replyto":"VRTCXYvPxc","id":"yDbUNOCoWu","forumContent":{"TLDR":{"value":"Propose a method for protecting videos from malicious editing"},"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Diffusion Models","Video Editing Protection"]},"supplementary_material":{"value":"/attachment/20e3d2c3ac3822028b1bbb174a98eadc21a786e4.zip"},"primary_area":{"value":"alignment, fairness, safety, privacy, and societal considerations"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"With the rapid development of generative technology, current generative models can generate high-fidelity digital content and edit it in a controlled manner. However, there is a risk that malicious individuals might misuse these capabilities for misleading or unlawful activities. Although existing research has attempted to shield photographic images from being manipulated by generative models, there remains a significant disparity in the protection offered to video content editing. To bridge the gap, we propose a protection method named VideoGuard, which can effectively protect videos from unauthorized malicious editing. This protection is achieved through the subtle introduction of nearly unnoticeable perturbation that interferes with the functioning of the intended generative diffusion models. Different from images, videos consist of sequential frames, containing not only visual content but also motion dynamics. Due to the redundancy between video frames, and inter-frame attention mechanism in video diffusion models, simply applying image-based protection methods separately to every video frame can not shield video from unauthorized editing. To tackle the above challenge, rather than optimize perturbation in a frame-wise manner like image-based methods, we adopt joint frame optimization, treating all the video frames as an optimization entity. Furthermore, we extract video motion information and fuse it into optimization objectives. Thereby, these alterations can effectively compel the models to produce outputs that are implausible and inconsistent. We provide a pipeline to optimize such a perturbation. Finally, we use both objective metrics and subjective metrics to demonstrate the efficacy of our method, and the results show that the protection performance of VideoGuard is superior to all the baseline methods."},"_bibtex":{"value":"@misc{\ncao2025videoguard,\ntitle={{VIDEOGUARD}: {PROTECTING} {VIDEO} {CONTENT} {FROM} {UNAUTHORIZED} {EDITING}},\nauthor={Junjie Cao and Hongxiang Li and Xinchun Yu and Jindong Gu and Li KaiZhou and Yansong Tang and Xiao-Ping Zhang},\nyear={2025},\nurl={https://openreview.net/forum?id=VRTCXYvPxc}\n}"},"title":{"value":"VIDEOGUARD: PROTECTING VIDEO CONTENT FROM UNAUTHORIZED EDITING"},"pdf":{"value":"/pdf/b987d0e5587d1a49ef55bbb3a965e4484b23d7a6.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"cao|videoguard_protecting_video_content_from_unauthorized_editing"},"authorids":{"value":["~Junjie_Cao5","~Hongxiang_Li3","~Xinchun_Yu1","~Jindong_Gu1","~Li_KaiZhou1","~Yansong_Tang1","~Xiao-Ping_Zhang1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Junjie Cao","Hongxiang Li","Xinchun Yu","Jindong Gu","Li KaiZhou","Yansong Tang","Xiao-Ping Zhang"]}},"version":2},{"content":{"summary":{"value":"This paper proposes TC-Bench, a new benchmark designed to assess the Temporal Compositionality of video generation models. TC-Bench is divided into two components: TC-Bench-T2V, which includes 150 prompts for evaluating Text-to-Video (T2V) models across a spectrum of attributes, actions, and objects, defining the initial and final states of scenes; and TC-Bench-I2V, comprising 120 prompt-video pairs that serve as ground truth videos and reference data for Image-to-Video (I2V) models. The metrics introduced in this study demonstrate a significant correlation with human judgments, enhancing the evaluation of temporal compositionality."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"n/a"},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The focus on temporal compositionality in video generation is both novel and important, given the rapid advancements in conditional video generation. The benchmarks and quantitative experiments presented are valuable additions to the field.\n\n2. The methodology is robust, featuring comprehensive experiments with detailed explanations, facilitating replication and further study.\n\n3. The paper is well-written, structured for clarity."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The sizes of the test sets for both T2V (150) and I2V (120) benchmarks are relatively small, which may limit their ability to comprehensively analyze the capability for temporal compositionality.\n\n2. There is a lack of analysis on the distribution of the test sets, raising concerns about their representativeness of real-world scenarios relevant to the tasks.\n\n3. The evaluation is limited to only two methods, SEINE and DynamiCrafter, within the I2V category. This might not provide a full perspective on the field, given the variety of available I2V methods.\n\n4. The influence of the structure and length of the input prompts on video generation quality is a critical aspect that remains unexamined, which could impact the effectiveness of the benchmarks."}},"nonreaders":[],"tmdate":1731427751830,"tcdate":1730694950271,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission12250/Reviewer_p1Js"],"signatures":["ICLR.cc/2025/Conference/Submission12250/Reviewer_p1Js"],"forum":"xSOl0s1u77","number":3,"license":"CC BY 4.0","cdate":1730694950271,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission12250/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427751830,"domain":"ICLR.cc/2025/Conference","replyto":"xSOl0s1u77","id":"8tNoHxvk0o","forumContent":{"TLDR":{"value":"We propose a new benchmark suite to evaluate temporal compositionality for conditional video generation"},"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video Generation Benchmark; Text-to-Video Generation; Compositional Video Generation"]},"supplementary_material":{"value":"/attachment/e03d3a33aca0cdba2996ed4a1ab1550bc6d03314.zip"},"primary_area":{"value":"datasets and benchmarks"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Video generation has many unique challenges beyond those of image generation. The temporal dimension introduces extensive possible variations across frames, over which consistency and continuity may be violated. In this study, we move beyond evaluating simple actions and argue that generated videos should incorporate the emergence of new concepts and their relation transitions like in real-world videos as time progresses. To assess the \\textbf{T}emporal \\textbf{C}ompositionality of video generation models, we propose TC-Bench, a benchmark of meticulously crafted text prompts, corresponding ground truth videos, and robust evaluation metrics. The prompts articulate the initial and final states of scenes, effectively reducing ambiguities for frame development and simplifying the assessment of transition completion. In addition, by collecting aligned real-world videos corresponding to the prompts, we expand TC-Bench's applicability from text-conditional models to image-conditional ones that can perform generative frame interpolation. We also develop new metrics to measure the completeness of component transitions in generated videos, which demonstrate significantly higher correlations with human judgments than existing metrics. Our comprehensive experimental results reveal that most video generators achieve less than ～20% of the compositional changes, highlighting enormous space for future improvement. Our analysis indicates that current video generation models struggle to interpret descriptions of compositional changes and dynamically map varied semantics across different time steps."},"_bibtex":{"value":"@misc{\nfeng2025tcbench,\ntitle={{TC}-Bench: Benchmarking Temporal Compositionality in Conditional Video Generation},\nauthor={Weixi Feng and Jiachen Li and Michael Saxon and Tsu-Jui Fu and Wenhu Chen and William Yang Wang},\nyear={2025},\nurl={https://openreview.net/forum?id=xSOl0s1u77}\n}"},"title":{"value":"TC-Bench: Benchmarking Temporal Compositionality in Conditional Video Generation"},"pdf":{"value":"/pdf/7f8f685e2a00166cd013ca39bb988d26a6d6b16f.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"feng|tcbench_benchmarking_temporal_compositionality_in_conditional_video_generation"},"authorids":{"value":["~Weixi_Feng2","~Jiachen_Li6","~Michael_Saxon1","~Tsu-Jui_Fu2","~Wenhu_Chen3","~William_Yang_Wang2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Weixi Feng","Jiachen Li","Michael Saxon","Tsu-Jui Fu","Wenhu Chen","William Yang Wang"]}},"version":2},{"content":{"venue":{"value":"CVPR Workshops 2021"},"pdf":{"value":"https://openaccess.thecvf.com/content/CVPR2021W/DynaVis/papers/Bridgeman_Dynamic_Appearance_Modelling_From_Minimal_Cameras_CVPRW_2021_paper.pdf"},"venueid":{"value":"dblp.org/conf/CVPR/2021"},"paperhash":{"value":"bridgeman|dynamic_appearance_modelling_from_minimal_cameras"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Lewis_Bridgeman:","https://dblp.org/search/pid/api?q=author:Jean-Yves_Guillemaut:","~Adrian_Hilton1"]},"html":{"value":"https://openaccess.thecvf.com/content/CVPR2021W/DynaVis/html/Bridgeman_Dynamic_Appearance_Modelling_From_Minimal_Cameras_CVPRW_2021_paper.html"},"_bibtex":{"value":"@inproceedings{DBLP:conf/cvpr/BridgemanGH21,\n  author={Lewis Bridgeman and Jean-Yves Guillemaut and Adrian Hilton},\n  title={Dynamic Appearance Modelling From Minimal Cameras},\n  year={2021},\n  cdate={1609459200000},\n  pages={1760-1769},\n  url={https://openaccess.thecvf.com/content/CVPR2021W/DynaVis/html/Bridgeman_Dynamic_Appearance_Modelling_From_Minimal_Cameras_CVPRW_2021_paper.html},\n  booktitle={CVPR Workshops},\n  crossref={conf/cvpr/2021w}\n}\n"},"abstract":{"value":"We present a novel method for modelling dynamic texture appearance from a minimal set of cameras. Previous methods to capture the dynamic appearance of a human from multi-view video have relied on large, expensive camera setups, and typically store texture on a frame-by-frame basis. We fit a parameterised human body model to multi-view video from minimal cameras (as few as 3), and combine the partial texture observations from multiple viewpoints and frames in a learned framework to generate full-body textures with dynamic details given an input pose. Key to our method are our multi-band loss functions, which apply separate blending functions to the high and low spatial frequencies to reduce texture artefacts. We evaluate our method on a range of multi-view datasets, and show that our model is able to accurately produce full-body dynamic textures, even with only partial camera coverage. We demonstrate that our method outperforms other texture generation methods on minimal camera setups."},"title":{"value":"Dynamic Appearance Modelling From Minimal Cameras"},"authors":{"value":["Lewis Bridgeman","Jean-Yves Guillemaut","Adrian Hilton"]}},"tmdate":1731487783214,"pdate":1609459200000,"tcdate":1731484489463,"writers":["~"],"signatures":["~Adrian_Hilton1"],"forum":"on9ksi4d1D","license":"CC BY-SA 4.0","number":214333,"cdate":1609459200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1731487783214,"domain":"DBLP.org","id":"on9ksi4d1D","version":2},{"content":{"venue":{"value":"IEEE Trans. Neural Networks Learn. Syst. 2024"},"pdf":{"value":"https://ieeexplore.ieee.org/iel7/5962385/10772360/10250856.pdf"},"venueid":{"value":"dblp.org/journals/TNN/2024"},"paperhash":{"value":"ma|rectify_vit_shortcut_learning_by_visual_saliency"},"authorids":{"value":["~Chong_Ma2","~Lin_Zhao6","https://dblp.org/search/pid/api?q=author:Yuzhong_Chen:","https://dblp.org/search/pid/api?q=author:Lei_Guo_0002:","https://dblp.org/search/pid/api?q=author:Tuo_Zhang:","~Xintao_Hu1","https://dblp.org/search/pid/api?q=author:Dinggang_Shen:","https://dblp.org/search/pid/api?q=author:Xi_Jiang_0001:","https://dblp.org/search/pid/api?q=author:Tianming_Liu_0001:"]},"html":{"value":"https://doi.org/10.1109/TNNLS.2023.3310531"},"_bibtex":{"value":"@article{DBLP:journals/tnn/MaZCGZHSJL24,\n  author={Chong Ma and Lin Zhao and Yuzhong Chen and Lei Guo and Tuo Zhang and Xintao Hu and Dinggang Shen and Xi Jiang and Tianming Liu},\n  title={Rectify ViT Shortcut Learning by Visual Saliency},\n  year={2024},\n  month={December},\n  cdate={1733011200000},\n  journal={IEEE Trans. Neural Networks Learn. Syst.},\n  volume={35},\n  number={12},\n  pages={18013-18025},\n  url={https://doi.org/10.1109/TNNLS.2023.3310531}\n}\n"},"abstract":{"value":"Shortcut learning in deep learning models occurs when unintended features are prioritized, resulting in degenerated feature representations and reduced generalizability and interpretability. However, shortcut learning in the widely used vision transformer (ViT) framework is largely unknown. Meanwhile, introducing domain-specific knowledge is a major approach to rectifying the shortcuts that are predominated by background-related factors. For example, eye-gaze data from radiologists are effective human visual prior knowledge that has the great potential to guide the deep learning models to focus on meaningful foreground regions. However, obtaining eye-gaze data can still sometimes be time-consuming, labor-intensive, and even impractical. In this work, we propose a novel and effective saliency-guided ViT (SGT) model to rectify shortcut learning in ViT with the absence of eye-gaze data. Specifically, a computational visual saliency model (either pretrained or fine-tuned) is adopted to predict saliency maps for input image samples. Then, the saliency maps are used to filter the most informative image patches. Considering that this filter operation may lead to global information loss, we further introduce a residual connection that calculates the self-attention across all the image patches. The experiment results on natural and medical image datasets show that our SGT framework can effectively learn and leverage human prior knowledge without eye-gaze data and achieves much better performance than baselines. Meanwhile, it successfully rectifies the harmful shortcut learning and significantly improves the interpretability of the ViT model, demonstrating the promise of transferring human prior knowledge derived visual saliency in rectifying shortcut learning."},"title":{"value":"Rectify ViT Shortcut Learning by Visual Saliency"},"authors":{"value":["Chong Ma","Lin Zhao","Yuzhong Chen","Lei Guo","Tuo Zhang","Xintao Hu","Dinggang Shen","Xi Jiang","Tianming Liu"]}},"tmdate":1762403799407,"pdate":1704067200000,"tcdate":1744724755238,"writers":["~"],"signatures":["~Lin_Zhao6"],"forum":"NCHTyPxC0w","license":"CC BY-SA 4.0","number":397761,"cdate":1733011200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1762403799407,"domain":"DBLP.org","id":"NCHTyPxC0w","version":2},{"content":{"summary":{"value":"This paper introduces a benchmark suite for evaluating contextual positional bias of large video language models (LVLMs) considering long contexts. The data collection involves three stages,\n1. QA Generation: They collect videos from existing video-text resources including ego-centric videos, and generate frame-wise captions using GPT-4o. These captions are combined with crowd-sourced task definitions, and then an LLM generates question-answer pairs.\n2. QA Refinement: The generated question-answer pairs are filtered by GPT-4o, considering hallucinations and cases where the question does not require visual information in order to be answered. Human validators filter out the invalid question-answer pairs before proceeding to next step.\n3. Distractors: LLMs generate distractors (i.e., incorrect answer choices) for the collected question-answer pairs, later, again refined by human validators.\n\nAdditionally, the benchmark is divided into 7 sub-categories, which are OCR, Attribute Perception, Object Reasoning, Count Problem, Relationship Recognition, Action Reasoning, and Instructed Description. The benchmarked models are tested under different conditions, which are denoted as customized contexts. The set of customized contexts within this work include multiple videos inputs, long videos, multimodal interleaved input and lastly template video with ImageNet mean pixel values. This work introduces 3 different metrics to measure the positional bias, where all 3 metrics are centered around the relative score (RS) measure. These 3 metrics are, average relative score ( $P_{mean}$ ), difference between maximum and minimum relative score ( $P_{ran}$ ), and the variance in relative scores ( $P_{var}$ ). The evaluations on the proposed benchmark include 27 LVLMs including both open-weight and proprietary models where the model scale is ranging up to 108 billion parameters. Further studies investigates the effect of context type, context length and model size on positional bias."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"- I think it would be good to show visual examples of customized contexts in the appendix part.\n- There is some typo on Fig2: GTP-4o-latest should be GPT-4o-latest.\n- There is also white text on the last figure (Fig. 22) which can be barely seen on the background: `\"Please output the questions and reference answers in the following JSON format: [ {'question': 'xxx', 'answer': 'xxx’}, {'question': 'xxx', 'answer': 'xxx’},\"`"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":1},"contribution":{"value":3},"strengths":{"value":"- A novel benchmark for assessing the positional bias in video-language models.\n- Focuses an important aspect which could help to assess robustness of LVLMs under different positional settings, which could help further developing more trusthworthy LVLMs.\n- Detailed experimentation: 27 models, further beneficial studies on the effect of different design choices."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The presentation needs to be improved,\n\t- In Figure 3, the entire set of sub-tasks should be illustrated. This is not feasible.\n\t- In Figure 3, the tasks should be renamed by following the terminology already exist in the literature. I don't understand why someone should pose these tasks as *reasoning* tasks, where they are *recognition* tasks in fact. For instance, please see [Fig 1](https://arxiv.org/pdf/2306.13394) in MMU work to see the difference between reasoning and recognition type of tasks. So, the tasks should be renamed as,\n\t\t- Object Reasoning -> Object Recognition\n\t\t- Count Problem -> Object Counting (because actions can be counted also as well)\n\t\t- Action Reasoning -> Action Recognition\n\t\t- Attribute Perception -> Attribute Recognition (for consistency)\n\t- In Fig1(a), what do the numbers next to the tick and x marks?\n\t- Fig 2 is cluttered, I think it would be better if the only metric is positional variance.\n- Human refinement process remains entirely opaque. There are no details on this process, even how many human validators participated in the refinement process.\n- The proposed metric is not a standard metric used to evaluate the models, so I think this part should be expanded in the main text. This is such an important element of this paper but it only takes up 20 lines in the main text currently.\n\t- For instance, why should one use the relative score metric, and why should the standalone accuracy should be in the denominator?\n\t- What is the position $i$ ? What is the unit, seconds, or frame? Is this position absolute or relative? Are these positions randomly sampled for each individual example? Do these positions guarantee that the question-answer pairs remain valid for the video with custom context?\n\t- The morphological recognition (MR) term is confusing because morphology is a term which exists in NLP literature. Additionally, MR term currently seems unclear in the main text, and it is not motivated well enough."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1764360535539,"tcdate":1761589689596,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission1486/Reviewer_TpSR"],"signatures":["ICLR.cc/2026/Conference/Submission1486/Reviewer_TpSR"],"forum":"0V0bQi24YC","number":2,"license":"CC BY 4.0","cdate":1761589689596,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission1486/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1764360535539,"domain":"ICLR.cc/2026/Conference","replyto":"0V0bQi24YC","id":"ahh9IJRD2a","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Contextual Positional Bais","Video Benchmark","Large Video Language Model"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Large video language models (LVLMs) have made notable progress in video understanding, spurring the development of corresponding evaluation benchmarks. However, existing benchmarks generally assess overall performance across entire video sequences, overlooking nuanced behaviors such as contextual positional bias, a critical yet under-explored aspect of LVLM performance. We present **Video-LevelGauge**, a dedicated benchmark designed to systematically assess positional bias in LVLMs. We employ standardized probes and customized contextual setups, allowing flexible control over context length, probe position, and contextual types to simulate diverse real-world scenarios. In addition, we introduce a comprehensive analysis method that combines statistical measures with bias pattern recognition to characterize bias. Our benchmark comprises 438 manually curated videos spanning multiple types, yielding 1,177 high-quality multiple-choice questions and 120 open-ended questions, validated for their effectiveness in exposing positional bias. Based on these, we evaluate 27 state-of-the-art LVLMs, including both commercial and open-source models. Our findings reveal significant positional biases in many leading open-source models, typically exhibiting head or neighbor-content preferences. In contrast, commercial models such as Gemini 2.5 Pro show impressive, consistent performance across entire video sequences. Further analyses on context variation, context length, model scale, and multi-modal reasoning provide insights for mitigating bias and guiding model enhancement."},"_bibtex":{"value":"@inproceedings{\nxia2026videolevelgauge,\ntitle={Video-LevelGauge: Investigating Contextual Positional Bias in Video Language Models.},\nauthor={Hou Xia and Zheren Fu and Fangcan Ling and Jiajun Li and Yi Tu and Zhendong Mao and Yongdong Zhang},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=0V0bQi24YC}\n}"},"title":{"value":"Video-LevelGauge: Investigating Contextual Positional Bias in Video Language Models."},"pdf":{"value":"/pdf/ea3d1d7c4dbc7ea9a98066476a511430d8dd7e57.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"xia|videolevelgauge_investigating_contextual_positional_bias_in_video_language_models"},"authorids":{"value":["~Hou_Xia1","~Zheren_Fu1","~Fangcan_Ling2","~Jiajun_Li2","~Yi_Tu5","~Zhendong_Mao1","~Yongdong_Zhang2"]},"authors":{"value":["Hou Xia","Zheren Fu","Fangcan Ling","Jiajun Li","Yi Tu","Zhendong Mao","Yongdong Zhang"]}},"version":2},{"content":{"summary":{"value":"This paper introduces MMCLIMA, a large-scale multimodal framework for climate QA. Addressing the limitations of existing text-only and small-scale benchmarks, MMCLIMA includes over 104k expert-validated QA pairs spanning text, video transcripts, and scientific figures across five core climate domains. Beyond the dataset, MMCLIMA serves as a reusable framework for multimodal QA evaluation. The authors also present MMCLIMA-70B-TXT, a domain-adapted model that surpasses both open- and closed-source baselines, demonstrating the value of specialized multimodal resources for climate science."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1.The paper lacks an ablation study isolating the contribution of each modality. Specifically, it would be valuable to evaluate models using only text, only transcripts, and only images, then compare these results with the full multimodal setup. Such an analysis would clarify whether multimodal fusion truly introduces new knowledge and reasoning capabilities beyond text alone, or if the benchmark essentially remains text-dominated with limited added value from visual and video modalities.\n\n2.Although the paper highlights the strong performance of the fine-tuned MMCLIMA-70B-TXT, it lacks ablation studies to explain the source of its improvement. Experiments varying the amount of training data and applying the same fine-tuning process to different base models would clarify whether the gains arise from data quality or sheer quantity, and whether the dataset is truly generalizable or mainly effective for the chosen LLaMA-3.3-70B architecture."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"This work establishes a systematic and extensible framework for multimodal climate QA, integrating rigorous pipeline design with comprehensive benchmarking. Its multi-stage approach—combining automated claim extraction, factual verification, and expert validation—ensures unprecedented scale and reliability, surpassing previous text-centric benchmarks. The evaluation of 28 LLMs and 8 VLMs reveals critical inter-modal performance gaps, while domain-adapted fine-tuning demonstrates significant gains, highlighting the value of specialized data. This foundational resource advances multimodal scientific reasoning through its scalable methodology and actionable insights."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Although the benchmark includes visual question answering (VQA), the dataset is overwhelmingly dominated by textual QA pairs (100k), while video transcripts (8k) and visual QA pairs (around 4k by estimation) constitute only a small fraction. This imbalance may cause the benchmark to emphasize textual understanding over genuine multimodal reasoning, offering limited stress testing of models’ visual interpretation capabilities."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762922281172,"tcdate":1761989435904,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission11109/Reviewer_epXe"],"signatures":["ICLR.cc/2026/Conference/Submission11109/Reviewer_epXe"],"forum":"j9TdFswuZ3","number":4,"license":"CC BY 4.0","cdate":1761989435904,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission11109/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762922281172,"domain":"ICLR.cc/2026/Conference","replyto":"j9TdFswuZ3","id":"mig0RVQuj6","forumContent":{"venue":{"value":"ICLR 2026 Conference Desk Rejected Submission"},"keywords":{"value":["Multimodal Climate Benchmark","Scientific Foundation Models","Scientific Question Answering","Large Language Models","Automated QA generation"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Climate change research increasingly requires AI systems that can operate across multiple modalities, including natural language, dynamic visual content, and scientific figures. Yet existing climate QA benchmarks remain limited: they include relatively small sets of questions, rely almost exclusively on text, and evaluate only a narrow range of models. As a result, they fail to reflect the multimodal and large-scale nature of climate knowledge. In this work, we introduce MMClima, a multimodal framework for climate question answering. MMClima contains over 104k expert-validated question–answer pairs spanning text, video transcriptions, and figures, alongside covering a diverse range of five core climate science domains. The dataset is constructed through automated claim extraction combined with human-in-the-loop validation to ensure both scale and reliability. Beyond serving as a dataset, MMClima provides a reusable framework for extending QA resources across modalities. Using MMClima, we evaluate state-of-the-art multimodal language models on tasks spanning factual recall, visual interpretation, and cross-modal synthesis. We further fine-tune on the textual split, yielding mmclima-70b-txt, a domain-adapted baseline that surpasses both open- and closed-source models. Finally, we release the dataset, evaluation pipeline, fine-tuned model weights, and data creation framework as open resources, establishing the first step toward standardized multimodal evaluation in climate science."},"_bibtex":{"value":"@misc{\nanonymous2026mmclima,\ntitle={{MMC}lima: A Framework for Multimodal Climate Science Data and Evaluation},\nauthor={Anonymous},\nyear={2026},\nurl={https://openreview.net/forum?id=j9TdFswuZ3}\n}"},"title":{"value":"MMClima: A Framework for Multimodal Climate Science Data and Evaluation"},"pdf":{"value":"/pdf/7552adda18ee392d7c3e7c9cab83f339bfa618d2.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Desk_Rejected_Submission"},"paperhash":{"value":"sheikh|mmclima_a_framework_for_multimodal_climate_science_data_and_evaluation"},"authorids":{"value":["~Muhammad_Umer_Sheikh1","~Hassan_Abid1","~Khawar_shehzad1","~Ufaq_Khan1","~Muhammad_Haris_Khan3"]},"authors":{"value":["Muhammad Umer Sheikh","Hassan Abid","Khawar shehzad","Ufaq Khan","Muhammad Haris Khan"]}},"version":2},{"content":{"comment":{"value":"Dear Reviewer,\n\nWe thank you for recognizing the value of our paper to the community, i.e. exposing and addressing shortcuts with extensive evaluation and a large-scale curation process. Your suggestions were also very helpful and we incorporated them!\n\nWe will address each comment, as well as our respective changes to the paper and experiments. To make it easy to spot what changed in the updated paper PDF, we temporarily put new addition between red brackets “[[...]]”. \nNote: if you look at the main results Table 4 you will find that we added additional shortcut baselines on MVP (after a great suggestion from you). For now these were run on MVP-mini only to speed up experiments but we will run on the full dataset for the final paper.\n\n### **Clear presentation of results/numbers:**\nThank you for this suggestion, we have now unified the mentions of MVP-mini vs MVP-full more.\n\n### **Include fine-grained sub-task performance in the appendix:**\nAs suggested, we now include several appendix tables that break down MVP into many fine-grained categories (Appendix Section E). In addition, we now also elaborate in more detail in Section 4.1 “VideoLLMs performance on dataset sub-tasks” on where models are failing using these taxonomies.\n\n### **Evaluate more shortcut baselines on MVP:**\nThis particular suggestion will make our paper stronger and more coherent, thank you! So we take the same shortcut baseline we used to analyze existing benchmarks (Table 1) and apply them to MVP (Table 4). Specifically we add the missing video-only and Socratic LLM baseline. As noted above, for now this shows MVP-mini numbers and we will run on MVP-full for the final paper (the trend should be the same as it is a representative subset). We discuss the findings in Section 4.1.\n\n### **Include more figures to demonstrate MVP and qualitative results of the models:**\nWe now include 18 examples (2 for each original source in MVP) in the appendix as well as some examples of intuitive physics in the main paper, where models are either failing or partially succeeding.\nWe are happy to include even more examples in the main paper if more are helpful to the reader.\n\nOverall, we hope this addressed your main concerns and suggestions, and that our edits in the paper reflect this adequately. We are happy to discuss further in the following days if you have any follow up questions.."},"title":{"value":"Addressing main suggestions such as running more shortcut baselines on MVP"}},"parentInvitations":"TMLR/-/Official_Comment","tmdate":1758579605640,"tcdate":1758579605640,"writers":["TMLR","TMLR/Paper5382/Authors"],"signatures":["TMLR/Paper5382/Authors"],"forum":"gvFgNJcSw1","number":9,"license":"CC BY 4.0","cdate":1758579605640,"readers":["everyone"],"invitations":["TMLR/Paper5382/-/Official_Comment"],"mdate":1758579605640,"domain":"TMLR","replyto":"pv5MTxZO78","id":"aK8xui3R52","forumContent":{"submission_length":{"value":"Regular submission (no more than 12 pages of main content)"},"venue":{"value":"Accepted by TMLR"},"abstract":{"value":"Existing benchmarks for assessing the spatio-temporal understanding and reasoning abilities of video language models are susceptible to score inflation due to the presence of shortcut solutions based on superficial visual or textual cues. This paper mitigates the challenges in accurately assessing model performance by introducing the Minimal Video Pairs (MVP) benchmark, a simple shortcut-aware video QA benchmark for assessing the physical understanding of video language models. The benchmark is comprised of 55K high-quality multiple-choice video QA examples focusing on physical world understanding. Examples are curated from nine video data sources, spanning first-person egocentric and exocentric videos, robotic interaction data, and cognitive science intuitive physics benchmarks. To mitigate shortcut solutions that rely on superficial visual or textual cues and biases, each sample in MVP has a minimal-change pair — a visually similar video accompanied by an identical question but an opposing answer. To answer a question correctly, a model must provide correct answers for both examples in the minimal-change pair; as such, models that solely rely on visual or textual biases would achieve below random performance. Human performance on MVP is 92.9%, while the best open-source state-of-the- art video-language model achieves 40.2% compared to random performance at 25%."},"_bibtex":{"value":"@article{\nkrojer2025a,\ntitle={A Shortcut-aware Video-{QA} Benchmark for Physical Understanding via Minimal Video Pairs},\nauthor={Benno Krojer and Mojtaba Komeili and Candace Ross and Quentin Garrido and Koustuv Sinha and Nicolas Ballas and Mido Assran},\njournal={Transactions on Machine Learning Research},\nissn={2835-8856},\nyear={2025},\nurl={https://openreview.net/forum?id=gvFgNJcSw1},\nnote={}\n}"},"title":{"value":"A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs"},"pdf":{"value":"/pdf/f01dd06430bcd7b2aba068243d47011dcc96be6e.pdf"},"venueid":{"value":"TMLR"},"paperhash":{"value":"krojer|a_shortcutaware_videoqa_benchmark_for_physical_understanding_via_minimal_video_pairs"},"authorids":{"value":["~Benno_Krojer1","~Mojtaba_Komeili1","~Candace_Ross1","~Quentin_Garrido1","~Koustuv_Sinha1","~Nicolas_Ballas1","~Mido_Assran1"]},"assigned_action_editor":{"value":"~Tae-Hyun_Oh3"},"authors":{"value":["Benno Krojer","Mojtaba Komeili","Candace Ross","Quentin Garrido","Koustuv Sinha","Nicolas Ballas","Mido Assran"]}},"version":2},{"content":{"summary":{"value":"This paper proposes STANCE, a controllable image-to-video framework that focuses on rigid object interactions and aims to improve motion coherence and interaction plausibility. The method converts sparse 2.5D user cues into dense instance-aligned motion fields, enhances them with a Dense RoPE tagging scheme, and jointly predicts RGB and structural maps to stabilize training. Experiments show improved temporal and physical consistency compared to baselines."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"(3.2.1 Training)  \nHow do you distinguish individual instances here? Are the instance masks generated by a segmentation model, or are they included in the dataset itself? If they are obtained from a model such as SAM during inference, mentioning this explicitly in the training section would make it clearer."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. **Clear problem focus**  \n   The paper identifies practical challenges in motion-coherent video generation, particularly the loss of control signal density and the difficulty of achieving temporal and physical consistency.\n\n2. **Practical design**  \n   The sparse-to-dense instance cue formulation is intuitive, human-editable, and easy to integrate into existing video diffusion pipelines.\n\n3. **Reasonable overall design**  \n   While the components are not entirely novel, the overall framework is coherent and the design choices are appropriate for improving controllability and motion coherence in rigid-object video generation."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. **Limitation of Using 2.5D Conditioning**  \n   The method relies only on 2.5D cues for motion conditioning, while recent approaches such as *Diffusion as Shader* [1] have demonstrated the feasibility of incorporating full 3D conditioning for controlled video generation. Since 3D information can potentially enhance both spatial and physical coherence, which are central to this work, the decision to limit conditioning to 2.5D appears restrictive. Moreover, for example, point-based 3D representations are inherently denser than the 2D control masks discussed in the paper and could naturally alleviate the sparsity issue that the proposed method aims to address. A clearer justification for this design choice would strengthen the paper.\n\n2. **Limitation of Dense RoPE Design**  \n   Dense RoPE is an effective heuristic for retaining sparse motion cues, but it does not fundamentally resolve the sparsity of the input control signals. The task itself focuses on representing object motion, yet the current RoPE design does not explicitly account for dynamic spatial changes or motion-aware positional encoding. Recent work on dynamic or trajectory-aligned positional embeddings [2] has explored such aspects, and considering this trend, the use of a static first-frame RoPE feels somewhat outdated and not particularly efficient, making it difficult to see clear advantages of this choice.\n\n3. **Limited Novelty in Joint Auxiliary Generation**  \n   The idea of jointly generating auxiliary structural cues alongside RGB has been explored in prior work (e.g., [3], [4]) and is not particularly novel. While this strategy can help stabilize training and improve motion coherence, similar multi-head or multi-stream supervision schemes have been widely adopted in recent video generation and 3D-aware diffusion models. Therefore, it is difficult to attribute clear novelty to this component. If the authors intend to highlight this as a contribution, it would be helpful to cover relevant prior work in the related works section and clarify how their formulation differs from existing approaches.\n\n[1] *Diffusion as Shader: 3D-aware Video Diffusion for Versatile Video Generation Control*  \n[2] *RoPECraft: Training-Free Motion Transfer with Trajectory-Guided RoPE Optimization on Diffusion Transformers*  \n[3] *World-consistent Video Diffusion with Explicit 3D Modeling*  \n[4] *JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers*"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762916464207,"tcdate":1761787258671,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission2961/Reviewer_97AZ"],"signatures":["ICLR.cc/2026/Conference/Submission2961/Reviewer_97AZ"],"forum":"FwtKMYHov7","number":1,"license":"CC BY 4.0","cdate":1761787258671,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission2961/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762916464207,"domain":"ICLR.cc/2026/Conference","replyto":"FwtKMYHov7","id":"FAJACqB9j6","forumContent":{"venue":{"value":"ICLR 2026 Conference Desk Rejected Submission"},"keywords":{"value":["Video Generation","Generative Model"]},"supplementary_material":{"value":"/attachment/415a700d351d8bd244411f2900432bea49e53253.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Video generation has recently made striking visual progress, but maintaining coherent object motion and interactions remains difficult. We trace two practical bottlenecks: (i) human-provided motion hints (e.g., small 2D maps) often collapse to too few effective tokens after encoding, weakening guidance; and (ii) optimizing for appearance and motion in a single head can favor texture over temporal consistency. We present STANCE, an image-to-video framework that addresses both issues with two simple components.\nFirst, we introduce Instance Cues—a pixel-aligned control signal that turns sparse, user-editable hints into a dense 2.5D (camera-relative) motion field by averaging per-instance flow and augmenting with monocular depth over the instance mask. This reduces depth ambiguity compared to 2D drag/arrow inputs while remaining easy to user. Second, we preserve the salience of these cues in token space with Dense RoPE, which tags a small set of motion tokens (anchored on the first frame) with time-addressable rotary embeddings. Paired with joint RGB + auxiliary-map prediction (segmentation or depth), our model anchors structure while RGB handles appearance, stabilizing optimization and improving temporal coherence without requiring per-frame trajectory scripts."},"_bibtex":{"value":"@misc{\nanonymous2026stance,\ntitle={{STANCE}: Motion Coherent Video Generation Via Sparse-To-dense Anchored Encoding},\nauthor={Anonymous},\nyear={2026},\nurl={https://openreview.net/forum?id=FwtKMYHov7}\n}"},"title":{"value":"STANCE: Motion Coherent Video Generation Via Sparse-To-dense Anchored Encoding"},"pdf":{"value":"/pdf/dcb00196385b99164d59c430fb8613633f2432c0.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Desk_Rejected_Submission"},"paperhash":{"value":"chen|stance_motion_coherent_video_generation_via_sparsetodense_anchored_encoding"},"authorids":{"value":["~ZhiFei_Chen1","~Tianshuo_Xu1","~Leyi_Wu1","~Luozhou_Wang2","~Dongyu_Yan1","~Zihan_You2","~Wenting_Luo1","~Guo_Zhang2","~Ying-Cong_Chen1"]},"authors":{"value":["ZhiFei Chen","Tianshuo Xu","Leyi Wu","Luozhou Wang","Dongyu Yan","Zihan You","Wenting Luo","Guo Zhang","Ying-Cong Chen"]}},"version":2},{"content":{"TLDR":{"value":"We propose a context-aware learning pipeline for occlusion handling in the VIS task."},"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video Instance Segmentation","Contrastive Learning","Instance Prototype","Context-aware Learning"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"In this paper, we introduce the Context-Aware Video Instance Segmentation (CAVIS), a novel framework designed to enhance instance association by integrating contextual information adjacent to each object. To efficiently extract and leverage this information, we propose the Context-Aware Instance Tracker (CAIT), which merges contextual data surrounding the instances with the core instance features to improve tracking accuracy. Additionally, we introduce the Prototypical Cross-frame Contrastive (PCC) loss, which ensures consistency in object-level features across frames, thereby significantly enhancing instance matching accuracy. CAVIS demonstrates superior performance over state-of-the-art methods on all benchmark datasets in video instance segmentation (VIS) and video panoptic segmentation (VPS). Notably, our method excels on the OVIS dataset, which is known for its particularly challenging videos."},"_bibtex":{"value":"@misc{\nlee2025contextaware,\ntitle={Context-Aware Video Instance Segmentation},\nauthor={Seunghun Lee and Jiwan Seo and Kiljoon Han and Minwoo Choi and Sunghoon Im},\nyear={2025},\nurl={https://openreview.net/forum?id=VhQelEo27A}\n}"},"title":{"value":"Context-Aware Video Instance Segmentation"},"pdf":{"value":"/pdf/11cd1c2e65f6b7f1e81d00edbbd7f162483f2af3.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"lee|contextaware_video_instance_segmentation"},"authorids":{"value":["~Seunghun_Lee3","~Jiwan_Seo2","~Kiljoon_Han1","~Minwoo_Choi1","~Sunghoon_Im1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Seunghun Lee","Jiwan Seo","Kiljoon Han","Minwoo Choi","Sunghoon Im"]}},"tmdate":1738735619662,"tcdate":1726285508371,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission637/Authors"],"signatures":["ICLR.cc/2025/Conference/Submission637/Authors"],"forum":"VhQelEo27A","license":"CC BY 4.0","number":637,"cdate":1726285508371,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/-/Submission","ICLR.cc/2025/Conference/-/Post_Submission","ICLR.cc/2025/Conference/Submission637/-/Full_Submission","ICLR.cc/2025/Conference/Submission637/-/Rebuttal_Revision","ICLR.cc/2025/Conference/-/Edit"],"mdate":1738735619662,"odate":1728008565725,"domain":"ICLR.cc/2025/Conference","id":"VhQelEo27A","version":2},{"content":{"venue":{"value":"CoRR 2024"},"pdf":{"value":"http://arxiv.org/pdf/2407.03010v1"},"venueid":{"value":"dblp.org/journals/CORR/2024"},"paperhash":{"value":"lee|contextaware_video_instance_segmentation"},"authorids":{"value":["~Seunghun_Lee3","~Jiwan_Seo2","https://dblp.org/search/pid/api?q=author:Kiljoon_Han:","~Minwoo_Choi1","https://dblp.org/search/pid/api?q=author:Sunghoon_Im:"]},"html":{"value":"https://doi.org/10.48550/arXiv.2407.03010"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2407-03010,\n  publtype={informal},\n  author={Seunghun Lee and Jiwan Seo and Kiljoon Han and Minwoo Choi and Sunghoon Im},\n  title={Context-Aware Video Instance Segmentation},\n  year={2024},\n  cdate={1704067200000},\n  journal={CoRR},\n  volume={abs/2407.03010},\n  url={https://doi.org/10.48550/arXiv.2407.03010}\n}\n"},"abstract":{"value":"In this paper, we introduce the Context-Aware Video Instance Segmentation (CAVIS), a novel framework designed to enhance instance association by integrating contextual information adjacent to each object. To efficiently extract and leverage this information, we propose the Context-Aware Instance Tracker (CAIT), which merges contextual data surrounding the instances with the core instance features to improve tracking accuracy. Additionally, we introduce the Prototypical Cross-frame Contrastive (PCC) loss, which ensures consistency in object-level features across frames, thereby significantly enhancing instance matching accuracy. CAVIS demonstrates superior performance over state-of-the-art methods on all benchmark datasets in video instance segmentation (VIS) and video panoptic segmentation (VPS). Notably, our method excels on the OVIS dataset, which is known for its particularly challenging videos."},"title":{"value":"Context-Aware Video Instance Segmentation"},"authors":{"value":["Seunghun Lee","Jiwan Seo","Kiljoon Han","Minwoo Choi","Sunghoon Im"]}},"tmdate":1731490333873,"pdate":1704067200000,"tcdate":1726243924300,"writers":["~"],"signatures":["~Seunghun_Lee3"],"forum":"ICh6WPMVp8","license":"CC BY-SA 4.0","number":80753,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1731490333873,"domain":"DBLP.org","id":"ICh6WPMVp8","version":2},{"content":{"summary":{"value":"This paper highlights the rapid advancements in language-model-based video understanding, particularly through the development of Large Language Models (LLMs). It critiques previous research for relying on a basic projection layer to map video features to tokens, which is inefficient. The authors introduce VaQuitA, a novel framework that enhances the integration of video and textual information by employing CLIP-score guided frame sampling and a trainable Video Perceiver with a Visual-Query Transformer. The study also finds that adding the prompt \"Please be critical.\" significantly improves LLM video comprehension, and VaQuitA sets new benchmarks for zero-shot video question-answering tasks while facilitating high-quality multi-turn dialogues."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"No any other questions"},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The proposed method is effective for video-text alignment in video-llms\n\n2. the experimental results are sufficient for verifying the effectiveness of the proposed method"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Recently, some video datasets have been released to test long or short video understanding and reasoning, such as video-mme, video-vista, and MVBench. These datasets can better test the overall performance of models. \n\n2. Some video-llms should be compared, such as Video-CCAM-v1.0, CogVLM2-Video-Chat, LLaMA-VID, MiniGPT4-Video, and others."}},"nonreaders":[],"tmdate":1731427647750,"tcdate":1730622120129,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission2739/Reviewer_vZuT"],"signatures":["ICLR.cc/2025/Conference/Submission2739/Reviewer_vZuT"],"forum":"LSq9ef8ANs","number":1,"license":"CC BY 4.0","cdate":1730622120129,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission2739/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427647750,"domain":"ICLR.cc/2025/Conference","replyto":"LSq9ef8ANs","id":"k1a1uqZmkz","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video Understanding","Large Language Model","Alignment"]},"supplementary_material":{"value":"/attachment/d8063bb38e274ca115c84f3167b17958c85d5165.zip"},"primary_area":{"value":"foundation or frontier models, including LLMs"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Recent advancements in language-model-based video understanding have been progressing at a remarkable pace, spurred by the introduction of Large Language Models (LLMs). However, the focus of prior research has been predominantly on devising a projection layer that maps video features to tokens, an approach that is both rudimentary and inefficient. In our study, we introduce a cutting-edge framework, VaQuitA, designed to refine the synergy between video and textual information. At the data level, instead of sampling frames uniformly, we implement a sampling method guided by CLIP-score rankings, which enables a more aligned selection of frames with the given question. At the feature level, we integrate a trainable Video Perceiver alongside a Visual-Query Transformer (abbreviated as VQ-Former), which bolsters the interplay between the input question and the video features. We also discover that incorporating a simple prompt, ``Please be critical.'', into the LLM input can substantially enhance its video comprehension capabilities. Our experimental results indicate that VaQuitA consistently sets a new benchmark for zero-shot video question-answering tasks and is adept at producing high-quality, multi-turn video dialogues with users. The code will be released."},"_bibtex":{"value":"@misc{\nwang2024vaquita,\ntitle={VaQuitA: Enhancing Alignment in {LLM}-Assisted Zero-Shot Video Understanding},\nauthor={Yizhou Wang and Ruiyi Zhang and Haoliang Wang and Uttaran Bhattacharya and Yun Fu and Gang Wu},\nyear={2024},\nurl={https://openreview.net/forum?id=LSq9ef8ANs}\n}"},"title":{"value":"VaQuitA: Enhancing Alignment in LLM-Assisted Zero-Shot Video Understanding"},"pdf":{"value":"/pdf/1b60e905d09fa516be63289c5298441b1797fec8.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"wang|vaquita_enhancing_alignment_in_llmassisted_zeroshot_video_understanding"},"authorids":{"value":["~Yizhou_Wang3","~Ruiyi_Zhang3","~Haoliang_Wang1","~Uttaran_Bhattacharya1","~Yun_Fu1","~Gang_Wu4"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Yizhou Wang","Ruiyi Zhang","Haoliang Wang","Uttaran Bhattacharya","Yun Fu","Gang Wu"]}},"version":2},{"content":{"summary":{"value":"This paper tackles the problem of long sequence LLMs. The authors point out that the challenge lies in : video encoder and modality alignment projector are fixed, preventing the integration of additional frames into Video-LLMs, and the LLM backbone is limited in its content length capabilities, which complicates the processing of an increased number of video tokens. The authors introduce a video token rearrangement technique that circumvents limitations imposed by the fixed video encoder and alignment projector. Furthermore, a training-free LLM context window extension method is proposed to enable Video-LLMs to understand a correspondingly increased number of visual tokens."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"See weakness."},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The paper proposes a video token rearrangement technique that bypasses the restrictions imposed by the fixed video encoder and alignment projector. \n2. A training-free Video-LLM context window extension method is proposed to ensure that the interpolated Video-LLM can handle any number of video frames.\n3. The presentations are good."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The chosed baselines are not complete. For example, PLLaVA, Video-LLaMA 2, Flash-VStream are not included.\n2. Table 2 should include the #frames of each model.\n3. Some others works about LLM sequence extension should be discussed and analysised. For example, LongVA: Long Context Transfer from Language to Vision.\n4. Lack some benchmark evaluations: VideoMME, MoVQA, MVBench, etc.\n\n[1] PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning\n[2] VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs\n[3] Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams\n[4] MoVQA: A Benchmark of Versatile Question-Answering for Long-Form Movie Understanding."}},"nonreaders":[],"tmdate":1731427986521,"tcdate":1730642369176,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission3931/Reviewer_C5Zx"],"signatures":["ICLR.cc/2025/Conference/Submission3931/Reviewer_C5Zx"],"forum":"QrTvFCa4nX","number":2,"license":"CC BY 4.0","cdate":1730642369176,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission3931/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427986521,"domain":"ICLR.cc/2025/Conference","replyto":"QrTvFCa4nX","id":"7f9ZFuOy7R","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"TLDR":{"value":"Obtain a Video-LLM which can process more frames in a totally training-free manner."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video Understanding"]},"primary_area":{"value":"foundation or frontier models, including LLMs"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Advancements in Large Language Models (LLMs) inspire various strategies for integrating video modalities. \nA key approach is Video-LLMs, which incorporate an optimizable interface linking sophisticated video encoders to LLMs. \nHowever, due to computation and data limitations, these Video-LLMs are typically pre-trained to process only short videos, limiting their broader application for understanding longer video content. Additionally, fine-tuning Video-LLMs to handle longer videos is cost-prohibitive.\nConsequently, it becomes essential to explore the interpolation of Video-LLMs under a completely training-free setting. In this paper, we first identify the primary challenges in interpolating Video-LLMs: (1) the video encoder and modality alignment projector are fixed, preventing the integration of additional frames into Video-LLMs, and (2) the LLM backbone is limited in its content length capabilities, which complicates the processing of an increased number of video tokens.\nTo address these challenges, we propose a specific INTerPolation method for Video-LLMs (INTP-Video-LLMs). We introduce an alternative video token rearrangement technique that circumvents limitations imposed by the fixed video encoder and alignment projector. Furthermore, we introduce a training-free LLM context window extension method to enable Video-LLMs to understand a correspondingly increased number of visual tokens."},"_bibtex":{"value":"@misc{\nshang2025interpolating,\ntitle={Interpolating Video-{LLM}s:  Toward Longer-sequence {LMM}s in a Training-free Manner},\nauthor={Yuzhang Shang and Bingxin Xu and Weitai Kang and Mu Cai and Yuheng Li and Zehao Wen and Zhen Dong and Kurt Keutzer and Yong Jae Lee and Yan Yan},\nyear={2025},\nurl={https://openreview.net/forum?id=QrTvFCa4nX}\n}"},"title":{"value":"Interpolating Video-LLMs:  Toward Longer-sequence LMMs in a Training-free Manner"},"pdf":{"value":"/pdf/610d4d569fba56bc840ff0a6d6108535129e04d6.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"shang|interpolating_videollms_toward_longersequence_lmms_in_a_trainingfree_manner"},"authorids":{"value":["~Yuzhang_Shang1","~Bingxin_Xu1","~Weitai_Kang1","~Mu_Cai1","~Yuheng_Li1","~Zehao_Wen1","~Zhen_Dong3","~Kurt_Keutzer1","~Yong_Jae_Lee2","~Yan_Yan6"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Yuzhang Shang","Bingxin Xu","Weitai Kang","Mu Cai","Yuheng Li","Zehao Wen","Zhen Dong","Kurt Keutzer","Yong Jae Lee","Yan Yan"]}},"version":2},{"content":{"venue":{"value":"ICML 2026 regular"},"keywords":{"value":["Long-Tailed Learning","Imbalanced Learning"]},"_bibtex":{"value":"@inproceedings{\nwan2026shortcutresistant,\ntitle={Shortcut-Resistant {CAM} Distillation for Long-Tailed Recognition},\nauthor={Wenhai Wan and Teng Zhang and Shao-Yuan Li and Xinrui Wang and Qiang-Sheng Hua and Songcan Chen},\nbooktitle={Forty-third International Conference on Machine Learning},\nyear={2026},\nurl={https://openreview.net/forum?id=pDrQ4YeSIr}\n}"},"title":{"value":"Shortcut-Resistant CAM Distillation for Long-Tailed Recognition"},"paperhash":{"value":"wan|shortcutresistant_cam_distillation_for_longtailed_recognition"},"originally_submitted_PDF":{"value":"/pdf/7bf353f96af08b0740a782b91bd07b470f532fbb.pdf"},"TLDR":{"value":"Addressing the long-tail problem from the perspective of mitigating shortcut learning"},"primary_area":{"value":"deep_learning"},"abstract":{"value":"Real-world datasets often follow a long-tailed distribution, making generalization to tail classes difficult. We revisit this problem through the lens of shortcut learning, where models prefer the easiest predictive cues (e.g., background or textures) over object-centric semantics, especially under scarce and biased supervision. We find that this tendency is amplified for tail classes: limited examples often share similar contexts, making non-semantic signals highly correlated and thus tempting shortcuts, whereas head classes with diverse appearances and environments encourage more stable object-focused representations. Motivated by this observation, we propose Shortcut-Resistant CAM Distillation (SRCD), a plug-and-play framework that transfers object-focused explanations from head to tail classes. SRCD operates in the Class Activation Map (CAM) space, where a CAM provides a class-specific spatial evidence map for a prediction. SRCD aggregates CAMs from a small set of head-class candidates into a shortcut-resistant teacher using an energy-model weighting based on coherence and concentration, and distills it to the tail-class CAM. We provide a theoretical analysis that quantifies shortcut reliance as shortcut-region evidence mass in CAM space and shows that SRCD suppresses tail shortcuts. Extensive experiments on long-tailed benchmarks consistently improve strong baselines. The code is available at \\url{https://github.com/Haifeng3/SRCD}."},"pdf":{"value":"/pdf/fb3b650da44124e4f4cff8d3c152c516edd9ccf4.pdf"},"lay_summary":{"value":"Real-world image datasets are often unbalanced: some categories have many examples, while others have only a few. This makes it hard for AI models to recognize rare categories well. When only limited examples are available, the model may rely on misleading clues, such as background, color, or texture, instead of learning the actual object. For example, if a rare animal often appears with a similar background in the training data, the model may mistakenly treat that background as an important clue. In this paper, we study this problem and find that rare categories are more likely to suffer from such shortcut learning, while common categories usually provide more diverse examples that help the model focus on the object itself. Based on this idea, we propose Shortcut-Resistant CAM Distillation (SRCD), a simple method that uses more reliable visual evidence learned from common categories to guide the learning of rare categories. SRCD helps the model pay more attention to object-related regions and reduce its dependence on misleading background cues, without changing the model architecture. We also provide theoretical analysis to explain why this method can reduce shortcut learning, and experiments on several long-tailed image recognition benchmarks show that SRCD consistently improves existing methods, especially for rare categories."},"venueid":{"value":"ICML.cc/2026/Conference"},"authorids":{"value":["~Wenhai_Wan1","~Teng_Zhang3","~Shao-Yuan_Li1","~Xinrui_Wang4","~Qiang-Sheng_Hua1","~Songcan_Chen1"]},"authors":{"value":["Wenhai Wan","Teng Zhang","Shao-Yuan Li","Xinrui Wang","Qiang-Sheng Hua","Songcan Chen"]}},"tmdate":1790067405237,"pdate":1777576686067,"tcdate":1769169095804,"writers":["ICML.cc/2026/Conference","ICML.cc/2026/Conference/Submission24030/Authors"],"signatures":["ICML.cc/2026/Conference/Submission24030/Authors"],"forum":"pDrQ4YeSIr","license":"CC BY 4.0","number":24030,"cdate":1769169095804,"readers":["everyone"],"invitations":["ICML.cc/2026/Conference/-/Submission","ICML.cc/2026/Conference/-/Edit","ICML.cc/2026/Conference/-/Post_Submission","ICML.cc/2026/Conference/Submission24030/-/Full_Submission","ICML.cc/2026/Conference/Submission24030/-/Camera_Ready_Revision"],"mdate":1790067405237,"odate":1782341991524,"domain":"ICML.cc/2026/Conference","id":"pDrQ4YeSIr","version":2},{"content":{"summary":{"value":"The paper introduces a novel method, Weight Space Correlation Analysis (WSCA), for detecting and measuring shortcut learning. While most studies focus on detecting metadata signals in embeddings, the authors address whether a model actually uses these metadata signals (e.g., scanner/site) for clinical prediction, rather than merely encoding them in its representations. The method relies on metadata heads, training linear probes on frozen embeddings to predict metadata, projecting layer weight vectors into a PCA subspace of the embedding manifold, and computing cosine similarities between primary-task head weights and metadata-task head weights as a proxy for shared feature utilization. The authors evaluate the framework on two tasks: fetal ultrasound plane classification and preterm birth prediction on the SA-SonoNet dataset to assess relationships between prediction and various clinical/technical factors. The first experiment serves as a proof-of-concept with two controlled scenarios: i) scanner is encoded but not used, and ii) an induced-bias variant showing higher WSCA when plane–scanner correlation is introduced."},"justification_of_final_rating":{"value":"The authors methodically addressed all my concerns, added the aforementioned experiments, and strengthened the rationale for their method. I therefore recommend acceptance and would like to warmly thank the authors for this contribution."},"confidence":{"value":3},"final_rating":{"value":5},"justification_of_the_preliminary_rating":{"value":"The paper addresses a highly relevant problem: diagnosing shortcut learning and distinguishing encodability from the utilization of confounders. The authors propose a simple, implementable diagnostic (WSCA) with intuitive visualizations and a reasonable sanity check for induced bias. This makes the work potentially useful for practitioners working with multi-site / multi-scanner medical imaging.\n\nHowever, the current evidence does not yet justify the central interpretation that WSCA reliably reflects the actual reliance on confounders. The method is only weakly tied to interventional or distribution-shift outcomes (e.g., scanner-held-out generalization or balanced-scanner evaluation), and the heatmaps remain largely qualitative due to missing calibration (null distributions, seed variability, and guidance on what constitutes a meaningful correlation). In addition, the approach may be fragile to implementation choices (PCA, multi-class head parameterization). I therefore rate this paper as borderline, adding one decisive experiment that links WSCA to performance vulnerability, plus basic calibration and a brief PCA sensitivity check, would likely move this to Weak Accept."},"confidentiality_llm_acknowledgment":{"value":"Yes"},"strengths":{"value":"* The paper addresses an important and practical research gap in shortcut-learning analysis: distinguishing situations in which a model can predict metadata from embeddings from situations in which the model relies on metadata for the main task, which is highly relevant in medical imaging. \n* The proposed method is operationally simple and yields intuitive visualizations. \n* Experiments include a reasonably controlled validation where induced bias via data curation increases the measured correlations, emphasizing the potential of the proposed method for detecting and measuring shortcut learning."},"weaknesses":{"value":"* The methodological novelty is moderate. WSCA largely reuses familiar building blocks (linear probing, multi-task heads, and weight/representation geometry) into a diagnostic measure. While the framing of “utilization vs. encodability” is clear, it reads more like an incremental analysis lens than a fundamentally new learning framework or substantially new interpretability method.\n* The main point of the method is to treat weight-vector alignment between task- and metadata linear heads as evidence of confounder utilization. Such a hypothesis is plausible but not guaranteed: two heads can be related without weight alignment (representation can rotate/redistribute features), alignment can happen because both tasks correlate with the same true factor (or because the tasks predict each other under bias). \n* For softmax classifiers, pairwise decision boundaries depend on (wᵢ − wⱼ), not raw wᵢ alone. Correlating raw class weight vectors can be misleading due to invariances.\n* For the PCA, retaining 99% variance (with a floor of 50 PCs) may discard low-variance directions that still drive classification (common in shortcut settings). How do the results vary with changes in the number of PCs or the percentage of variance retained? \n* The paper would benefit from a clearer discussion of limitations/failure modes. For instance, would it be possible to observe cases where metadata is encoded, WSCA is low, yet we still observe performance collapse during scanner shifts? What should we conclude?"},"detailed_comments":{"value":"See weaknesses."},"questions_to_address_in_the_rebuttal":{"value":"1/ Add experiments whose results show that WSCA links to real shortcut reliance. For instance, perform an experiment on a balanced-scanner set (each class has the same number of each scanner) in which we would expect high encoding but low WSCA. \n\n2/ Provide a null WSCA distribution, allowing users to have an idea of when WSCA becomes significant and truly leads to shortcut learning.\n\n3/ Add experiments on robustness to the PCA/implementation choices."},"preliminary_rating":{"value":3}},"parentInvitations":"MIDL.io/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1771530043175,"tcdate":1767795533847,"writers":["MIDL.io/2026/Conference","MIDL.io/2026/Conference/Submission234/Reviewer_PjDp"],"signatures":["MIDL.io/2026/Conference/Submission234/Reviewer_PjDp"],"forum":"Os1ua32K56","number":2,"license":"CC BY 4.0","cdate":1767795533847,"readers":["everyone"],"invitations":["MIDL.io/2026/Conference/Submission234/-/Official_Review","MIDL.io/2026/Conference/-/Edit"],"mdate":1771530043175,"domain":"MIDL.io/2026/Conference","replyto":"Os1ua32K56","id":"Vy967Tz10C","forumContent":{"TLDR":{"value":"A method to quantify how much does a trained classifier model rely on irrelevant features in its prediction."},"venue":{"value":"MIDL 2026 Poster"},"midl_latex_submission_checklist":{"value":["The paper compiles correctly using the pdflatex compiler.","Created a single midl26_NNN.zip file with midl26_NNN.tex, midl26_NNN.bib, all necessary figures and files.","The LaTeX file includes the correct header commands before \\title.","Hyperref is preloaded; do not modify it.","The times package is not used.","All co-authors are listed with correct firstname, lastname, and any suffix/prefix.","Math in the title and abstract uses valid LaTeX.","References are provided through the .bib file only.","Tables and figures stay within the page margins.","The zip archive contains all necessary figures and no unused files.","Special formatting from rebuttal has been removed.","All special characters use LaTeX commands.","Appendices and supplementary material are included in the same PDF after references.","The main paper does not exceed 12 pages."]},"keywords":{"value":["shortcut learning","feature utilization","obstetric ultrasound"]},"reproducibility":{"value":"https://github.com/wong-ck/wsc-analysis"},"read_cfp_and_author_instructions":{"value":"Yes"},"originality_policy":{"value":"Yes"},"abstract":{"value":"Deep learning models in medical imaging are susceptible to shortcut learning, relying on confounding metadata (e.g. scanner model) that is often encoded in image embeddings. The crucial question is whether the model actively utilizes this encoded information for its final prediction. We introduce Weight Space Correlation analysis, an interpretable methodology that quantifies feature utilization by measuring the alignment between the classification heads of a primary clinical task and auxiliary metadata tasks. We first validate our method by successfully detecting artificially induced shortcut learning. We then apply it to probe the feature utilization of an SA-SonoNet model trained for Spontaneous Preterm Birth (sPTB) prediction. Our analysis confirmed that while the embeddings contain substantial metadata, the sPTB classifier's weight vectors were highly correlated with clinically relevant factors (e.g. cervical length) but decoupled from clinically irrelevant acquisition factors (e.g. scanner). Our methodology provides a tool for verifying model trustworthiness, by inspecting whether it utilizes features unrelated to the genuine clinical signal. Code available at https://github.com/wong-ck/wsc-analysis"},"_bibtex":{"value":"@inproceedings{\nwong2026weight,\ntitle={Weight Space Correlation Analysis: Quantifying Feature Utilization in Deep Learning Models},\nauthor={Chun Kit Wong and Paraskevas Pegios and Nina Weng and Emilie Pi Fogtmann Sejer and Martin Gr{\\o}nneb{\\ae}k Tolsgaard and Anders Nymark Christensen and Aasa Feragen},\nbooktitle={Medical Imaging with Deep Learning},\nyear={2026},\nurl={https://openreview.net/forum?id=Os1ua32K56}\n}"},"title":{"value":"Weight Space Correlation Analysis: Quantifying Feature Utilization in Deep Learning Models"},"latex_code":{"value":"/attachment/f0d19d5ca19899663f0003a6f73ab0e9f29d648b.zip"},"secondary_subject_area":{"value":"Integration of Imaging and Clinical Data"},"pdf":{"value":"/pdf/800af085a90aae8f9da5b0a8f0144eae85aa7c75.pdf"},"copyright_form":{"value":"/attachment/068ac4af1df9743d3f284cec1fa4b58b3db34226.pdf"},"visa":{"value":"No"},"single_blind_notice":{"value":"Yes"},"venueid":{"value":"MIDL.io/2026/Conference"},"paperhash":{"value":"wong|weight_space_correlation_analysis_quantifying_feature_utilization_in_deep_learning_models"},"primary_subject_area":{"value":"Fairness and Bias"},"authorids":{"value":["~Chun_Kit_Wong1","~Paraskevas_Pegios1","~Nina_Weng1","~Emilie_Pi_Fogtmann_Sejer1","~Martin_Grønnebæk_Tolsgaard1","~Anders_Nymark_Christensen1","~Aasa_Feragen2"]},"registration":{"value":"Yes"},"authors":{"value":["Chun Kit Wong","Paraskevas Pegios","Nina Weng","Emilie Pi Fogtmann Sejer","Martin Grønnebæk Tolsgaard","Anders Nymark Christensen","Aasa Feragen"]},"llm_policy_acknowledgment":{"value":"Yes"}},"version":2},{"content":{"summary":{"value":"This paper proposes a novel method for video scene segmentation based on the Minimum Description Length (MDL) principle. By leveraging the MDL property, the proposed approach achieves an elegant balance between two conflicting objectives in scene segmentation: ensuring intra-scene frame similarity while controlling the total number of scenes. Notably, this balance is attained without any manual parameter tuning or training. Experiments demonstrate that our method yields more accurate scene boundaries, particularly in long video settings, and further enhances the performance of two downstream tasks—long video summarization and long video question answering (long video QA)."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"Considering the weaknesses mentioned above, it would be appreciated if the authors could address the following questions: \n\n1.\tPlease clarify the motivation of the study and specify the particular problems it aims to address. \n\n2.\tWhen processing individual segments in isolation, does the method neglect essential global dependencies and potentially disrupt contextual coherence? \n\n3.\tPlease clarify the relationships among frames, scenes, partitions, and videos within the proposed framework. \n\n4.\tPlease provide an explicit definition or formulation of the final cost for a given partition. \n\n5.\tCan the proposed method accurately segment scenes that share similar statistical characteristics but differ in semantic content (e.g., transitional shots)? \n\n6.\tPlease clarify the demonstrations of the two downstream tasks discussed in Section 4, as mentioned in the fifth weakness item. \n\n7.\tLines 367–369 indicate that the proposed method achieves a similar inference speed to more complex methods such as BASSL and NeighborNet. However, Table 1 shows that MDLSeg performs similarly to these methods on the OVSD dataset, but significantly outperforms them on the BBC dataset (approximately 24× faster than BASSL and 2.5× faster than NeighborNet). Please explain the reason for this discrepancy. \n\n8.\tPlease clarify the “w/o names” setting mentioned in Table 2. \n\n9.\tPlease clarify the purpose of the “w/o input” setting in Table 3, given that the proposed method does not involve a training process."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":1},"contribution":{"value":2},"strengths":{"value":"1.\tThe paper proposes a novel MDL-based method for video scene segmentation that is both parameter-free and training-free. This design elegantly balances two conflicting objectives in clustering-based problems—preserving intra-scene frame similarity while controlling the overall number of scenes. The simplicity and conceptual clarity of the approach are commendable.  \n\n2.\tThe exploration of downstream tasks, including long video summarization and question answering, demonstrates improved performance facilitated by the proposed scene segmentation method. This indicates that better scene segmentation can effectively enhance long-video understanding."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Despite the strengths demonstrated above, this paper should be rejected for the following main reasons: \n\n1.\tThe motivation and the specific problems addressed in this study are not clearly articulated, which raises questions about the significance and necessity of the proposed method. \n\n2.\tThe effectiveness of independent chunk understanding in video summarization, as presented in lines 35–37, is questionable. Effective video summarization requires a holistic comprehension of the overall narrative and structural flow. Processing individual segments in isolation may neglect essential global dependencies and disrupt contextual coherence. \n\n3.\tThe method is presented in a confusing manner without sufficient preliminaries or clear equations to define the cost of a partition. The unclear relationships among frames, scenes, partitions, and videos further obscure the overall design.  \n\n4.\tThere is a concern regarding the fine-grained segmentation and generalization capabilities of the method while dealing with scenes that share similar statistical characteristics but differ in semantic content, which often occurs in transitional shots. The MDL-based approach may fail to correctly segment such scenes due to the MDL principle’s emphasis on minimizing description length. \n\n5.\tThe demonstrations of the two downstream tasks in Section 4 are unclear. The claim in line 243 that the system “requires only video input” neglects the character bank from IMDB. In line 250, the phrase “timestamps in the transcript” is ambiguous, raising the question of whether these timestamps are aligned with the boundaries of the segmented scenes. Moreover, it is unclear in line 254 how cosine similarity is computed between a multi-frame scene and a single question. \n\n6.\tThe experimental section is insufficient in several aspects:\n\n(1) Lines 194–196 introduce a hyperparameter L, which is stated to achieve the best performance when set to approximately 10 minutes; however, no explicit ablation study is provided to support this claim.\n\n(2) Commonly used evaluation metrics for video scene segmentation—such as AP, mIoU, AUC-ROC, and F1 [1, 2, 3]—should also be considered and compared in the experiments.\n\n(3) The conclusion presented in lines 456–459 lacks experimental validation. It would be beneficial to include comparative experiments on different sets of videos to substantiate this conclusion. \n\n7.\tThe analysis of the “w/o names” setting in paragraph Summarization of Section 6 is confusing, as this setting is discussed in the text but not shown in Table 2. \n\n8.\tThe purpose of the “w/o input” setting in Table 3 is debatable, as the proposed method does not involve any training process. Therefore, it seems unnecessary to discuss the potential contamination issue within the proposed system. \n\nBesides, the paper is imprecise and unpolished: \n\n1.\tThe title “Video Understanding” in the second paragraph of the Related Work section is broader than the actual content of the paragraph. As outlined in [4], video understanding generally includes three major categories: video content understanding, descriptive understanding, and video content generation and manipulation. However, the paragraph primarily discusses video captioning, video summarization, and video question answering, which belong specifically to descriptive understanding tasks. \n\n2.\tThe target task is ambiguously defined, being inconsistently referred to as video segmentation, scene segmentation, and scene detection. \n\n3.\tThe coverage of prior work is incomplete. Despite the claim in Section 6 that all existing scene segmentation methods are included in Section 2, relevant studies such as [5, 6] are omitted. \n\n4.\tLines 370–377 focus on video summarization rather than scene detection and should be relocated to the subsequent paragraph for clarity. \n\nPresentation and Formatting Issues:\n\n1.\tThere are extensive citation formatting errors that need to be corrected. For example, author-year references are sometimes written without parentheses.\n\n2.\tThe frame size in footnote on Page 2 can be better written as “1024\\times 1024”. \n\n3.\t“scaped” in line 211 should be “scraped”. \n\n4.\tThere should be a space after the full stop symbol in line 101 and line 316. \n\n5.\tThe table in lines 378–394 lacks a caption and the whole table appears to have been deleted. \n\n6.\t“eTable 2” in line 426 should be “Table 2”. \n\n7.\tCitation “Ataallah et al.” in line 464 is repeated. \n\n8.\tMethod qwen-vl in line 468 should be cited. \n\n[1] Mun, Jonghwan, et al. \"Bassl: Boundary-aware self-supervised learning for video scene segmentation.\" Proceedings of the Asian Conference on Computer Vision. 2022.\n\n[2] Wu, Haoqian, et al. \"Scene consistency representation learning for video scene segmentation.\" Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2022.\n\n[3] Tan, Jiawei, et al. \"Neighbor Relations Matter in Video Scene Detection.\" Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2024.\n\n[4] Madan, Neelu, et al. \"Foundation models for video understanding: A survey.\" arXiv preprint arXiv:2405.03770 (2024).\n\n[5] Chen, Shixing, et al. \"Shot contrastive self-supervised learning for scene boundary detection.\" Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2021.\n\n[6] Islam, Md Mohaiminul, et al. \"Efficient movie scene detection using state-space transformers.\" Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2023."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762926951503,"tcdate":1761143591320,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission16927/Reviewer_2KJT"],"signatures":["ICLR.cc/2026/Conference/Submission16927/Reviewer_2KJT"],"forum":"uh6aDR1jlw","number":1,"license":"CC BY 4.0","cdate":1761143591320,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission16927/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762926951503,"domain":"ICLR.cc/2026/Conference","replyto":"uh6aDR1jlw","id":"hlEndp6H8y","forumContent":{"TLDR":{"value":"An algorithm for splitting a video into scenes that requires no set parameters, is more accurate than existing methods and movie summarisation and VQA."},"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["minimum description length","scene break detection","summarisation"]},"primary_area":{"value":"other topics in machine learning (i.e., none of the above)"},"abstract":{"value":"The proliferation of creative video content has driven demand for adapting language models to handle video input and enable multimodal understanding. However, end-to-end models struggle to process long videos due to their size and complexity. An effective alternative is to divide them into smaller chunks to be processed separately, and this motivates a method for choosing where the chunk boundaries should be. In this paper, we propose an algorithm for segmenting videos into contiguous chunks, based on the minimum description length principle, coupled with a dynamic programming search. The algorithm is entirely parameter-free, given feature vectors, not requiring a set threshold or the number or size of chunks to be specified. We show empirically that the breakpoints it produces more accurately approximate scene boundaries in long videos, compared with existing methods for scene detection, even when such methods have access to the true number of scenes. We then showcase this algorithm in two tasks: long video summarization, and retrieval-augmented video question answering. In both cases, scene breaks produced by our algorithm lead to better downstream performance than existing methods for video segmentation."},"_bibtex":{"value":"@misc{\nmahon2025a,\ntitle={A Parameter-free Scene Detection Algorithm for Vision and Language Understanding},\nauthor={Louis Mahon and Mirella Lapata},\nyear={2025},\nurl={https://openreview.net/forum?id=uh6aDR1jlw}\n}"},"title":{"value":"A Parameter-free Scene Detection Algorithm for Vision and Language Understanding"},"pdf":{"value":"/pdf/9a30b8346ca2cc33162d5c11eef070ff9239b014.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"mahon|a_parameterfree_scene_detection_algorithm_for_vision_and_language_understanding"},"authorids":{"value":["~Louis_Mahon1","~Mirella_Lapata1"]},"authors":{"value":["Louis Mahon","Mirella Lapata"]}},"version":2},{"content":{"summary":{"value":"This paper extends GRPO for Video QA by proposing a spatiotemporal grouping strategy. The method, ST-GRPO, aims to solve the low reward variance problem in standard GRPO by creating groups across different spatiotemporal video variants, which is shown to stabilize training and improve performance on six VQA benchmarks."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"1. There appears to be a notational gap; how does the \"importance-based grouping\" strategy, which generates video variants, connect to the final GxM responses used in the objective function?\n\n2. The paper hypothesizes that \"importance-based\" sampling improves temporal reasoning, yet experimental results show it performs negligibly better than \"stochastic\" sampling, questioning its unique benefit.\n\n3. The ST-GRPO objective function does not seem to include any explicit reward for spatio-temporal correctness, so how does the model learn this, as suggested in the qualitative examples?\n\n4. The reward analysis only plots the reward standard deviation; could the authors also provide the mean reward curve over training steps to better understand the learning dynamics?\n\n5. Does the proposed spatio-temporal transformation (e.g., temporal cropping or importance sampling) risk omitting critical frames, leading to a mismatch where the video variant no longer contains the answer to the question?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. The proposed ST-GRPO method achieves consistent, albeit modest, performance gains over the standard GRPO baseline across all six evaluated benchmarks.\n\n2. The method's effectiveness is validated across a wide range of six different video understanding benchmarks."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The performance improvement from ST-GRPO is limited.\n\n2. The method's novelty is somewhat incremental, as it combines existing techniques like GRPO, video augmentation, and question-aware frame relevance scoring.\n\n3. The paper's core claims lack sufficient experimental support; for instance, the \"Importance-Based\" grouping performs almost identically to \"Stochastic\" grouping, questioning its specific contribution."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762927951294,"tcdate":1761594550015,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission18206/Reviewer_PXxo"],"signatures":["ICLR.cc/2026/Conference/Submission18206/Reviewer_PXxo"],"forum":"UVhAIvnmbz","number":2,"license":"CC BY 4.0","cdate":1761594550015,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission18206/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762927951294,"domain":"ICLR.cc/2026/Conference","replyto":"UVhAIvnmbz","id":"VvESkgcGwv","forumContent":{"venue":{"value":"ICLR 2026 Conference Desk Rejected Submission"},"keywords":{"value":["Video Question Answering","Large Multimodal Models","Post-Training"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"We introduce SpatioTemporal-GRPO (ST-GRPO), a novel extension of the GRPO algorithm for video question answering. ST-GRPO addresses a limitation of standard GRPO: when all responses in a group have similar correctness, the low reward variance gives the model an uninformative signal for improvement. Our method overcomes this by generating multiple spatiotemporal variants of a video to serve as complementary inputs. Unlike standard GRPO, which only groups textual responses, ST-GRPO forms groups across both textual and spatiotemporal variants. This increases reward variance within each group, providing a more informative signal for learning. To ensure these visual variations are meaningful, we propose an importance-based grouping strategy. This approach computes per-frame relevance scores using cross-modal embeddings, prioritizing frames that carry higher semantic weight relative to the question. This question-aware method ensures our spatiotemporal groups are informed by the relevant visual cues for each query. Our experiments demonstrate consistent improvements across six challenging video understanding benchmarks, including VideoMME, TempCompass, VideoMMMU, MMVU, VSI-Bench, and PerceptionTest, showing that incorporating structured visual diversity into reinforcement learning provides a more effective approach for learning from spatiotemporal cues in video question answering."},"_bibtex":{"value":"@misc{\nanonymous2026spatiotemporalgrpo,\ntitle={SpatioTemporal-{GRPO}: Post-Training Large Multimodal Models for Video {QA}},\nauthor={Anonymous},\nyear={2026},\nurl={https://openreview.net/forum?id=UVhAIvnmbz}\n}"},"title":{"value":"SpatioTemporal-GRPO: Post-Training Large Multimodal Models for Video QA"},"pdf":{"value":"/pdf/f08af00877b1f7935b435b9b2ec41c89911f80b4.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Desk_Rejected_Submission"},"paperhash":{"value":"bahrami|spatiotemporalgrpo_posttraining_large_multimodal_models_for_video_qa"},"authorids":{"value":["~Emad_Bahrami1","~Olga_Zatsarynna1","~Parth_Pathak3","~Sunando_Sengupta1","~Juergen_Gall1","~Mohsen_Fayyaz2"]},"authors":{"value":["Emad Bahrami","Olga Zatsarynna","Parth Pathak","Sunando Sengupta","Juergen Gall","Mohsen Fayyaz"]}},"version":2},{"content":{"venue":{"value":"TrustCom/BigDataSE 2018"},"pdf":{"value":"https://ieeexplore.ieee.org/iel7/8454845/8455868/08456000.pdf"},"venueid":{"value":"dblp.org/conf/TRUSTCOM/2018"},"paperhash":{"value":"wang|pairs_privacyaware_identification_and_recommendation_of_spatiofriends"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Shuo_Wang_0012:","~Richard_O._Sinnott1","https://dblp.org/search/pid/api?q=author:Surya_Nepal:"]},"html":{"value":"https://doi.org/10.1109/TrustCom/BigDataSE.2018.00131"},"_bibtex":{"value":"@inproceedings{DBLP:conf/trustcom/WangSN18,\n  author={Shuo Wang and Richard O. Sinnott and Surya Nepal},\n  title={PAIRS: Privacy-Aware Identification and Recommendation of Spatio-Friends},\n  year={2018},\n  cdate={1514764800000},\n  pages={920-931},\n  url={https://doi.org/10.1109/TrustCom/BigDataSE.2018.00131},\n  booktitle={TrustCom/BigDataSE},\n  crossref={conf/trustcom/2018}\n}\n"},"abstract":{"value":"Due to the prevalence of location-based services, it has now become possible to infer social connections between people by observing their spatial behaviors over time. Such spatial behaviors, if shared, can be utilized to identify and recommend friends for web-based social service users. However, such approaches cannot be implemented without solving two key challenges: (a) guaranteeing an individual's privacy in the shared spatiotemporal data, and (b) addressing the inherent sparseness of shared spatiotemporal data. In this paper, we propose a Privacy-Aware Identification and Recommendation of Spatio-Friends (PAIRS) approach, that can infer and recommend potential social connections by analyzing spatiotemporal information of social media users using robust privacy guarantee mechanisms. To achieve this, PAIRS constructs co-occurrence profiles using a cluster-based anchor representation to alleviate the sparseness of shared spatiotemporal information. It utilizes the diversity, time and weighted frequency-based inference to efficiently infer the strength of potential social connections from co-occurrence profile by reducing the negative impact of coincidences and thereby enhances accuracy. To tackle the privacy concerns, PAIRS sanitizes the cluster-based anchors, the location entropy values as well as the co-occurrence profile under differential privacy, including optimization mechanisms to handle trade-offs in utility and privacy. Extensive experiments are conducted with real-world datasets including both individuals' spatiotemporal data and their actual social connections. We confirm that our approach can achieve two often contradictory goals: a provable robust privacy protection for sharing data and an efficient social strength inference and spatio-friend identification mechanism. Specifically, PAIRS remains approximately 70% accuracy (precision) and 80% efficiency (recommendation potential) after perturbation."},"title":{"value":"PAIRS: Privacy-Aware Identification and Recommendation of Spatio-Friends"},"authors":{"value":["Shuo Wang","Richard O. Sinnott","Surya Nepal"]}},"tmdate":1747798334208,"pdate":1514764800000,"tcdate":1747798265451,"writers":["~"],"signatures":["~Richard_Sinnott1"],"forum":"QLt3LFcIqE","license":"CC BY-SA 4.0","number":545940,"cdate":1514764800000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1747798334208,"domain":"DBLP.org","id":"QLt3LFcIqE","version":2},{"content":{"summary":{"value":"This paper proposes a novel method termed InfVSR, aiming to achieve efficient and temporally-scalable diffusion-based video super-resolution. InfVSR firstly adopts a pretrained DiT into causal structure and maintaining local and global coherence with rolling KV-cache, then distills the model with distribution matching to achieve one-step diffusion inference. This paper also proposes MovieLQ, a long-sequence video benchmark to evaluate the VSR in long-term semantic-level consistency, fidelity and efficiency."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Refer to the weaknesses. The visualization is not very persuasive. The novelty should be clearified with detailed discussion. How about using prompt to boost the SR performance and guide the semantic content?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1.\tThe idea is easy to follow. It adapts a pretrained T2V DiT-based diffusion model to causal form. Then applies distribution matching distillation to achieve one-step diffusion inference. This is effective and efficient for the model to perform unlimited length super-resolution video prediction together with the rolling KV-cache design.\n2.\tA good and straight-forward solution on long-sequence video super-resolution. Both semantic and pixel consistencies are maintained through DMD loss and pixel-level reconstruction loss.\n3.\tA new benchmark VideoLQ is proposed to evaluate the long-sequence video super-resolution task."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1.\tThe long-term consistency is not thoroughly discussed in the paper, e.g., what is the visual results and comparison between the SOTA methods after 10, 100 or 1000 frames? Also the visual results in the supplement seems not fidel to the GT in some text images, this seems to be brought by the generative instability. \n2.\tThe main idea of improving the efficiency of video model using DMD and causal structure is trivial since it has been adopted in many other video methods [1][2]. Some more discussions are needed to specify the novelty of this paper.\n3.\tThe InfVSR is built upon the Wan T2V model, also there are models like SeeSR[3] using text to enhance the super-resolution results. Can model achieve better result with proper text prompt or guidance?\n4.\tThe parameter compared with SOTA methods should be provided to better validate the efficiency of the proposed method.\n\n[1] Self-forcing: bridging the train-test gap in autoregressive video diffusion. arXiv preprint arXiv:2506.08009\n[2] Matrix-Game 2.0: An Open-Source, Real-Time, and Streaming Interactive World Model arXiv preprint arXiv:2508.13009\n[3] Seesr: Towards semantics-aware real-world image super-resolution CVPR 2024"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762915788995,"tcdate":1761837778905,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission1517/Reviewer_Uppd"],"signatures":["ICLR.cc/2026/Conference/Submission1517/Reviewer_Uppd"],"forum":"fZi8HxJbMO","number":4,"license":"CC BY 4.0","cdate":1761837778905,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission1517/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762915788995,"domain":"ICLR.cc/2026/Conference","replyto":"fZi8HxJbMO","id":"EcmZJIZ5uq","forumContent":{"TLDR":{"value":"A one-step diffusion based AR method for video super-resolution"},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Video Super-Resolution","One-Step Diffusion","Auto-Regression"]},"supplementary_material":{"value":"/attachment/f5a4ca8f80f5f3be9f486af6b3a2129044fa1302.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Real-world videos often extend over thousands of frames, posing unique demands far beyond current short benchmarks. Existing video super-resolution (VSR) approaches, however, face two persistent challenges when processing long sequences: (1) Efficiency due to the heavy cost of multi-step denoising for full-length sequences; and (2) Scalability hindered by temporal decomposition that causes artifacts and discontinuities. To break these limits, we propose InfVSR, which novelly reformulate VSR as an autoregressive-one-step-diffusion paradigm. This enables streaming inference while fully leveraging pre-trained video diffusion priors. First, we adapt the pre-trained DiT into a causal structure, maintaining  both local and global coherence via rolling KV-cache and joint visual guidance. Second, we distill diffusion process into a single step efficiently, with patch-wise pixel supervision and cross-chunk distribution matching. Together, these designs enable efficient and scalable VSR for unbounded-length videos. To fill the gap in long-form video evaluation, we build a new benchmark tailored for extended sequences, and further introduce semantic-level metrics to comprehensively assess temporal consistency. Our method pushes the frontier of long-form VSR, achieves state-of-the-art quality with enhanced semantic consistency, and delivers up to 58x speed-up over existing methods such as MGLD-VSR. Code will be released soon."},"_bibtex":{"value":"@misc{\nzhang2026infvsr,\ntitle={Inf{VSR}: Breaking Length Limits of Generic Video Super-Resolution},\nauthor={Ziqing Zhang and Kai Liu and Zheng Chen and Xi Li and Yucong Chen and Bingnan Duan and Linghe Kong and Yulun Zhang},\nyear={2026},\nurl={https://openreview.net/forum?id=fZi8HxJbMO}\n}"},"title":{"value":"InfVSR: Breaking Length Limits of Generic Video Super-Resolution"},"pdf":{"value":"/pdf/e812c789a62dc7257bfe4b5756e0c9a11f67fd28.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"zhang|infvsr_breaking_length_limits_of_generic_video_superresolution"},"authorids":{"value":["~Ziqing_Zhang2","~Kai_Liu21","~Zheng_Chen11","~Xi_Li6","~Yucong_Chen3","~Bingnan_Duan2","~Linghe_Kong1","~Yulun_Zhang1"]},"authors":{"value":["Ziqing Zhang","Kai Liu","Zheng Chen","Xi Li","Yucong Chen","Bingnan Duan","Linghe Kong","Yulun Zhang"]}},"version":2},{"content":{"venue":{"value":"CoRR 2017"},"pdf":{"value":"https://arxiv.org/pdf/1710.00974v1"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"li|a_concatenating_framework_of_shortcut_convolutional_neural_networks"},"html":{"value":"http://arxiv.org/abs/1710.00974"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-1710-00974,\n  publtype={informal},\n  author={Yujian Li and Ting Zhang and Zhaoying Liu and Haihe Hu},\n  title={A concatenating framework of shortcut convolutional neural networks},\n  year={2017},\n  cdate={1483228800000},\n  journal={CoRR},\n  volume={abs/1710.00974},\n  url={http://arxiv.org/abs/1710.00974}\n}\n"},"abstract":{"value":"It is well accepted that convolutional neural networks play an important role in learning excellent features for image classification and recognition. However, in tradition they only allow adjacent layers connected, limiting integration of multi-scale information. To further improve their performance, we present a concatenating framework of shortcut convolutional neural networks. This framework can concatenate multi-scale features by shortcut connections to the fully-connected layer that is directly fed to the output layer. We do a large number of experiments to investigate performance of the shortcut convolutional neural networks on many benchmark visual datasets for different tasks. The datasets include AR, FERET, FaceScrub, CelebA for gender classification, CUReT for texture classification, MNIST for digit recognition, and CIFAR-10 for object recognition. Experimental results show that the shortcut convolutional neural networks can achieve better results than the traditional ones on these tasks, with more stability in different settings of pooling schemes, activation functions, optimizations, initializations, kernel numbers and kernel sizes."},"title":{"value":"A concatenating framework of shortcut convolutional neural networks"},"authors":{"value":[{"fullname":"Yujian Li","username":""},{"fullname":"Ting Zhang","username":""},{"fullname":"Zhaoying Liu","username":"~Zhaoying_Liu1"},{"fullname":"Haihe Hu","username":""}]}},"tmdate":1784721017638,"pdate":1514678400000,"externalIds":["dblp:journals/corr/abs-1710-00974"],"tcdate":1784721012140,"writers":["~","OpenReview.net/Public_Article/DBLP.org","OpenReview.net/Support"],"signatures":["~Zhaoying_Liu1"],"forum":"H0rGZkPcpj","license":"CC BY-SA 4.0","number":91381,"cdate":1483228800000,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/DBLP.org/-/Record","OpenReview.net/Public_Article/-/Edit"],"mdate":1784721017638,"domain":"OpenReview.net/Public_Article","id":"H0rGZkPcpj","version":2},{"content":{"venue":{"value":"IEEE Trans. Circuits Syst. Video Technol. 2024"},"pdf":{"value":"https://ieeexplore.ieee.org/iel7/76/10700029/10505173.pdf"},"venueid":{"value":"dblp.org/journals/TCSV/2024"},"paperhash":{"value":"xie|video_question_generation_for_dynamic_changes"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Jiayuan_Xie:","~Jiali_Chen1","https://dblp.org/search/pid/api?q=author:Zhenghao_Liu:","https://dblp.org/search/pid/api?q=author:Yi_Cai_0001:","https://dblp.org/search/pid/api?q=author:Qingbao_Huang:","https://dblp.org/search/pid/api?q=author:Qing_Li_0001:"]},"html":{"value":"https://doi.org/10.1109/TCSVT.2024.3391415"},"_bibtex":{"value":"@article{DBLP:journals/tcsv/XieCLCHL24,\n  author={Jiayuan Xie and Jiali Chen and Zhenghao Liu and Yi Cai and Qingbao Huang and Qing Li},\n  title={Video Question Generation for Dynamic Changes},\n  year={2024},\n  month={September},\n  cdate={1725148800000},\n  journal={IEEE Trans. Circuits Syst. Video Technol.},\n  volume={34},\n  number={9},\n  pages={8710-8721},\n  url={https://doi.org/10.1109/TCSVT.2024.3391415}\n}\n"},"abstract":{"value":"Video question generation task aims to generate meaningful questions about a video targeting an answer. Existing methods merely focus on the static appearance features in the image frames or simply identify a motion in the video to ask general questions. However, a video contains dynamically changing visual content that deserves to be questioned, e.g., changes in object motions, object states and relationships among objects, which is more practical and closer to the dynamic world we live in. In this paper, we propose a difference-aware video question generation model that aims to generate questions about temporal differences in the video, i.e., capturing the dynamic changes between image frames of a video to ask questions. To capture the dynamic changes between image frames, we utilize a temporal difference extractor to localize the differences for each frame pair of a video through an attention mechanism. Then, we introduce an answer-aware module to capture the answer-related image frame pair containing their differences for question generation, which aims to guide our model to focus on answer-related content for questioning. Finally, the output of the answer-aware module is sent to a decoder module to generate questions. Extensive experiments on SVQA and MSVD-QA datasets show that the proposed model outperforms state-of-the-art models, e.g., our model achieves at least 17.1% improvement over existing models in the SVQA dataset. This is because our model can generate questions similar to ground truths that involve changes between image frames in videos. Our code is available at https://github.com/Gary-code/D-VQG."},"title":{"value":"Video Question Generation for Dynamic Changes"},"authors":{"value":["Jiayuan Xie","Jiali Chen","Zhenghao Liu","Yi Cai","Qingbao Huang","Qing Li"]}},"tmdate":1745034063521,"pdate":1704067200000,"tcdate":1745034055190,"writers":["~"],"signatures":["~Jiali_Chen1"],"forum":"A6nsHMfbXb","license":"CC BY-SA 4.0","number":404632,"cdate":1725148800000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1745034063521,"domain":"DBLP.org","id":"A6nsHMfbXb","version":2},{"content":{"venue":{"value":"IEEE Trans. Circuits Syst. Video Technol. 2024"},"pdf":{"value":"https://ieeexplore.ieee.org/iel7/76/10700029/10505173.pdf"},"venueid":{"value":"dblp.org/journals/TCSV/2024"},"paperhash":{"value":"xie|video_question_generation_for_dynamic_changes"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Jiayuan_Xie:","~Jiali_Chen1","https://dblp.org/search/pid/api?q=author:Zhenghao_Liu:","https://dblp.org/search/pid/api?q=author:Yi_Cai_0001:","~Qingbao_Huang1","~Qing_Li5"]},"html":{"value":"https://doi.org/10.1109/TCSVT.2024.3391415"},"_bibtex":{"value":"@article{DBLP:journals/tcsv/XieCLCHL24,\n  author={Jiayuan Xie and Jiali Chen and Zhenghao Liu and Yi Cai and Qingbao Huang and Qing Li},\n  title={Video Question Generation for Dynamic Changes},\n  year={2024},\n  month={September},\n  cdate={1725148800000},\n  journal={IEEE Trans. Circuits Syst. Video Technol.},\n  volume={34},\n  number={9},\n  pages={8710-8721},\n  url={https://doi.org/10.1109/TCSVT.2024.3391415}\n}\n"},"abstract":{"value":"Video question generation task aims to generate meaningful questions about a video targeting an answer. Existing methods merely focus on the static appearance features in the image frames or simply identify a motion in the video to ask general questions. However, a video contains dynamically changing visual content that deserves to be questioned, e.g., changes in object motions, object states and relationships among objects, which is more practical and closer to the dynamic world we live in. In this paper, we propose a difference-aware video question generation model that aims to generate questions about temporal differences in the video, i.e., capturing the dynamic changes between image frames of a video to ask questions. To capture the dynamic changes between image frames, we utilize a temporal difference extractor to localize the differences for each frame pair of a video through an attention mechanism. Then, we introduce an answer-aware module to capture the answer-related image frame pair containing their differences for question generation, which aims to guide our model to focus on answer-related content for questioning. Finally, the output of the answer-aware module is sent to a decoder module to generate questions. Extensive experiments on SVQA and MSVD-QA datasets show that the proposed model outperforms state-of-the-art models, e.g., our model achieves at least 17.1% improvement over existing models in the SVQA dataset. This is because our model can generate questions similar to ground truths that involve changes between image frames in videos. Our code is available at https://github.com/Gary-code/D-VQG ."},"title":{"value":"Video Question Generation for Dynamic Changes"},"authors":{"value":["Jiayuan Xie","Jiali Chen","Zhenghao Liu","Yi Cai","Qingbao Huang","Qing Li"]}},"tmdate":1739942412695,"pdate":1704067200000,"tcdate":1731464087763,"writers":["~"],"signatures":["~Qingbao_Huang1"],"forum":"iT2OQr3ZJv","license":"CC BY-SA 4.0","number":189655,"cdate":1725148800000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1739942412695,"domain":"DBLP.org","id":"iT2OQr3ZJv","version":2},{"content":{"summary":{"value":"This paper focuses on the challenges of maintaining consistency and action continuity in the task of Video Storytelling. The authors propose a series of methods, including Time-weighted Blending, which preserves historical information; Dynamics-Informed Prompt Weighting, which adjusts the influence ratio of the double-prompt mechanism; and Semantic Action Representation. The proposed approach demonstrates promising results on the Video Storytelling task, enabling the generation of multi-scene long-form videos while effectively maintaining both subject consistency and action continuity."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"I have several questions regarding the task definition, design choices, and evaluation results, which may help clarify the authors’ approach.\n- How do the authors define the scope of the Video Storytelling task? Based on the examples provided in the supplementary material, the notion of consistency appears to focus mainly on subject and action. Could the authors clarify whether scene consistency is also considered within the task definition?\n- Regarding the examples involving multiple character sequences, maintaining consistency across multiple subjects may introduce additional complexity compared to single-subject scenarios. A brief discussion on whether multi-subject consistency poses distinct challenges could strengthen the presentation.\n- The proposed approach introduces discrete prompts that separately describe the main content and actions. Have the authors explored using more detailed or multi-aspect prompts, e.g., generating multiple samples with enriched textual descriptions and then concatenating them into a long video, for comparison? How different do the authors expect such detailed prompts to be from the proposed discrete prompt formulation?\n- In Table 1, the proposed method performs lower on LPIPS compared with other methods. What might be the reason for this? Additionally, when action representations are removed, the LPIPS score improves. Could the authors provide insights into why this might occur?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"- The paper focuses on the long-form video storytelling task and highlights the challenges of subject consistency and action continuity, which remain pressing issues in the field.\n- The proposed Time-weighted Blending mechanism offers a novel perspective on maintaining consistency in long-form video generation."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- **Ambiguous problem definition**\n    - The Abstract states that the paper focuses on long-form storytelling while also mentioning the goal of preventing abrupt transitions, which makes the problem formulation somewhat unclear. It is not entirely evident whether the focus is on pixel-level scene transitions or multi-shot video generation.\n    - The authors identify temporal coherence, semantic meaning, and action continuity as the main challenges. However, if I understand correctly, semantic meaning and action continuity can be considered components of temporal coherence. Since the work primarily aims to ensure consistency across multiple video segments, it might be helpful for the authors to clarify whether temporal coherence refers to high-level semantic consistency rather than pixel-level smoothness. A clearer statement of the task definition and key challenges would strengthen the paper’s focus.\n- **Lack of clarity in presentation**\n    - The Abstract and Introduction introduce specific module names (e.g., Time-weighted Blending, Dynamics-Informed Prompt Weighting) without prior explanation, which may make it difficult for readers to follow at first.\n    - Section 3 would benefit from more detailed descriptions. For example:\n        - It is unclear whether the generated frame refers to historical or current frames (L240).\n        - The difference between the subscripts $T$ in Eq. 3 and $N−1$ in Eq. 1 is not well explained.\n        - The hyperparameter $\\gamma$ in Eq. 4 is not specified.\n        - The role of $\\alpha'$ in Section 3.3 and the motivation behind the design of the SAR mechanism are not clearly described. Providing further clarification on these aspects would make the method section easier to follow.\n- **Limited evaluation scope**\n    - For the task of maintaining consistency across multiple video segments, it would be beneficial to include comparisons with additional baselines such as reference-to-video methods [1] and multi-shot video generation models [2–3].\n    - The evaluation currently relies mainly on DINO and LPIPS, which capture only part of the intended objectives. The authors are encouraged to include more comprehensive metrics, such as background consistency (as in VBench [4]) and ViCLIP feature similarity, to provide a fuller assessment of video consistency and continuity.\n- **Limited qualitative diversity**\n    - The qualitative results mostly feature similar subject pairs (e.g., A Woman & A Man, Dog & Cat). Including a wider range of examples would better demonstrate the generalization and robustness of the proposed method.\n\n[1] Liu L, Ma T, Li B, et al. Phantom: Subject-consistent video generation via cross-modal alignment[J]. arXiv preprint arXiv:2502.11079, 2025.(ICCV 2025)\n\n[2] Long F, Qiu Z, Yao T, et al. Videostudio: Generating consistent-content and multi-scene videos[C]//European Conference on Computer Vision. Cham: Springer Nature Switzerland, 2024: 468-485. (ECCV 2024)\n\n[3] Zhou Y, Zhou D, Cheng M M, et al. Storydiffusion: Consistent self-attention for long-range image and video generation[J]. Advances in Neural Information Processing Systems, 2024, 37: 110315-110340. (NIPS 2024)\n\n[4] Huang Z, He Y, Yu J, et al. Vbench: Comprehensive benchmark suite for video generative models[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2024: 21807-21818."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762918619885,"tcdate":1761907202747,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission6326/Reviewer_MAwd"],"signatures":["ICLR.cc/2026/Conference/Submission6326/Reviewer_MAwd"],"forum":"pSwlegpXZ0","number":3,"license":"CC BY 4.0","cdate":1761907202747,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission6326/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762918619885,"domain":"ICLR.cc/2026/Conference","replyto":"pSwlegpXZ0","id":"xdbOINhB5L","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Long-form Story Generation","Dynamics-Informed Prompt Weighting","Time-Weighted Blending","Semantic Action Representation"]},"supplementary_material":{"value":"/attachment/7e442395fc8a654b728aa132857caffae74e95c3.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Generating coherent long-form video sequences from discrete input using only text prompts is a critical task in content creation. While diffusion-based models excel at short video synthesis, long-form storytelling from text remains largely unexplored and a challenge due to difficulties in temporal coherency, preserving semantic meaning, and maintaining both scene context and action continuity across the video. We introduce a novel storytelling framework that achieves this by integrating scene and action prompts through dynamics-inspired prompt mixing. Specifically, we first present a bidirectional time-weighted latent blending strategy to ensure temporal consistency between segments of the long-form video being generated. We then propose a dynamics-informed prompt weighting (DIPW) mechanism that adaptively balances the influence of scene and action prompts at each diffusion timestep by jointly considering CLIP-based alignment, narrative continuity, and temporal smoothness. To further enhance motion continuity, we incorporate a semantic action representation to encode high-level action semantics into the blending process, dynamically adjusting transitions based on action similarity and ensuring smooth yet adaptable motion changes. Latent space blending maintains spatial coherence between objects in a scene, while time-weighted blending enforces bidirectional constraints for temporal consistency. The resulting integrative system prevents abrupt transitions while ensuring fluid storytelling that faithfully reflects both scene and action cues. Extensive experiments demonstrate significant improvements over baselines, achieving temporally consistent and visually compelling video narratives without any additional training. This approach bridges the gap between short clips and extended video to establish a new paradigm in GenAI-driven video synthesis from text."},"_bibtex":{"value":"@misc{\nkang2025dynamicsinspired,\ntitle={Dynamics-Inspired Text-Guided Video Storytelling},\nauthor={Taewon Kang and Divya Kothandaraman and Ming Lin},\nyear={2025},\nurl={https://openreview.net/forum?id=pSwlegpXZ0}\n}"},"title":{"value":"Dynamics-Inspired Text-Guided Video Storytelling"},"pdf":{"value":"/pdf/df8f20e018ec433d2d3ea26b9f59db8b88179bf3.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"kang|dynamicsinspired_textguided_video_storytelling"},"authorids":{"value":["~Taewon_Kang2","~Divya_Kothandaraman1","~Ming_Lin2"]},"authors":{"value":["Taewon Kang","Divya Kothandaraman","Ming Lin"]}},"version":2},{"content":{"summary":{"value":"SpeakerVid-5M îs a large-scale dyadic talking humans dataset. The dataset contains automatically extracted annotations of 2D kpts, audio, audio-to-text, and scene descriptions. The dataset is accompanied by an extensive video benchmark."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"Are the person ids per video or are they consistent across multiple videos, i.e. same speaker, different day?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"Dyadic human videos are an important mode of human-centric video generation. Having a large dataset with paired ASR, Audio and video is highly valuable. The paper is easy to follow and the authors provide an extensive ethics statement. VidChatBench is a reasonable benchmark suite for the proposed dataset."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"It seems that the dataset only contains pairs of videos - initiator -> respond. However, to future-prove this, I wonder if the authors could make available extended back-and-forth sequences of their data as well, i.e. where the initiator and responder engage in a back and forth way.\n\nThe authors claim that the video resolution is 1080P - however, their sample videos in the supplementary material are crops which are significantly smaller than 1080P. Could the authors clarify if they will release the full resolution? \n\nI find Figure 1 a bit unclear: what are the multi-modal annotations? I think the figure should make more clear in what the model inputs and outputs are for the generation process. \n\nIt seems some of the example videos contain jump cuts at the end (Body Composition/full_body/3.mp4 ) - I wonder how many videos contain those “”incorrect cuts - could the authors comment if they are planning to post-process the dataset to identify those instances or if they believe those to not be a problem?\n\n\nSuggested additional citations: TalkCuts [1].\n\n\n[1] A Large-Scale Dataset for Multi-Shot Human Speech Video Generation; NeurIPS DBT 2025"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762920552961,"tcdate":1761946929784,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission8778/Reviewer_a726"],"signatures":["ICLR.cc/2026/Conference/Submission8778/Reviewer_a726"],"forum":"U004uqALWl","number":3,"license":"CC BY 4.0","cdate":1761946929784,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission8778/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762920552961,"domain":"ICLR.cc/2026/Conference","replyto":"U004uqALWl","id":"fPjJoGrVXq","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"TLDR":{"value":"This paper propose a large-scale high-quality dataset for audio-visual dyadic interactive human generation."},"keywords":{"value":["Video generation","Digital human","Human-Centric dataset"]},"supplementary_material":{"value":"/attachment/1ab92ded039da7fc197cb948b5bc4b4cd2de8325.zip"},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"The rapid development of large-scale models has catalyzed significant breakthroughs in the digital human domain. These advanced methodologies offer high-fidelity solutions for avatar driving and rendering, leading academia to focus on the next major challenge: audio-visual dyadic interactive virtual human. To facilitate research in this emerging area, we present SpeakerVid-5M dataset, the first large-scale, high-quality dataset designed for audio-visual dyadic interactive virtual human generation. Totaling over $8,743$ hours, SpeakerVid-5M contains more than $5.2$ million video clips of human portraits. It covers diverse scales and interaction types, including monadic talking, listening, and dyadic conversations. Crucially, the dataset is structured along two key dimensions: interaction type and data quality. First, it is categorized into four types (dialogue branch, single branch, listening branch and multi-turn branch) based on the interaction scenario. Second, it is stratified into a large-scale pre-training subset and a curated, high-quality subset for Supervised Fine-Tuning (SFT). This dual structure accommodates a wide array of 2D virtual human tasks. In addition, we provide an autoregressive (AR)-based video chat baseline trained on this data, accompanied by a dedicated set of metrics and test data to serve as a benchmark (VidChatBench) for future work. Both the dataset and the corresponding data processing code will be publicly released."},"_bibtex":{"value":"@inproceedings{\nzhang2026speakervidm,\ntitle={SpeakerVid-5M: A Large-Scale High-Quality Dataset for Audio-Visual Dyadic Interactive Human Generation},\nauthor={Youliang Zhang and Zhaoyang Li and Duomin Wang and jiahe zhang and Deyu Zhou and Zixin Yin and Xili Dai and Gang YU and Xiu Li},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=U004uqALWl}\n}"},"title":{"value":"SpeakerVid-5M: A Large-Scale High-Quality Dataset for Audio-Visual Dyadic Interactive Human Generation"},"pdf":{"value":"/pdf/eef6eb90d54ce7b258514d6cd84028daaa892a43.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"zhang|speakervid5m_a_largescale_highquality_dataset_for_audiovisual_dyadic_interactive_human_generation"},"authorids":{"value":["~Youliang_Zhang2","~Zhaoyang_Li7","~Duomin_Wang1","~jiahe_zhang3","~Deyu_Zhou2","~Zixin_Yin2","~Xili_Dai2","~Gang_YU2","~Xiu_Li1"]},"authors":{"value":["Youliang Zhang","Zhaoyang Li","Duomin Wang","jiahe zhang","Deyu Zhou","Zixin Yin","Xili Dai","Gang YU","Xiu Li"]}},"version":2},{"content":{"summary":{"value":"This paper aims to enhance ViT for long-term video understanding. The authors design a memory bank to store historical information and develop input-aware adaptive memory selection to retrieve the relevant information to assist long-term analysis. The experiments show that the architecture demonstrates satisfactory performance with high efficiency."},"soundness":{"value":2},"confidence":{"value":5},"questions":{"value":"1. Does the KV Cache in this paper retain the gradient?\n2. This paper focuses on a pure vision model with enhanced memory design. However, the ViT-only architecture is capable of a limited range of video-related tasks. Is it possible to integrate it with video-language models to achieve wider range of video tasks to exert more impact on the community?"},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. The analysis of the limited temporal receptive field in long-term video understanding makes sense, and the motivation is clear.\n2. The method is simple and intuitive."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The experiments are limited. Only AVA and Epic-Kitchens are reported. Results on more video datasets are required to verify the effectiveness of the adaptive memory design. Besides, the performance improvements are marginal.\n2. The memory bank is recurrently updated by adaptive selection. Is it possible that in a long video, the content in the middle of the video is not closely related to the beginning, and only relevant content appears towards the end? However, during the memory bank update process, the tokens of the earlier video content were already discarded."}},"nonreaders":[],"tmdate":1731427670929,"tcdate":1730696005681,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission2872/Reviewer_CdCg"],"signatures":["ICLR.cc/2025/Conference/Submission2872/Reviewer_CdCg"],"forum":"1DEHVMDBaO","number":3,"license":"CC BY 4.0","cdate":1730696005681,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission2872/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427670929,"domain":"ICLR.cc/2025/Conference","replyto":"1DEHVMDBaO","id":"605yvNt4hJ","forumContent":{"TLDR":{"value":"We introduce the Adaptive Memory Vision Transformer (AMViT), which dynamically adjusts its Temporal Receptive Field using an Adaptive Memory Mechanism for effectiveness and efficiency improvement in long-form video understanding."},"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Key-Value Cache","Vision Transformer","Video Understanding"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"In long-form video understanding, selecting an optimal Temporal Receptive Field (TRF) is crucial for Vision Transformer (ViT) models due to the dynamic nature of diverse video motion contents, which varies in duration and velocity. A short TRF can result in loss of critical information, while a long TRF may decrease ViT's performance and computational efficiency caused by the unrelated contents in videos and the quadratic complexity of the attention mechanism. To tackle this issue, we introduce Adaptive Memory Mechanism (AMM) that enables ViT to adjust its TRF dynamically in response to the video's dynamic contents. Instead of discarding Key-Value (KV) Cache from the earliest inference when the settings limit is reached, our approach uses a Memory Bank (MB) to retain the most important embeddings from the Key-Value Cache that would otherwise be discarded in memory-augmented methods. The selection is based on the attention score calculated between the Class Token (CLS) in current iteration and the KV Cache in previous iterations. We demonstrate that Adaptive Memory Vision Transformer (AMViT) outperforms existing methods across a diverse array of tasks (action recognition, action anticipation, and action detection)."},"_bibtex":{"value":"@misc{\nliu2025adaptive,\ntitle={Adaptive Memory Mechanism in Vision Transformer for Long-form Video Understanding},\nauthor={Zhenshun Liu and Zijian Lei and Kejing Yin and William K. Cheung},\nyear={2025},\nurl={https://openreview.net/forum?id=1DEHVMDBaO}\n}"},"title":{"value":"Adaptive Memory Mechanism in Vision Transformer for Long-form Video Understanding"},"pdf":{"value":"/pdf/89a290cc5b5541967c05c287db67008301f18a46.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"liu|adaptive_memory_mechanism_in_vision_transformer_for_longform_video_understanding"},"authorids":{"value":["~Zhenshun_Liu1","~Zijian_Lei1","~Kejing_Yin1","~William_K._Cheung1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Zhenshun Liu","Zijian Lei","Kejing Yin","William K. Cheung"]}},"version":2},{"content":{"venue":{"value":"WCSP 2010"},"pdf":{"value":"https://ieeexplore.ieee.org/iel5/5620887/5629524/05632618.pdf"},"venueid":{"value":"dblp.org/conf/WCSP/2010"},"paperhash":{"value":"ma|distributed_linkaware_rate_allocation_for_rd_optimal_multiple_video_streaming_over_wireless_networks"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Zejun_Ma:","~Li_Song3","https://dblp.org/search/pid/api?q=author:Cheng_Zhi:","https://dblp.org/search/pid/api?q=author:Libo_Yang:"]},"html":{"value":"https://doi.org/10.1109/WCSP.2010.5632618"},"_bibtex":{"value":"@inproceedings{DBLP:conf/wcsp/MaSZY10,\n  author={Zejun Ma and Li Song and Cheng Zhi and Libo Yang},\n  title={Distributed link-aware rate allocation for R-D optimal multiple video streaming over wireless networks},\n  year={2010},\n  cdate={1262304000000},\n  pages={1-6},\n  url={https://doi.org/10.1109/WCSP.2010.5632618},\n  booktitle={WCSP},\n  crossref={conf/wcsp/2010}\n}\n"},"abstract":{"value":"When multiple video streams are delivered in wireless networks, careful rate allocation is desirable to share the available bandwidth without overflowing the whole network or any individual link, as well as to maximize the total received quality among the competing streams. This paper presents a distributed rate allocation for Rate Distortion (R-D) optimal multiple video streaming in a video- and link-aware fashion. We take into account video R-D characteristics, as well as network and link congestion in the optimization framework, to strive for the minimal total video distortion of all participating streams while limiting network and individual link channel time utilization. Initial Start scheme addresses the initialization problem for video streaming effectively, and Fast Recovery scheme eliminates network and any individual link congestion as fast as possible. Global optimal rates can be obtained in a distributed manner. Results from numerical experiments confirm excellent performance of this distributed rate allocation."},"title":{"value":"Distributed link-aware rate allocation for R-D optimal multiple video streaming over wireless networks"},"authors":{"value":["Zejun Ma","Li Song","Cheng Zhi","Libo Yang"]}},"tmdate":1753241513982,"pdate":1262304000000,"tcdate":1753241495334,"writers":["~"],"signatures":["~Li_Song3"],"forum":"7ovscsGsym","license":"CC BY-SA 4.0","number":583198,"cdate":1262304000000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1753241513982,"domain":"DBLP.org","id":"7ovscsGsym","version":2},{"content":{"summary":{"value":"The paper presents a benchmark, called InfiniBench, for the evaluation of long video understanding. The dataset consists of 1219 videos. The average length of the videos is 52.59 minutes. There are 108.2K (video, question) pairs. The questions are divided into 9 categories. Some categories require the ability to make associations across a longtime span. Some categories require in-depth understanding and reasoning capabilities. It is a very interesting new benchmark for long video understanding."},"soundness":{"value":4},"confidence":{"value":5},"questions":{"value":"Can the authors comment on the limited variety of TV shows? What about sports events like NBA, NFL, Tennis, etc. \n\nOn GPT-4o evaluation, only 250 frames are selected. Are the 250 frames selected uniformly? Have you tried to reduce the frame size and squeeze more frames into GPT-4o?\n\nWill all the videos be released to public? Are there any legal issues?\n\nAre there text scripts (screenplay) associated with all the videos (movies and TV shows)?"},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":4},"strengths":{"value":"The videos are very long with an average length 52 minutes.\n\nThe number of (question, answer) pairs is large (108k)\n\nSome of the questions are unique such as spoiler questions, global appearance, and scene transitions. \n\nCompared to the existing benchmarks, this benchmark contains much longer videos and contains some new interesting types of questions. It'll be very useful to the researchers who work on long video understanding."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The variety of TV show sources is limited since there are only 6 different TV shows."}},"nonreaders":[],"tmdate":1731427565772,"tcdate":1730010230918,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission2250/Reviewer_dcjq"],"signatures":["ICLR.cc/2025/Conference/Submission2250/Reviewer_dcjq"],"forum":"2D0uXQbntW","number":1,"license":"CC BY 4.0","cdate":1730010230918,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission2250/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427565772,"domain":"ICLR.cc/2025/Conference","replyto":"2D0uXQbntW","id":"Zfj5mXf9BG","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["video understanding","benchmark","long video benchmark","long video understanding"]},"supplementary_material":{"value":"/attachment/e4286990af860ccb3f16a8fce15b1da8635ef0b5.zip"},"primary_area":{"value":"datasets and benchmarks"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Understanding long videos, ranging from tens of minutes to several hours, presents unique challenges in video comprehension. Despite the increasing importance of long-form video content, existing benchmarks primarily focus on shorter clips. To address this gap, we introduce InfiniBench a comprehensive benchmark for very long video understanding which presents 1)very long video duration, averaging 52.59 minutes per video 2)The largest number of question-answer pairs, 108.2K 3) Diversity in questions that examine nine different skills and include both multiple-choice questions and open-ended questions 4) Memory questions, such as Global Appearance that require remembering and tracking the visual aspects through the video. Using InfiniBench, we comprehensively evaluate existing Large Multi-Modality Models (LMMs) on each skill, including the commercial models such as GPT-4o and Gemini 1.5 Flash and the recent open-source models. \nThe evaluation shows significant challenges in our benchmark.\nOur findings reveal that even leading AI models like GPT-4o and Gemini 1.5 Flash face challenges in achieving high performance in long video understanding, with average accuracies of just 56.01 % and 43.32 %, and average scores of 3.25 and 2.79 out of 5, respectively.\nQwen2-VL matches Gemini's performance in the MCQ skills but lags significantly in open-ended question tasks.\nWe hope this benchmark will stimulate the LMMs community towards long video and human-level understanding."},"_bibtex":{"value":"@misc{\nataallah2025infinibench,\ntitle={InfiniBench: A Comprehensive Benchmark for Large Multimodal Models in Very Long Video Understanding},\nauthor={Kirolos Ataallah and Chenhui Gou and Eslam Mohamed BAKR and Khushbu Pahwa and Jian Ding and Mohamed Elhoseiny},\nyear={2025},\nurl={https://openreview.net/forum?id=2D0uXQbntW}\n}"},"title":{"value":"InfiniBench: A Comprehensive Benchmark for Large Multimodal Models in Very Long Video Understanding"},"pdf":{"value":"/pdf/032fb6c70b723929924e625fabf7cebb80b48405.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"ataallah|infinibench_a_comprehensive_benchmark_for_large_multimodal_models_in_very_long_video_understanding"},"authorids":{"value":["~Kirolos_Ataallah1","~Chenhui_Gou1","~Eslam_Mohamed_BAKR1","~Khushbu_Pahwa1","~Jian_Ding3","~Mohamed_Elhoseiny1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Kirolos Ataallah","Chenhui Gou","Eslam Mohamed BAKR","Khushbu Pahwa","Jian Ding","Mohamed Elhoseiny"]}},"version":2},{"content":{"summary":{"value":"This paper addresses a gap in storytelling video generation: the lack of coherent, character-driven dialogue and speech. The authors propose Action2Dialogue, a modular, training-free pipeline that synthesizes character-specific dialogue and expressive speech directly from scene-level text prompts.\nThe system takes a sequence of prompt pairs as input, where each pair defines a scene's setting and a character's action. The pipeline then performs three main functions:\nScene Visualization: An existing story generation model (e.g., Text2Story) creates the visual video clip.\nDialogue Generation: A large language model (LLM) generates dialogue. This is the core contribution. The LLM is conditioned on the input prompts, visual features from a representative keyframe (extracted via a model like BLIP), and a novel dialogue history mechanism.\nSpeech Synthesis: A reference-driven voice synthesis model renders the generated text into expressive speech using a character's voice sample.\nThe paper's primary component is the Recursive Narrative Bank (RNB), a speaker-aware, temporally-structured memory. This mechanism provides the LLM with the dialogue history, enabling it to generate utterances that are consistent with a character's persona and evolving narrative context across multiple scenes. The authors demonstrate through automated metrics and a human subject study that their full pipeline, particularly the visual grounding and the RNB, produces significantly more natural, coherent, and character-consistent narratives than ablated baselines."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"On Lip-Sync: The lack of lip-sync is the most apparent gap. See weaknesses"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The paper tackles a clear and currently underserved problem in generative AI. While video generation has advanced, most outputs are \"silent films.\" This work provides a concrete framework for adding a crucial dimension of emotional depth and narrative realism through character-driven dialogue.\n2. The proposed method is training-free and does not introduce new model architecture, but propose modular composition of existing SOTA components (video generation, VLM, LLM, TTS). \n3. The paper is well-written and easy to follow. The problem statement is clear, the proposed method is described logically, and Figure provides an excellent overview of the entire pipeline. The authors clearly delineate their contributions."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Missing Lip-Sync / Audiovisual Synchronization: The most significant weakness is the lack of lip synchronization. The paper generates a video clip and a separate audio track. While it calls this a \"multimodal video narrative,\" the result is effectively a video with a character-specific voice-over, not a video of a character speaking. This omission feels like an incomplete solution to the problem of \"character-driven dialogue\" in a visual medium. A truly integrated system would need to either generate the video with the correct mouth movements or apply a post-processing lip-sync model. This limitation should be more prominently discussed. Video and dialogue are generated by two separate, unaligned systems. The video model generates visuals from (p1, p2), while the LLM generates text from (p1, p2, c, Ht). There is no mechanism to enforce that the generated dialogue precisely matches the specific generated visuals.\n2. Weak Visual Grounding: The system's entire visual understanding of a scene hinges on a single, representative middle frame. This is a very lossy representation of a dynamic video clip. If a key action described in the prompt occurs at the beginning or end of the clip, the middle frame might be uninformative (e.g., just showing Donkey on the ground), leading to poorly-grounded dialogue. A more robust approach might use a video-language model to encode the entire clip's temporal dynamics.\n3. The paper's related work section is missing a critical line of research: speech-driven video generation (or talking head/portrait animation). This field directly addresses the paper's (unsolved) final step of creating a speaking character.\nHighly relevant works like\n\nEMO: Emote Portrait Alive -- Generating Expressive Portrait Videos with Audio2Video Diffusion Model under Weak Conditions\n\nHallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer\n\nomnihuman-1: rethinking the scaling-up of one-stage conditioned human animation models\n\nmocha: towards movie-grade talking character synthesis"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762918622667,"tcdate":1761949769961,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission6327/Reviewer_STLo"],"signatures":["ICLR.cc/2026/Conference/Submission6327/Reviewer_STLo"],"forum":"bVsVbmI0qC","number":2,"license":"CC BY 4.0","cdate":1761949769961,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission6327/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762918622667,"domain":"ICLR.cc/2026/Conference","replyto":"bVsVbmI0qC","id":"ID6exqqgNq","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Narrative Generation","Video Synthesis"]},"supplementary_material":{"value":"/attachment/8b4e302a0fe4900b49ba035f95cfd8577cb0fa77.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Recent advances in scene-based video generation have enabled systems to synthesize coherent visual narratives from structured prompts. However, a crucial dimension of storytelling—character-driven dialogue and speech—remains underexplored. In this paper, we present a modular pipeline that transforms action-level prompts into visually and auditorily grounded narrative dialogue, enriching visual storytelling with natural voice and character expression. Our method takes as input a pair of prompts per scene, where the first defines the setting and the second specifies a character’s behavior. While a story generation model such as Text2Story produces the corresponding visual scene, we focus on generating expressive, character-consistent utterances grounded in both the prompts and the scene image. A pretrained vision-language encoder extracts high-level semantic features from a representative frame, capturing salient visual context. These features are then integrated with structured prompts to guide a large language model in synthesizing natural dialogue. To ensure contextual and emotional consistency across scenes, we introduce a \\textit{Recursive Narrative Bank}—a speaker-aware, temporally structured memory that recursively accumulates each character’s dialogue history. Inspired by Script Theory in cognitive psychology, this design enables characters to speak in ways that reflect their evolving goals, social context, and narrative roles throughout the story. Finally, we render each utterance as expressive, character-conditioned speech, resulting in fully-voiced, multimodal video narratives. Our training-free framework generalizes across diverse story settings—from fantasy adventures to slice-of-life episodes—offering a scalable solution for coherent, character-grounded audiovisual storytelling."},"_bibtex":{"value":"@misc{\nkang2025characterdriven,\ntitle={Character-Driven Narrative Generation for Scene-Based Video Synthesis},\nauthor={Taewon Kang and Ming Lin},\nyear={2025},\nurl={https://openreview.net/forum?id=bVsVbmI0qC}\n}"},"title":{"value":"Character-Driven Narrative Generation for Scene-Based Video Synthesis"},"pdf":{"value":"/pdf/97d8ad5f734782287d3e0337488d04ad6e837ccb.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"kang|characterdriven_narrative_generation_for_scenebased_video_synthesis"},"authorids":{"value":["~Taewon_Kang2","~Ming_Lin2"]},"authors":{"value":["Taewon Kang","Ming Lin"]}},"version":2},{"content":{"venue":{"value":"ICANN (2) 2017"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-319-68612-7_4.pdf"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"zhang|shortcut_convolutional_neural_networks_for_classification_of_gender_and_texture"},"html":{"value":"https://doi.org/10.1007/978-3-319-68612-7_4"},"_bibtex":{"value":"@inproceedings{DBLP:conf/icann/ZhangLL17,\n  author={Ting Zhang and Yujian Li and Zhaoying Liu},\n  title={Shortcut Convolutional Neural Networks for Classification of Gender and Texture},\n  year={2017},\n  cdate={1483228800000},\n  pages={30-39},\n  url={https://doi.org/10.1007/978-3-319-68612-7_4},\n  booktitle={ICANN (2)},\n  crossref={conf/icann/2017-2}\n}\n"},"abstract":{"value":"Convolutional neural networks are global trainable multi-stage architectures that automatically learn translation invariant features from raw input images. However, in tradition they only allow adjacent layers connected, limiting integration of multi-scale information. To further improve their performance in classification, we present a new architecture called shortcut convolutional neural networks. This architecture can concatenate multi-scale feature maps by shortcut connections to form the fully-connected layer that is directly fed to the output layer. We give an investigation of the proposed shortcut convolutional neural networks on gender classification and texture classification. Experimental results show that shortcut convolutional neural networks have better performances than those without shortcut connections, and it is more robust to different settings of pooling schemes, activation functions, initializations, and optimizations."},"title":{"value":"Shortcut Convolutional Neural Networks for Classification of Gender and Texture"},"authors":{"value":[{"fullname":"Ting Zhang","username":""},{"fullname":"Yujian Li","username":""},{"fullname":"Zhaoying Liu","username":"~Zhaoying_Liu1"}]}},"tmdate":1784721018792,"pdate":1514678400000,"externalIds":["dblp:conf/icann/ZhangLL17"],"tcdate":1784721012341,"writers":["~","OpenReview.net/Public_Article/DBLP.org","OpenReview.net/Support"],"signatures":["~Zhaoying_Liu1"],"forum":"MNmCx8wZx6","license":"CC BY-SA 4.0","number":91399,"cdate":1483228800000,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/DBLP.org/-/Record","OpenReview.net/Public_Article/-/Edit"],"mdate":1784721018792,"domain":"OpenReview.net/Public_Article","id":"MNmCx8wZx6","version":2},{"content":{"summary":{"value":"The author introduces a zero-training video refinement pipeline that leverages neuro-symbolic feedback to automatically enhance video generation."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"see weakness"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"The author introduces an interesting zero-training video refinement pipeline that leverages neuro-symbolic feedback to automatically enhance video generation."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1.The author uses video editing to correct incorrect video generation. However, video editing is a more difficult task than video generation, with lower accuracy and higher computational resource requirements. Therefore, I think this approach doesn't make sense in practice. Relying on video editing for correction is less effective than improving the success rate of video generation in the first place.\n\n\n2. This pipeline is too idealistic and complex. If the video generation involves complex object changes, this pipeline is likely to fail."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762920050166,"tcdate":1762008391115,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission8061/Reviewer_FLya"],"signatures":["ICLR.cc/2026/Conference/Submission8061/Reviewer_FLya"],"forum":"ifJ91JSLhq","number":4,"license":"CC BY 4.0","cdate":1762008391115,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission8061/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762920050166,"domain":"ICLR.cc/2026/Conference","replyto":"ifJ91JSLhq","id":"Vc7MVfeRTy","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"TLDR":{"value":"We propose a training free method to improve the temporal fidelity text to video generation."},"keywords":{"value":["Neuro-symbolic AI","Text-to-Video Synthesis and Generation","Explainable Computer Vision"]},"supplementary_material":{"value":"/attachment/b22f935dcfed4ca204abdfb8b0254521cd2af223.zip"},"primary_area":{"value":"neurosymbolic & hybrid AI systems (physics-informed, logic & formal reasoning, etc.)"},"abstract":{"value":"Current text-to-video (T2V) generation models are increasingly popular due to their ability to produce coherent videos from textual prompts. However, these models often struggle to generate semantically and temporally consistent videos when dealing with longer, more complex prompts involving multiple objects or sequential events. Additionally, the high computational costs associated with training or fine-tuning make direct improvements impractical. To overcome these limitations, we introduce NeuS-E, a novel zero-training video refinement pipeline that leverages neuro-symbolic feedback to automatically enhance video generation, achieving superior alignment with the prompts. Our approach first derives the neuro-symbolic feedback by analyzing a formal video representation and pinpoints semantically inconsistent events, objects, and their corresponding frames. This feedback then guides targeted edits to the original video. Extensive empirical evaluations on both open-source and proprietary T2V models demonstrate that NeuS-E significantly enhances temporal and logical alignment across diverse prompts by almost 40%."},"_bibtex":{"value":"@misc{\nchoi2026well,\ntitle={We'll Fix it in Post: Improving Text-to-Video Generation with Zero Training},\nauthor={Minkyu Choi and S P Sharan and Harsh Goel and Sahil Shah and Sandeep P. Chinchali},\nyear={2026},\nurl={https://openreview.net/forum?id=ifJ91JSLhq}\n}"},"title":{"value":"We'll Fix it in Post: Improving Text-to-Video Generation with Zero Training"},"pdf":{"value":"/pdf/70da27db733ed9eb78e57fb621cb21d3cad59875.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"choi|well_fix_it_in_post_improving_texttovideo_generation_with_zero_training"},"authorids":{"value":["~Minkyu_Choi2","~S_P_Sharan1","~Harsh_Goel1","~Sahil_Shah3","~Sandeep_P._Chinchali1"]},"authors":{"value":["Minkyu Choi","S P Sharan","Harsh Goel","Sahil Shah","Sandeep P. Chinchali"]}},"version":2},{"content":{"summary":{"value":"The author addresses the challenge of long-horizon, spatially-consistent panoramic video generation by building on the insight of incorportin explicit 3D memory over generated contents. Approach-wise, EvoWorld starts from sinegle panoramic image input, and then leverage video generator with camera control to generate video (assuming it's all static contents, since they finetuned on their static data). They use VGGT to localize the generated frames, and distill changes to 3D for later video extrapolation steps. The key insight, it levages 3D reconstraction along with video generation for better spatial guidance. For dense view control, we provide their own spherical plucker coordinate encoding for extrinsincs in the panoramic setup. \n\nIn their experiments, they compare with approach on other video generator only baselines to showcase theirs are the best in terms of 2d and 3d quality and consistencies. Through their ablation, they demonstrate SpherePlücker + 3D memory gives the best video results. They also further evaluate EvoWorld in several downstream tasks."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. Could you compare your results with existing work that leverage explicit 3D memory? For example, DiffDreamer or other Panoramic scene generation work? \n\n2. What prevent the current model from generating long-horizon video contents? \n\n3. How is the EvoWorld differentiate itself from existing work?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The comparison on video quality against baselines are carefully conducted. Both metrics in 2d and 3d are reported, including MEt3R. \n\n2. The results are validated across synthetic and real-world data benchmarks. \n\n3. Results are also evaluated carefully in several downstream tasks. \n\nOverall, I appreciate the authors efforts on thorought evaluation and comparison on the proposed method against baseline across different benchmarks."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Missing Key Related Work and Unclear Novelty Positioning\nThe related work section overlooks several highly relevant research threads, which makes it difficult to clearly identify the contribution’s novelty. In particular, prior efforts in perpetual video / world generation (e.g., Infinite Nature, LoTR, DiffDreamer) and panoramic view synthesis (e.g., works summarized in Sec. 3.3.1 of this survey: arXiv:2505.05474) are not adequately discussed. These works share similar goals of long-horizon view consistency and world exploration, so a discussion is necessary to contextualize what is new in this paper.\n\n2. Claims Not Fully Supported by Results\nAlthough introducing a 3D memory representation to video generation is a reasonable idea, the experimental evidence does not convincingly demonstrate the claimed benefits:\n-- Spatial inconsistency: Multi-view frames in Fig. 4 display noticeable inconsistencies, particularly in building facades, which become more evident when inspecting the provided video examples.\n-- Limited trajectory length: The generated camera motion appears to span only short distances (≈10 meters), which does not align with the paper’s claims of enabling long-horizon generation. The author also mentions this in limitation section. \nAs a result, the key claims of “long-range, spatially consistent video generation” presented in the abstract are not sufficiently validated.\n\n3. No Strategy for Loop Closure or Revisitation\nHowever, the proposed approach does not articulate any mechanism to support scene generation after revisiting as is shown in the teaser image. Additionally, the paper provides no qualitative examples where a camera trajectory completes a loop (e.g., around a city block) and returns to an earlier location without accumulating significant drift or scene errors.\n\n4. Incremental Insight Relative to Prior 3D-Aware Generative Models\nThe central idea—that explicit 3D memory improves view consistency—has been explored in earlier works such as SynSin, DiffDreamer, and others in the literature on 3D-aware generation. The contribution therefore feels somewhat incremental unless stronger experimental evidence or new theoretical insights can be demonstrated. \n\n5. VGGT localization necessary?\nIt is a bit unclear in terms of why the author rely on pose-conditioned video generator, but still rely on VGGT for localization."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762918530668,"tcdate":1761533540556,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission6188/Reviewer_4M6D"],"signatures":["ICLR.cc/2026/Conference/Submission6188/Reviewer_4M6D"],"forum":"Pwag2BlEZC","number":2,"license":"CC BY 4.0","cdate":1761533540556,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission6188/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762918530668,"domain":"ICLR.cc/2026/Conference","replyto":"Pwag2BlEZC","id":"cuYV3nfe4z","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"TLDR":{"value":"We propose a framework for evolving world generation that builds and updates an explicit 3D memory from generated panoramic videos. Conditioning on this reconstructed geometry mitigates error accumulation, ensuring spatial consistency over time."},"keywords":{"value":["World Model","Video Generation","Embodied AI","Panorama","Spatial Consistency"]},"supplementary_material":{"value":"/attachment/4f4b25e2d62529ab2af4cd4878ec6606eaa1eac4.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Humans possess a remarkable ability to mentally explore and replay 3D environments they have previously experienced. Inspired by this mental process, we present EvoWorld: a world model that bridges panoramic video generation with evolving 3D memory to enable spatially consistent long-horizon exploration. Given a single panoramic image as input, EvoWorld first generates future video frames by leveraging a video generator with fine-grained view control, then evolves the scene's 3D reconstruction using a feedforward plug-and-play transformer, and finally synthesizes futures by conditioning on geometric reprojections from this evolving explicit 3D memory. Unlike prior state-of-the-arts that synthesize videos only, our key insight lies in exploiting this evolving 3D reconstruction as explicit spatial guidance for the video generation process, projecting the reconstructed geometry onto target viewpoints to provide rich spatial cues that significantly enhance both visual realism and geometric consistency. To evaluate long-range exploration capabilities, we introduce the first comprehensive benchmark spanning synthetic outdoor environments, Habitat indoor scenes, and challenging real-world scenarios, with particular emphasis on loop-closure detection and spatial coherence over extended trajectories. Extensive experiments demonstrate that our evolving 3D memory substantially improves visual fidelity and maintains spatial scene coherence compared to existing approaches, representing a significant advance toward practical long-horizon spatially consistent world modeling."},"_bibtex":{"value":"@misc{\nwang2025evoworld,\ntitle={EvoWorld: Evolving Panoramic World Generation with Explicit 3D Memory},\nauthor={Jiahao Wang and Luoxin Ye and TaiMing Lu and Junfei Xiao and Jiahan Zhang and Yuxiang Guo and Xijun Liu and Rama Chellappa and Cheng Peng and Alan Yuille and Jieneng Chen},\nyear={2025},\nurl={https://openreview.net/forum?id=Pwag2BlEZC}\n}"},"title":{"value":"EvoWorld: Evolving Panoramic World Generation with Explicit 3D Memory"},"pdf":{"value":"/pdf/dbcbe1995af724382fb8f16c156932c6dfa35547.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"wang|evoworld_evolving_panoramic_world_generation_with_explicit_3d_memory"},"authorids":{"value":["~Jiahao_Wang5","~Luoxin_Ye1","~TaiMing_Lu1","~Junfei_Xiao1","~Jiahan_Zhang1","~Yuxiang_Guo1","~Xijun_Liu1","~Rama_Chellappa1","~Cheng_Peng2","~Alan_Yuille1","~Jieneng_Chen1"]},"authors":{"value":["Jiahao Wang","Luoxin Ye","TaiMing Lu","Junfei Xiao","Jiahan Zhang","Yuxiang Guo","Xijun Liu","Rama Chellappa","Cheng Peng","Alan Yuille","Jieneng Chen"]}},"version":2},{"content":{"summary":{"value":"**Update after discussion period**:\n\nMy biggest concern of this paper is about its audio reconstruction pipeline. The original pipeline consists of an RGB VAE and a non-DNN mel2wave conversion module, resulting in terrible sound quality. Moreover, such a pipeline has been used in video2audio community widely, which is a pity.\n\nDuring the rebuttal period, the authors not only added the experiment of using a vocoder for the mel2wave conversion, but also re-trained their models with an audio VAE. As a result of their efforts, the FAD score improved from 4.37 to 2.16 and eventually 1.34. Congrats on the achievement!\n\nI listened to the latest samples in the Google Drive, and found a distinct advantage of audio VAE in samples that contain non-noise or musical sources (sample027__SrU3mfTPYg_000032.mp4, sample045_2es7oZzwLWM_000030.mp4). Although the AV-align score is slightly worse with an audio VAE, the score is still higher than conventional methods by a large margin.\n\nSince the AV-align score is computed by a pretrained DNN model, moreover, considering the gap between v2a community and audio generation community, the metric might not be very reliable. For example, in samples such as \"sample007_-2sOH8XovEE_000484.mp4\", I don't feel the alignment is as good as the scores indicate. Perhaps some future works can be done to further improve the metric itself.\n\nI increased my ratings to **encourage further collaboration between video2audio and audio communities**.\n\n---------------\n\nThe paper proposes \"MDSGen\", an efficient model based on Masked Diffusion Transformer, for video-to-audio generation.\n\nThe challenges of video-to-audio generation are mainly:\n1. Heavy computation and memory usage;\n2. Requirements for the audio quality;\n3. Requirements for the audio-video alignment;\n\nMDSGen reduces the resource consumption by using very light-weight Transformer coupled with fast diffusion samplers such as DPM solver, as well as a dimension reduction module to reduce the size of the video conditioning embeddings.\n\nMDSGen improves audio quality and audio-video quality by introducing a time-aware masking strategy into the mask DiT framework, together with other efforts.\n\nConceptualy, MDSGen looks like a framework that replaces the \"text prompt\" in text-to-audio DiT [StableAudioOpen],[MakeAnAudio2] by a video feature embedding. Hence the technical contributions are more in micro aspects.\n\nHowever, some design choices may have severely affected the audio quality, making the work less solid or reusable to the community. Audio quality observed in the supplementary files is far from the level in modern text-to-audio models such as [AudioLDM], [MakeAnAudio2], [SpecMaskGIT], [StableAudioOpen]."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"Three models are presented in the paper (Tiny, Small and Base) except the overfitting large model. What is the model presented in the supplementary files?\n\nIs there any reason to use an image VAE instead of audio VAEs? \n\nDid the authors observe any advantage of GLA over a neural vocoder?\n\nIs it possible to run a subjective listening test, and see the consistency between human evaluation and the audio-video alignment accuracy measured by a DNN model?\n\nCould the authors measure the FAD scores on top of the current FID scores?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":3},"strengths":{"value":"Most contributions of MDSGen are about micro design aspects.\n1. Channel selection of Mel-spec\n2. Time-aware masking strategy for generative models\n3. Reduced dimension of the video features\n4. Small model size and fast inference speed"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"## Major issues\n### 1. Improper audio reconstrution pipeline\nMDSGen ustilizes the VAE from Stable Diffusion, which is not trained for Mel-spec. Although the authors carefully discussed how to take the most advantage of this image VAE, the discussion itself is **NOT** reusable for the audio community, as there have been plenty of audio VAE designs, some of which are publicly available such as [AudioLDM], [MakeAnAudio2], [StableAudioOpen], [DAC].\n\nAnother improper choice is that, MDSGen utilizes the Griffin-Lim Algorithm (GLA) to convert mel-spec back to wave forms. GLA has almost been abandoned by audio community, due to the recent advance in neural vocoder, e.g., [HiFiGAN], [UnivNet], [BigVGAN]. I believe the apparent phase distortion in the supplementary files might have been caused by GLA.\n\nThere is a rough comparison on the quality of audio reconstruction pipeline in a recent paper [SpecMaskGIT], I hope it could be useful for the improvement of MDSGen.\n\nI strongly recommend the authors to consider audio-specified reconstruction pipelines for improved audio quality. Even in audio-visual generation community, we can see the usage of such audio VAE for excellent audio quality, e.g., [VisualEchoes]\n### 2. Invalid claims on the result\nBecause the audio quality is far from the baseline in audio generation community, it is improper to claim that MDSGen is better in \"audio-video alignment\".\n\nI believe, the audio-video alignment can be evaluated only when the audio quality is sufficiently good. Given the current audio quality, I don't think the model is ready for further evaluation.\n## Minor issues \n### 1. Evaluation metrics\nThe FID used in this paper comes from the implementation of SpecVQGAN, a pioneer of audio generation. However, the FAD implementation ([AudioLDM],[FAD_github]) has been more widely accepted in audio community. Evaluating with the widely adopted FAD metric can also help the readers to compre the audio quality with other audio generation models.\n### 2. Insufficient ablation study\nMDSGen trains a learnable module to reduce the video feature sequence into a single vector. From Figure 6, we can observe that the learned weights are quite evenly distributed (except the beginning and ending frames).\n\nThe observation posts a question: How much improvement can the learnable reducer bring compared to a naive average pooling?\n\n[StableAudioOpen]: https://arxiv.org/abs/2407.14358\n[MakeAnAudio2]: https://arxiv.org/abs/2305.18474\n[AudioLDM]: https://audioldm.github.io/\n[SpecMaskGIT]: https://arxiv.org/abs/2406.17672\n[DAC]: https://github.com/descriptinc/descript-audio-codec\n[HiFiGAN]: https://github.com/jik876/hifi-gan\n[UnivNet]: https://github.com/rishikksh20/UnivNet-pytorch\n[BigVGAN]: https://github.com/NVIDIA/BigVGAN\n[VisualEchoes]: https://arxiv.org/abs/2405.14598\n[AudioLDMEval]: https://github.com/haoheliu/audioldm_eval\n[FAD_github]: https://github.com/gudgud96/frechet-audio-distance"}},"nonreaders":[],"tmdate":1733244810252,"tcdate":1730683351783,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission265/Reviewer_GqkQ"],"signatures":["ICLR.cc/2025/Conference/Submission265/Reviewer_GqkQ"],"forum":"yFEqYwgttJ","number":4,"license":"CC BY 4.0","cdate":1730683351783,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission265/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1733244810252,"domain":"ICLR.cc/2025/Conference","replyto":"yFEqYwgttJ","id":"ZvfOoUvtYv","forumContent":{"TLDR":{"value":"A novel approach is presented for highly efficient vision-guided sound synthesis using masked diffusion models"},"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["vision-guided audio generation","fast inference","open-domain sound synthesis","masked diffusion models","temporal learning","visual sound source localization","generative AI"]},"supplementary_material":{"value":"/attachment/540764b582c3bd4ee182a390db04bfe032eb2087.zip"},"primary_area":{"value":"generative models"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"We introduce MDSGen, a novel framework for vision-guided open-domain sound generation optimized for model parameter size, memory consumption, and inference speed. This framework incorporates two key innovations: (1) a redundant video feature removal module that filters out unnecessary visual information, and (2) a temporal-aware masking strategy that leverages temporal context for enhanced audio generation accuracy. In contrast to existing resource-heavy Unet-based models, MDSGen employs denoising masked diffusion transformers,  facilitating efficient generation without reliance on pre-trained diffusion models. Evaluated on the benchmark VGGSound dataset, our smallest model (5M parameters) achieves 97.9% alignment accuracy, using 172x fewer parameters, 371% less memory, and offering 36x faster inference than the current 860M-parameter state-of-the-art model (93.9% accuracy). The larger model (131M parameters) reaches nearly 99% accuracy while requiring 6.5x fewer parameters. These results highlight the scalability and effectiveness of our approach. The code is available at https://bit.ly/mdsgen."},"_bibtex":{"value":"@inproceedings{\npham2025mdsgen,\ntitle={{MDSG}en: Fast and Efficient Masked Diffusion Temporal-Aware Transformers for Open-Domain Sound Generation},\nauthor={Trung X. Pham and Tri Ton and Chang D. Yoo},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=yFEqYwgttJ}\n}"},"title":{"value":"MDSGen: Fast and Efficient Masked Diffusion Temporal-Aware Transformers for Open-Domain Sound Generation"},"pdf":{"value":"/pdf/ac9a020becd94dd25f980950f2ac0b59550f2139.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"pham|mdsgen_fast_and_efficient_masked_diffusion_temporalaware_transformers_for_opendomain_sound_generation"},"authorids":{"value":["~Trung_X._Pham1","~Tri_Ton1","~Chang_D._Yoo1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Trung X. Pham","Tri Ton","Chang D. Yoo"]}},"version":2},{"content":{"summary":{"value":"This paper studies temporal dependence schemes that decouple content and semantic information, establishing semantic-enhanced Dual Temporal Adjacent Maps for video moment retrieval, conferred as DTAM. Specifically, DTAM designs two branches to encode visual appearance and semantic knowledge from video clips respectively, where knowledge from the appearance branch is distilled into the semantic branch to help DTAM distinguish features with the same visual content but different semantics with a well-designed semantic-aware contrastive loss. Besides, a moment-aware mechanism is also developed to assist temporal adjacent maps' learning for better video grounding. Finally, extensive experimental results and analysis demonstrate the superiority of the proposed DTAM over existing state-of-the-art approaches on three challenging video moment retrieval benchmarks, i.e., TACoS, Charades\u0002STA, and ActivityNet Captions."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"Please refer to the weakness part, especially the novelty issue."},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. This paper tries to address the sophisticated temporal dependency among moments, which leads to an inability to encode temporal dependencies between moments that are crucial for moment retrieval. Therefore, it proposes semantic-enhanced Dual Temporal Adjacent Maps (DTAM) for effective video grounding.\n2. The proposed DTAM achieves satisfactory performance on three public datasets."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"0. The novelty of this paper is somewhat limited. Temporal Adjacent Map strategy was proposed in AAAI 2019 and became a popular trick for video moment retrieval. The proposed Dual Temporal Adjacent Maps method seems to incrementally modify the previous work, without any substantive theoretical analysis. Besides, the contrastive learning strategy is a normal and commonly-sued tricks.\n\n1. Why don't the authors leverage the pre-trained semantic space provided by popular VLMs for feature alignment? Such visual and textual features from VLMs offer rich semantic knowledge and, compared to I3D/C3D and LSTM features, provide better semantic alignment. It is important for matching appropriate video clips with queries. \n\n2. Moreover, the paper states, \"Previous video moment retrieval method ignore the flexibility and complexity of moment description, resulting in an inability to encode temporal dependencies between moments that are crucial for moment retrieval.\" The viewpoint needs to be supported in the experiments or with reference to other published results. Specifically, how are complex queries (e.g., those containing \"again\" as mentioned before) defined, and for samples containing these complex queries, what are the evaluation results of the current methods compared to the proposed DTAM?\n\n3. What distinguishes the appearance temporal adjacent map from the semantic temporal adjacent map, and what are their respective roles? And how is temporal and semantic decoupling accomplished?\n\n4. A previous study [1] suggests that current video moment retrieval approaches, influenced by dataset distribution, lead models to learn dataset-specific biases. What are the similarities and differences between the motivations of this work and those of the proposed DTAM? Additionally, it would be beneficial to evaluate the model's performance on Charades-CD and ActivityNet-CD.\n\n[1] Hao J, Sun H, Ren P, et al. Can shuffling video benefit temporal bias problem: A novel training framework for temporal grounding[C]//European Conference on Computer Vision. Cham: Springer Nature Switzerland, 2022: 130-147."}},"nonreaders":[],"tmdate":1731428020779,"tcdate":1730530532617,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission4047/Reviewer_ygoQ"],"signatures":["ICLR.cc/2025/Conference/Submission4047/Reviewer_ygoQ"],"forum":"l3CSCOnGPB","number":4,"license":"CC BY 4.0","cdate":1730530532617,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission4047/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428020779,"domain":"ICLR.cc/2025/Conference","replyto":"l3CSCOnGPB","id":"ygeZlyOH3U","forumContent":{"TLDR":{"value":"a novel semantic-enhanced dual temporal adjacent maps (DTAM) for effective video moment retrieval, which models temporal dependencies between moments in an appearance-semantic decoupled fashion."},"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Computer Vision","Muylti-modal Understanding","Video Grounding"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Retrieving a specific moment from an untrimmed video via a text description is a central problem in vision-language learning. It is a challenging task due to the sophisticated temporal dependency among moments. Existing methods fail to deal with this issue well since they establish temporal relations of moments in a way that visual content and semantics are coupled. This paper studies temporal dependence schemes that decouple content and semantic information, establishing semantic-enhanced Dual Temporal Adjacent Maps for video moment retrieval, conferred as DTAM. Specifically, DTAM designs two branches to encode visual appearance and semantic knowledge from video clips respectively, where knowledge from the appearance branch is distilled into the semantic branch to help DTAM distinguish features with the same visual content but different semantics with a well-designed semantic-aware contrastive loss. Besides, we also develop a moment-aware mechanism to assist temporal adjacent maps' learning for better video grounding. Finally, extensive experimental results and analysis demonstrate the superiority of the proposed DTAM over existing state-of-the-art approaches on three challenging video moment retrieval benchmarks, i.e., TACoS, Charades-STA, and ActivityNet Captions."},"_bibtex":{"value":"@misc{\nwang2025learning,\ntitle={Learning Semantic-Enhanced Dual Temporal Adjacent Maps for Video Moment Retrieval},\nauthor={Yu Wang and Shengjie Zhao and Shiwei Chen},\nyear={2025},\nurl={https://openreview.net/forum?id=l3CSCOnGPB}\n}"},"title":{"value":"Learning Semantic-Enhanced Dual Temporal Adjacent Maps for Video Moment Retrieval"},"pdf":{"value":"/pdf/83e3b8d8b0191f0c4f665fc641a4a682d0f67e99.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"wang|learning_semanticenhanced_dual_temporal_adjacent_maps_for_video_moment_retrieval"},"authorids":{"value":["~Yu_Wang32","~Shengjie_Zhao1","~Shiwei_Chen3"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Yu Wang","Shengjie Zhao","Shiwei Chen"]}},"version":2},{"content":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["medical video analysis","one-shot video object segmentation","test-time training","self-distillation"]},"primary_area":{"value":"unsupervised, self-supervised, semi-supervised, and supervised representation learning"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"This paper introduces a novel task and approach for one-shot medical video object segmentation using static image datasets. We address the critical challenge of limited annotated video data in medical imaging by proposing a framework that leverages readily available labeled static images to segment objects in medical videos with minimal annotation---specifically, a ground truth mask for only the first frame. Our method comprises training a one-shot segmentation model exclusively on images, followed by adapting it to medical videos through a test-time training strategy. This strategy incorporates a memory mechanism to utilize spatiotemporal context and employs self-distillation to maintain generalization capabilities. To facilitate research in this domain, we present OS-I2V-Seg, a comprehensive dataset comprising 28 categories in images and 4 categories in videos, totaling 68,416 image/frame-mask pairs. Extensive experiments demonstrate the efficacy of our approach in this extremely low-data regime for video object segmentation, establishing baseline performance on OS-I2V-Seg. The code and data will be made publicly available."},"_bibtex":{"value":"@misc{\nzheng2024temporalaware,\ntitle={Temporal-Aware Test-Time Training via Self-Distillation for One-Shot Image-to-Video Segmentation},\nauthor={Qicong Wang and Yilei Shi and Jingliang Hu and Xiao Xiang Zhu and Lichao Mou},\nyear={2024},\nurl={https://openreview.net/forum?id=BhECSDSkAE}\n}"},"title":{"value":"Temporal-Aware Test-Time Training via Self-Distillation for One-Shot Image-to-Video Segmentation"},"pdf":{"value":"/pdf/545f3a422f20d41a30e003f7fa3165c198d4e8d7.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"wang|temporalaware_testtime_training_via_selfdistillation_for_oneshot_imagetovideo_segmentation"},"authorids":{"value":["~Qicong_Wang2","~Yilei_Shi1","~Jingliang_Hu1","~Xiao_Xiang_Zhu1","~Lichao_Mou3"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Qicong Wang","Yilei Shi","Jingliang Hu","Xiao Xiang Zhu","Lichao Mou"]}},"tmdate":1753983459540,"tcdate":1727492931844,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission13118/Authors"],"signatures":["ICLR.cc/2025/Conference/Submission13118/Authors"],"forum":"BhECSDSkAE","license":"CC BY 4.0","number":13118,"cdate":1727492931844,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/-/Submission","ICLR.cc/2025/Conference/-/Post_Submission","ICLR.cc/2025/Conference/Submission13118/-/Full_Submission","ICLR.cc/2025/Conference/-/Withdrawn_Submission","ICLR.cc/2025/Conference/-/Edit"],"mdate":1753983459540,"odate":1728008565725,"domain":"ICLR.cc/2025/Conference","id":"BhECSDSkAE","version":2},{"content":{"summary":{"value":"This work proposes a trajectory-aware spatial-temporal graph for video salient object ranking. The proposed model includes a spatial correlation graph and a temporal correlation graph. Unlike previous VSOR methods, this work suggests to modeling instance-level temporal relations. They conduct experiments to demonstrate the advantage of their method, and also show the effectiveness of their method in video retargeting task."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"1. How does the model benefit from \"Trajectory-wise Contrast\" if the motion is too dynamic? For example, in Figure 2, the persons are moving away from their original position and go out of the expanded bounding box at t+1 frame. Then, $f_1$ cannot track the future trajectory by comparing $f_1$ and $f_1^{t+1}$, as the person goes outside of the bounding box. If the model fails to track the object, how does the model benefit from \"Trajectory-wise Contrast\"?\n\n2. It would be beneficial if the authors included studies to investigate how the bounding box area impacts VSOR performance. Whether a larger bounding box is more friendly for faster motion?\n\n3. What does the \"||\" mean in Eq. 1, Eq. 2, Eq. 3 and Eq. 4? Does it mean concatenation? It would be better to specify this just after Eq. 1. \n\n4. Why the numerical results in table 1 are significantly different from those reported in [1]?\n\n5. The Eq. 2 and Eq. 3 are confusing and inconsistent with Figure 2 ($T^i, T^t$). $f_{j}^{t}$ means $j^{th}$ instance at frame $t$, and $f_j$ mean $j^{th}$ instance at current frame. If so, why $f_j$ in Eq. 2 doesn't include a superscript? The $f_*$ in Figure 2 ($T^i$) also misses some superscripts. Does $h_{N_{i}^{t_i}}$ aggregate information from $f_i$ in Eq. 2?\n\n6. How long does it take to train the model?"},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":1},"strengths":{"value":"1. The idea of this work is clear. Unlike previous VSOR methods, which focus on frame-level temporal relations, this work proposes to explore the instance-level temporal relations.\n\n2. They apply their VSOR model to the video retargeting task and achieved good results."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The authors claim that they **explicitly** model the motion trajectories of each instance (L8-9,L42,L69). However, the proposed method doesn't support this claim as the proposed method neither track each instance nor re-identify the same instance. Instead, the proposed method aggregates local contextual features from the same position across frames to **approximates** the instance trajectory. This claim might be not accurate.\n\n2. Lack comparison with latest SOR methods. VSOR is highly related to image-based SOR. In table 1, the compared image-based methods are [2,5] (2021), however, some latest methos are available, such as, \"Bi-directional Object-context Prioritization Learning for Saliency Ranking\" (2022-CVPR), \"Partitioned Saliency Ranking with Dense Pyramid Transformers\" (2023-ACM MM) and \"SeqRank: Sequential Ranking of Salient Objects\" (2024-AAAI).\n\n3. There is a real lack of experimental results to substantiate its advantages. In Table 1, only three methods are presented: two image-based salient object recognition (SOR) methods from 2021 and one video salient object recognition (VSOR) method from 2022. To strengthen the evaluation, it would be prudent to compare the proposed method with more approaches from related fields. For instance, consider video salient object detection methods like “Shifting More Attention to Video Salient Object Detection” (2019 CVPR) and “Dynamic Context-Sensitive Filtering Network for Video Salient Object Detection” (2021 ICCV). Additionally, exploring video shadow detection methods such as “SCOTCH and SODA: A Transformer Video Shadow Detection Framework” (2023 CVPR) could provide valuable insights.\n\n4. Why the SA-SOR scores of DAVSOD is much worse than those in RVSOD? Need explanation. In addition, it would be better to include a naive solution as comparison, since the SA-SOR score is very low. For example, a naive solution is to rank the instances according to their sizes to see how well the model outperforms such naive solution."},"limitations":{"value":"The authors have discussed the limitations, and I agree that there is no potential negative societal impact of their work."}},"nonreaders":[],"tmdate":1730879956982,"tcdate":1718340307805,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission18167/Reviewer_TupH"],"signatures":["NeurIPS.cc/2024/Conference/Submission18167/Reviewer_TupH"],"forum":"VUBtAcQN44","number":1,"license":"CC BY 4.0","cdate":1718340307805,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission18167/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730879956982,"domain":"NeurIPS.cc/2024/Conference","replyto":"VUBtAcQN44","id":"kmGlfUIUM3","forumContent":{"venue":{"value":"NeurIPS 2024 poster"},"keywords":{"value":["Video Salient Object Ranking","Spatio-temporal Graph","Video Retargeting"]},"primary_area":{"value":"machine_vision"},"abstract":{"value":"Video salient object ranking aims to simulate the human attention mechanism by dynamically prioritizing the visual attraction of objects in a scene over time. Despite its numerous practical applications, this area remains underexplored. In this work, we propose a graph model for video salient object ranking. This graph simultaneously explores multi-scale spatial contrasts and intra-/inter-instance temporal correlations across frames to extract diverse spatio-temporal saliency cues. It has two advantages: 1. Unlike previous methods that only perform global inter-frame contrast or compare all proposals across frames globally, we explicitly model the motion of each instance by comparing its features with those in the same spatial region in adjacent frames, thus obtaining more accurate motion saliency cues. 2. We synchronize the spatio-temporal saliency cues in a single graph for joint optimization, which exhibits better dynamics compared to the previous stage-wise methods that prioritize spatial cues followed by temporal cues. Additionally, we propose a simple yet effective video retargeting method based on video saliency ranking. Extensive experiments demonstrate the superiority of our model in video salient object ranking and the effectiveness of the video retargeting method. Our codes/models are released at [https://github.com/zyf-815/VSOR/tree/main](https://github.com/zyf-815/VSOR/tree/main)."},"_bibtex":{"value":"@inproceedings{\nchen2024a,\ntitle={A Motion-aware Spatio-temporal Graph for Video Salient Object Ranking},\nauthor={Hao Chen and Yufei Zhu and Yongjian Deng},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=VUBtAcQN44}\n}"},"title":{"value":"A Motion-aware Spatio-temporal Graph for Video Salient Object Ranking"},"pdf":{"value":"/pdf/0d92c407daaca6fbfa64bbeac6b9fbed7ebdd3b6.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"chen|a_motionaware_spatiotemporal_graph_for_video_salient_object_ranking"},"authorids":{"value":["~Hao_Chen39","~Yufei_Zhu4","~Yongjian_Deng1"]},"authors":{"value":["Hao Chen","Yufei Zhu","Yongjian Deng"]}},"version":2},{"content":{"venue":{"value":"AAAI 2026"},"pdf":{"value":"https://ojs.aaai.org/index.php/AAAI/article/download/42281/46242"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"suhail|shortcut_learning_susceptibility_in_vision_classifiers_student_abstract"},"html":{"value":"https://doi.org/10.1609/aaai.v40i48.42281"},"_bibtex":{"value":"@inproceedings{DBLP:conf/aaai/SuhailGS26,\n  author={Pirzada Suhail and Vrinda Goel and Amit Sethi},\n  title={Shortcut Learning Susceptibility in Vision Classifiers (Student Abstract)},\n  year={2026},\n  cdate={1767225600000},\n  pages={41390-41392},\n  url={https://doi.org/10.1609/aaai.v40i48.42281},\n  booktitle={AAAI},\n  crossref={conf/aaai/2026}\n}\n"},"abstract":{"value":"Shortcut learning, where machine learning models exploit spurious correlations in data instead of capturing meaningful features, poses a significant challenge to building generalizable models. Vision classifiers based on Convolutional Neural Networks (CNNs), Multi-Layer Perceptrons (MLPs), and Vision Transformers (ViTs) leverage distinct architectural principles to process spatial and structural information, making them differently susceptible to shortcut learning. In this study, we systematically evaluate these architectures by introducing deliberate shortcuts into the dataset that are correlated with class labels both positionally and via intensity, creating a controlled setup to assess whether models rely on these artificial cues or learn actual distinguishing features. We perform both quantitative evaluation by training on the shortcut-modified dataset and testing on two different test sets—one containing the same shortcuts and another without them—to determine the extent of reliance on shortcuts. Additionally, qualitative evaluation is performed using network inversion-based reconstruction techniques to analyze what the models internalize in their weights, aiming to reconstruct the training data as perceived by the classifiers. Further, we evaluate susceptibility to shortcut learning across different learning rates. Our analysis reveals that CNNs at lower learning rates tend to be more reserved against entirely picking up shortcut features, while ViTs, particularly those without positional encodings, almost entirely ignore the distinctive image features in the presence of shortcuts."},"title":{"value":"Shortcut Learning Susceptibility in Vision Classifiers (Student Abstract)"},"authors":{"value":[{"fullname":"Pirzada Suhail","username":"~Pirzada_Suhail1"},{"fullname":"Vrinda Goel","username":""},{"fullname":"Amit Sethi","username":"~Amit_Sethi2"}]}},"tmdate":1785934426872,"pdate":1798675200000,"externalIds":["dblp:conf/aaai/SuhailGS26"],"tcdate":1782152567170,"writers":["~","OpenReview.net/Public_Article/DBLP.org","OpenReview.net/Support"],"signatures":["~Amit_Sethi2"],"forum":"OSB06uk2MP","license":"CC BY-SA 4.0","number":45769,"cdate":1767225600000,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/DBLP.org/-/Record","OpenReview.net/Public_Article/-/Edit","OpenReview.net/Public_Article/-/Authorship_Claim"],"mdate":1785934426872,"domain":"OpenReview.net/Public_Article","id":"OSB06uk2MP","version":2},{"content":{"summary":{"value":"GeoDrag is a geometry-aware drag-based image editing framework combining 3D structural cues and 2D spatial priors. It predicts displacement fields in diffusion latent space, incorporating depth-based modulation and spatial plane adjustment, along with a conflict-free partitioning scheme for multi-point editing."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"1.\tHow sensitive is the method to depth map errors?\n2.\tCan GeoDrag operate with estimated depth from monocular models?\n3.\tIs the diffusion model re-trained or used as a frozen backbone?"},"rating":{"value":4},"details_of_ethics_concerns":{"value":"None"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. Addresses a concrete weakness: Most drag methods ignore geometry, causing distortions. GeoDrag effectively solves this.\n\n2. Principled integration: The combination of 3D depth modulation and 2D plane adjustment is physically intuitive and well-designed.\n\n3. Practical contribution: One-step editing drastically improves usability and speed over iterative gradient-based methods.\n\n4. Strong empirical results: Quantitative improvements (1.4× DAI, 1.1× MD) and excellent qualitative visuals demonstrate real impact.\n\n5. Multi-point conflict resolution: - The conflict-free partitioning module is a simple yet powerful idea that improves robustness.\n6. Extensive evaluation: Diverse examples (faces, landscapes, multi-object scenes) show generality."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1.\tDepth dependency: Accuracy hinges on reliable depth maps; the paper assumes ideal depth input without addressing noise robustness.\n2.\tLimited theoretical grounding: The interaction between geometry and diffusion latent space lacks mathematical explanation.\n3.\tAblation insufficiency:The individual contribution of geometry-aware vs. plane-aware modules is unclear.\n4..\tEfficiency discussion shallow: Though claimed efficient, actual GPU/latency data is minimal.\n5.\tApplicability scope: Focused on still images; no experiments on video or non-diffusion backbones.\n6..\tMinor clarity issues: Some figures are over-compressed, making displacement fields hard to interpret."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762923246963,"tcdate":1761401361889,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission12321/Reviewer_m8ds"],"signatures":["ICLR.cc/2026/Conference/Submission12321/Reviewer_m8ds"],"forum":"MBiMt3wp8M","number":1,"license":"CC BY 4.0","cdate":1761401361889,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission12321/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762923246963,"domain":"ICLR.cc/2026/Conference","replyto":"MBiMt3wp8M","id":"lDUJBnn6lT","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Diffusion Model; Drag-based Image Editing"]},"supplementary_material":{"value":"/attachment/22d353a68c01660ad65f7e7d4dfee26915559a38.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Interactive point-based image editing serves as a controllable editor, enabling precise and flexible manipulation of image content. However, most drag-based methods operate primarily on the 2D pixel plane with limited use of 3D cues. As a result, they often produce imprecise and inconsistent edits, particularly in geometry-intensive scenarios such as rotations and perspective transformations. To address these limitations, we propose a novel geometry-guided drag-based image editing method—GeoDrag, which addresses three key challenges: 1) incorporating 3D geometric cues into pixel-level editing, 2) mitigating discontinuities caused by geometry-only guidance, and 3) resolving conflicts arising from multi-point dragging. Built upon a unified displacement field that jointly encodes 3D geometry and 2D spatial priors, GeoDrag enables coherent, high-fidelity, and structure-consistent editing in a single forward pass. In addition, a conflict-free partitioning strategy is introduced to isolate editing regions, effectively preventing interference and ensuring consistency. Extensive experiments across various editing scenarios validate the effectiveness of our method, showing superior precision, structural consistency, and reliable multi-point editability. Project page: https://xinyu-pu.github.io/projects/geodrag."},"_bibtex":{"value":"@inproceedings{\npu2026dragging,\ntitle={Dragging with Geometry: From Pixels to Geometry-Guided Image Editing},\nauthor={Xinyu Pu and Hongsong Wang and Jie Gui and Pan Zhou},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=MBiMt3wp8M}\n}"},"title":{"value":"Dragging with Geometry: From Pixels to Geometry-Guided Image Editing"},"pdf":{"value":"/pdf/57deef21f397fac9a8baaa0d0230792ac82d9137.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"pu|dragging_with_geometry_from_pixels_to_geometryguided_image_editing"},"authorids":{"value":["~Xinyu_Pu1","~Hongsong_Wang2","~Jie_Gui1","~Pan_Zhou3"]},"authors":{"value":["Xinyu Pu","Hongsong Wang","Jie Gui","Pan Zhou"]}},"version":2},{"content":{"TLDR":{"value":"This paper measures and mitigates shortcut leaning in the misinformation detection."},"venue":{"value":"NeurIPS 2025 poster"},"keywords":{"value":["Shortcut Learning","Injection Attack","Misinformation Detection"]},"supplementary_material":{"value":"/attachment/32eb4c62e2bfba393e15c799d60f50adffae790a.zip"},"primary_area":{"value":"applications"},"abstract":{"value":"Misinformation detectors often rely on superficial cues (i.e., shortcuts) that correlate with misinformation in training data but fail to generalize to the diverse and evolving nature of real-world misinformation. This issue is exacerbated by large language models (LLMs), which can easily generate convincing misinformation using simple prompts. We introduce TruthOverTricks, a unified evaluation paradigm for measuring shortcut learning in misinformation detection. TruthOverTricks categorizes shortcut behaviors into intrinsic shortcut induction and extrinsic shortcut injection, and evaluates seven representative detectors across 14 popular benchmarks, along with two new factual misinformation datasets, NQ-Misinfo and Streaming-Misinfo. Empirical results reveal that existing detectors suffer severe performance degradation when exposed to both naturally occurring and adversarially crafted shortcuts. To address this, we propose the Shortcut Mitigation Framework (SMF), an LLM-augmented data augmentation framework that mitigates shortcut reliance through paraphrasing, factual summarization, and sentiment normalization. SMF consistently enhances robustness across 16 benchmarks, forcing models to rely on deeper semantic understanding rather than shortcut cues."},"_bibtex":{"value":"@inproceedings{\nwan2025truth,\ntitle={Truth over Tricks: Measuring and Mitigating Shortcut Learning in Misinformation Detection},\nauthor={Herun Wan and Jiaying Wu and Minnan Luo and Zhi Zeng and Zhixiong Su},\nbooktitle={The Thirty-ninth Annual Conference on Neural Information Processing Systems},\nyear={2025},\nurl={https://openreview.net/forum?id=ngxGNQE1M2}\n}"},"title":{"value":"Truth over Tricks: Measuring and Mitigating Shortcut Learning in Misinformation Detection"},"pdf":{"value":"/pdf/9c68e8baf3c47644094e482abcc69d6b27b23902.pdf"},"venueid":{"value":"NeurIPS.cc/2025/Conference"},"paperhash":{"value":"wan|truth_over_tricks_measuring_and_mitigating_shortcut_learning_in_misinformation_detection"},"authorids":{"value":["~Herun_Wan1","~Jiaying_Wu2","~Minnan_Luo1","~Zhi_Zeng4","~Zhixiong_Su1"]},"authors":{"value":["Herun Wan","Jiaying Wu","Minnan Luo","Zhi Zeng","Zhixiong Su"]}},"tmdate":1783627313391,"pdate":1758217316188,"tcdate":1746969848452,"writers":["NeurIPS.cc/2025/Conference","NeurIPS.cc/2025/Conference/Submission21440/Authors"],"signatures":["NeurIPS.cc/2025/Conference/Submission21440/Authors"],"forum":"ngxGNQE1M2","license":"CC BY 4.0","number":21440,"cdate":1746969848452,"readers":["everyone"],"invitations":["NeurIPS.cc/2025/Conference/-/Submission","NeurIPS.cc/2025/Conference/-/Post_Submission","NeurIPS.cc/2025/Conference/Submission21440/-/Full_Submission","NeurIPS.cc/2025/Conference/-/Edit","NeurIPS.cc/2025/Conference/Submission21440/-/Camera_Ready_Revision"],"mdate":1783627313391,"odate":1761705009128,"domain":"NeurIPS.cc/2025/Conference","id":"ngxGNQE1M2","version":2},{"content":{"research_area_keywords":{"value":"optimization methods, reinforcement learning, generalization"},"keywords":{"value":["Direct Preference Optimization","Iterative Preference Optimization","Verifiable Mathematical Reasoning","Preference Pair Filtering","Support-Aware Filtering","On-policy and Fallback Pairs","Implicit Reward Margin"]},"B_use_or_create_scientific_artifacts":{"value":"Yes"},"languages_studied":{"value":"English"},"D4_ethics_review_board_approval":{"value":"N/A"},"C4_parameters_for_packages":{"value":"No"},"B6_elaboration":{"value":"section 4, appendix D"},"B1_cite_creators_of_artifacts":{"value":"Yes"},"B2_discuss_the_license_for_artifacts":{"value":"No"},"B4_elaboration":{"value":"The paper uses existing mathematical reasoning benchmarks rather than collecting new user-generated or personal data."},"B5_elaboration":{"value":"section 4"},"D_human_subjects_including_annotators":{"value":"No"},"B2_elaboration":{"value":"The paper uses existing public benchmarks, GSM8K and MATH-500, but does not explicitly discuss their licenses or terms of use. We will add this information in the final version."},"B3_elaboration":{"value":"Section 4"},"E_ai_assistants_in_research_or_writing":{"value":"Yes"},"C2_experimental_setup_and_hyperparameters":{"value":"Yes"},"C1_model_size_and_budget":{"value":"No"},"venue":{"value":"ACL ARR 2026 May Submission"},"D3_data_consent":{"value":"N/A"},"_bibtex":{"value":"@inproceedings{\nanonymous2026not,\ntitle={Not All Verifiable Preference Pairs Are Equal: Support-Aware Filtering for Stable Iterative {DPO}},\nauthor={Anonymous},\nbooktitle={Submitted to ACL Rolling Review - May 2026},\nyear={2026},\nurl={https://openreview.net/forum?id=Ms8fv561vJ},\nnote={under review}\n}"},"title":{"value":"Not All Verifiable Preference Pairs Are Equal: Support-Aware Filtering for Stable Iterative DPO"},"C3_descriptive_statistics":{"value":"Yes"},"contribution_types":{"value":["NLP engineering experiment","Data analysis"]},"A1_limitations_section":{"value":"This paper has a limitations section."},"B6_statistics_for_data":{"value":"Yes"},"B4_data_contains_personally_identifying_info_or_offensive_content":{"value":"No"},"B3_artifact_use_consistent_with_intended_use":{"value":"Yes"},"A2_potential_risks":{"value":"No"},"E1_information_about_use_of_ai_assistants":{"value":"No"},"abstract":{"value":"Direct Preference Optimization (DPO) is well suited to verifiable reasoning tasks because preference pairs can be constructed automatically from answer correctness. However, in iterative DPO, correctness alone does not determine whether a pair provides a useful training signal. We distinguish two regimes according to the origin of the chosen response: On-policy pairs, where both the chosen and rejected responses are sampled from the current policy, and Fallback pairs, where the policy fails to generate a correct response and the chosen completion is supplied by a dataset-provided reference solution. Based on this distinction, we propose support-aware filtering, which applies margin-based filtering to On-policy pairs and a likelihood-based support gate to Fallback pairs. Experiments on GSM8K and MATH-500 with Qwen2.5-style 0.5B, 1.5B, and 3B models show that support-aware filtering yields its clearest gains on MATH-500, while GSM8K results are more mixed. Overall, our findings suggest that preference margins in iterative verifiable DPO are pair-origin dependent, and that chosen-response support is a useful signal for stabilizing fallback supervision."},"D1_instructions_given_to_participants":{"value":"N/A"},"paper_type":{"value":"Short"},"C_computational_experiments":{"value":"Yes"},"pdf":{"value":"/pdf/246af7827e18106dde15a3ead091ff687de89bf7.pdf"},"research_area":{"value":"Machine Learning for NLP"},"B5_documentation_of_artifacts":{"value":"Yes"},"EMNLP_2026_AI_Reviewing_Experiment":{"value":"no"},"D2_recruitment_and_payment":{"value":"N/A"},"venueid":{"value":"aclweb.org/ACL/ARR/2026/May/Submission"}},"tmdate":1790696955776,"tcdate":1779690571813,"writers":["aclweb.org/ACL/ARR/2026/May","aclweb.org/ACL/ARR/2026/May/Submission6721/Authors"],"signatures":["aclweb.org/ACL/ARR/2026/May/Submission6721/Authors"],"forum":"Ms8fv561vJ","license":"CC BY 4.0","number":6721,"cdate":1779690571813,"readers":["everyone"],"invitations":["aclweb.org/ACL/ARR/2026/May/-/Submission","aclweb.org/ACL/ARR/2026/May/-/Edit","aclweb.org/ACL/ARR/2026/May/-/Post_Submission","aclweb.org/ACL/ARR/2026/May/-/Preprint_Post_Submission"],"mdate":1790696955776,"odate":1780385598983,"domain":"aclweb.org/ACL/ARR/2026/May","id":"Ms8fv561vJ","version":2},{"content":{"summary":{"value":"This paper presents a zero-shot video moment retrieval method. It specifically addresses the dependence of existing zero-shot methods on query-to-context similarity which suffer from modality and language-style gaps. To overcome this, the paper proposes Self-Similarity-based Moment proposal and Scoring to generate consistent candidates. It also introduces a query-aware MLLM-based reasoning stage to further refine the alignment between text and video. SOTA results are reported on multiple zero-shot video moment retrieval benchmarks."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"See the Weaknesses section."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"The paper is well written. The structure is easy to follow, and illustrations are clear.  \n\nThe idea of self-similarity to address the modality gap in zero-shot video moment retrieval is interesting and plausible. \n\nBoundary scores calculation to generate candidate spans from self-similarity of frame/caption features is likely to work well.  \n\nBesides improvement of the event boundary detection, adding caption information in addition to the visual features is an interesting direction. \n\nThe idea of using MLLM to re-rank the candidate spans is also interesting and should lead to improved performance. \n\nThe overall method seems to be well designed and convincing."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The idea of self-similarity to address modality gap is not new. If fact, it is one of the earliest methods I can recall in this line of work. I am not sure whether it has been used in the context of video-moment retrieval or not. Never-the-less, the novelty of the proposed method remains low as this would be just another application. I am keen so see how the other reviewers perceive this and remain open to being convinced otherwise during the rebuttal period. \n\nVision-language models such as CLIP are extensively trained to reduce the modality gap. How is self-similarity better? \n\nIn the Introduction, the idea of event boundaries is linked with shot/scene clustering only. However, later on in the methods, information about additionally using captions is presented. For completeness, this information should also be mentioned in the Introduction, as it is a vital part of the method. \n\nThe candidate spans will vary significantly based on the chosen threshold. Did you do an experiment for this?  \n\nThere are other adhoc choices and hyper-parameters e.g. top-$k_S$ frame-level scores, SMS calculation, and $\\alpha$, $\\beta$ hyper-parameters. How sensitive is the method to these choices? \n\nThe framework's ability to be scaled for practical settings remains unclear. The added complexity and costs seem to outweigh the performance gains compared to existing works. For example, the gains on R1@0.5 metric over existing works are not that high in general and much low in the case of ActivityNet-Captions dataset.\n\nThe ablations in Table 4 show minimal improvement provided by adding SMS and Re-ranking modules. Similarly, in Table 5, Baseline is as almost as good as MLLM Re-ranking. The significant overhead of using an MLLM does not seem justified.\n\nTypo on page 4: “Then, The TSMs ...” -> Then, the TSMs..."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762922332186,"tcdate":1761873334683,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission11175/Reviewer_hLgp"],"signatures":["ICLR.cc/2026/Conference/Submission11175/Reviewer_hLgp"],"forum":"o6Msp3XZGz","number":3,"license":"CC BY 4.0","cdate":1761873334683,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission11175/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762922332186,"domain":"ICLR.cc/2026/Conference","replyto":"o6Msp3XZGz","id":"eanG66HIWp","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Video Moment Retrieval","Zero-shot Video Moment Retrieval"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Zero-shot video moment retrieval (ZMR) aims to overcome the limitations of traditional approaches that require large-scale datasets annotated with text and its relevant temporal spans. Despite advances in pre-trained vision–language models (VLMs) and multimodal large language models (MLLMs), existing ZMR methods still heavily depend on query-to-context similarity, making them vulnerable to modality and language-style gaps.  These gaps lead to unreliable span proposals and unstable moment retrieval results. To address this issue, we propose Self-Similarity-based Moment proposal and Scoring (Self-SiMS) that instead exploits intrinsic relationships within videos, enabling consistent candidate generation and scoring. By deriving self-similarity only from the video content, we circumvent the noisy and mismatched patterns of query–frame or query–caption similarities, thereby mitigating both modality and language-style gaps. Furthermore, we introduce a query-aware MLLM-based reasoning stage to further sharpen alignment between text and video by mitigating modality and language-style gaps. Extensive experiments demonstrate that Self-SiMS achieves the state-of-the-art performance across multiple ZMR benchmarks."},"_bibtex":{"value":"@misc{\nlee2025mitigating,\ntitle={Mitigating Modality and Language-Style Gaps for Zero-Shot Video Moment Retrieval},\nauthor={Jihyun Lee and Cheol-Ho Cho and Woojin Jun and Woojin Jeong and Jae-Pil Heo},\nyear={2025},\nurl={https://openreview.net/forum?id=o6Msp3XZGz}\n}"},"title":{"value":"Mitigating Modality and Language-Style Gaps for Zero-Shot Video Moment Retrieval"},"pdf":{"value":"/pdf/b42e7543e1427ee870fb239aac408cdf686cb007.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"lee|mitigating_modality_and_languagestyle_gaps_for_zeroshot_video_moment_retrieval"},"authorids":{"value":["~Jihyun_Lee7","~Cheol-Ho_Cho1","~Woojin_Jun1","~Woojin_Jeong1","~Jae-Pil_Heo3"]},"authors":{"value":["Jihyun Lee","Cheol-Ho Cho","Woojin Jun","Woojin Jeong","Jae-Pil Heo"]}},"version":2},{"content":{"TLDR":{"value":"We propose a multimodal multi-agent Theory of Mind benchmark, MuMA-ToM"},"venue":{"value":"Video-Langauge Models Poster"},"keywords":{"value":["Multimodal Theory of Mind","Multi-agent Interactions","Bayesian Inverse Planning"]},"supplementary_material":{"value":"/attachment/8b9aebc271649a11a599bf3ee35792c2f5d2222d.pdf"},"abstract":{"value":"Understanding people’s social interactions in complex real-\nworld scenarios often relies on intricate mental reasoning.\nTo truly understand how and why people interact with one\nanother, we must infer the underlying mental states that give\nrise to the social interactions, i.e., Theory of Mind reason-\ning in multi-agent interactions. Additionally, social interac-\ntions are often multi-modal – we can watch people’s actions,\nhear their conversations, and/or read about their past behav-\niors. For AI systems to successfully and safely interact with\npeople in real-world environments, they also need to under-\nstand people’s mental states as well as their inferences about\neach other’s mental states based on multi-modal information\nabout their interactions. For this, we introduce MuMA-ToM,\na Multi-modal Multi-Agent Theory of Mind benchmark.\nMuMA-ToM is the first multi-modal Theory of Mind bench-\nmark that evaluates mental reasoning in embodied multi-\nagent interactions. In MuMA-ToM, we provide video and\ntext descriptions of people’s multi-modal behavior in realis-\ntic household environments. Based on the context, we then\nask questions about people’s goals, beliefs, and beliefs about\nothers’ goals. We validated MuMA-ToM in a human ex-\nperiment and provided a human baseline. We also proposed\na novel multi-modal, multi-agent ToM model, LIMP (Lan-\nguage model-based Inverse Multi-agent Planning). Our ex-\nperimental results show that LIMP significantly outperforms\nstate-of-the-art methods, including large multi-modal mod-\nels (e.g., GPT-4o, Gemini-1.5 Pro) and a recent multi-modal\nToM model, BIP-ALM."},"_bibtex":{"value":"@inproceedings{\nshi2025mumatom,\ntitle={Mu{MA}-ToM: Multi-modal Multi-Agent Theory of Mind},\nauthor={Haojun Shi and Suyu Ye and Xinyu Fang and Chuanyang Jin and Leyla Isik and Yen-Ling Kuo and Tianmin Shu},\nbooktitle={Workshop on Video-Language Models @ NeurIPS 2024},\nyear={2025},\nurl={https://openreview.net/forum?id=9WfmsKmXF9}\n}"},"title":{"value":"MuMA-ToM: Multi-modal Multi-Agent Theory of Mind"},"pdf":{"value":"/pdf/c65e2cbe9b5af9e284f07e01e0d3d653f9d038b7.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Workshop/Video-Langauge_Models"},"paperhash":{"value":"shi|mumatom_multimodal_multiagent_theory_of_mind"},"authorids":{"value":["~Haojun_Shi1","~Suyu_Ye1","~Xinyu_Fang2","~Chuanyang_Jin2","~Leyla_Isik1","~Yen-Ling_Kuo1","~Tianmin_Shu1"]},"track":{"value":"Long Paper Track (up to 9 pages)"},"authors":{"value":["Haojun Shi","Suyu Ye","Xinyu Fang","Chuanyang Jin","Leyla Isik","Yen-Ling Kuo","Tianmin Shu"]}},"tmdate":1736861079962,"pdate":1730081751974,"tcdate":1725322271803,"writers":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission7/Authors"],"signatures":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission7/Authors"],"forum":"9WfmsKmXF9","license":"CC BY 4.0","number":7,"cdate":1725322271803,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Post_Submission","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/-/Edit","NeurIPS.cc/2024/Workshop/Video-Langauge_Models/Submission7/-/Camera-Ready_Revision"],"mdate":1736861079962,"odate":1736861079933,"domain":"NeurIPS.cc/2024/Workshop/Video-Langauge_Models","id":"9WfmsKmXF9","version":2},{"content":{"venue":{"value":"ICME 2023"},"pdf":{"value":"https://ieeexplore.ieee.org/iel7/10219544/10219545/10219770.pdf"},"venueid":{"value":"dblp.org/conf/ICMCS/2023"},"paperhash":{"value":"menon|just_noticeable_differenceaware_perscene_bitrateladdering_for_adaptive_video_streaming"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Vignesh_V._Menon:","https://dblp.org/search/pid/api?q=author:Jingwen_Zhu:","https://dblp.org/search/pid/api?q=author:Prajit_T._Rajendran:","https://dblp.org/search/pid/api?q=author:Hadi_Amirpour:","https://dblp.org/search/pid/api?q=author:Patrick_Le_Callet:","~Christian_Timmerer1"]},"html":{"value":"https://doi.org/10.1109/ICME55011.2023.00288"},"_bibtex":{"value":"@inproceedings{DBLP:conf/icmcs/MenonZRACT23,\n  author={Vignesh V. Menon and Jingwen Zhu and Prajit T. Rajendran and Hadi Amirpour and Patrick Le Callet and Christian Timmerer},\n  title={Just Noticeable Difference-Aware Per-Scene Bitrate-Laddering for Adaptive Video Streaming},\n  year={2023},\n  cdate={1672531200000},\n  pages={1673-1678},\n  url={https://doi.org/10.1109/ICME55011.2023.00288},\n  booktitle={ICME},\n  crossref={conf/icmcs/2023}\n}\n"},"abstract":{"value":"In video streaming applications, a fixed set of bitrate-resolution pairs (known as a bitrate ladder) is typically used during the entire streaming session. However, an optimized bitrate ladder per scene may result in (i) decreased storage or delivery costs or/and (ii) increased Quality of Experience. This paper introduces a Just Noticeable Difference (JND)-aware perscene bitrate ladder prediction scheme (JASLA) for adaptive video-on-demand streaming applications. JASLA predicts jointly optimized resolutions and corresponding constant rate factors (CRFs) using spatial and temporal complexity features for a given set of target bitrates for every scene, which yields an efficient constrained Variable Bitrate encoding. Moreover, bitrate-resolution pairs that yield distortion lower than one JND are eliminated. Experimental results show that, on average, JASLA yields bitrate savings of 34.42% and 42.67% to maintain the same PSNR and VMAF, respectively, compared to the reference HTTP Live Streaming (HLS) bitrate ladder Constant Bitrate encoding using x265 HEVC encoder, where the maximum resolution of streaming is Full HD (1080p). Moreover, a 54.34% average cumulative decrease in storage space is observed."},"title":{"value":"Just Noticeable Difference-Aware Per-Scene Bitrate-Laddering for Adaptive Video Streaming"},"authors":{"value":["Vignesh V. Menon","Jingwen Zhu","Prajit T. Rajendran","Hadi Amirpour","Patrick Le Callet","Christian Timmerer"]}},"tmdate":1741250367991,"pdate":1672531200000,"tcdate":1741250323035,"writers":["~"],"signatures":["~Christian_Timmerer1"],"forum":"BgchGy6exe","license":"CC BY-SA 4.0","number":359805,"cdate":1672531200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1741250367991,"domain":"DBLP.org","id":"BgchGy6exe","version":2},{"content":{"summary":{"value":"The paper introduces Listen to Motion (LiMo), a new approach that incorporates motion information to refine the correlation between audio and dynamic visual elements. LiMo works by extracting temporal visual semantics to foster interaction between video frames while preserving the static visual-audio correlations of prior models. It improves audio-visual alignment by differentiating between different video clips and introducing a method to identify and reweight false positive or multiple positive instances. Extensive experiments are conducted on several retrieval and motion-specific tasks."},"presentation":{"value":"2 fair"},"contribution":{"value":"2 fair"},"soundness":{"value":"1 poor"},"strengths":{"value":"+ The proposed adaptive reweighted contrastive loss addresses the noisy data issue in audio-visual self-supervised learning.\n\n+ The experiments show that the proposed method achieves superior performance over compared approaches on several tasks, including Audio-visual retrieval, audio-base video grounding, and Lip-speech retrieval."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The contributions of the work are overclaimed, and some very important relevant works are missing. A lot of claims and statements are clearly wrong. \n\n1. In the paper, the authors claimed that \n\n(a) \"Previous methods mainly model the static “object” information from a few frames while lacking the ability to capture and align the important temporal “motion” information\";\n\n(b) \"Although these methods achieve promising performance on different audio-visual downstream tasks, they either lack modeling of visual motion information or explicit learning of motion audio alignment, which significantly constrains their upper bound in capturing audio-visual correlations\";\n\n(c) \"Previous audio-visual representation methods mainly follow a video-level contrastive loss, which pulls features from the same video close while pushing features of different videos away. To acquire a more precise understanding of the correlation between motion and audio, we further propose cliplevel contrastive loss, clips in different videos and the same video are both used for calculating contrastive loss.\" \n\nHowever, self-supervised audio-visual learning using temporal synchronization beyond semantic correspondence matching is not new. In 2018, Andrew and Alexei [1] proposed a method to learn audio and visual representations by training a neural network to predict whether video frames and audio are temporally aligned. They paired temporally shifted audio with visual frames from the same video to build negative samples. Concurrently, Korbar et al. [2] adopted contrastive learning for audio-visual self-supervised learning, using both easy and hard negative samples. \"Easy negatives\" are pairs where the video frames and audio come from two different videos. \"Hard negatives\" are pairs taken from the same video, but with at least a half-second time gap between the audio sample and the visual clip. These two pioneering works are well-known in the field, and the first work is probably the most cited self-supervised audio-visual learning work. I was surprised that the authors ignored these works and made these invalid statements. \n\n[1] Owens, Andrew, and Alexei A. Efros. \"Audio-visual scene analysis with self-supervised multisensory features.\" Proceedings of the European conference on computer vision (ECCV). 2018.\n\n[2] Korbar, Bruno, Du Tran, and Lorenzo Torresani. \"Cooperative learning of audio and video models from self-supervised synchronization.\" Advances in Neural Information Processing Systems 31 (2018).\n\n2. About the audio-visual noisy data issue, the authors claimed that \n\n(a) \"Despite their impressive performance, two key issues still limit the further development of audio-visual representations: ... 2) The unlabeled web-scale video data are noisy. In a video, the visual information is limited in the camera perspective, while the audio can originate from all directions. Consequently, not all visual objects make sounds, and not all sound sources are visible in the video.\nThis unavoidable noisy data compromises the quality of the learned representations.\"\n\n(b) \"the web-scale video data for training is very noisy. However, previous audio-visual representation methods lack the analysis and design to alleviate the adverse effect of the noisy data.\" \n\nMorgado et al. [3] have already explored the issue of false positive noisy data. They proposed a weighted contrastive learning loss to down-weigh the contribution of false positives to the overall loss. The main idea of the proposed method is very similar to what they did.\nThe authors have also ignored this work, which clearly contradicts their claim of novelty.\n\n[3] Morgado, Pedro, Ishan Misra, and Nuno Vasconcelos. \"Robust audio-visual instance discrimination.\" Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2021.\n\nConsidering these facts, I do not think the work can be accepted.\n\nI wonder why the authors did not evaluate their method on video action recognition (see [3]) and audio-visual classification (see CAV-MAE paper), which are standard evaluation tasks for audio-visual learning."},"confidence":{"value":"1: You are unable to assess this paper and have alerted the ACs to seek an opinion from different reviewers."},"questions":{"value":"Please see the Weaknesses."},"rating":{"value":"1: strong reject"},"details_of_ethics_concerns":{"value":"N/A"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636321940,"tcdate":1699073863424,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission3658/Reviewer_5gkS"],"signatures":["ICLR.cc/2024/Conference/Submission3658/Reviewer_5gkS"],"forum":"mOFACpjXe2","number":3,"license":"CC BY 4.0","cdate":1699073863424,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission3658/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636321940,"domain":"ICLR.cc/2024/Conference","replyto":"mOFACpjXe2","id":"KysdOn885G","forumContent":{"venue":{"value":"ICLR 2024 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["audio-visual representation learning"]},"primary_area":{"value":"representation learning for computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Audio-visual correlation learning has many applications and is pivotal in broader multimodal understanding and generation. Recently, many existing methods try to learn audio-visual contrastive representations from web-scale videos and show impressive performance. However, these methods mainly focus on learning the correlation between audio and static visual information (such as objects and background) while ignoring the crucial role of motion information in determining sounds in videos. Besides, the widespread presence of false and multiple positive audio-visual pairs in web-scale unlabeled videos also limits the performance of audio-visual representations. In this paper, we propose \\textbf{Li}sten to \\textbf{Mo}tion (LiMo) to capture motion information explicitly and align motion and audio robustly. Specifically, for modeling the motion in video, we extract the temporal visual semantic by facilitating the interaction between frames, while retaining static visual-audio correlation knowledge acquired in previous models. To prompt a more robust audio-visual alignment, we propose learning motion-audio alignment more specifically by distinguishing different clips within the same video. And we quantitatively measure the likelihood of each sample being false positive or containing multiple positive instances, then adaptively reweight samples in the final learning objective. Our extensive experiments demonstrate the effectiveness of LiMo on various audio-visual downstream tasks. On audio-visual retrieval, LiMo achieves absolute improvements of at least 15\\% top1 accuracy on AudioSet and VGGSound. On our newly proposed motion-specific tasks, LiMo exhibits much better performance. Moreover, LiMo also achieves advanced accuracy on audio event recognition, demonstrating enhanced discriminability of audio representations."},"_bibtex":{"value":"@misc{\nwang2024listen,\ntitle={Listen to Motion: Robustly Learning Correlated Audio-Visual Representations},\nauthor={Zehan Wang and Xize Cheng and Li Tang and Luping Liu and Yang Zhao and Tao Jin and Chengfei Cai and WANG HongFa and Wei Liu and Zhou Zhao},\nyear={2024},\nurl={https://openreview.net/forum?id=mOFACpjXe2}\n}"},"title":{"value":"Listen to Motion: Robustly Learning Correlated Audio-Visual Representations"},"pdf":{"value":"/pdf/c2d570ec1cd2295d627e36b1e3024bf06b17d7e3.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Withdrawn_Submission"},"paperhash":{"value":"wang|listen_to_motion_robustly_learning_correlated_audiovisual_representations"},"authorids":{"value":["~Zehan_Wang2","~Xize_Cheng1","~Li_Tang3","~Luping_Liu2","~Yang_Zhao14","~Tao_Jin2","~Chengfei_Cai1","~WANG_HongFa1","~Wei_Liu3","~Zhou_Zhao3"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Zehan Wang","Xize Cheng","Li Tang","Luping Liu","Yang Zhao","Tao Jin","Chengfei Cai","WANG HongFa","Wei Liu","Zhou Zhao"]}},"version":2},{"content":{"summary":{"value":"This paper presents PanoWorld-X, a framework for generating controllable 360-degree panoramic videos with trajectory-guided exploration. The main contributions are: (1) PanoExplorer dataset with 116,759 panoramic video-trajectory pairs generated via Unreal Engine, (2) Explorable Sphere-Aware DiT architecture featuring Exploration-Aware Attention for trajectory control and Sphere-Aware Attention that leverages spherical geometry to improve consistency, and (3) experiments showing improvements over existing methods in visual quality, motion range, and control precision. The work enables immersive world generation for VR, embodied AI, and autonomous driving applications."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. How to define the yaw, pitch, and roll angles (Line 252)?\n2. Can the author give more detailed description regarding the scale normalization strategy? (Line 263)\n3. Can the author compare PanoWorld-X with other baseline methods using the datasets comes from other sources?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The PanoExplorer dataset addresses a critical gap in panoramic video research with 116,759 high-quality samples featuring diverse scenes (10 categories), extensive motion patterns, and precise trajectory annotations. The data curation pipeline is well-designed with collision detection and multi-stage quality filtering.\n2. The sphere-aware attention mechanism is theoretically well-motivated, properly accounting for the spherical topology of panoramic data. \n3.  The paper provides extensive comparisons with both panoramic generation methods (360DVD, Imagine360, GenEX) and camera-controllable methods (CameraCtrl, AC3D), demonstrating clear improvements across multiple metrics. Ablation studies systematically validate each component."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The entire dataset is synthetic from Unreal Engine. While Section D.1 shows some zero-shot generalization to in-the-wild images, more rigorous evaluation on real panoramic videos would strengthen the claims about practical applicability. The domain gap could limit real-world deployment.\n2. There lacks some analyze regarding the proposed datasets. For example, this paper does not give the camera trajectory distribution (move forward, move forward and turn left / right, etc), which is important to understand the proposed datasets. And keeping a balanced camera trajectory is important for learning a model with good trajectory generalization ability.\n3. The comparison between PanoWorld-X and other methods is unfair, since other methods are not trained using the proposed datasets, while the evaluation dataset comes from the proposed dataset."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1764358737012,"tcdate":1761706296638,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission11906/Reviewer_kKFp"],"signatures":["ICLR.cc/2026/Conference/Submission11906/Reviewer_kKFp"],"forum":"iZyBEbq6jR","number":1,"license":"CC BY 4.0","cdate":1761706296638,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission11906/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1764358737012,"domain":"ICLR.cc/2026/Conference","replyto":"iZyBEbq6jR","id":"dk5o08mpv7","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["immersive content"]},"supplementary_material":{"value":"/attachment/5773686dc6b17b2632d3dc874878b8b3fbb061b7.zip"},"primary_area":{"value":"other topics in machine learning (i.e., none of the above)"},"abstract":{"value":"Generating a complete and explorable 360-degree visual world enables a wide range of downstream applications. While prior works have advanced the field, they remain constrained by either narrow field-of-view limitations, which hinder the synthesis of continuous and holistic scenes, or insufficient controllability that restricts free exploration by users or autonomous agents. To address this, we propose PanoWorld-X, a novel framework for high-fidelity and controllable panoramic video generation with diverse camera trajectories. \n First, we propose a novel pipeline for synthesizing panoramic video-trajectory dataset pairs in virtual 3D environments via Unreal Engine. This pipeline consists of four main steps and enables the collection of a large-scale dataset with rich scene diversity and accurate trajectory annotations.\n To achieve precise panoramic video generation, we identify that the bottleneck arises from the misalignment between the spherical geometry of panoramic data and the inductive priors of conventional video diffusion models. To address this, we leverage the spherical connectivity characteristics of panorama data, and propose a Sphere-Aware Diffusion Transformer that reprojects equirectangular features onto the spherical surface, thereby capturing geometric adjacency in the latent space. This design significantly improves both visual fidelity and spatiotemporal continuity.\n   Extensive experiments demonstrate that our PanoWorld-X achieves superior performance in various aspects, including motion range, control precision, and visual quality, underscoring its potential for real-world applications."},"_bibtex":{"value":"@misc{\nyin2026panoworldx,\ntitle={PanoWorld-X: Generating Explorable Panoramic Worlds via Sphere-Aware Video Diffusion},\nauthor={Yuyang Yin and Hao-Xiang Guo and Fangfu Liu and Mengyu Wang and Hanwen Liang and Eric Li and Yikai Wang and Xiaojie Jin and Yao Zhao and Yunchao Wei},\nyear={2026},\nurl={https://openreview.net/forum?id=iZyBEbq6jR}\n}"},"title":{"value":"PanoWorld-X: Generating Explorable Panoramic Worlds via Sphere-Aware Video Diffusion"},"pdf":{"value":"/pdf/8d2464ff9acb82ee4bdbbfda2c8b3031ed0755c7.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"yin|panoworldx_generating_explorable_panoramic_worlds_via_sphereaware_video_diffusion"},"authorids":{"value":["~Yuyang_Yin1","~Hao-Xiang_Guo1","~Fangfu_Liu2","~Mengyu_Wang5","~Hanwen_Liang2","~Eric_Li8","~Yikai_Wang2","~Xiaojie_Jin1","~Yao_Zhao1","~Yunchao_Wei1"]},"authors":{"value":["Yuyang Yin","Hao-Xiang Guo","Fangfu Liu","Mengyu Wang","Hanwen Liang","Eric Li","Yikai Wang","Xiaojie Jin","Yao Zhao","Yunchao Wei"]}},"version":2},{"content":{"summary":{"value":"This paper propose to detect shortcut by evaluating the mutual information between the latent variable Z of a network and the input X. The intuition is that the latent variable Z of a network learning shortcuts would have a lower mutual information with the input X. Based on the intuition, the paper propose to detect shortcut by calculating mutual information using neural tangent kernel (NTK)."},"presentation":{"value":"2 fair"},"contribution":{"value":"2 fair"},"soundness":{"value":"2 fair"},"strengths":{"value":"1). Using mutual information to detect whether networks is an interesting topic.\n\n2). This paper provide a detailed related work introduction."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The proposed method seems infeasible.\n\n1). In algorithm 1, a dataset is said to contain shortcut if $I(X_{test};Z)<I(X_{tr};Z)$. However, since $X_{test}$ have different distribution than $X_{tr}$, it seems natural to have $I(X_{test};Z)<I(X_{tr};Z)$. Therefore an important question need to be answered: are there any datasets satisfying $I(X_{test};Z)\\geq I(X_{tr};Z)$?  There lacks empirical evidences in this paper to answer the question and I do not think the algorithm 1 could effectively detect shortcuts.\n\n2). Unlike the proposed algorithm 1, the experiments in this paper, on the other hand, mainly compare networks trained on two datasets \"with\" and \"without\" shortcut. This approach also have problems since it requires comparing with a network trained on dataset that is \"without\" shortcut. By defining \"without shortcut\", it also involves domain knowledge and human expertise to detect shortcuts.\n\nBased on the above two points, I think that the propose method have major flaws.\n\nMinor issues:\n\n1). the introduction of estimating mutual information using NTK is vague. For example, what is the definition of $\\Theta(x,X)$? $\\sigma$ is used in Eq.5 but it is written as $\\Sigma$ in Eq.7.\n\n2). Figure legends (e.g. Fig.4) could be more detailed."},"confidence":{"value":"4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work."},"questions":{"value":"As mentioned in the weakness section, I have concerns over the propose method.\n\n1). For those datasets without shortcut, are they satisfy $I(X_{test};Z)=I(X_{tr};Z)$? Please provide empirical evidences to show that algorithm 1 is feasible.\n\n2). I notice that at the begining of the training, before the mutual information $I(X;Z)$ starts to be different between \"with shortcut\" and \"without shortcut\", the test loss has become different (e.g. 100 epoch in Fig.6 and 1000 epoch in Fig.7).  Why is it?"},"rating":{"value":"3: reject, not good enough"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636989670,"tcdate":1698570681044,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission8022/Reviewer_XNRX"],"signatures":["ICLR.cc/2024/Conference/Submission8022/Reviewer_XNRX"],"forum":"hr4HTShC6l","number":1,"license":"CC BY 4.0","cdate":1698570681044,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission8022/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636989670,"domain":"ICLR.cc/2024/Conference","replyto":"hr4HTShC6l","id":"hTCaReNxqe","forumContent":{"TLDR":{"value":"Proposed a mutual-information based method to detect shortcuts/spurious correlations."},"venue":{"value":"Submitted to ICLR 2024"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["shortcuts","spurious correlation","mutual information","information theory","neural tangent kernel"]},"supplementary_material":{"value":"/attachment/51f14e81aa32b370b97ad25ffc7becbdabd22c25.pdf"},"primary_area":{"value":"societal considerations including fairness, safety, privacy"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"The failure of deep neural networks to generalize to out-of-distribution (OOD) data is a well-known problem that raises concerns about the deployment of trained networks in safety-critical domains such as healthcare and autonomous vehicles. We study a particular kind of distribution shift — shortcuts or spurious correlations in the training data. These correlations are not present in real-world test data, so there\nis a performance drop due to distribution shift, also referred to as shortcut learning. Shortcut learning is often only exposed when models are evaluated in carefully controlled experimental settings, posing a serious dilemma for AI practitioners to properly assess the effectiveness of a trained model for real-world applications. In this work, we try to understand shortcut learning using information-theoretic tools and propose to use the mutual information (MI) between the learned representation and the input space as a domain-agnostic metric for detecting shortcuts in the training datasets. For studying the training dynamics of shortcut learning, we develop a Neural Tangent Kernel (NTK) based framework, which can be used to detect shortcuts and spurious correlations in the training data without requiring class labels\nof the test data. We empirically demonstrate on multiple datasets, such as MNIST, CelebA, NICO, Waterbirds, and BenchMD, that MI can effectively detect shortcuts. We benchmark against multiple OOD detection baselines to show that OOD detectors cannot detect shortcuts, and our method can be used in complementary with OOD detectors to identify all types of distribution shifts in the datasets, including\nshortcuts."},"_bibtex":{"value":"@misc{\nadnan2024detecting,\ntitle={Detecting Shortcuts using Mutual Information},\nauthor={Mohammed Adnan and Yani Ioannou and Kenyon Tsai and Angus Galloway and Hamid Tizhoosh and Rahul G Krishnan and Graham W. Taylor},\nyear={2024},\nurl={https://openreview.net/forum?id=hr4HTShC6l}\n}"},"title":{"value":"Detecting Shortcuts using Mutual Information"},"pdf":{"value":"/pdf/ca51045eef9d7fcc21c7c83eaf9e2cd5f0cd86f2.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Rejected_Submission"},"paperhash":{"value":"adnan|detecting_shortcuts_using_mutual_information"},"authorids":{"value":["~Mohammed_Adnan1","~Yani_Ioannou1","~Kenyon_Tsai1","~Angus_Galloway1","~Hamid_Tizhoosh1","~Rahul_G_Krishnan1","~Graham_W._Taylor1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Mohammed Adnan","Yani Ioannou","Kenyon Tsai","Angus Galloway","Hamid Tizhoosh","Rahul G Krishnan","Graham W. Taylor"]}},"version":2},{"content":{"summary":{"value":"This paper addresses the issue of noisy pairs in vision-language pre-training by proposing a weighted contrastive loss that estimates the noisiness of data pairs using Bayesian techniques. The proposed method is evaluated on the CLIP model and demonstrates reasonable improvements compared to vanilla CLIP."},"presentation":{"value":"2 fair"},"contribution":{"value":"2 fair"},"soundness":{"value":"2 fair"},"strengths":{"value":"1. The problem of addressing noisy pairs in vision-language pre-training has practical implications and is valuable.\n2. The proposed method takes into account both the false positive and false negative problems in contrastive learning."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. There is a concern regarding the originality of the problem studied in this paper. The claim of being the first to consider the noise problem in contrastive learning is incorrect, as several prior works have already addressed this issue.: Robust Cross-Modal Representation Learning with Progressive Self-Distillation, Ranking info noise contrastive estimation: Boosting contrastive learning via ranked positives, Cross-Modal Retrieval with Partially Mismatched Pairs and Debiased Contrastive Learning, to name a few.\n\n2. Given the focus on noisy pairs, the related works section should provide a comprehensive discussion covering noisy labels, noisy correspondence, and robust vision-language pre-training. Moreover, the approach of incorporating soft weights into contrastive learning is not novel in either the noisy label or vision-language learning community.\n\n3. The proposed method does not show notable improvements compared to vanilla CLIP, particularly according to Table 2.\n\n4. The estimation of the confidence of noisy pairs through the joint posterior distribution over the global model parameter and local random weight (Eq. 2) is still unclear. It lacks sufficient details and intuitive explanation, making it difficult to understand how the confidence of noisy pairs is estimated."},"confidence":{"value":"5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully."},"questions":{"value":"In the implementation details, the authors manually simulate around 10% noisy pairs. The rationale behind this choice is unclear. It would be more informative to evaluate the method on naturally occurring partially-matched pairs, such as the 3% to 22% range found in CC. Additionally, exploring performance on even noisier datasets with 30% or 50% noisy pairs would provide a more challenging scenario for evaluation if the simulate noisy pair is needed."},"rating":{"value":"3: reject, not good enough"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636227025,"tcdate":1698806417168,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission2836/Reviewer_g84h"],"signatures":["ICLR.cc/2024/Conference/Submission2836/Reviewer_g84h"],"forum":"CvxcWCDX0h","number":4,"license":"CC BY 4.0","cdate":1698806417168,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission2836/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636227025,"domain":"ICLR.cc/2024/Conference","replyto":"CvxcWCDX0h","id":"jzwGtifte3","forumContent":{"venue":{"value":"ICLR 2024 Conference Withdrawn Submission"},"TLDR":{"value":"we propose a novel solution by reformulating the standard CL into a probability framework, and introducing learnable random weights to associate with data pairs, so as to allow automatic inference of the degree of noisiness for each data pair."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["multi-modal learning","contrastive learning","foundation models"]},"primary_area":{"value":"representation learning for computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Contrastive learning~(CL) represents one of the most successful paradigms for self-supervised representation learning, which has been applied to SOTA multi-modal learning applications. One overlooked limitation of standard contrastive learning, however, is that it is not designed for robust learning in the presence of noisy data pairs. For example, not all negative samples are truly negative, {\\it e.g.}, within a mini-batch there can be negative samples that are semantically as positive as the positive sample. This is common in most web-sourced multi-modal datasets such as CC3M and YFCC that are frequently used for CL, due to the noisy nature when crawling the datasets. Consequently, the noise in the datasets could significantly impair the power of CL. To remedy this issue, we propose a novel solution by reformulating the standard CL into a probability framework, and introducing learnable random weights to associate with data pairs, so as to allow automatic inference of the degree of noisiness for each data pair. Within our probability framework, posterior inference of the random weights can be done efficiently with Bayesian data augmentation. Consequently, the model can be effectively optimized by a novel learning algorithm based on stochastic expectation maximization. We demonstrate the effectiveness of our approach on several standard multi-modal contrastive learning benchmarks, which significantly outperforms standard contrastive learning."},"_bibtex":{"value":"@misc{\njiang2024learning,\ntitle={Learning Multi-Modal Representation Alignments from Noisy Data-Pairs},\nauthor={Qian Jiang and Jingjing Meng and Alireza Bagheri Garakani and Yang Jiao and Yetian Chen and Yikai Ni and Yan Gao and Yi Sun and Changyou Chen},\nyear={2024},\nurl={https://openreview.net/forum?id=CvxcWCDX0h}\n}"},"title":{"value":"Learning Multi-Modal Representation Alignments from Noisy Data-Pairs"},"pdf":{"value":"/pdf/5b5dc724edb76920a4b1bd79e791974ea61f81aa.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Withdrawn_Submission"},"paperhash":{"value":"jiang|learning_multimodal_representation_alignments_from_noisy_datapairs"},"authorids":{"value":["~Qian_Jiang1","~Jingjing_Meng1","~Alireza_Bagheri_Garakani1","~Yang_Jiao6","~Yetian_Chen1","yika@amazon.com","~Yan_Gao8","~Yi_Sun13","~Changyou_Chen1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Qian Jiang","Jingjing Meng","Alireza Bagheri Garakani","Yang Jiao","Yetian Chen","Yikai Ni","Yan Gao","Yi Sun","Changyou Chen"]}},"version":2},{"content":{"venue":{"value":"ICML 2026 regular"},"keywords":{"value":["AI-Generated Video Detection","AI Security."]},"_bibtex":{"value":"@inproceedings{\nfang2026real,\ntitle={Real Data Lies: Unveiling and Closing the Quality Shortcut in Generalizable {AI}-Generated Video Detection},\nauthor={Ziyuan Fang and Tianyi Wei and Guanjie Wang and Weiming Zhang and Nenghai Yu and Wenbo Zhou},\nbooktitle={Forty-third International Conference on Machine Learning},\nyear={2026},\nurl={https://openreview.net/forum?id=sYzf4xhthY}\n}"},"title":{"value":"Real Data Lies: Unveiling and Closing the Quality Shortcut in Generalizable AI-Generated Video Detection"},"paperhash":{"value":"fang|real_data_lies_unveiling_and_closing_the_quality_shortcut_in_generalizable_aigenerated_video_detection"},"originally_submitted_PDF":{"value":"/pdf/72a4202ef78ece911d9b22d66d02fd970d8c7311.pdf"},"primary_area":{"value":"social_aspects->security"},"abstract":{"value":"Recent advances in video generation have enabled highly realistic synthetic content, raising concerns about the integrity of digital media and motivating the development of benchmarks and detection methods for generated videos. Prior works have largely prioritized bolstering model generalization against unseen generators. However, we uncover a neglected factor: the quality distribution of real videos plays a pivotal role. Current training protocols suffer from a clear quality bias between real and fake data, prone to shortcut learning. Compounded by testing on similar real data distributions, this creates an illusion of generalization. In reality, these models fail to generalize when exposed to real data with significantly different quality profiles. To address this, we propose training with quality-matched real and fake data to mitigate bias. Building on this, we introduce a data expansion strategy that broadens the training set to comprehensively cover the full quality spectrum. This approach enables the model to learn quality-agnostic features for detection, thereby achieving generalization across real data of varying qualities and enhancing real-world applicability. Extensive experiments demonstrate that our method scales well across diverse backbones, consistently enhancing the generalization capability of existing models."},"link_to_code":{"value":"https://github.com/Umbreller-F/Real-Data-Lies"},"pdf":{"value":"/pdf/f94a59d00aec1c860d09f5f5fa84cdc51b6c807e.pdf"},"lay_summary":{"value":"AI-generated videos are becoming incredibly realistic, making it hard to tell what is real and what is fake. Researchers have built detection tools that claim to spot these fake videos even when faced with new types of AI generators they have never seen before. However, our study reveals a hidden flaw in how these tools are trained and tested.\n\nWe found that current detectors often cheat by relying on video quality. During training, the real videos used are typically lower in quality than the AI-generated ones. As a result, the detectors learn a simple but wrong rule: \"high quality means fake.\" This works when testing on similar low-quality real videos, but fails dramatically when encountering high-quality real footage, such as professional or ultra-HD videos.\n\nTo fix this, we propose matching the quality of real and fake training videos so the model cannot use quality as a shortcut. We also expand training data to cover the full range of video quality found in the real world, from blurry clips to crystal-clear footage. This forces the detector to learn genuine signs of forgery rather than simple quality cues.\n\nOur experiments show this approach significantly improves reliability across different video qualities. We also highlight that current benchmark tests are incomplete because they only use narrow quality ranges, and we suggest a more comprehensive evaluation protocol to better reflect real-world conditions."},"venueid":{"value":"ICML.cc/2026/Conference"},"authorids":{"value":["~Ziyuan_Fang1","~Tianyi_Wei1","~Guanjie_Wang2","~Weiming_Zhang2","~Nenghai_Yu1","~Wenbo_Zhou1"]},"authors":{"value":["Ziyuan Fang","Tianyi Wei","Guanjie Wang","Weiming Zhang","Nenghai Yu","Wenbo Zhou"]}},"tmdate":1790067330765,"pdate":1777576335789,"tcdate":1768909303635,"writers":["ICML.cc/2026/Conference","ICML.cc/2026/Conference/Submission9683/Authors"],"signatures":["ICML.cc/2026/Conference/Submission9683/Authors"],"forum":"sYzf4xhthY","license":"CC BY-NC 4.0","number":9683,"cdate":1768909303635,"readers":["everyone"],"invitations":["ICML.cc/2026/Conference/-/Submission","ICML.cc/2026/Conference/-/Post_Submission","ICML.cc/2026/Conference/Submission9683/-/Full_Submission","ICML.cc/2026/Conference/Submission9683/-/Reciprocal_Reviewing_Correction","ICML.cc/2026/Conference/-/Edit","ICML.cc/2026/Conference/Submission9683/-/Camera_Ready_Revision"],"mdate":1790067330765,"odate":1782341935111,"domain":"ICML.cc/2026/Conference","id":"sYzf4xhthY","version":2},{"content":{"summary":{"value":"This paper proposes Video-GPT, a concise large video foundation model that unifies autoregressive modeling and diffusion via a novel \"next clip diffusion\" paradigm. Inspired by GPT’s next token prediction, the model treats video clips as \"visual words\" to model spatial-temporal details in the visual world—addressing the limitation of language sequences in capturing such details. The key design involves constructing interleaved sequences of noisy and clean clips, using hierarchical attention masking to leverage historical clean clips for denoising future noisy clips."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. Could you clarify the core differences between Video-GPT’s \"next clip diffusion\" and prior hybrid diffusion-autoregressive models? Specifically, how does your clip-level autoregressive design and hierarchical masking outperform their frame-level or pixel-level combinations?\n2. The paper compares the performance of Video GPT with other models in video prediction, achieving SOTA in Physics-IQ. However, it doesn't examine its performance in other aspects that need to be considered in video generation, such as motion quality and subject consistency.\n3. How does Video-GPT perform on longer videos (e.g., 1 minute or more) in terms of temporal coherence and content consistency? Have you tested it against long-video generation models (e.g., Flexifilm, Open-Sora-Plan) and observed any performance degradation?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The \"next clip diffusion\" paradigm is a creative combination of autoregressive modeling (from GPT) and diffusion (for high-quality generation). Treating clips as visual words and using historical clean clips as context for denoising is a novel adaptation of language modeling to video, filling the gap between discrete text tokens and continuous video data. This hybrid design effectively unifies short-term generation and long-term prediction.\n2. As a unified video foundation model, Video-GPT bridges video generation and understanding tasks, advancing the goal of visual world modeling. Its strong performance on physics-aware prediction (Physics-IQ) indicates progress in learning world knowledge from video.\n3. Its generalization to six downstream tasks highlights its potential as a backbone for diverse video applications."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Insufficient comparison with hybrid baselines: The paper mentions prior works that combine diffusion and autoregressive modeling but lacks a detailed comparison of their core differences. \n2. There is a lack of comparisons with some newer autoregressive + diffusion video generation models, such as self-forcing, apt2.\n3. Limited analysis of architectural choices: The model inherits Phi-3-mini’s architecture and SDXL’s VAE without justifying these choices. There is no comparison with other architectures (e.g., DiT, U-Net) or VAEs (e.g., 3D VAE vs. 2D VAE) to show whether these selections are critical to performance. Additionally, the progressive training strategy’s effectiveness is only validated via frame count ablation, without analyzing how frame interval or clip number affects convergence."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762915836265,"tcdate":1761980426099,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission1629/Reviewer_vBbq"],"signatures":["ICLR.cc/2026/Conference/Submission1629/Reviewer_vBbq"],"forum":"E0ZAcqy9TB","number":3,"license":"CC BY 4.0","cdate":1761980426099,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission1629/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762915836265,"domain":"ICLR.cc/2026/Conference","replyto":"E0ZAcqy9TB","id":"8qGmDsDzw6","forumContent":{"TLDR":{"value":"Video-GPT treats video as new language for visual world modeling."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Video; Diffusion; LLM"]},"primary_area":{"value":"foundation or frontier models, including LLMs"},"abstract":{"value":"GPT has shown its remarkable success in natural language processing. However, the language sequence is not sufficient to describe spatial-temporal details in the visual world. Alternatively, the video sequence is good at capturing such details. Motivated by this fact, we propose a concise Video-GPT in this paper by treating video as new language for visual world modeling. By analogy to next token prediction in GPT, we introduce a novel next clip diffusion paradigm for pretraining Video-GPT. Different from the previous works, this distinct paradigm allows Video-GPT to tackle both short-term generation and long-term prediction, by autoregressively denoising the noisy clip according to the clean clips in the history. Extensive experiments show our Video-GPT achieves the state-of-the-art performance on video prediction, which is the key factor towards world modeling (Physics-IQ Benchmark: Video-GPT 34.97 vs. Kling 23.64 vs. Wan 20.89). Moreover, it can be well adapted on 6 mainstream video tasks in both video generation and understanding, showing its great generalization capacity in downstream."},"_bibtex":{"value":"@inproceedings{\nzhuang2026videogpt,\ntitle={Video-{GPT} via Next Clip Diffusion},\nauthor={Shaobin Zhuang and Zhipeng Huang and Ying Zhang and Fangyikang Wang and Canmiao Fu and Binxin Yang and Chong Sun and Chen Li and Yali Wang},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=E0ZAcqy9TB}\n}"},"title":{"value":"Video-GPT via Next Clip Diffusion"},"pdf":{"value":"/pdf/717aa186ca36c0bc98678fe5f11097effe05f8bf.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"zhuang|videogpt_via_next_clip_diffusion"},"authorids":{"value":["~Shaobin_Zhuang1","~Zhipeng_Huang3","~Ying_Zhang9","~Fangyikang_Wang1","~Canmiao_Fu1","~Binxin_Yang1","~Chong_Sun1","~Chen_Li11","~Yali_Wang1"]},"authors":{"value":["Shaobin Zhuang","Zhipeng Huang","Ying Zhang","Fangyikang Wang","Canmiao Fu","Binxin Yang","Chong Sun","Chen Li","Yali Wang"]}},"version":2},{"content":{"venue":{"value":"CoRR 2025"},"pdf":{"value":"https://arxiv.org/pdf/2503.18950v2"},"venueid":{"value":"dblp.org/journals/CORR/2025"},"paperhash":{"value":"kim|targetaware_video_diffusion_models"},"authorids":{"value":["~Taeksoo_Kim2","https://dblp.org/search/pid/api?q=author:Hanbyul_Joo:"]},"html":{"value":"https://doi.org/10.48550/arXiv.2503.18950"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2503-18950,\n  publtype={informal},\n  author={Taeksoo Kim and Hanbyul Joo},\n  title={Target-Aware Video Diffusion Models},\n  year={2025},\n  month={March},\n  cdate={1740787200000},\n  journal={CoRR},\n  volume={abs/2503.18950},\n  url={https://doi.org/10.48550/arXiv.2503.18950}\n}\n"},"abstract":{"value":"We present a target-aware video diffusion model that generates videos from an input image in which an actor interacts with a specified target while performing a desired action. The target is defined by a segmentation mask and the desired action is described via a text prompt. Unlike existing controllable image-to-video diffusion models that often rely on dense structural or motion cues to guide the actor's movements toward the target, our target-aware model requires only a simple mask to indicate the target, leveraging the generalization capabilities of pretrained models to produce plausible actions. This makes our method particularly effective for human-object interaction (HOI) scenarios, where providing precise action guidance is challenging, and further enables the use of video diffusion models for high-level action planning in applications such as robotics. We build our target-aware model by extending a baseline model to incorporate the target mask as an additional input. To enforce target awareness, we introduce a special token that encodes the target's spatial information within the text prompt. We then fine-tune the model with our curated dataset using a novel cross-attention loss that aligns the cross-attention maps associated with this token with the input target mask. To further improve performance, we selectively apply this loss to the most semantically relevant transformer blocks and attention regions. Experimental results show that our target-aware model outperforms existing solutions in generating videos where actors interact accurately with the specified targets. We further demonstrate its efficacy in two downstream applications: video content creation and zero-shot 3D HOI motion synthesis."},"title":{"value":"Target-Aware Video Diffusion Models"},"authors":{"value":["Taeksoo Kim","Hanbyul Joo"]}},"tmdate":1769241358663,"pdate":1767139200000,"externalIds":["dblp:journals/corr/abs-2503-18950"],"tcdate":1769241356494,"writers":["~"],"signatures":["~Taeksoo_Kim2"],"forum":"NpuPTjWzxc","license":"CC BY-SA 4.0","number":793917,"cdate":1740787200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1769241358663,"domain":"DBLP.org","id":"NpuPTjWzxc","version":2},{"content":{"summary":{"value":"This paper proposes a new concept called “align before adapt” in contrast to the previous “adapt then align” when a pre-trained image encoder is extended for video.\nWhen using VLP, prior works usually construct a video embedding from video frames (this is the “adapt” from image to video), then match it with a text embedding (this is “align”).\nIn contrast, this work first “aligns” frame features with text, then aggregates them to a video-level feature.\n\n- The visual encoder shown in Fig.3(a) is a ViT with ToMe to produce a frame feature X_t for frame t.\n- Align: Then X_t is used as a query to encode the semantic embedding Q_t from a text corpus (Fig.3(b)).\n- Adapt: These frame features X_1, …, X_T are 1D convolved and attended by Q_1, …, Q_T to produce a video-level embedding z (Fig.3(c)).\n\nExperiments of fully supervised setting demonstrate that the proposed method works better than or comparable to other state-of-the-art models, with much fewer computation cost.\nFor zero-shot and few-shot, the proposed method outperform others."},"presentation":{"value":"2 fair"},"contribution":{"value":"3 good"},"soundness":{"value":"3 good"},"strengths":{"value":"The proposed concept, “align before adapt”, is reasonable and interesting. This kind of approach is not so popular because 3D patches are usually better for video representation such as ViViT and VideoMAE. Recent CLIP-based video models uses and combines features of CLIP applied to each frame, work well for zero-shot and few-shot settings, while the main approach is first construct a video-level representation.\n\nThe approach of this work transforms CLIP features of frames first, before temporal aggregation. It might inspire other potential following works for video representation by using pre-trained image encoders."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The alignment of the “align before adapt” concept in this work heavily relies on the text corpus (p.4, sec.3.1). It seems to be versatile and reusable, however, hand-crafted and the performance may be affected by the quality of the corpus and the vocabulary size K. There are not discussions on the corpus quality, and K is not shown.\n\nThe corpus may help the model by providing cues of “objects” and “scenes”, and the appendix shows show to construct the corpus by asking “What are identifying object/body parts/scenes/roles of an action {label}?” to an LLM model. However it may boost the notorious representation bias (background bias or scene bias). If it is true, it should be discussed as a limitation of the approach using a text corpus constructed in this way.\nAlso datasets with less bias (such as SSv2 or Diving48) should be used in the experiments.\n\nThere is a trend using learnable prompts when a frozen encoder, and there are some works for action recognition, such as ViFi CLIP. This approach shows a resemblance to the proposed method in terms of transforming frame-wise features of a pre-trained image encoder before aggregation. This prompt learning approach is simpler and does not require the preparation of a hand-crafted text corpus, which could be an advantage over the proposed method.\n- (ViFi CLIP) Fine-tuned CLIP Models are Efficient Video Learners, CVPR2023\n\n\nThe proposed “region-aware” image encoder (Fig3a) uses ToMe, however the meaning of “region-aware” is not well explained and hence not clear. It is simply using ToMe, but \nin the experiments ToMe and region-aware tokens are separately used in Table 5. It should be clarify more about the token merging and the term “region-aware”.\n\n\n\n\nMinor comments:\n\n- ALT, the name of the proposed method, is not explained anywhere.\n- ToMe appears p2, but not explained (full-spelled) until p4.\n- Fig3(c), the final \\hat{Q}_M looks the text semantic embeddings because of the symbol and the connection from Q in the left of Fig3(c). But it is actually the visual embedding as \\hat{X} is used as value. It might be better to change the notation and figure."},"confidence":{"value":"4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work."},"questions":{"value":"How much does the quality of the text corpus affect the results, and how much is the representation bias in this approach ?\n\nWhat is the similarity with prompt learning approaches to modify frame-wise features before temporal aggregation ?"},"rating":{"value":"6: marginally above the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636419796,"tcdate":1698113379178,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission4446/Reviewer_X9sb"],"signatures":["ICLR.cc/2024/Conference/Submission4446/Reviewer_X9sb"],"forum":"dyHn2MAYxM","number":1,"license":"CC BY 4.0","cdate":1698113379178,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission4446/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636419796,"domain":"ICLR.cc/2024/Conference","replyto":"dyHn2MAYxM","id":"xqBRlBr9L3","forumContent":{"venue":{"value":"ICLR 2024 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Action and event understanding","Vision and language","Multimodal learning"]},"primary_area":{"value":"representation learning for computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Large-scale pre-trained visual-language models have achieved significant success in various video tasks. However, most existing methods follow an 'adapt then align' paradigm, where pre-trained image encoders are adapted to model video-level representations, which are then aligned to the semantics or one-hot labels of target actions. This paradigm overlooks the challenge of mapping from static images to complicated activity concepts. In this paper, we propose a novel and efficient 'align before adapt' paradigm. We introduce a token-merging strategy to the pre-trained image model, generating region-aware embeddings in a hierarchical manner. This enhances the visual-semantic alignment at a fine-grained level. Additionally, we align the region-aware embeddings with the text corpus of action-related entities, such as objects, body parts, primitive motions, and scenes. The embeddings of the aligned text entities serve as queries for the transformer-based video adapter, better aligning with the activity concepts in a video sequence. Our proposed framework achieves competitive performance and superior generalizability while significantly reducing computational costs. In fully-supervised scenarios, our method achieves 87.9% top-1 accuracy on Kinetics-400, using only 4947 GFLOPs. Furthermore, in 2-shot experiments, our method outperforms the previous state-of-the-art by 13.0% and 12.0% on HMDB-51 and UCF-101, respectively."},"_bibtex":{"value":"@misc{\nchen2024align,\ntitle={Align before Adapt: Efficient and Generalizable Video Action Recognition with Text Corpus},\nauthor={Yifei Chen and Dapeng Chen and Ruijin Liu and SHIXIANG TANG and Sai Zhou and Wenyuan Xue and Wei Peng},\nyear={2024},\nurl={https://openreview.net/forum?id=dyHn2MAYxM}\n}"},"title":{"value":"Align before Adapt: Efficient and Generalizable Video Action Recognition with Text Corpus"},"pdf":{"value":"/pdf/f4b07a54b515c37c6734662ca8392de174bc04ab.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Withdrawn_Submission"},"paperhash":{"value":"chen|align_before_adapt_efficient_and_generalizable_video_action_recognition_with_text_corpus"},"authorids":{"value":["~Yifei_Chen5","~Dapeng_Chen4","~Ruijin_Liu1","~SHIXIANG_TANG1","~Sai_Zhou1","~Wenyuan_Xue1","~Wei_Peng6"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Yifei Chen","Dapeng Chen","Ruijin Liu","SHIXIANG TANG","Sai Zhou","Wenyuan Xue","Wei Peng"]}},"version":2},{"content":{"summary":{"value":"This paper introduces Vinoground, a temporal counterfactual benchmark designed to investigate the dense temporal reasoning abilities of video large language models (VidLLMs). The authors propose an elaborate data curation and categorization process to ensure that the counterfactual videos and captions are derived from a natural data distribution, thus guaranteeing data quality. The evaluation results indicate that existing video understanding models struggle significantly with temporal reasoning, even for short videos (less than 10 seconds)."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. What specific Chain of Thought (COT) prompts were used in the evaluation?\n\n2. Do any time-sensitive VidLLMs (e.g., [3][4]) perform better on Vinoground?\n    \n[3] TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding\n\n[4] VTimeLLM: Empower LLM to Grasp Video Moments"},"rating":{"value":5},"details_of_ethics_concerns":{"value":"NA"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":2},"strengths":{"value":"1. The paper is well-written, with a clear presentation of the proposed method that is easy to follow.\n\n2. The benchmark design is sound, and the curated data used for evaluation is of high quality. I believe Vinoground will serve as a valuable benchmark for assessing the true performance of VidLLMs in understanding various actions, background transitions, and object transformations.\n\n3. The authors provide a thorough evaluation of current prevalent VidLLMs, revealing their poor performance on dense temporal reasoning tasks, even with short video clips."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. **Limited Novelty**: \n\nThis is my primary concern. First, the evaluation protocol closely follows that of Winoground. Second, similar concepts for constructing counterfactual caption/video pairs have been proposed in [1][2]. While I acknowledge that the data quality of Vinoground may surpass that of [1][2] due to the curated videos being sourced from natural data distributions, the overall novelty appears to be incremental.\n    \n[1] TempCompass: Do video LLMs really understand videos?\n    \n[2] Paxion: Patching Action Knowledge in Video-Language Foundation Models.\n    \n2. **Lack of Insight into Model Design**: \n\nThe evaluation results indicate that existing video understanding models are inadequate in terms of temporal reasoning, even for short videos (less than 10 seconds). However, the authors do not provide in-depth insights into model design. For example, which types of modules could better model temporal information? What training data might enhance performance on temporal reasoning tasks? Such analyses would significantly improve the quality of the paper."}},"nonreaders":[],"tmdate":1731427456149,"tcdate":1730360638887,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission1636/Reviewer_LzWR"],"signatures":["ICLR.cc/2025/Conference/Submission1636/Reviewer_LzWR"],"forum":"a1P5kh2oo8","number":3,"license":"CC BY 4.0","cdate":1730360638887,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission1636/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427456149,"domain":"ICLR.cc/2025/Conference","replyto":"a1P5kh2oo8","id":"nvlSCQ33m6","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"TLDR":{"value":"Modern SoTA LMMs still demonstrates subpar performance at temporal reasoning with our temporal counterfactual benchmark composed of natural videos."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["temporal reasoning; counterfactual reasoning; short video comprehension"]},"supplementary_material":{"value":"/attachment/fd093e5ac3c36911b68e297a2932dcf6ce1ed019.zip"},"primary_area":{"value":"datasets and benchmarks"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"There has been growing sentiment recently that modern large multimodal models (LMMs) have addressed most of the key challenges related to short video comprehension. As a result, both academia and industry are gradually shifting their attention towards the more complex challenges posed by understanding long-form videos. \nHowever, is this really the case?  Our studies indicate that LMMs still lack many fundamental reasoning capabilities even when dealing with short videos.  We introduce Vinoground, a temporal counterfactual LMM evaluation benchmark encompassing 1000 short and natural video-caption pairs. We demonstrate that existing LMMs severely struggle to distinguish temporal differences between different actions and object transformations.  For example, the best model GPT-4o only obtains $\\sim$50\\% on our text and video scores, showing a large gap compared to the human baseline of $\\sim$90\\%. All open-source multimodal models and CLIP-based models perform much worse, producing mostly random chance performance. Through this work, we shed light onto the fact that temporal reasoning in short videos is a problem yet to be fully solved. We will make our benchmark publicly available."},"_bibtex":{"value":"@misc{\nzhang2025vinoground,\ntitle={Vinoground: Scrutinizing {LMM}s over Dense Temporal Reasoning with Short Videos},\nauthor={Jianrui Zhang and Mu Cai and Yong Jae Lee},\nyear={2025},\nurl={https://openreview.net/forum?id=a1P5kh2oo8}\n}"},"title":{"value":"Vinoground: Scrutinizing LMMs over Dense Temporal Reasoning with Short Videos"},"pdf":{"value":"/pdf/2e0dd590d0c11d11d7071fa71d7eaea83ded91ac.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"zhang|vinoground_scrutinizing_lmms_over_dense_temporal_reasoning_with_short_videos"},"authorids":{"value":["~Jianrui_Zhang1","~Mu_Cai1","~Yong_Jae_Lee2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Jianrui Zhang","Mu Cai","Yong Jae Lee"]}},"version":2},{"content":{"venue":{"value":"ICME 2024"},"pdf":{"value":"https://ieeexplore.ieee.org/iel8/10685847/10687354/10687384.pdf"},"venueid":{"value":"dblp.org/conf/ICMCS/2024"},"paperhash":{"value":"pu|msn_adaptive_and_dynamic_multimodal_shortcut_network_architecture_for_latencyaware_applications"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Yifei_Pu:","https://dblp.org/search/pid/api?q=author:Chi_Wang:","https://dblp.org/search/pid/api?q=author:Xiaofeng_Hou:","https://dblp.org/search/pid/api?q=author:Cheng_Xu:","https://dblp.org/search/pid/api?q=author:Jiacheng_Liu_0001:","~Jing_Wang112","https://dblp.org/search/pid/api?q=author:Minyi_Guo:","https://dblp.org/search/pid/api?q=author:Chao_Li_0009:"]},"html":{"value":"https://doi.org/10.1109/ICME57554.2024.10687384"},"_bibtex":{"value":"@inproceedings{DBLP:conf/icmcs/PuWHX0WGL24,\n  author={Yifei Pu and Chi Wang and Xiaofeng Hou and Cheng Xu and Jiacheng Liu and Jing Wang and Minyi Guo and Chao Li},\n  title={MSN: Adaptive and Dynamic Multi-modal Shortcut Network Architecture for Latency-Aware Applications},\n  year={2024},\n  cdate={1704067200000},\n  pages={1-6},\n  url={https://doi.org/10.1109/ICME57554.2024.10687384},\n  booktitle={ICME},\n  crossref={conf/icmcs/2024}\n}\n"},"abstract":{"value":"Multi-modal neural networks have demonstrated exceptional performance by merging information across modalities, surpassing the state-of-the-art uni-modal DNNs. However, this accuracy improvement comes at the cost of increased computation, leading to higher inference latency. This defect significantly limits the practical value of multi-modal DNNs, especially for latency-aware applications. Therefore, we propose an adaptive and efficient multi-modal shortcut architecture called M2SN to reduce the execution latency with accuracy guarantees. It skips ineffective network layers to reduce computational costs as well as alleviate the overfitting problem adaptive to specific models and scenarios. The key contributions of M2SN are twofold: 1) We design and insert shortcuts into each uni-modal network to perform adaptive computing. 2) We design a navigator to dynamically choose the optimal shortcuts. Unlike previous approaches, M2SN features high generality as it does not rely on any prior knowledge. The experimental results show that M2SN can reduce 28.3% average latency while obtaining the same or higher accuracy compared with SOTA baselines."},"title":{"value":"MSN: Adaptive and Dynamic Multi-modal Shortcut Network Architecture for Latency-Aware Applications"},"authors":{"value":["Yifei Pu","Chi Wang","Xiaofeng Hou","Cheng Xu","Jiacheng Liu","Jing Wang","Minyi Guo","Chao Li"]}},"tmdate":1768024458949,"pdate":1704067200000,"externalIds":["dblp:conf/icmcs/PuWHX0WGL24"],"tcdate":1768024449634,"writers":["~"],"signatures":["~Jing_Wang112"],"forum":"K6hu6Ltgg1","license":"CC BY-SA 4.0","number":742220,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1768024458949,"domain":"DBLP.org","id":"K6hu6Ltgg1","version":2},{"content":{"venue":{"value":"ICME 2024"},"pdf":{"value":"https://ieeexplore.ieee.org/iel8/10685847/10687354/10687384.pdf"},"venueid":{"value":"dblp.org/conf/ICMCS/2024"},"paperhash":{"value":"pu|msn_adaptive_and_dynamic_multimodal_shortcut_network_architecture_for_latencyaware_applications"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Yifei_Pu:","https://dblp.org/search/pid/api?q=author:Chi_Wang:","https://dblp.org/search/pid/api?q=author:Xiaofeng_Hou:","https://dblp.org/search/pid/api?q=author:Cheng_Xu:","https://dblp.org/search/pid/api?q=author:Jiacheng_Liu_0001:","https://dblp.org/search/pid/api?q=author:Jing_Wang_0055:","~Minyi_Guo1","https://dblp.org/search/pid/api?q=author:Chao_Li_0009:"]},"html":{"value":"https://doi.org/10.1109/ICME57554.2024.10687384"},"_bibtex":{"value":"@inproceedings{DBLP:conf/icmcs/PuWHX0WGL24,\n  author={Yifei Pu and Chi Wang and Xiaofeng Hou and Cheng Xu and Jiacheng Liu and Jing Wang and Minyi Guo and Chao Li},\n  title={MSN: Adaptive and Dynamic Multi-modal Shortcut Network Architecture for Latency-Aware Applications},\n  year={2024},\n  cdate={1704067200000},\n  pages={1-6},\n  url={https://doi.org/10.1109/ICME57554.2024.10687384},\n  booktitle={ICME},\n  crossref={conf/icmcs/2024}\n}\n"},"abstract":{"value":"Multi-modal neural networks have demonstrated exceptional performance by merging information across modalities, surpassing the state-of-the-art uni-modal DNNs. However, this accuracy improvement comes at the cost of increased computation, leading to higher inference latency. This defect significantly limits the practical value of multi-modal DNNs, especially for latency-aware applications. Therefore, we propose an adaptive and efficient multi-modal shortcut architecture called M2SN to reduce the execution latency with accuracy guarantees. It skips ineffective network layers to reduce computational costs as well as alleviate the overfitting problem adaptive to specific models and scenarios. The key contributions of M2SN are twofold: 1) We design and insert shortcuts into each uni-modal network to perform adaptive computing. 2) We design a navigator to dynamically choose the optimal shortcuts. Unlike previous approaches, M2SN features high generality as it does not rely on any prior knowledge. The experimental results show that M2SN can reduce 28.3% average latency while obtaining the same or higher accuracy compared with SOTA baselines."},"title":{"value":"MSN: Adaptive and Dynamic Multi-modal Shortcut Network Architecture for Latency-Aware Applications"},"authors":{"value":["Yifei Pu","Chi Wang","Xiaofeng Hou","Cheng Xu","Jiacheng Liu","Jing Wang","Minyi Guo","Chao Li"]}},"tmdate":1741136591568,"pdate":1704067200000,"tcdate":1741136586393,"writers":["~"],"signatures":["~Minyi_Guo1"],"forum":"yOARWcjf9t","license":"CC BY-SA 4.0","number":350964,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1741136591568,"domain":"DBLP.org","id":"yOARWcjf9t","version":2},{"content":{"venue":{"value":"ICME 2024"},"pdf":{"value":"https://ieeexplore.ieee.org/iel8/10685847/10687354/10687384.pdf"},"venueid":{"value":"dblp.org/conf/ICMCS/2024"},"paperhash":{"value":"pu|msn_adaptive_and_dynamic_multimodal_shortcut_network_architecture_for_latencyaware_applications"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Yifei_Pu:","https://dblp.org/search/pid/api?q=author:Chi_Wang:","~Xiaofeng_Hou1","https://dblp.org/search/pid/api?q=author:Cheng_Xu:","https://dblp.org/search/pid/api?q=author:Jiacheng_Liu_0001:","https://dblp.org/search/pid/api?q=author:Jing_Wang_0055:","https://dblp.org/search/pid/api?q=author:Minyi_Guo:","https://dblp.org/search/pid/api?q=author:Chao_Li_0009:"]},"html":{"value":"https://doi.org/10.1109/ICME57554.2024.10687384"},"_bibtex":{"value":"@inproceedings{DBLP:conf/icmcs/PuWHX0WGL24,\n  author={Yifei Pu and Chi Wang and Xiaofeng Hou and Cheng Xu and Jiacheng Liu and Jing Wang and Minyi Guo and Chao Li},\n  title={MSN: Adaptive and Dynamic Multi-modal Shortcut Network Architecture for Latency-Aware Applications},\n  year={2024},\n  cdate={1704067200000},\n  pages={1-6},\n  url={https://doi.org/10.1109/ICME57554.2024.10687384},\n  booktitle={ICME},\n  crossref={conf/icmcs/2024}\n}\n"},"abstract":{"value":"Multi-modal neural networks have demonstrated exceptional performance by merging information across modalities, surpassing the state-of-the-art uni-modal DNNs. However, this accuracy improvement comes at the cost of increased computation, leading to higher inference latency. This defect significantly limits the practical value of multi-modal DNNs, especially for latency-aware applications. Therefore, we propose an adaptive and efficient multi-modal shortcut architecture called M2SN to reduce the execution latency with accuracy guarantees. It skips ineffective network layers to reduce computational costs as well as alleviate the overfitting problem adaptive to specific models and scenarios. The key contributions of M2SN are twofold: 1) We design and insert shortcuts into each uni-modal network to perform adaptive computing. 2) We design a navigator to dynamically choose the optimal shortcuts. Unlike previous approaches, M2SN features high generality as it does not rely on any prior knowledge. The experimental results show that M2SN can reduce 28.3% average latency while obtaining the same or higher accuracy compared with SOTA baselines."},"title":{"value":"MSN: Adaptive and Dynamic Multi-modal Shortcut Network Architecture for Latency-Aware Applications"},"authors":{"value":["Yifei Pu","Chi Wang","Xiaofeng Hou","Cheng Xu","Jiacheng Liu","Jing Wang","Minyi Guo","Chao Li"]}},"tmdate":1739262085541,"pdate":1704067200000,"tcdate":1739262082526,"writers":["~"],"signatures":["~Xiaofeng_Hou1"],"forum":"Kc9VJ67FE1","license":"CC BY-SA 4.0","number":318357,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1739262085541,"domain":"DBLP.org","id":"Kc9VJ67FE1","version":2},{"content":{"summary":{"value":"* Introduces CinePile, a large-scale long-form video understanding benchmark: ~305k MCQ QAs across 9,396 videos (train/test).\n\n* Targets temporal reasoning, human–object interactions, and narrative/plot understanding—explicitly aiming to defeat single-frame shortcuts.\n\n* Built via an LLM-driven, human-in-the-loop pipeline that leverages human audio descriptions aligned to public movie clips; questions are generated without direct video input.\n\n* At evaluation time, models get only raw video + dialogue.\n\n* Benchmarks 24 models; humans outperform best commercial models by ~25% and open-source by ~37%; instruction-tuning on CinePile yields large gains (e.g., Video-LLaVA +71% relative)."},"soundness":{"value":2},"confidence":{"value":5},"questions":{"value":"1- Compare against InfiniBench in detail, as it seems to offer more skills and longer videos.\n\n2- Please detail all the missing vague numbers in Section 2.\n\n3- Add more recent benchmarks and methods to polish the paper.\n\n4- What are the insights of the benchmark? A beneficial benchmark should shed light on the current limitations to guide future research. Therefore, highlighting interesting conclusions is essential."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"* Clear problem motivation: Existing “long-video” datasets permit frame-based shortcuts; CinePile stresses genuine temporal/narrative comprehension.\n\n* Scale & utility: Large enough for both instruction-tuning and evaluation; already seeing community adoption.\n\n* Generalizable pipeline: Automated QA generation/verification shown to extend to longer (≤30 min) and non-movie videos with minimal prompt changes.\n\n* Question diversity: Broad coverage (temporal, perceptual, reasoning); includes quantitative diversity metrics and a hard split analysis."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1.\tOrganization & clarity of exposition\nThe manuscript would benefit from tighter structure and clearer analysis. For example, Section 3 could be integrated as a subsection of Section 2. Section 2 currently reads as procedural rather than analytical; please quantify key elements (e.g., the proportion of “hard” questions detected and the share of weak Q&A filtered). Statements such as “we observed that some questions…” (L213) should be replaced with concrete statistics and definitions of “trivial/basic.” Reporting these ratios will turn the section from a workflow description into evidence-backed analysis.\n\n2.\tVisual reliance policy\nIf visual grounding is a primary objective, consider filtering out non–vision-centric items rather than merely down-weighting them. Alternatively, justify the scoring approach with ablations showing how the visual-reliance score correlates with model performance and dataset difficulty.\n\n3.\tRelated work coverage & comparative positioning\nThe related work and Table 1 under-represent recent benchmarks. Please update with stronger, more current baselines and include a detailed, side-by-side comparison with InfiniBench (released >1 year ago), covering scale, clip duration, domains, question types, and construction pipeline. If CinePile is smaller/shorter, clarify the intended complementary role; if larger/longer in some dimensions, make that explicit.\n\n4.\tTemplate-based generation: motivation & impact\nThe rationale for using templates needs clarification. What benefits do templates provide over free-form generation (e.g., control, coverage, difficulty)? Quantify any trade-offs: linguistic diversity, semantic variety, and question naturalness. If templates risk constraining the benchmark, consider expanding template sets or adding a free-form subset and report diversity metrics across both.\n\n5.\tEvaluation format (MCQ-only)\nLimiting the evaluation to MCQ may mask real reasoning ability. Adding an open-ended evaluation regime (short-answer/free-form with calibrated automatic or human grading) would substantially strengthen claims about long-form understanding and reduce the risk of distractor-driven shortcuts.\n\n6.\tInput-modality ablations\nPlease include a systematic analysis of modality contributions: (i) video-only (no subtitles), (ii) subtitles/dialogue-only (no frames), and (iii) full multimodal input. Reporting the deltas will clarify visual vs. textual reliance and help assess potential leakage or superficial cue exploitation.\n\n7.\tBaselines and recency in comparisons\nSeveral compared methods in Table 2 appear outdated. Please include major video(-LLM) releases from the past year and ensure consistent, transparent evaluation settings (frame rate, sampling strategy, input modalities, prompt formats) to support fair comparisons."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762924749656,"tcdate":1762826170257,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission14313/Reviewer_yLyE"],"signatures":["ICLR.cc/2026/Conference/Submission14313/Reviewer_yLyE"],"forum":"gF51sPe5Yc","number":4,"license":"CC BY 4.0","cdate":1762826170257,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission14313/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762924749656,"domain":"ICLR.cc/2026/Conference","replyto":"gF51sPe5Yc","id":"tRAO1VBfPu","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Datasets and benchmarking","Video understanding","Multi-modal learning","Visual question answering","Long-form video","Metrics and benchmarks"]},"supplementary_material":{"value":"/attachment/0b48d3dee81adb3c113688e15c73d4b55479e82a.zip"},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Current datasets for long-form video understanding often fall short of providing genuine long-form comprehension challenges, as many tasks derived from these datasets can be successfully tackled by analyzing just one or a few random frames from a video. To address this issue, we present a novel dataset and benchmark, CinePile, specifically designed for authentic long-form video understanding. This paper details our innovative approach for creating a question-answer dataset, utilizing advanced LLMs with human-in-the-loop and building upon human-generated raw data. Our comprehensive dataset comprises 305,000 multiple-choice questions (MCQs), covering various visual and multimodal aspects, including temporal comprehension, understanding human-object interactions, and reasoning about events or actions within a scene. Additionally, we fine-tuned open-source Video-LLMs on the training split and evaluated both open-source and proprietary video-centric LLMs on the test split of our dataset. The findings indicate that although current models underperform compared to humans, fine-tuning these models can lead to significant improvements in their performance."},"_bibtex":{"value":"@misc{\nrawal2025cinepile,\ntitle={CinePile: A Large-Scale Video Question Answering Dataset and Benchmark},\nauthor={Ruchit Rawal and Khalid Saifullah and Miquel Farr{\\'e} and Ronen Basri and David Jacobs and Gowthami Somepalli and Tom Goldstein},\nyear={2025},\nurl={https://openreview.net/forum?id=gF51sPe5Yc}\n}"},"title":{"value":"CinePile: A Large-Scale Video Question Answering Dataset and Benchmark"},"pdf":{"value":"/pdf/857e99107f45569d53616139b21709b66bdebd4c.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"rawal|cinepile_a_largescale_video_question_answering_dataset_and_benchmark"},"authorids":{"value":["~Ruchit_Rawal1","~Khalid_Saifullah1","~Miquel_Farré1","~Ronen_Basri1","~David_Jacobs1","~Gowthami_Somepalli1","~Tom_Goldstein1"]},"authors":{"value":["Ruchit Rawal","Khalid Saifullah","Miquel Farré","Ronen Basri","David Jacobs","Gowthami Somepalli","Tom Goldstein"]}},"version":2},{"content":{"summary":{"value":"The paper proposes LOVE-R1, a large video-language model designed for efficient and accurate long video understanding. The approach combines dense, low-resolution frame sampling for global temporal context and selectively zooms in on high-resolution segments when finer spatial details are necessary, as determined by the model’s reasoning process. The method is trained with a three-stage post-training paradigm. Experiments on several long video benchmarks show consistent improvements over strong baselines such as Qwen2.5-VL, achieving state-of-the-art performance with a 3.1% average gain. Overall, the paper contributes a novel adaptive sampling framework and a reasoning-based training pipeline that effectively balance temporal coverage and spatial detail for long video understanding."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. Given the marginal 0.5% gain from decoupled reinforcement finetuning, under what conditions does this method significantly outperform standard GRPO, justifying its added complexity?\n2. Can the authors provide error analysis on zoom-in failures, such as mislocalization rates or incorrect decision-to-zoom vs. answer choices?"},"rating":{"value":6},"details_of_ethics_concerns":{"value":"No ethics review needed."},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"**1. Originality:**\nThe paper addresses a fundamental limitation of current video-language models, the trade-off between temporal coverage and spatial fidelity under constrained context length. By introducing an adaptive, query-driven zoom-in mechanism, it reframes long video understanding as a progressive multi-step reasoning process (decision, zoom-in, answer). This formulation is both intuitive and impactful, demonstrating clear conceptual originality. \n\n**2. Technical Soundness:**\nThe proposed three-stage post-training framework, comprising slow–fast template finetuning, chain-of-thought (CoT) initialization, and decoupled reinforcement optimization, is logically structured and empirically validated. Each stage contributes a distinct function.\n\n**3. Empirical Validation:**\nThe experimental evaluation is extensive and convincing. The proposed LOVE-R1 model consistently outperforms strong baseline across multiple long video understanding benchmarks (Video-MME, LongVideoBench, LVBench, MLVU). The ablation studies (Tables 2–3) clearly attribute gains to specific model components, confirming the utility of adaptive sampling and multi-step reasoning."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"**1. Marginal Gain from Decoupled RL:**\nA core innovation presented for the training process is the \"decoupled reinforcement finetuning\" (Stage 3). The ablation study in Table 2b shows that this decoupled optimization (labeled \"Single-Step Optimization\") achieves an overall score of 66.2%, only a marginal 0.5% improvement over the standard multi-step GRPO baseline (\"Multi-Step Optimization\"), which scored 65.7%. This minimal gain does not strongly justify the added complexity of the decoupled approach, and the paper does not clearly demonstrate the specific conditions or failure cases under which this method becomes essential.\n\n**2. Missing Error Analysis and Efficiency Metrics:**\nThe paper lacks a crucial discussion of the model's specific failure modes. While ablation studies compare against baselines like \"random zoom-in\", there is no qualitative or quantitative analysis of when the LOVE-R1 model itself fails. For instance, the analysis does not explore the frequency or nature of incorrect zoom-in localizations, errors in the decision-making step (i.e., deciding to zoom vs. to answer), or misinterpretations of the detailed \"slow video\" frames.\n\n**3. Unknown Computation Cost:**\nFurthermore, the paper omits any reporting on computational cost. A primary motivation for an adaptive \"slow-fast\" mechanism is to balance performance with efficiency. However, the paper provides no metrics on inference time, memory footprint, or a clear discussion of the practical efficiency trade-offs between sampling density and spatial resolution."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762917814118,"tcdate":1760944656708,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission4998/Reviewer_yZYH"],"signatures":["ICLR.cc/2026/Conference/Submission4998/Reviewer_yZYH"],"forum":"IlmqwQtY20","number":1,"license":"CC BY 4.0","cdate":1760944656708,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission4998/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762917814118,"domain":"ICLR.cc/2026/Conference","replyto":"IlmqwQtY20","id":"POxMMvCef4","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Long video understanding","multimodal reasoning"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Long video understanding is still challenging for recent Large Video-Language Models (LVLMs) due to the conflict between long-form temporal understanding and detailed spatial perception. LVLMs with a uniform frame sampling mechanism, which samples frames with an equal frame size and fixed sampling rate, inevitably sacrifice either temporal clues or spatial details, resulting in suboptimal solutions. To mitigate this dilemma, we propose LOVE-R1, a model that can adaptively zoom in on a video clip. The model is first provided with densely sampled frames but in a small resolution. If some spatial details are needed, the model can zoom in on a clip of interest with a large frame resolution based on its reasoning until key visual information is obtained. The whole process is implemented as a multi-step reasoning process. To train the reasoning ability, we first finetune the model on our collected 38k high-quality CoT data and enhance it with decoupled reinforcement finetuning. As outcome rewards can not provide fine-grained process supervision, we decouple multi-step reasoning into multiple single-step reasoning and optimize the internal zoom-in ability explicitly. Experiments on long video understanding benchmarks show that our model with the slow-fast adaptive frame sampling mechanism achieves a great trade-off between sampling density and frame resolutions, and LOVE-R1 outperforms our baseline Qwen2.5-VL by an average of 3.1\\% points across 4 common long video understanding benchmarks."},"_bibtex":{"value":"@misc{\nfu2026lover,\ntitle={{LOVE}-R1: Advancing Long Video Understanding with Adaptive Zoom-in Mechanism via Multi-Step Reasoning},\nauthor={Shenghao Fu and Qize Yang and Yuan-Ming Li and Xihan Wei and Xiaohua Xie and Wei-Shi Zheng},\nyear={2026},\nurl={https://openreview.net/forum?id=IlmqwQtY20}\n}"},"title":{"value":"LOVE-R1: Advancing Long Video Understanding with Adaptive Zoom-in Mechanism via Multi-Step Reasoning"},"pdf":{"value":"/pdf/4300a6fc45108f2fed82916f7bb731eebf9cba7f.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"fu|lover1_advancing_long_video_understanding_with_adaptive_zoomin_mechanism_via_multistep_reasoning"},"authorids":{"value":["~Shenghao_Fu1","~Qize_Yang1","~Yuan-Ming_Li7","~Xihan_Wei1","~Xiaohua_Xie1","~Wei-Shi_Zheng3"]},"authors":{"value":["Shenghao Fu","Qize Yang","Yuan-Ming Li","Xihan Wei","Xiaohua Xie","Wei-Shi Zheng"]}},"version":2},{"content":{"summary":{"value":"The paper reframes preference-data quality as model-dependent and introduces a Truncated Influence Function (TIF) lens showing that medium-IF pairs—not very small or very large ones—drive the most stable alignment gains. To avoid the high cost of exact IF on LLMs, the authors propose two forward-pass proxies—Loss Difference (LossDiff) and Implicit Reward Margin (IRM)—and a combined LossDiff–IRM selector that closely tracks TIF."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"See weaknesses."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. **Principled, model-aware data valuation.** The paper grounds selection in a TIF-motivated view of *model-dependent* data value and shows that combining **LossDiff** and **IRM** outperforms either proxy alone; moreover, pairs discarded by the selector are empirically low-value for alignment, confirming the criterion’s discriminative power.  \n\n2. **Strong empirical generality.** Across models, objectives (DPO/SLiC), and benchmarks, the method achieves higher win rates using reduced data (e.g., 64% subset with top performance)."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. **Incomplete positioning vs. recent work.** The paper discusses Filtered DPO, but does not engage with several highly relevant baselines:\n\n   * [1] margin-based preference selection for alignment quality (ICML 2025),\n   * [2] RS-DPO (rejection sampling + DPO for cleaner preference data),\n   * [3] active preference learning for LLMs (querying informative pairs instead of passively filtering).\n     These works target the same core problem — selecting / curating high-value preference pairs — and should be compared both conceptually and empirically.\n\n2. **Practicality.** The proposed selectors still require rescoring large volumes of pairs with both the current model and an auxiliary model. The paper does not report the real cost (GPU hours, throughput) or show that this is cheaper/more scalable than RS-DPO-style sampling or active preference acquisition.\n\n3. **Robustness.** The “medium-IF is best” claim is convincing on the reported setups but is not stress-tested across broader domains, model sizes, or stages of alignment; it is unclear how stable this curriculum is outside the presented benchmarks.\n\n[1] Larger or Smaller Reward Margins to Select Preferences for Alignment? ICML 2025.  \n[2] Rs-dpo: A hybrid rejection sampling and direct preference optimization method for alignment of large language models. NAACL 2024.  \n[3] Active preference learning for large language models. ICML 2024."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1764359919368,"tcdate":1761988113590,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission13211/Reviewer_8ymy"],"signatures":["ICLR.cc/2026/Conference/Submission13211/Reviewer_8ymy"],"forum":"FUp0KeEEBs","number":3,"license":"CC BY 4.0","cdate":1761988113590,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission13211/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1764359919368,"domain":"ICLR.cc/2026/Conference","replyto":"FUp0KeEEBs","id":"0qyKcj2SMA","forumContent":{"TLDR":{"value":"We assess preference data quality through our newly proposed truncated influence function (TIF), and then we propose a set of candidate scoring functions that are positive correlated with TIF to select valuable preference data."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Large language model alignment","preference data","influence function"]},"primary_area":{"value":"foundation or frontier models, including LLMs"},"abstract":{"value":"Large language model (LLM) alignment is typically achieved through learning from human preference comparisons, making the quality of preference data critical to its success. Existing studies often pre-process raw training datasets to identify valuable preference pairs using external reward models or off-the-shelf LLMs, achieving improved overall performance but rarely examining whether individual, selected data point is genuinely beneficial. We assess data quality through individual influence on validation data using our newly proposed truncated influence function (TIF), which mitigates the over-scoring present in traditional measures and reveals that preference data quality is inherently a property of the model. In other words, a data pair that benefits one model may harm another. This leaves the need to improve the preference data selection approaches to be adapting to specific models. To this end, we introduce two candidate scoring functions (SFs) that are computationally simpler than TIF and positively correlated with it. They are also model dependent and can serve as potential indicators of individual data quality for preference data selection. Furthermore, we observe that these SFs inherently exhibit errors when compared to TIF. To this end, we combine them to offset their diverse error sources, resulting in a simple yet effective data selection rule that enables the models to achieve a more precise selection of valuable preference data. We conduct experiments across diverse alignment benchmarks and various LLM families, with results demonstrating that better alignment performance can be achieved using less data, showing the generality of our findings and new methods. Our code is publicly available at~\\url{https://github.com/tmlr-group/TIF_LossDiff-IRM}."},"_bibtex":{"value":"@inproceedings{\nzhang2026towards,\ntitle={Towards Understanding Valuable Preference Data for Large Language Model Alignment},\nauthor={Zizhuo Zhang and Qizhou Wang and Shanshan Ye and Jianing Zhu and Jiangchao Yao and Bo Han and Masashi Sugiyama},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=FUp0KeEEBs}\n}"},"title":{"value":"Towards Understanding Valuable Preference Data for Large Language Model Alignment"},"pdf":{"value":"/pdf/57226567930a062c2cbc65ab18cfbcb06d576705.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"zhang|towards_understanding_valuable_preference_data_for_large_language_model_alignment"},"authorids":{"value":["~Zizhuo_Zhang1","~Qizhou_Wang1","~Shanshan_Ye1","~Jianing_Zhu2","~Jiangchao_Yao1","~Bo_Han1","~Masashi_Sugiyama1"]},"authors":{"value":["Zizhuo Zhang","Qizhou Wang","Shanshan Ye","Jianing Zhu","Jiangchao Yao","Bo Han","Masashi Sugiyama"]}},"version":2},{"content":{"summary":{"value":"The paper proposes a method or 3D scene reconstruction from as few as two views via a two-stage approach. First, a finetuned video diffusion model generates multiple view that interpolate the camera trajectory between the given input views. In order to improve 3D consistency of these generated images, the authors introduce a 3D structure conditioning by encoding a point cloud obtained from an existing learned stereo reconstruction method (DUSt3R). In a second stage, the generated views are fused into a 3D representation, for which the authors extend 3D Gaussian Splatting by a weighting based on confidence of DUSt3R.\nExperimental evaluation results show strong performance for both in training distribution data as well as cross-dataset generalization, outperforming state-of-the-art baselines. Ablation studies validate the effectiveness of the technical contributions."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"- Is the mentioned counterexample for the proof correct or am I missing something?\n- What exactly is the implementation of the learnable embedding function PosEmb in equation 3?\n- What is meant with \"uncertainty of unconstrained images\" (line 265)? Does this mean that generated images are not fully 3D-consistent?\n- Regarding the confidence-aware 3DGS optimization in section 4.4, how is the loss in equation 6 still used / relevant for the final loss in equation 7, which does not use equation 6 anymore?\n- Is there any specific reason why you limit yourself to the case of view interpolation and do not evaluate extrapolation?"},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"- The paper tackles an important and very challenging task of reconstructing general scenes from very sparse views.\n- The two-stage approach leverages the strong prior of video diffusion models and the technical contributions fit well into this framework.\n- The experimental evaluation shows strong results consistently outperforming state-of-the-art baselines on both\n   - in training distribution data \n     - significant quantitative and qualitative advantage for small angle variance in input views (table 1, figure 3)\n     - even larger gap for large angle variance in input views (table 2, figure 4)\n  - and cross-dataset generalization\n    - baselines fail completely, while the proposed method achieves reasonable metric scores (table 4)\n    - impressive qualitative results (figure 6)\n- The ablation study validates the effectiveness of the proposed 3D structure conditioning and the confidence-aware 3DGS optimization (especially convincing qualitative ablation in figure 5).\n- The paper is mostly well-written and easy to follow.\n- The supplemental video provides additional convincing qualitative reconstruction results."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The lack of precise mathematical formulation raises doubt about proposition 1 and its proof:\n  - The main inequality in equation 17 (appendix) is justified only verbally and is not obvious to me.\n  - A counterexample for the inequality in line 841 could be a dataset mainly consisting of dark rooms with all ground truth renderings being black except for one where the lights are on (with possibly very complex geometry or even transparent and reflective materials). For this case, fitting the marginal distribution can be easier than fitting the conditional distribution, if the given scene s happens to be the one with lights on.\n  - Moreover, if the 3D structure (point cloud) is obtained from 2D images only via stereo reconstruction (DUSt3R [1]), then an image encoder should be theoretically able to encode the images in the same way as the composition of DUSt3R and the point cloud encoder used in the method. Following this argument, a strict inequality as in proposition 1 is not reasonable.\n\n- Some lack of clarity:\n  - What exactly is the implementation of the learnable embedding function PosEmb in equation 3?\n  - What is meant with \"uncertainty of unconstrained images\" (line 265)? Does this mean that generated images are not fully 3D-consistent?\n  - Regarding the confidence-aware 3DGS optimization in section 4.4, it is unclear how the loss in equation 6 is still used / relevant for the final loss in equation 7, which does not use equation 6 anymore.\n\n- Somewhat unfair comparison with baselines:\n  - The proposed method relies heavily on the strong priors learned by DUSt3R [1] and the video diffusion model.\n  - All baselines are trained more or less from scratch on small-scale datasets compared to the large-scale datasets for training video diffusion models.\n  - Therefore, it is not that surprising that the proposed method generalizes much better to other datasets out of the training/finetuning data distribution.\n\n- Restriction to view interpolation:\n  - Given that the method makes use of a pre-trained video diffusion model, the restriction to interpolation of the camera trajectory between two input views is unsatisfactory.\n  - In many cases, two views as conditioning already eliminate uncertainty in the reconstruction task entirely such that deterministic approaches like pixelSplat [4] and MVSplat [5] already produce detailed and sharp reconstructions.\n  - The proposed generative approach with a strong prior trained on large-scale single view data should be able to extrapolate from input views to some extent, which the paper misses to evaluate.\n  - This is especially critical, if the paper motivates the method with the limitation of generalizable 3D reconstruction methods that \"struggle to generate high-quality images in areas not visible from the input perspectives\" (lines 143f.).\n\n- Missing relevant related work and possible baseline:\n  - latentSplat [2] aims to bridge the gap between regression models for generalizable NVS and generative models for 3D reconstruction and is therefore a very relevant work and potential baseline.\n  - GeNVS [3] was one of the first approaches to generative novel views with a diffusion model conditioned on 3D-aware (pixelNeRF) features and is therefore important related work.\n\nMinor comments:\n- The paper sometimes uses the terms \"conditioning\" and \"guidance\" interchangeably, although they have different meanings to me (conditioning: giving something as additional input to the diffusion model; guidance: using CFG or also gradients (e.g. classifier guidance) to affect the sampling process). A more precise use of these terms would be helpful.\n- The ablation study in table 3 could be extended to contain all different combinations of leaving out individual components in order to find possible dependencies between them.\n\nReferences:\n- [1] DUSt3R: Geometric 3D Vision Made Easy. CVPR 2024\n- [2] latentSplat: Autoencoding Variational Gaussians for Fast Generalizable 3D Reconstruction. ECCV 2024\n- [3] Generative Novel View Synthesis with 3D-Aware Diffusion Models. ICCV 2023\n- [4] pixelSplat: 3D Gaussian Splats from Image Pairs for Scalable Generalizable 3D Reconstruction. CVPR 2024\n- [5] MVSplat: Efficient 3D Gaussian Splatting from Sparse Multi-View Images. ECCV 2024"}},"nonreaders":[],"tmdate":1733089181023,"tcdate":1730637829219,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission9317/Reviewer_Vp2m"],"signatures":["ICLR.cc/2025/Conference/Submission9317/Reviewer_Vp2m"],"forum":"Z30Mdbv5jO","number":5,"license":"CC BY 4.0","cdate":1730637829219,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission9317/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1733089181023,"domain":"ICLR.cc/2025/Conference","replyto":"Z30Mdbv5jO","id":"r3BOYPJI88","forumContent":{"TLDR":{"value":"ReconX is a novel sparse-view 3D scene reconstruction paradigm that reframes the ambiguous reconstruction challenge as a temporal generation task."},"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Sparse-view Reconstruction","Video Diffusion","Gaussian Splatting"]},"supplementary_material":{"value":"/attachment/85bb44579fb246807dc36f1fcaffbce8ea01e93d.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Advancements in 3D scene reconstruction have transformed 2D images from the real world into 3D models, producing realistic 3D results from hundreds of input photos. Despite great success in dense-view reconstruction scenarios, rendering a detailed scene from sparse views is still an ill-posed optimization problem, often resulting in artifacts and distortions in unseen areas. In this paper, we propose ReconX, a novel 3D scene reconstruction paradigm that reframes the ambiguous reconstruction problem as a temporal generation task. The key insight is to unleash the strong generative prior of large pre-trained video diffusion models for sparse-view reconstruction. Nevertheless, it is challenging to preserve 3D view consistency when directly generating video frames from pre-trained models. To address this issue, given limited input views, the proposed ReconX first constructs a global point cloud and encodes it into a contextual space as the 3D structure condition. Guided by the condition, the video diffusion model then synthesizes video frames that are detail-preserved and exhibit a high degree of 3D consistency, ensuring the coherence of the scene from various perspectives. Finally, we recover the 3D scene from the generated video through a confidence-aware 3D Gaussian Splatting optimization scheme. Extensive experiments on various real-world datasets show the superiority of ReconX over state-of-the-art methods in terms of quality and generalizability."},"_bibtex":{"value":"@misc{\nliu2025reconx,\ntitle={ReconX: Reconstruct Any Scene from Sparse Views with Video Diffusion Model},\nauthor={Fangfu Liu and Wenqiang Sun and Hanyang Wang and Yikai Wang and Haowen Sun and Junliang Ye and Jun Zhang and Yueqi Duan},\nyear={2025},\nurl={https://openreview.net/forum?id=Z30Mdbv5jO}\n}"},"title":{"value":"ReconX: Reconstruct Any Scene from Sparse Views with Video Diffusion Model"},"pdf":{"value":"/pdf/9364ea472d15f95d4ba49d9d3ab667616846ed10.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"liu|reconx_reconstruct_any_scene_from_sparse_views_with_video_diffusion_model"},"authorids":{"value":["~Fangfu_Liu2","~Wenqiang_Sun1","~Hanyang_Wang4","~Yikai_Wang2","~Haowen_Sun1","~Junliang_Ye1","~Jun_Zhang25","~Yueqi_Duan1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Fangfu Liu","Wenqiang Sun","Hanyang Wang","Yikai Wang","Haowen Sun","Junliang Ye","Jun Zhang","Yueqi Duan"]}},"version":2},{"content":{"summary":{"value":"This paper identifies a key failure point in autoregressive video generation: static sampling strategies like top-k/top-p, which work well for language, are ill-suited for the low-density, high-redundancy nature of video tokens, leading to error accumulation. To address this, the authors propose Entropy-Guided k-Guard (ENKG) sampling, a training-free, inference-time method. ENKG dynamically adjusts the nucleus sampling probability for each token based on its predictive entropy—using a smaller, more conservative candidate pool for low-entropy (high-confidence) regions and a larger, more diverse pool for high-entropy (uncertain) regions. This is combined with a \"k-guard\" mechanism that ensures a minimal number of top candidates are always included to prevent the model from becoming overly greedy and causing artifacts like frozen frames."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"While the related work section mentions some uses of entropy in LLMs, such as for model switching or retrieval, it doesn't discuss entropy-guided sampling for text generation itself. Have similar adaptive, entropy-guided sampling strategies already been investigated in the LLM literature? For example, there is a very similar paper[1], which explores most of the concepts in the proposed method, but in LLM settings.\n\n[1] EDT: Improving Large Language Models’ Generation by Entropy-based Dynamic Temperature Sampling"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"The paper's primary strength is its clear and insightful diagnosis of the problem, effectively distinguishing the statistical properties of video versus language tokens and identifying the \"entropy collapse\" phenomenon. The proposed ENKG method is simple and highly practical, as it can be applied as a plug-and-play module to existing models without any retraining. The experimental results are convincing, showing consistent and significant improvements across multiple state-of-the-art video models and datasets, well-supported by thorough ablation studies that validate each component of the proposed method"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The main weakness lies in the potentially limited novelty of the core concepts. While their application to video generation is new and insightful, entropy-guided adaptation and hybrid sampling methods have been explored in other domains. Additionally, the evaluation is heavily focused on autonomous driving scenarios. While effective here, it remains an open question how well this strategy would generalize to more open-domain or creative video generation tasks, which may exhibit different uncertainty structures and benefit from different sampling trade-offs."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762922862312,"tcdate":1761904968089,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission11846/Reviewer_esDH"],"signatures":["ICLR.cc/2026/Conference/Submission11846/Reviewer_esDH"],"forum":"8yKSnpZjyd","number":3,"license":"CC BY 4.0","cdate":1761904968089,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission11846/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762922862312,"domain":"ICLR.cc/2026/Conference","replyto":"8yKSnpZjyd","id":"wKrdb7p7kc","forumContent":{"TLDR":{"value":"We propose ENkG, an entropy-guided sampling policy with a k-guard that mitigates error accumulation and preserves structure in long-horizon autoregressive video generation—plug-and-play at inference, no retraining."},"venue":{"value":"ICLR 2026 Conference Desk Rejected Submission"},"keywords":{"value":["Autoregressive video generation","discrete tokens","VQ-VAE","entropy","top-p / nucleus sampling","top-k","adaptive decoding","error accumulation","uncertainty-aware sampling"]},"supplementary_material":{"value":"/attachment/2149f2c43ada90cc09d71b6cd0b02b25292f32b7.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Autoregressive (AR) architectures have achieved significant successes in LLM, inspiring explorations for video generation.  In LLMs, top-$p$/top-$k$ sampling strategies work exceptionally well: language tokens have high semantic density and low redundancy, so a fixed size of token candidates already strike a balance between semantic accuracy and generation diversity. In contrast, video tokens have low semantic density and high spatio-temporal redundancy. This mismatch makes static top-k/top-p strategies ineffective for video decoders: they either introduce unnecessary randomness for low-uncertainty regions (static backgrounds) or stuck in early errors for high-uncertainty regions (foreground objects). Prediction errors will accumulate as more frames are generated and eventually\nseverely degrade long-horizon quality. \nTo address this, we propose Entropy-Guided $k$-Guard (ENkG) sampling, a simple yet effective strategy that adapts sampling to token-wise dispersion, quantified by the entropy of each token’s predicted distribution. \nENkG uses adaptive token candidate sizes: for low-entropy regions, it employs fewer candidates to suppress redundant noise and preserve structural integrity; for high-entropy regions, it uses more candidates to mitigate error compounding.\nENkG is model-agnostic, training-free, and adds negligible overhead. Experiments demonstrate consistent improvements in perceptual quality and structural stability compared to static top-k/top-p strategies."},"_bibtex":{"value":"@misc{\nanonymous2026entropyguided,\ntitle={Entropy-Guided k-Guard Sampling for Long-Horizon Autoregressive Video Generation},\nauthor={Anonymous},\nyear={2026},\nurl={https://openreview.net/forum?id=8yKSnpZjyd}\n}"},"title":{"value":"Entropy-Guided k-Guard Sampling for Long-Horizon Autoregressive Video Generation"},"pdf":{"value":"/pdf/0110e66be73c396e503bc7cf10101f77c2b52ab5.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Desk_Rejected_Submission"},"paperhash":{"value":"han|entropyguided_kguard_sampling_for_longhorizon_autoregressive_video_generation"},"authorids":{"value":["~Yizhao_Han1","~Tianxing_Shi1","~Zhao_Wang11","~Zhiyuan_Pu1","~Mingxiao_Li4","~Qian_Zhang7","~Wei_Yin2","~Xiao-Xiao_Long1"]},"authors":{"value":["Yizhao Han","Tianxing Shi","Zhao Wang","Zhiyuan Pu","Mingxiao Li","Qian Zhang","Wei Yin","Xiao-Xiao Long"]}},"version":2},{"content":{"summary":{"value":"This paper investigates issues in time series contrastive learning caused by \"bad\" positive pairs violating the assumption of shared semantic information between views. It identifies two types of bad pairs: noisy positive pairs mainly sharing noise patterns when original signal is noisy, and faulty positive pairs where data augmentation alters important temporal patterns, resulting in views no longer sharing meaning. Through analysis and observing a simulated dataset, the paper shows these bad pairs degrade representations via noisy alignment or faulty alignment during training. To address this, the authors propose a Dynamic Bad Pair Mining (DBPM) algorithm that tracks training loss of each pair over time to identify potential bad pairs based on statistics of historical losses. DBPM estimates weights to suppress identified bad pairs, mitigating their detrimental impacts. Experiments show integrating DBPM into various contrastive learning frameworks consistently improves linear evaluation performance on time series benchmarks. Controlled tests also demonstrate DBPM provides greater robustness against injected bad pairs versus baseline contrastive learning. Overall, this paper provides first investigation of the bad pair problem in time series contrastive learning and proposes the DBPM solution to reliably identify and suppress bad pairs as a simple plug-in enhancing existing methods."},"presentation":{"value":"3 good"},"contribution":{"value":"3 good"},"soundness":{"value":"3 good"},"strengths":{"value":"- Identifies and provides the first study of an important issue in time series contrastive learning - the problem of bad positive pairs - that has not been thoroughly explored before. Provides both theoretical analysis and empirical observations on simulated data to demonstrate the detrimental effects of noisy and faulty positive pairs on representation learning.\n\n- Proposes a simple solution (DBPM) that is model-agnostic and can work as a plug-in to boost various existing contrastive learning methods for time series.\n\n- Evaluation on multiple real-world time series benchmarks demonstrates clear and consistent improvements from integrating DBPM across different datasets and models. Additional experiments validate DBPM's ability to confer greater robustness against injected bad pairs compared to baseline contrastive learning.\n\n- Thorough ablation studies analyze the effects of different design choices and hyperparameters\n\n- The code and simulated dataset will be made publicly available to facilitate future research"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The identification of bad pairs is based on heuristically set thresholds. Could more principled statistical methods could be explored to automatically determine optimal thresholds?\n2. All evaluations were studied with the InfoNCE loss, does this approach generally work for other losses, for e.g., the supervised contrastive (SupCon) loss Khosla et. al, NeurIPS 2020."},"confidence":{"value":"3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked."},"questions":{"value":"1. Did the authors conduct ablation experiments to distinguish the individual impacts of noisy versus faulty positive pairs on the model's performance? It would be insightful to determine if specific datasets are predominantly influenced by one over the other. Is it possible to isolate certain features and then forecast based on them?"},"rating":{"value":"6: marginally above the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636193212,"tcdate":1699046225572,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission2562/Reviewer_Cugc"],"signatures":["ICLR.cc/2024/Conference/Submission2562/Reviewer_Cugc"],"forum":"K2c04ulKXn","number":2,"license":"CC BY 4.0","cdate":1699046225572,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission2562/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636193212,"domain":"ICLR.cc/2024/Conference","replyto":"K2c04ulKXn","id":"umk3ZBspWP","forumContent":{"venue":{"value":"ICLR 2024 poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Time Series Contrastive Learning","Healthcare","Self-Supervised Representation Learning"]},"primary_area":{"value":"representation learning for computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"*Not all positive pairs are beneficial to time series contrastive learning*. In this paper, we study two types of bad positive pairs that can impair the quality of time series representation learned through contrastive learning: the noisy positive pair and the faulty positive pair. We observe that, with the presence of noisy positive pairs, the model tends to simply learn the pattern of noise (Noisy Alignment). Meanwhile, when faulty positive pairs arise, the model wastes considerable amount of effort aligning non-representative patterns (Faulty Alignment). To address this problem, we propose a Dynamic Bad Pair Mining (DBPM) algorithm, which reliably identifies and suppresses bad positive pairs in time series contrastive learning. Specifically, DBPM utilizes a memory module to dynamically track the training behavior of each positive pair along training process. This allows us to identify potential bad positive pairs at each epoch based on their historical training behaviors. The identified bad pairs are subsequently down-weighted through a transformation module, thereby mitigating their negative impact on the representation learning process. DBPM is a simple algorithm designed as a lightweight **plug-in** without learnable parameters to enhance the performance of existing state-of-the-art methods. Through extensive experiments conducted on four large-scale, real-world time series datasets, we demonstrate DBPM's efficacy in mitigating the adverse effects of bad positive pairs."},"_bibtex":{"value":"@inproceedings{\nlan2024towards,\ntitle={Towards Enhancing Time Series Contrastive Learning: A Dynamic Bad Pair Mining Approach},\nauthor={Xiang Lan and Hanshu Yan and Shenda Hong and Mengling Feng},\nbooktitle={The Twelfth International Conference on Learning Representations},\nyear={2024},\nurl={https://openreview.net/forum?id=K2c04ulKXn}\n}"},"title":{"value":"Towards Enhancing Time Series Contrastive Learning: A Dynamic Bad Pair Mining Approach"},"pdf":{"value":"/pdf/99d93a2891889c6677593a50a92480572438f057.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference"},"paperhash":{"value":"lan|towards_enhancing_time_series_contrastive_learning_a_dynamic_bad_pair_mining_approach"},"authorids":{"value":["~Xiang_Lan2","~Hanshu_Yan1","~Shenda_Hong1","~Mengling_Feng1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Xiang Lan","Hanshu Yan","Shenda Hong","Mengling Feng"]}},"version":2},{"content":{"venue":{"value":"AISTATS 2022"},"pdf":{"value":"https://proceedings.mlr.press/v151/makar22a/makar22a.pdf"},"venueid":{"value":"dblp.org/conf/AISTATS/2022"},"paperhash":{"value":"makar|causally_motivated_shortcut_removal_using_auxiliary_labels"},"authorids":{"value":["~Maggie_Makar1","https://dblp.org/search/pid/api?q=author:Ben_Packer:","https://dblp.org/search/pid/api?q=author:Dan_Moldovan:","https://dblp.org/search/pid/api?q=author:Davis_W._Blalock:","~Yoni_Halpern1","https://dblp.org/search/pid/api?q=author:Alexander_D'Amour:"]},"html":{"value":"https://proceedings.mlr.press/v151/makar22a.html"},"_bibtex":{"value":"@inproceedings{DBLP:conf/aistats/MakarPMBHD22,\n  author={Maggie Makar and Ben Packer and Dan Moldovan and Davis W. Blalock and Yoni Halpern and Alexander D'Amour},\n  title={Causally motivated shortcut removal using auxiliary labels},\n  year={2022},\n  cdate={1640995200000},\n  pages={739-766},\n  url={https://proceedings.mlr.press/v151/makar22a.html},\n  booktitle={AISTATS},\n  crossref={conf/aistats/2022}\n}\n"},"abstract":{"value":"Shortcut learning, in which models make use of easy-to-represent but unstable associations, is a major failure mode for robust machine learning. We study a flexible, causally-motivated approach to training robust predictors by discouraging the use of specific shortcuts, focusing on a common setting where a robust predictor could achieve optimal i.i.d generalization in principle, but is overshadowed by a shortcut predictor in practice. Our approach uses auxiliary labels, typically available at training time, to enforce conditional independences implied by the causal graph. We show both theoretically and empirically that causally-motivated regularization schemes (a) lead to more robust estimators that generalize well under distribution shift, and (b) have better finite sample efficiency compared to usual regularization schemes, even when no shortcut is present. Our analysis highlights important theoretical properties of training techniques commonly used in the causal inference, fairness, and disentanglement literatures. Our code is available at github.com/mymakar/causally_motivated_shortcut_removal"},"title":{"value":"Causally motivated shortcut removal using auxiliary labels"},"authors":{"value":["Maggie Makar","Ben Packer","Dan Moldovan","Davis W. Blalock","Yoni Halpern","Alexander D'Amour"]}},"tmdate":1739952032783,"pdate":1640995200000,"tcdate":1727781684630,"writers":["~"],"signatures":["~Maggie_Makar1"],"forum":"9UZcyhf0wN","license":"CC BY-SA 4.0","number":131083,"cdate":1640995200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1739952032783,"domain":"DBLP.org","id":"9UZcyhf0wN","version":2},{"content":{"venue":{"value":"CoRR 2021"},"pdf":{"value":"http://arxiv.org/pdf/2105.06422v3"},"venueid":{"value":"dblp.org/journals/CORR/2021"},"paperhash":{"value":"makar|causallymotivated_shortcut_removal_using_auxiliary_labels"},"authorids":{"value":["~Maggie_Makar1","https://dblp.org/search/pid/api?q=author:Ben_Packer:","https://dblp.org/search/pid/api?q=author:Dan_Moldovan:","https://dblp.org/search/pid/api?q=author:Davis_W._Blalock:","https://dblp.org/search/pid/api?q=author:Yoni_Halpern:","https://dblp.org/search/pid/api?q=author:Alexander_D'Amour:"]},"html":{"value":"https://arxiv.org/abs/2105.06422"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2105-06422,\n  publtype={informal},\n  author={Maggie Makar and Ben Packer and Dan Moldovan and Davis W. Blalock and Yoni Halpern and Alexander D'Amour},\n  title={Causally-motivated Shortcut Removal Using Auxiliary Labels},\n  year={2021},\n  cdate={1609459200000},\n  journal={CoRR},\n  volume={abs/2105.06422},\n  url={https://arxiv.org/abs/2105.06422}\n}\n"},"abstract":{"value":"Shortcut learning, in which models make use of easy-to-represent but unstable associations, is a major failure mode for robust machine learning. We study a flexible, causally-motivated approach to training robust predictors by discouraging the use of specific shortcuts, focusing on a common setting where a robust predictor could achieve optimal \\emph{iid} generalization in principle, but is overshadowed by a shortcut predictor in practice. Our approach uses auxiliary labels, typically available at training time, to enforce conditional independences implied by the causal graph. We show both theoretically and empirically that causally-motivated regularization schemes (a) lead to more robust estimators that generalize well under distribution shift, and (b) have better finite sample efficiency compared to usual regularization schemes, even when no shortcut is present. Our analysis highlights important theoretical properties of training techniques commonly used in the causal inference, fairness, and disentanglement literatures. Our code is available at https://github.com/mymakar/causally_motivated_shortcut_removal"},"title":{"value":"Causally-motivated Shortcut Removal Using Auxiliary Labels"},"authors":{"value":["Maggie Makar","Ben Packer","Dan Moldovan","Davis W. Blalock","Yoni Halpern","Alexander D'Amour"]}},"tmdate":1727781689650,"pdate":1609459200000,"tcdate":1727781684414,"writers":["~"],"signatures":["~Maggie_Makar1"],"forum":"CEWaZefk0X","license":"CC BY-SA 4.0","number":131075,"cdate":1609459200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1727781689650,"domain":"DBLP.org","id":"CEWaZefk0X","version":2},{"content":{"venue":{"value":"ICML 2023 Poster"},"abstract":{"value":"In this study, we propose Shortcut Fine-Tuning (SFT), a new approach for addressing the challenge of fast sampling of pretrained Denoising Diffusion Probabilistic Models (DDPMs). SFT advocates for the fine-tuning of DDPM samplers through the direct minimization of Integral Probability Metrics (IPM), instead of learning the backward diffusion process. This enables samplers to discover an alternative and more efficient sampling shortcut, deviating from the backward diffusion process. Inspired by a control perspective, we propose a new algorithm SFT-PG: Shortcut Fine-Tuning with Policy Gradient, and prove that under certain assumptions, gradient descent of diffusion models with respect to IPM is equivalent to performing policy gradient. To our best knowledge, this is the first attempt to utilize reinforcement learning (RL) methods to train diffusion models. Through empirical evaluation, we demonstrate that our fine-tuning method can further enhance existing fast DDPM samplers, resulting in sample quality comparable to or even surpassing that of the full-step model across various datasets."},"title":{"value":"Optimizing DDPM Sampling with Shortcut Fine-Tuning"},"pdf":{"value":"/pdf/91556fadf755af970e93bdce63b8d4c9da8e0c35.pdf"},"venueid":{"value":"ICML.cc/2023/Conference"},"paperhash":{"value":"fan|optimizing_ddpm_sampling_with_shortcut_finetuning"},"authorids":{"value":["~Ying_Fan2","~Kangwook_Lee1"]},"authors":{"value":["Ying Fan","Kangwook Lee"]}},"tmdate":1686841448061,"pdate":1682373197265,"tcdate":1673305187609,"writers":["ICML.cc/2023/Conference","ICML.cc/2023/Conference/Submission76/Authors"],"signatures":["ICML.cc/2023/Conference/Submission76/Authors"],"forum":"3SiELE1Wzl","number":76,"cdate":1673305187609,"mdate":1686841448061,"readers":["everyone"],"invitations":["ICML.cc/2023/Conference/-/Submission","ICML.cc/2023/Conference/-/Edit","ICML.cc/2023/Conference/Submission76/-/Camera_Ready_Revision"],"odate":1686841448049,"domain":"ICML.cc/2023/Conference","id":"3SiELE1Wzl","version":2},{"content":{"summary":{"value":"This paper presents Q-Bench-Video, a benchmark specifically designed to evaluate the video quality understanding capabilities of Large Multi-modal Models (LMMs). Recognizing a gap in existing benchmarks that focus on high-level video comprehension rather than quality assessment, Q-Bench-Video targets four key dimensions of video quality: technical, aesthetic, temporal, and AI-generated content (AIGC) distortions. The benchmark includes a diverse set of videos from natural scenes, AI-generated content, and computer graphics, ensuring a balanced distribution across quality levels. To comprehensively assess LMMs, Q-Bench-Video employs various question types—Yes-or-No, What-How, open-ended, and video pair comparisons—that capture the models’ performance across straightforward and complex tasks. Validated on 17 LMMs (12 open-source and 5 proprietary), the benchmark reveals that, while LMMs demonstrate foundational capabilities in video quality assessment, they fall significantly short of human-level performance, particularly in open-ended and AIGC-specific queries. Q-Bench-Video provides a new standard for evaluating video quality understanding and aims to drive future research on enhancing LMMs’ video quality perception."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"1. Could you provide more detailed insights into the specific error patterns observed in LMMs’ open-ended responses?\n\n2. Have you considered testing LMMs on additional specialized domains, such as gaming videos, animation, or screen content, to better assess cross-domain generalization?\n\n3. Given the potential bias in GPT-assisted evaluation, did you consider incorporating human evaluations as a baseline for open-ended responses? \n\n4. Would you consider adding a more challenging subset focused specifically on low-quality or noisy videos?"},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. Q-Bench-Video is the first benchmark specifically focused on assessing video quality understanding in Large Multi-modal Models (LMMs), addressing a unique and underexplored aspect of LMMs that goes beyond typical video comprehension. By evaluating quality-related distortions—including technical, aesthetic, temporal, and AIGC-specific aspects—it provides a novel and comprehensive perspective on video quality assessment.\n\n2. The benchmark is meticulously designed with diverse video sources (natural, AIGC, CG) and a balanced quality distribution through uniform sampling from quality-annotated datasets. The use of various question types (Yes-or-No, What-How, open-ended, and video pair comparisons) creates a robust framework that evaluates LMMs across both straightforward and complex scenarios, ensuring a thorough assessment of model capabilities.\n\n3. Q-Bench-Video addresses an important gap in LMM evaluation by focusing on video quality understanding, which is essential for applications in video generation, content moderation, and quality control. The findings—highlighting substantial performance gaps between LMMs and human evaluators, especially in AIGC distortions—offer valuable insights for advancing LMMs and guide future research on improving their quality perception."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. While Q-Bench-Video includes diverse video types, there is limited exploration of how LMMs generalize across these domains. Testing LMMs on specific domains, such as medical or surveillance videos, would enhance the benchmark’s relevance by showing how well models handle domain-specific quality variations, which are critical in many real-world applications.\n\n2. The benchmark results indicate that LMMs struggle significantly with open-ended questions, but the paper lacks a detailed breakdown of common error types or patterns in these responses. A more granular analysis could provide clearer insights into where models fail in nuanced video quality understanding and offer guidance on specific areas for model improvement.\n\n3. The reliance on GPT-assisted evaluation for scoring open-ended responses may introduce subjectivity or bias, as the evaluation depends on the alignment of the assistant model’s judgments with human interpretation. Including human evaluations as a baseline for open-ended questions would provide a stronger reference and help validate the reliability of the GPT-assisted scores.\n\n4. The benchmark primarily focuses on a balanced quality distribution but lacks an emphasis on evaluating LMMs in scenarios with challenging low-quality or noisy data, which are common in real-world settings. Incorporating more degraded video samples could better test the robustness of LMMs and highlight areas where models may require improvement for real-world deployment.\n\n5. Some of the highly relevant video quality assessment works should be mentioned in the related work section, like [R1][R2][R3].\n\n[R1] UGC-VQA: Benchmarking Blind Video Quality Assessment for User Generated Content, TIP 2021\n[R2] RAPIQUE: Rapid and accurate video quality prediction of user generated content, OJSP 2021\n[R3] FAVER: Blind Quality Prediction of Variable Frame Rate Videos, SPIC 202"}},"nonreaders":[],"tmdate":1731427490028,"tcdate":1730711887687,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission1786/Reviewer_W4DH"],"signatures":["ICLR.cc/2025/Conference/Submission1786/Reviewer_W4DH"],"forum":"VaUy5GZO3f","number":4,"license":"CC BY 4.0","cdate":1730711887687,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission1786/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427490028,"domain":"ICLR.cc/2025/Conference","replyto":"VaUy5GZO3f","id":"wbNquaDtCg","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Large multi-modal model","benchmark","video quality assessment"]},"primary_area":{"value":"datasets and benchmarks"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"With the rising interest in research on Large Multi-modal Models (LMMs) for video understanding, many studies have emphasized general video comprehension capabilities, neglecting the **systematic exploration into video quality understanding**. To address this oversight, we introduce **Q-Bench-Video** in this paper, a new benchmark specifically designed to evaluate LMMs' proficiency in discerning video quality. **a)** To ensure the diversity of video sources, Q-Bench-Video encompasses videos from natural scenes, computer graphics (CG), and AI-generated content (AIGC). **b)** Building on the traditional multiple-choice questions format with the *Yes-or-No* and *What-How* categories, we include *Open-ended* questions to better evaluate complex scenarios. Additionally, we incorporate the **video pair quality comparison** question to enhance comprehensiveness. **c)** Beyond the traditional *Technical*, *Aesthetic*, and *Temporal* distortions, we have expanded our evaluation aspects to include the dimension of *AIGC* distortions, which addresses the increasing demand for video generation. Finally, we collect a total of 2,378 question-answer pairs and test them on 12 open-source & 5 proprietary LMMs. Our findings indicate that while LMMs have a foundational understanding of video quality, their performance remains incomplete and imprecise, with a notable discrepancy compared to human-level performance. Through **Q-Bench-Video**, we seek to catalyze community interest, stimulate further research, and unlock the untapped potential of LMMs to close the gap in video quality understanding."},"_bibtex":{"value":"@misc{\nzhang2024qbenchvideo,\ntitle={Q-Bench-Video: Benchmarking the Video Quality Understanding of {LMM}s},\nauthor={Zicheng Zhang and Ziheng Jia and Haoning Wu and Chunyi Li and Zijian Chen and Yingjie Zhou and Wei Sun and Xiaohong Liu and Xiongkuo Min and Weisi Lin and Guangtao Zhai},\nyear={2024},\nurl={https://openreview.net/forum?id=VaUy5GZO3f}\n}"},"title":{"value":"Q-Bench-Video: Benchmarking the Video Quality Understanding of LMMs"},"pdf":{"value":"/pdf/b6665a67a0e193ca939b3b34f488e5e0b380c738.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"zhang|qbenchvideo_benchmarking_the_video_quality_understanding_of_lmms"},"authorids":{"value":["~Zicheng_Zhang7","~Ziheng_Jia1","~Haoning_Wu1","~Chunyi_Li1","~Zijian_Chen1","~Yingjie_Zhou1","~Wei_Sun12","~Xiaohong_Liu2","~Xiongkuo_Min1","~Weisi_Lin1","~Guangtao_Zhai1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Zicheng Zhang","Ziheng Jia","Haoning Wu","Chunyi Li","Zijian Chen","Yingjie Zhou","Wei Sun","Xiaohong Liu","Xiongkuo Min","Weisi Lin","Guangtao Zhai"]}},"version":2},{"content":{"TLDR":{"value":"Hawkeye scales the test-time compute of coding agents over a minimal taxonomy of unit tests to port hardware-aware GPU kernels across architectures, vendors, and precisions."},"venue":{"value":"ICML 2026 Workshop DL4C Oral"},"pdf":{"value":"/pdf/b635c08f50c0b6b899b8847fd4078ce6d74699ec.pdf"},"keywords":{"value":["GPU","kernel optimization","porting","hardware aware"]},"venueid":{"value":"ICML.cc/2026/Workshop/DL4C"},"paperhash":{"value":"tschand|hawkeye_hardwareaware_gpu_kernel_optimization_with_minimal_supervision"},"authorids":{"value":["~Arya_Tschand1","~Kesavan_Ramakrishnan1","~Alexander_Ingare1","~Simon_Guo1","~Jeffrey_Jian_Ma1","~Zishen_Wan1","~Simran_Arora1","~Azalia_Mirhoseini3","~Vijay_Janapa_Reddi1"]},"abstract":{"value":"Achieving peak GPU kernel performance increasingly relies on architecture-specific optimizations targeting new hardware features. While AI coding agents show promise in generating performant kernels, they lack the necessary context to effectively implement and stack hardware-specific optimizations, especially on newer GPU architectures. We propose Hawkeye (Hardware-Aware Kernel Optimization), an open-source framework that grounds autonomous kernel generation in a minimal and comprehensive taxonomy with only one unit test per optimization strategy per target architecture. Supporting a new accelerator therefore requires only 10 expert-written unit tests per architecture (one per recurring optimization strategy) that generalize across downstream workloads, rather than hand writing a new kernel for each workload and precision. Hawkeye effectively scales the test-time compute of coding agents with this minimal expert supervision to enable kernel generation that consistently leverages hardware-specific features, approaching and even surpassing expert-written PyTorch or Triton in BF16 and emerging low precision (FP8, NVFP4, MXFP4) across Ampere, Hopper, Blackwell, and MI350 GPUs. Hawkeye demonstrates that minimally supervised coding agents can exploit architecture-specific hardware features and reduce the overhead of supporting emerging hardware accelerators."},"_bibtex":{"value":"@inproceedings{\ntschand2026hawkeye,\ntitle={Hawkeye: Hardware-Aware {GPU} Kernel Optimization with Minimal Supervision},\nauthor={Arya Tschand and Kesavan Ramakrishnan and Alexander Ingare and Simon Guo and Jeffrey Jian Ma and Zishen Wan and Simran Arora and Azalia Mirhoseini and Vijay Janapa Reddi},\nbooktitle={Deep Learning for Code: Towards Human-Centered Coding Agents},\nyear={2026},\nurl={https://openreview.net/forum?id=e3pxJbBRBk}\n}"},"title":{"value":"Hawkeye: Hardware-Aware GPU Kernel Optimization with Minimal Supervision"},"authors":{"value":["Arya Tschand","Kesavan Ramakrishnan","Alexander Ingare","Simon Guo","Jeffrey Jian Ma","Zishen Wan","Simran Arora","Azalia Mirhoseini","Vijay Janapa Reddi"]}},"tmdate":1782331000723,"pdate":1781618626052,"tcdate":1778789859754,"writers":["ICML.cc/2026/Workshop/DL4C","ICML.cc/2026/Workshop/DL4C/Submission62/Authors"],"signatures":["ICML.cc/2026/Workshop/DL4C/Submission62/Authors"],"forum":"e3pxJbBRBk","license":"CC BY-NC 4.0","number":62,"cdate":1778789859754,"readers":["everyone"],"invitations":["ICML.cc/2026/Workshop/DL4C/-/Submission","ICML.cc/2026/Workshop/DL4C/-/Post_Submission","ICML.cc/2026/Workshop/DL4C/-/Edit"],"mdate":1782331000723,"odate":1781618626052,"domain":"ICML.cc/2026/Workshop/DL4C","id":"e3pxJbBRBk","version":2},{"content":{"venue":{"value":"CoRR 2025"},"pdf":{"value":"http://arxiv.org/pdf/2502.00688v1"},"venueid":{"value":"dblp.org/journals/CORR/2025"},"paperhash":{"value":"chen|highorder_matching_for_onestep_shortcut_diffusion_models"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Bo_Chen_0029:","~Chengyue_Gong1","https://dblp.org/search/pid/api?q=author:Xiaoyu_Li_0001:","https://dblp.org/search/pid/api?q=author:Yingyu_Liang:","https://dblp.org/search/pid/api?q=author:Zhizhou_Sha:","~Zhenmei_Shi1","~Zhao_Song2",""]},"html":{"value":"https://doi.org/10.48550/arXiv.2502.00688"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2502-00688,\n  publtype={informal},\n  author={Bo Chen and Chengyue Gong and Xiaoyu Li and Yingyu Liang and Zhizhou Sha and Zhenmei Shi and Zhao Song and Mingda Wan},\n  title={High-Order Matching for One-Step Shortcut Diffusion Models},\n  year={2025},\n  month={February},\n  cdate={1738368000000},\n  journal={CoRR},\n  volume={abs/2502.00688},\n  url={https://doi.org/10.48550/arXiv.2502.00688}\n}\n"},"abstract":{"value":"One-step shortcut diffusion models [Frans, Hafner, Levine and Abbeel, ICLR 2025] have shown potential in vision generation, but their reliance on first-order trajectory supervision is fundamentally limited. The Shortcut model's simplistic velocity-only approach fails to capture intrinsic manifold geometry, leading to erratic trajectories, poor geometric alignment, and instability-especially in high-curvature regions. These shortcomings stem from its inability to model mid-horizon dependencies or complex distributional features, leaving it ill-equipped for robust generative modeling. In this work, we introduce HOMO (High-Order Matching for One-Step Shortcut Diffusion), a game-changing framework that leverages high-order supervision to revolutionize distribution transportation. By incorporating acceleration, jerk, and beyond, HOMO not only fixes the flaws of the Shortcut model but also achieves unprecedented smoothness, stability, and geometric precision. Theoretically, we prove that HOMO's high-order supervision ensures superior approximation accuracy, outperforming first-order methods. Empirically, HOMO dominates in complex settings, particularly in high-curvature regions where the Shortcut model struggles. Our experiments show that HOMO delivers smoother trajectories and better distributional alignment, setting a new standard for one-step generative models."},"title":{"value":"High-Order Matching for One-Step Shortcut Diffusion Models"},"authors":{"value":["Bo Chen","Chengyue Gong","Xiaoyu Li","Yingyu Liang","Zhizhou Sha","Zhenmei Shi","Zhao Song","Mingda Wan"]}},"tmdate":1782468077500,"pdate":1735689600000,"tcdate":1742444493035,"writers":["~"],"signatures":["~Zhenmei_Shi1"],"forum":"3fgMsA4bgU","license":"CC BY-SA 4.0","number":370175,"cdate":1738368000000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1782468077500,"domain":"DBLP.org","id":"3fgMsA4bgU","version":2},{"content":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["High-Order Matching","Diffusion Model","One-Step Shortcut"]},"primary_area":{"value":"generative models"},"abstract":{"value":"One-step shortcut diffusion models [Frans, Hafner, Levine and Abbeel, ICLR 2025] have shown potential in vision generation, but their reliance on first-order trajectory supervision is fundamentally limited. The Shortcut model's simplistic velocity-only approach fails to capture intrinsic manifold geometry, leading to erratic trajectories, poor geometric alignment, and instability-especially in high-curvature regions. These shortcomings stem from its inability to model mid-horizon dependencies or complex distributional features, leaving it ill-equipped for robust generative modeling. In this work, we introduce HOMO (High-Order Matching for One-Step Shortcut Diffusion), a game-changing framework that leverages high-order supervision to revolutionize distribution transportation. By incorporating acceleration, jerk, and beyond, HOMO not only fixes the flaws of the Shortcut model but also achieves unprecedented smoothness, stability, and geometric precision. Theoretically, we prove that HOMO's high-order supervision ensures superior approximation accuracy, outperforming first-order methods. Empirically, HOMO dominates in complex settings, particularly in high-curvature regions where the Shortcut model struggles. Our experiments show that HOMO delivers smoother trajectories and better distributional alignment, setting a new standard for one-step generative models."},"_bibtex":{"value":"@misc{\nchen2025highorder,\ntitle={High-Order Matching for One-Step Shortcut Diffusion Models},\nauthor={Yubin Chen and Chengyue Gong and Xiaoyu Li and Yingyu Liang and Zhizhou Sha and Zhenmei Shi and Zhao Song},\nyear={2025},\nurl={https://openreview.net/forum?id=Sv5Ubt3dFi}\n}"},"title":{"value":"High-Order Matching for One-Step Shortcut Diffusion Models"},"pdf":{"value":"/pdf/c52172f317c0fcbf437a819feb095714446d4c8b.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"chen|highorder_matching_for_onestep_shortcut_diffusion_models"},"authorids":{"value":["~Yubin_Chen4","~Chengyue_Gong1","~Xiaoyu_Li12","~Yingyu_Liang1","~Zhizhou_Sha1","~Zhenmei_Shi1","~Zhao_Song3"]},"authors":{"value":["Yubin Chen","Chengyue Gong","Xiaoyu Li","Yingyu Liang","Zhizhou Sha","Zhenmei Shi","Zhao Song"]}},"tmdate":1764200549124,"tcdate":1758223392706,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission13834/Authors"],"signatures":["ICLR.cc/2026/Conference/Submission13834/Authors"],"forum":"Sv5Ubt3dFi","license":"CC BY 4.0","number":13834,"cdate":1758223392706,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/-/Submission","ICLR.cc/2026/Conference/-/Post_Submission","ICLR.cc/2026/Conference/Submission13834/-/Full_Submission","ICLR.cc/2026/Conference/-/Withdrawn_Submission"],"mdate":1764200549124,"odate":1759896705795,"domain":"ICLR.cc/2026/Conference","id":"Sv5Ubt3dFi","version":2},{"content":{"venue":{"value":"ICLR 2025 DeLTa Workshop Poster"},"keywords":{"value":["High-Order Matching","Diffusion Model","One-Step Shortcut"]},"abstract":{"value":"One-step shortcut diffusion models [Frans, Hafner, Levine and Abbeel, ICLR 2025] have shown potential in vision generation, but their reliance on first-order trajectory supervision is fundamentally limited. The Shortcut model's simplistic velocity-only approach fails to capture intrinsic manifold geometry, leading to erratic trajectories, poor geometric alignment, and instability-especially in high-curvature regions. These shortcomings stem from its inability to model mid-horizon dependencies or complex distributional features, leaving it ill-equipped for robust generative modeling.\nIn this work, we introduce *HOMO* (**H**igh-**O**rder **M**atching for **O**ne-Step Shortcut Diffusion), a game-changing framework that leverages high-order supervision to revolutionize distribution transportation. By incorporating acceleration, jerk, and beyond, HOMO not only fixes the flaws of the Shortcut model but also achieves unprecedented smoothness, stability, and geometric precision. Theoretically, we prove that HOMO's high-order supervision ensures superior approximation accuracy, outperforming first-order methods.\nEmpirically, HOMO dominates in complex settings, particularly in high-curvature regions where the Shortcut model struggles. Our experiments show that HOMO delivers smoother trajectories and better distributional alignment, setting a new standard for one-step generative models."},"_bibtex":{"value":"@inproceedings{\nchen2025highorder,\ntitle={High-Order Matching for One-Step Shortcut Diffusion Models},\nauthor={Bo Chen and Chengyue Gong and Xiaoyu Li and Yingyu Liang and Zhizhou Sha and Zhenmei Shi and Zhao Song and Mingda Wan},\nbooktitle={ICLR 2025 Workshop on Deep Generative Model in Machine Learning: Theory, Principle and Efficacy},\nyear={2025},\nurl={https://openreview.net/forum?id=be8vS1Dcix}\n}"},"title":{"value":"High-Order Matching for One-Step Shortcut Diffusion Models"},"pdf":{"value":"/pdf/46c256c15cb239250d90e46030cfc75de7255577.pdf"},"venueid":{"value":"ICLR.cc/2025/Workshop/DeLTa"},"paperhash":{"value":"chen|highorder_matching_for_onestep_shortcut_diffusion_models"},"authorids":{"value":["~Bo_Chen22","~Chengyue_Gong1","~Xiaoyu_Li12","~Yingyu_Liang1","~Zhizhou_Sha1","~Zhenmei_Shi1","~Zhao_Song3","~Mingda_Wan1"]},"Track":{"value":"long paper (up to 8 pages)"},"authors":{"value":["Bo Chen","Chengyue Gong","Xiaoyu Li","Yingyu Liang","Zhizhou Sha","Zhenmei Shi","Zhao Song","Mingda Wan"]}},"tmdate":1744680858943,"pdate":1741244033965,"tcdate":1738859841549,"writers":["ICLR.cc/2025/Workshop/DeLTa","ICLR.cc/2025/Workshop/DeLTa/Submission50/Authors"],"signatures":["ICLR.cc/2025/Workshop/DeLTa/Submission50/Authors"],"forum":"be8vS1Dcix","license":"CC BY 4.0","number":50,"cdate":1738859841549,"readers":["everyone"],"invitations":["ICLR.cc/2025/Workshop/DeLTa/-/Submission","ICLR.cc/2025/Workshop/DeLTa/-/Post_Submission","ICLR.cc/2025/Workshop/DeLTa/-/Edit","ICLR.cc/2025/Workshop/DeLTa/Submission50/-/Camera_Ready"],"mdate":1744680858943,"odate":1741949086703,"domain":"ICLR.cc/2025/Workshop/DeLTa","id":"be8vS1Dcix","version":2},{"content":{"summary":{"value":"This paper explores the challenges faced by contrastive learning in the context of time series data due to the presence of noisy and faulty positive pairs. It introduces a novel Dynamic Bad Pair Mining (DBPM) algorithm that dynamically identifies and mitigates the influence of these bad positive pairs, aiming to enhance the performance of existing contrastive learning methods. The paper demonstrates the effectiveness of DBPM through extensive experiments on a suite of real-world time series datasets."},"presentation":{"value":"3 good"},"contribution":{"value":"3 good"},"soundness":{"value":"3 good"},"strengths":{"value":"1. **Novel Problem Identification:** The paper introduces a previously unrecognized issue in time series contrastive learning—the challenge of 'bad' positive pairs. The authors convincingly illustrate two variants of such pairs: noisy and faulty, which arise due to augmentation on certain data inputs. The paper provides a persuasive and logical explanation for the emergence of these pairs, marking a significant observation for the time series contrastive learning community.\n\n2. **Empirical Observations:** Through empirical research on a simulated dataset, the authors observed that noisy positive pairs exhibit very low training loss, whereas faulty positive pairs show high training loss. Remarkably, both types display consistently low loss variances. Hence, on a plane plotting mean against variance of training loss, these two classes stand out as outliers from the bulk of good positive pairs. This finding is not only intriguing and insightful but also interpretable and plausible.\n\n3. **Algorithm Development:** Building upon these observations, the authors suggest utilizing the statistics (mean and variance) of training losses of positive pairs to flag the bad ones. This process includes a memory module that records the loss of all positive pairs over several epochs, followed by an algorithm that then downweights the bad pairs identified through this mechanism.\n\n4. **Presentation Clarity:** The paper is well-structured and articulate in detailing the problem, the proposed resolution, and the experimental framework. The inclusion of illustrative figures and tables effectively contributes to the reader's comprehension of the concepts and validates the significance of the proposed approach."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. **Ablation Study Limitations:** The presented ablation studies are notably limited, which constrains the understanding of each component's unique effects on the overall system performance. Particularly, the roles of the memory module and the transformation module have not been analyzed independently. Distinct investigations into how each module contributes to the suppression of bad pairs, the efficiency of the learning process, and the final model accuracy would greatly enhance the transparency of the results. \n\n2. **Algorithmic Complexity:** The discussion on the algorithm’s additional computational complexity is mentioned but lacks depth. A detailed exploration of the computational costs, including an assessment of trade-offs, would be instructive. Further analysis on the scaling behavior, particularly concerning large datasets and the associated runtime and memory implications in comparison to baseline methods, would greatly benefit the paper.\n\n3. **Parameter Sensitivity Analysis:** The method for selecting the hyperparameters critical for identifying bad positive pairs (specifically, the two $\\beta$ parameters) seems to be heuristic-based. A thorough sensitivity analysis of these hyperparameters, with an emphasis on how they influence model performance, could substantiate the robustness of the paper. Although the preset heuristics for setting thresholds perform adequately, an adaptive strategy for threshold selection could potentially be more effective than manual tuning of the β values."},"confidence":{"value":"3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked."},"questions":{"value":"What are other potential applications (e.g. forecasting) where DBPM could be beneficial? Would be good to see evaluation beyond just classification."},"rating":{"value":"6: marginally above the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636192739,"tcdate":1699513701170,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission2562/Reviewer_gEVT"],"signatures":["ICLR.cc/2024/Conference/Submission2562/Reviewer_gEVT"],"forum":"K2c04ulKXn","number":4,"license":"CC BY 4.0","cdate":1699513701170,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission2562/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636192739,"domain":"ICLR.cc/2024/Conference","replyto":"K2c04ulKXn","id":"EzCGb7c13o","forumContent":{"venue":{"value":"ICLR 2024 poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Time Series Contrastive Learning","Healthcare","Self-Supervised Representation Learning"]},"primary_area":{"value":"representation learning for computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"*Not all positive pairs are beneficial to time series contrastive learning*. In this paper, we study two types of bad positive pairs that can impair the quality of time series representation learned through contrastive learning: the noisy positive pair and the faulty positive pair. We observe that, with the presence of noisy positive pairs, the model tends to simply learn the pattern of noise (Noisy Alignment). Meanwhile, when faulty positive pairs arise, the model wastes considerable amount of effort aligning non-representative patterns (Faulty Alignment). To address this problem, we propose a Dynamic Bad Pair Mining (DBPM) algorithm, which reliably identifies and suppresses bad positive pairs in time series contrastive learning. Specifically, DBPM utilizes a memory module to dynamically track the training behavior of each positive pair along training process. This allows us to identify potential bad positive pairs at each epoch based on their historical training behaviors. The identified bad pairs are subsequently down-weighted through a transformation module, thereby mitigating their negative impact on the representation learning process. DBPM is a simple algorithm designed as a lightweight **plug-in** without learnable parameters to enhance the performance of existing state-of-the-art methods. Through extensive experiments conducted on four large-scale, real-world time series datasets, we demonstrate DBPM's efficacy in mitigating the adverse effects of bad positive pairs."},"_bibtex":{"value":"@inproceedings{\nlan2024towards,\ntitle={Towards Enhancing Time Series Contrastive Learning: A Dynamic Bad Pair Mining Approach},\nauthor={Xiang Lan and Hanshu Yan and Shenda Hong and Mengling Feng},\nbooktitle={The Twelfth International Conference on Learning Representations},\nyear={2024},\nurl={https://openreview.net/forum?id=K2c04ulKXn}\n}"},"title":{"value":"Towards Enhancing Time Series Contrastive Learning: A Dynamic Bad Pair Mining Approach"},"pdf":{"value":"/pdf/99d93a2891889c6677593a50a92480572438f057.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference"},"paperhash":{"value":"lan|towards_enhancing_time_series_contrastive_learning_a_dynamic_bad_pair_mining_approach"},"authorids":{"value":["~Xiang_Lan2","~Hanshu_Yan1","~Shenda_Hong1","~Mengling_Feng1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Xiang Lan","Hanshu Yan","Shenda Hong","Mengling Feng"]}},"version":2},{"content":{"summary":{"value":"This work targets promptable segmentation across multi-view images or videos with 3D consistency.\nPrior methods like SAM2-Video lack 3D awareness and produce inconsistent masks across views, while optimization-based methods (SA3D, OmniSeg3D) require costly per-scene fitting.\n\nTheir approach uses Pi3 to produce pointmaps (pixel-to-3D correspondences) from unposed images.\nThe key insight is that pointmaps naturally bridge 2D prompts and 3D geometry without rendering or projection.\nTheir pipeline\n(1) runs Pi3 to get pointmaps with confidences, extracts image embeddings using frozen SAM2-Video.\n(2) They then encode the 3d information into positional embeddings for the prompt embedding and their proposed \"confidence embedding\".\n(4) Finally they feed the enhanced embeddings through a mask decoder (standard)\n\nFor training, they use only the SA-1B (single-view images) dataset, with no multi-view data or 3D annotations required.\nThey evaluate on their method on the NVOS and SPIn-NeRF benchmarks, measuring mIoU and mean accuracy for segmentation.\nResults show consistent improvement over SAM2-Video (generalization / \"online\" baseline), while achieving competitive performance with per-scene optimization methods."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"Figure 5 compares single-view vs full-view attention, but to my understanding the model was never trained with this full-view attention scenario.\nCould the authors elaborate on why this is informative?\n\nIs there any difference between the training/evaluation settings in the ablation studies (Table 3) versus the main results (Table 1)?\nI'm trying to make sure I understand the slight difference in results."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"The technical idea is well-motivated and easy to follow, building off existing work.\nEnhancing SAM2-Video embeddings with 3D positional information improves consistency across views.\n\nResults show significant improvement over SAM2-Video across NVOS and SPIn-NeRF benchmarks.\nThey achieve competitive performance with optimization-based methods without requiring per-scene fitting.\n\nThe ablations are informative, particularly Table 3a which evaluates key decoder design choices.\nThey systematically compare attention scope, positional embeddings, and confidence embeddings.\n\nThe Appendix is thorough and I found many questions answered in there."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The comparison against generalization baselines is limited to SAM2-Video alone.\nIt would strengthen the paper to include other video or multi-view segmentation methods that don't require per-scene optimization.\n\nThe method relies heavily on Pi3 for pointmap generation, but the technical sections provide limited detail on how Pi3 works.\nGiven that Pi3 appears to do much of the heavy lifting, it's unclear how much of the contribution is genuinely novel versus simply combining existing components (Pi3 + SAM2-Video).\nIt would be helpful to either describe Pi3 in a preliminary section or compare it against other visual geometry models for generating pointmaps to understand the design choices.\n\nGiven that the proposed method slightly underperforms optimization-based methods, it would be useful to show an extra column or so that measures the inference time / number of frames to let their method shine more.\n\n### Minor Weaknesses\n\nThe model is trained on single-view object-image pairs from SA-1B (Section 3.4), which seems mismatched with the multi-view/video task at test time.\nA bit of justification/analysis on why not just train with multi-view would be helpful.\n\nIn general, some of the formatting and the placement of the results could be improved to better help flow for the reader.\nThis is very minor, but some of the tables are often quite \"far\" from the reference text.\n\nTable 3b is useful but could be better motivated by introducing some of those methods earlier in the related work."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762919876600,"tcdate":1762009303333,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission7831/Reviewer_nx6A"],"signatures":["ICLR.cc/2026/Conference/Submission7831/Reviewer_nx6A"],"forum":"xnQTMZAbZ0","number":3,"license":"CC BY 4.0","cdate":1762009303333,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission7831/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762919876600,"domain":"ICLR.cc/2026/Conference","replyto":"xnQTMZAbZ0","id":"M5kyyaTew8","forumContent":{"TLDR":{"value":"Solving 3D promotable segmentation using pointmap guidance."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["SAM","3D Vision","VGGT"]},"supplementary_material":{"value":"/attachment/719076ce8cd4f8ab439e58fc7ec72ebdd77069c8.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Promptable segmentation has emerged as a powerful paradigm in computer vision, enabling users to guide models in parsing complex scenes with prompts such as clicks, boxes, or textual cues. Recent advances, exemplified by the Segment Anything Model (SAM), have extended this paradigm to videos and multi-view images. However, the lack of 3D awareness often leads to inconsistent results, necessitating costly per-scene optimization to enforce 3D consistency. In this work, we introduce MV-SAM, a framework for multi-view segmentation that achieves 3D consistency using pointmaps—3D points reconstructed from unposed images by recent visual geometry models. Leveraging the pixel–point one-to-one correspondence of pointmaps, MV-SAM lifts images and prompts into 3D space, eliminating the need for explicit 3D networks or annotated 3D data. Specifically, MV-SAM extends SAM by lifting image embeddings from its pretrained encoder into 3D point embeddings, which are decoded by a transformer using cross-attention with 3D prompt embeddings. This design aligns 2D interactions with 3D geometry, enabling the model to implicitly learn consistent masks across views through 3D positional embeddings. Trained on the SA-1B dataset, our method generalizes well across domains, outperforming SAM2-Video and achieving comparable performance with per-scene optimization baselines on NVOS, SPIn-NeRF, ScanNet++, uCo3D, and DL3DV benchmarks. Code will be released."},"_bibtex":{"value":"@misc{\njeong2026mvsam,\ntitle={{MV}-{SAM}: Multi-view Promptable Segmentation using Pointmap Guidance},\nauthor={Yoonwoo Jeong and Cheng Sun and Yu-Chiang Frank Wang and Minsu Cho and Jaesung Choe},\nyear={2026},\nurl={https://openreview.net/forum?id=xnQTMZAbZ0}\n}"},"title":{"value":"MV-SAM: Multi-view Promptable Segmentation using Pointmap Guidance"},"pdf":{"value":"/pdf/80ac3e7b1ab928bf1c39159d24be2d96c9499b89.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"jeong|mvsam_multiview_promptable_segmentation_using_pointmap_guidance"},"authorids":{"value":["~Yoonwoo_Jeong1","~Cheng_Sun3","~Yu-Chiang_Frank_Wang2","~Minsu_Cho1","~Jaesung_Choe1"]},"authors":{"value":["Yoonwoo Jeong","Cheng Sun","Yu-Chiang Frank Wang","Minsu Cho","Jaesung Choe"]}},"version":2},{"content":{"summary":{"value":"The paper presents an innovative approach Norton for video-language pre-training that learns long-term temporal dependencies between short video clips and captions. The paper addresses two challenges: 1) the high computational cost of modeling long videos and 2) the noisy correspondence between clips and captions due to asynchronous and irrelevant pairs. The paper uses a modified optimal transport framework to measure the sequence similarity between clips and captions, and to filter out the noisy pairs. The paper also exploits the faulty negative samples in contrastive learning to improve clip representation. The paper evaluates the method on video-paragraph retrieval, text-to-video retrieval, videoQA, and action segmentation tasks, and shows that it outperforms existing methods on various metrics. The paper also conducts ablation studies to analyze the impact of different components."},"presentation":{"value":"4 excellent"},"contribution":{"value":"4 excellent"},"soundness":{"value":"4 excellent"},"strengths":{"value":"1.This paper is well-written, well-organized, and easy to follow.\n2.The problem tackled by the authors is an interesting one - noisy correspondence, almost all video harvested from the web would include such cases and hinders the performance of large-scale multi-modal video learning. This challenge is inherent in the data and has received little attention in the literature. The paper explores this novel and important problem in depth.\n3.The paper presents a unified OT framework that can efficiently and effectively solve the MNC problem at different levels of granularity. I commend the authors for providing a time cost table in the appendix which shows the efficiency of their method with various settings and verifies their claims.\nThis method's unique strength lies in using existing pre-training model and works in a self-bootstrapping capability. Note that directly pretraining a large multi-modal model is unpractical for the researchers not in company. This paper significantly improves the temporal ability of VideoCLIP without the need for additional models like DecemBert or TAN. This self-sufficiency significantly contributes to its enhanced scalability."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1.This paper does not provide specific numerical results on how well the proposed method handles noisy correspondence. It would be beneficial if the authors could provide more detailed numerical results or analysis on this aspect in their rebuttal or future work. For example, they could conduct experiments on synthetic noisy datasets to show the performance of their method. They could also compare their method with other methods that are designed to handle noisy correspondence and show how their method performs in comparison. This would provide more concrete evidence on the effectiveness of their method in handling noisy correspondence.\n2.In certain scenarios, such as video-paragraph retrieval on YouCookII (as shown in Table 1) and in the context of ablation experiments (as indicated in Table 7), we observe encouraging indications of potential improvements. While in other experiments (as depicted in Table 2), there are mixed results across various metrics. Can you shed light on the reasons behind this disparity?\n3.How does the proposed method compare with other OT-based methods like action sequence matching? What are the benefits or drawbacks of using OT for video-text learning? Please provide some comparisons and discussions."},"confidence":{"value":"4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work."},"questions":{"value":"The primary queries for the rebuttal are predominantly derived from the \"weaknesses\" section outlined earlier. For instance, it would be highly appreciated if the authors could augment their experiments regarding noisy correspondence and offer more clarification on how their approach differs from other OT-based methods. Resolving these raised concerns will make the submission stronger and I vote for accepting this paper."},"rating":{"value":"8: accept, good paper"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699635927899,"tcdate":1698749708192,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission39/Reviewer_TtfV"],"signatures":["ICLR.cc/2024/Conference/Submission39/Reviewer_TtfV"],"forum":"9Cu8MRmhq2","number":3,"license":"CC BY 4.0","cdate":1698749708192,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission39/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699635927899,"domain":"ICLR.cc/2024/Conference","replyto":"9Cu8MRmhq2","id":"qe4eJIxX7b","forumContent":{"venue":{"value":"ICLR 2024 oral"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video-language pre-training","Noisy correspondence"]},"primary_area":{"value":"representation learning for computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Existing video-language studies mainly focus on learning short video clips, leaving long-term temporal dependencies rarely explored due to over-high computational cost of modeling long videos. To address this issue, one feasible solution is learning the correspondence between video clips and captions, which however inevitably encounters the multi-granularity noisy correspondence (MNC) problem. To be specific, MNC refers to the clip-caption misalignment (coarse-grained) and frame-word misalignment (fine-grained), hindering temporal learning and video understanding. In this paper, we propose NOise Robust Temporal Optimal traNsport (Norton) that addresses MNC in a unified optimal transport (OT) framework. In brief, Norton employs video-paragraph and clip-caption contrastive losses to capture long-term dependencies based on OT. To address coarse-grained misalignment in video-paragraph contrast, Norton filters out the irrelevant clips and captions through an alignable prompt bucket and realigns asynchronous clip-caption pairs based on transport distance. To address the fine-grained misalignment, Norton incorporates a soft-maximum operator to identify crucial words and key frames. Additionally, Norton exploits the potential faulty negative samples in clip-caption contrast by rectifying the alignment target with OT assignment to ensure precise temporal modeling. Extensive experiments on video retrieval, videoQA, and action segmentation verify the effectiveness of our method. \nCode is available at https://lin-yijie.github.io/projects/Norton."},"_bibtex":{"value":"@inproceedings{\nlin2024multigranularity,\ntitle={Multi-granularity Correspondence Learning from Long-term Noisy Videos},\nauthor={Yijie Lin and Jie Zhang and Zhenyu Huang and Jia Liu and zujie wen and Xi Peng},\nbooktitle={The Twelfth International Conference on Learning Representations},\nyear={2024},\nurl={https://openreview.net/forum?id=9Cu8MRmhq2}\n}"},"title":{"value":"Multi-granularity Correspondence Learning from Long-term Noisy Videos"},"pdf":{"value":"/pdf/578b0930059c165430921bc67cd65b6a0657e518.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference"},"paperhash":{"value":"lin|multigranularity_correspondence_learning_from_longterm_noisy_videos"},"authorids":{"value":["~Yijie_Lin1","~Jie_Zhang42","~Zhenyu_Huang1","~Jia_Liu4","~zujie_wen1","~Xi_Peng3"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Yijie Lin","Jie Zhang","Zhenyu Huang","Jia Liu","zujie wen","Xi Peng"]}},"version":2},{"content":{"research_area_keywords":{"value":"Alignment, Direct Preference Optimization, Interpretability, Legal NLP, Data Generation/Augmentation"},"keywords":{"value":["LLM Alignment","Direct Preference Optimization","Linguistic Empathy","Syntactic Minimal Pairs","Legal Question Answering"]},"languages_studied":{"value":"English"},"venue":{"value":"ACL ARR 2026 January Submission"},"_bibtex":{"value":"@inproceedings{\nanonymous2026parservalidated,\ntitle={Parser-Validated Minimal-Pair Preferences for {DPO}: A Case Study in Legal ''Empathy''},\nauthor={Anonymous},\nbooktitle={Submitted to ACL Rolling Review - January 2026},\nyear={2026},\nurl={https://openreview.net/forum?id=mRyusakBE4},\nnote={under review}\n}"},"title":{"value":"Parser-Validated Minimal-Pair Preferences for DPO: A Case Study in Legal \"Empathy\""},"contribution_types":{"value":["Model analysis & interpretability","Data resources"]},"abstract":{"value":"Aligning LLMs to produce responses perceived as empathetic typically relies on costly human preference data and offers limited insight into which linguistic cues drive those preferences. We study legal question answering and introduce an automatic preference-data pipeline based on parser-validated syntactic minimal pairs. Grounded in linguistic accounts of perspective-taking, we generate rule-labeled minimal pairs along five dimensions (pronouns, voice, tense, polite imperatives, evaluative adverbs) and validate the intended contrast with dependency parsing (82.7% success). From 1,785 questions, we produce 7,378 minimal pairs and fine-tune three 7--8B model families (LLaMA-3, Mistral, Gemma) with DPO. Human evaluation (3,309 judgments, 35 raters) prefers DPO over an SFT baseline (68.8% vs. 31.2\\%, $p<0.001$, $h$=0.40), robust to length controls. Feature validation shows voice is the dominant above-chance contributor (80%, $h$=0.64), while other edits are register-sensitive. Overall, parser-validated minimal pairs provide an interpretable, scalable route to preference optimization and identify which cues align with human judgments in-domain."},"paper_type":{"value":"Long"},"pdf":{"value":"/pdf/3f7452e9f292d2e9bca34e298752247e8e5fc215.pdf"},"research_area":{"value":"Safety and Alignment in LLMs"},"venueid":{"value":"aclweb.org/ACL/ARR/2026/January/Submission"}},"tmdate":1780805121800,"tcdate":1767637554247,"writers":["aclweb.org/ACL/ARR/2026/January","aclweb.org/ACL/ARR/2026/January/Submission6206/Authors"],"signatures":["aclweb.org/ACL/ARR/2026/January/Submission6206/Authors"],"forum":"mRyusakBE4","license":"CC BY 4.0","number":6206,"cdate":1767637554247,"readers":["everyone"],"invitations":["aclweb.org/ACL/ARR/2026/January/-/Submission","aclweb.org/ACL/ARR/2026/January/-/Edit","aclweb.org/ACL/ARR/2026/January/-/Preprint_Release_Submission","aclweb.org/ACL/ARR/2026/January/-/Post_Submission","aclweb.org/ACL/ARR/2026/January/-/Preprint_Post_Submission"],"mdate":1780805121800,"odate":1774007940523,"domain":"aclweb.org/ACL/ARR/2026/January","id":"mRyusakBE4","version":2},{"content":{"TLDR":{"value":"We introduce CAT-Video, a corruption-aware training framework that improves robustness and temporal coherence in video diffusion models through structured, data-aligned noise injection."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["video diffusion","corruption-aware training","robust video generation","structured noise injection","multimodal robustness","temporal coherence"]},"supplementary_material":{"value":"/attachment/c99cdbeea034fc1e74ad38310569a2906228cc13.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Latent Video Diffusion Models (LVDMs) have achieved state-of-the-art generative quality for image and video generation; however, they remain brittle under noisy conditioning, where small perturbations in text or multimodal embeddings can cascade over timesteps and cause semantic drift. Existing corruption strategies from image diffusion (Gaussian, Uniform) fail in video settings because static noise disrupts temporal fidelity. In this paper, we propose **CAT-Video**, a corruption-aware training framework with structured, data-aligned noise injection tailored for video diffusion. Our two operators—*Batch-Centered Noise Injection (BCNI)* and *Spectrum-Aware Contextual Noise (SACN)*—align perturbations with batch semantics or spectral dynamics to preserve coherence. CAT-Video yields substantial gains: BCNI reduces FVD by **31.9%** on WebVid-2M, MSR-VTT, and MSVD, while SACN improves UCF-101 by **12.3%**, outperforming Gaussian, Uniform, and even large diffusion baselines like DEMO (2.3B) and Lavie (3B) despite training on $\\mathbf{5}\\times$ less data. Ablations confirm the unique value of low-rank, data-aligned noise, and theory establishes why these operators tighten robustness and generalization bounds. CAT-Video thus sets a new framework for robust video diffusion, and our experiments show that it can also be extended to autoregressive generation and multimodal video understanding LLMs."},"_bibtex":{"value":"@misc{\nmaduabuchi2026catvideo,\ntitle={{CAT}-{VIDEO}: {CORRUPTION}-{AWARE} {TRAINING} {FOR} {ROBUST} {VIDEO} {DIFFUSION} {MODELS}},\nauthor={Chika Maduabuchi and Hao Chen and Yujin Han and Jindong Wang},\nyear={2026},\nurl={https://openreview.net/forum?id=unZhwukf0T}\n}"},"title":{"value":"CAT-VIDEO: CORRUPTION-AWARE TRAINING FOR ROBUST VIDEO DIFFUSION MODELS"},"pdf":{"value":"/pdf/36b779e3109b1503084b84e3c1c30d2bc4985918.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"maduabuchi|catvideo_corruptionaware_training_for_robust_video_diffusion_models"},"authorids":{"value":["~Chika_Maduabuchi1","~Hao_Chen15","~Yujin_Han1","~Jindong_Wang4"]},"authors":{"value":["Chika Maduabuchi","Hao Chen","Yujin Han","Jindong Wang"]}},"tmdate":1770805109321,"tcdate":1758298523914,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission19699/Authors"],"signatures":["ICLR.cc/2026/Conference/Submission19699/Authors"],"forum":"unZhwukf0T","license":"CC BY 4.0","number":19699,"cdate":1758298523914,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/-/Submission","ICLR.cc/2026/Conference/-/Post_Submission","ICLR.cc/2026/Conference/Submission19699/-/Full_Submission","ICLR.cc/2026/Conference/-/Edit"],"mdate":1770805109321,"odate":1759896705795,"domain":"ICLR.cc/2026/Conference","id":"unZhwukf0T","version":2},{"content":{"summary":{"value":"This paper identifies the two main problems in automated video editing (lack of coherence in long-form narratives and inability to handle diverse tasks), and proposes VideoAgent, an all-in-one agentic framework with these main contributions:\n\n* a method for video shot creation with cross-model and context on the full narrative of a target output video\n* an orchestration framework that can coordinate a large set of agents to generate a final video edit out of the created video shots\n* a new video edit benchmark (VideoEdit)\n\nComplete evaluation follows using both VideoEdit and Shot2Story, which display high performance against the proposed baselines (87-98% success rate), reduced API costs (60% lower costs), and comparable results to those produced by human editors (just 4% below human edits).\n\nThe paper also provides detailed prompts and pseudocode for each of the agents used by VideoAgent, and includes source code for this task."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"* Have you considered using video generation models too, as an agent, for adding shots not included in the input materials?\n\n* Could a much simpler version of this work approach the same quality? For example, could the graph be non-dynamic but fixed, with each node gated on a selector for whether the node agent needs to be applied or not? (this has been explored to some extent in Table 3 which removes Intent Parsing and Agent Graph elements, but I wonder specifically about an existing but fixed graph)\n\n* Have you considered ablating the specific list of agents, to measure how they rank compared to each other? this could inform which other agents are possible future additions for the system"},"rating":{"value":4},"details_of_ethics_concerns":{"value":"* The paper displays copyrighted work (e.g. SpongeBob in page 16)\n\n* The voice synthesis and voice cloning features could be used for impersonation\n\n* Possible impact on employment is not discussed; could this work be done in such a way that it allows human video editors to collaborate with the system?"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"The main strengths of this work are:\n\n* Sound engineering, combining a large amount of agents with a solid orchestration method\n* Great quantitative results, generally superior to those of the baselines and, especially, very close to human-created videos (with the caveats discussed in the weaknesses section)\n* Novel approach that combines multiple-agents with a long-narrative aware video shot process, and an orchestration framework with self-aware elements\n* Clear descriptions of the methods used\n* Exhaustive details on prompts and pseudocode used in the agents, and open source code, both of which allow for high reproducibility of the work"},"flag_for_ethics_review":{"value":["Yes, Privacy, security and safety","Yes, Legal compliance (e.g., GDPR, copyright, terms of use, web crawling policies)","Yes, Other reasons (please specify below)"]},"weaknesses":{"value":"The main weaknesses of this work are:\n\n* Lack of details about the human baselines. Just 4% below human performance is a very impressive result, however, which is greatly diminished by lack on details about how this human performance is measured. For example, who are the humans, what is their expertise, which tools have they used, for how long, etc.\n  * this point is really critical because the non-human baselines are based on systems that aren't designed to handle multi-modal video editing. So it's difficult to understand what the quality of this system is without a valid human comparison.\n  * (minor suggestion): Besides an ad-hoc human baseline of videos edited just for the purpose of this evaluation, one can also wonder how the system would compare against video edits seen on the wild, for different categories. What is the success rate against, for example, against fan edits of existing works.\n\n* Lack of actual examples in video format. The paper displays many examples as frame sequences, but given the nature of this work, the addition of examples in video format would be beneficial to this work. Being able to actually watch and listen to the videos produced by the system (and to compare them to the raw input materials) would provide a better understanding of the system quality.\n\n* While this work details the creation of a significant engineering system, with many agents and a solid orchestration method, the research contributions appear more incremental. To make the research contributions clearer, the paper could describe in more detail how novel aspects in the introduced orchestration system differ from those in other multi agent systems based on LLMs that also deal with graphs. The Related Work section at this time merely describes the application to multimodal video editing workflows which I don't think is enough novelty. \n\n* (minor) Lack of details regarding system latency (though API costs are provided), especially when compared against other baselines. Ideally the paper would include a plot with an axis for latency and another for each key metric, such that the tradeoffs between quality and performance can be better understood. (A similar plot for API costs would be interesting too, and given the reduced API costs of this system, good evidence in favor of this work)\n\n* (minor): Lack of mentions of key downsides for this approach, and potential future work. What do the failure modes look like? What is insightful about them?\n\n* (minor): The appendix may be excessive. The paper could improve by just listing a short summary of each agent behavior, and pointing to the supplementary materials for details."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1764358989896,"tcdate":1761239291987,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission6731/Reviewer_pQcp"],"signatures":["ICLR.cc/2026/Conference/Submission6731/Reviewer_pQcp"],"forum":"cTqGsLYkRl","number":1,"license":"CC BY 4.0","cdate":1761239291987,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission6731/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1764358989896,"domain":"ICLR.cc/2026/Conference","replyto":"cTqGsLYkRl","id":"hriGuqzCrR","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Multi-Agent Systems; Multimodal Content Editing; Agentic AI"]},"supplementary_material":{"value":"/attachment/026381dc502ed23824895d4016100d97ae3b3431.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Video editing has become essential in digital media creation, yet existing automated systems are restricted to short segment processing and domain-specific tasks. They face two critical limitations: i) inability to handle diverse video comprehension and editing operations, and ii) lack of long-video understanding for coherent narrative creation. We propose VideoAgent, an all-in-one agentic framework addressing these challenges through two key innovations. First, we develop automated video shot creation with shot planning agents for coherent narratives and cross-modal retrieval for aligned visual content. Second, we design a multi-agent orchestration framework integrating over thirty specialized editing agents. Intent parsing filters relevant tools while self-reflective graph orchestration assembles complex editing pipelines. Extensive experiments on our newly-proposed VideoEdit benchmark and public datasets demonstrate VideoAgent's superiority over existing multimodal LLMs and agentic systems. VideoAgent achieves 87-98% orchestration success rates while reducing API costs by 60%. Human evaluation across six video categories shows VideoAgent produces professional-quality content approaching human-level performance, with ratings only 4% below human-created videos."},"_bibtex":{"value":"@misc{\nzhou2026videoagent,\ntitle={VideoAgent: All-in-One Agentic Framework for Video Understanding and Editing},\nauthor={Hengji Zhou and Lingxuan Huang and KunpengTan and Si Wu and Lianghao Xia and Chao Huang},\nyear={2026},\nurl={https://openreview.net/forum?id=cTqGsLYkRl}\n}"},"title":{"value":"VideoAgent: All-in-One Agentic Framework for Video Understanding and Editing"},"pdf":{"value":"/pdf/d2174487eb6853747d5a8f5d7a4743a4ef1fd6b8.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"zhou|videoagent_allinone_agentic_framework_for_video_understanding_and_editing"},"authorids":{"value":["~Hengji_Zhou1","~Lingxuan_Huang1","~KunpengTan1","~Si_Wu5","~Lianghao_Xia1","~Chao_Huang4"]},"authors":{"value":["Hengji Zhou","Lingxuan Huang","KunpengTan","Si Wu","Lianghao Xia","Chao Huang"]}},"version":2},{"content":{"summary":{"value":"Zoom-Zero proposes a reinforced coarse-to-fine framework for grounded video question answering (GVQA), where the model first localizes query-relevant temporal segments and then “zooms in” for fine-grained visual verification. By introducing a zoom-in accuracy reward and token-selective credit assignment, it enhances evidence faithful temporal grounding and answer accuracy, outperforming prior GRPO-based LVLMs across GVQA and long-video benchmarks."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Please refer to weaknesses section."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. **Strong Results.** The proposed model achieves state-of-the-art performance on grounded video QA and long video understanding benchmarks, clearly demonstrating its effectiveness.\n2. **Clarity of Writing.** The paper is well-organized and clearly written, allowing readers to easily follow the technical details and rationale of the proposed approach.\n3. **Motivation for Token-Level Advantage Estimation.** The paper insightfully identifies and addresses a key limitation in prior GRPO-based methods of collapsing multiple rewards into a single scalar, offering a well-motivated solution through token-level advantage estimation."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. **Novelty compared to VideoChat-R1.** The method largely resembles those of VideoChat-R1.\n    1. How is the temporal grounding reward different from the IoU reward used in VideoChat-R1?\n    2. Similarly, how does the zoom accuracy reward differ from the accuracy or recall reward in VideoChat-R1?\n2. **Novelty compared to Qwen2.5-VL + frame selection methods.** Since Qwen2.5-VL inherently supports dynamic frame sampling, the zoom-in capability seems intrinsic to the base model rather than a novel contribution of Zoom-Zero. Therefore, frame selection methods applied to Qwen2.5-VL could achieve similar spatial zoom-in behavior as Zoom-Zero.\n3. **Comparison with related works.** Frame selection for long video understanding has been extensively studied [1, 2, 3, 4]. A more comprehensive comparison with these works should be provided.\n    \n    [1] Hu et al, M-LLM Based Video Frame Selection for Efficient Video Understanding, CVPR 2025\n    \n    [2] Zhang et al., Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs, ICCV 2025\n    \n    [3] Wu et al., AdaFrame: Adaptive Frame Selection for Fast Video Recognition, CVPR 2019\n    \n    [4] Tang et al., Adaptive Keyframe Sampling for Long Video Understanding, CVPR 2025\n    \n4. **Validation of ‘verifiability’.** The authors claim that the zoom-in accuracy reward “verifies that grounded segments contain requisite evidence to answer the query.” However, no concrete mechanism ensures that the grounded segments indeed contain sufficient information.\n    1. How does the presence of answer in zoomed region verified?\n    2. What are the actual statistics of zoomed regions that contain versus lack the required evidence?"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762919495149,"tcdate":1761894706473,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission7362/Reviewer_u5TQ"],"signatures":["ICLR.cc/2026/Conference/Submission7362/Reviewer_u5TQ"],"forum":"OWLcXnCFdB","number":3,"license":"CC BY 4.0","cdate":1761894706473,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission7362/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762919495149,"domain":"ICLR.cc/2026/Conference","replyto":"OWLcXnCFdB","id":"cKljfpuG1b","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Grounded Video Question Answering; Reinforcement Learning"]},"supplementary_material":{"value":"/attachment/87e1ec9cd230e5946b37949778649027278ba66a.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Grounded video question answering (GVQA) aims to localize relevant temporal segments in videos and generate accurate answers to a given question; however, large video-language models (LVLMs) exhibit limited temporal awareness. Although existing approaches based on Group Relative Policy Optimization (GRPO) attempt to improve temporal grounding, they still struggle to faithfully ground their answers in the relevant video evidence, leading to temporal mislocalization and hallucinations. In this work, we present **Zoom-Zero**, a coarse-to-fine framework that first localizes query-relevant segments and then temporally zooms into the most salient frames for finer-grained visual verification. Our method addresses the limits of GRPO for the GVQA task with *two key innovations*: **(i)** a zoom-in accuracy reward that validates the fidelity of temporal grounding prediction and facilitates fine-grained visual verification on grounded frames; **(ii)** token-selective credit assignment, which attributes rewards to the tokens responsible for temporal localization or answer generation, mitigating GRPO’s issue in handling multi-faceted reward signals. Our proposed method advances grounded video question answering, improving temporal grounding by 5.2\\% on NExT-GQA and 4.6\\% on ReXTime, while also enhancing average answer accuracy by 2.4\\%. Additionally, the coarse-to-fine zoom-in during inference further benefits long-form video understanding by preserving critical visual details without compromising global context, yielding an average improvement of 6.4\\% on long-video benchmarks. Our code will be publicly available."},"_bibtex":{"value":"@misc{\nshen2026zoomzero,\ntitle={Zoom-Zero: Reinforced Coarse-to-Fine Video Understanding via Temporal Zoom-in},\nauthor={Xiaoqian Shen and Min-Hung Chen and Yu-Chiang Frank Wang and Mohamed Elhoseiny and Ryo Hachiuma},\nyear={2026},\nurl={https://openreview.net/forum?id=OWLcXnCFdB}\n}"},"title":{"value":"Zoom-Zero: Reinforced Coarse-to-Fine Video Understanding via Temporal Zoom-in"},"pdf":{"value":"/pdf/0415e6d1e424b145784091af201341ad6ce66761.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"shen|zoomzero_reinforced_coarsetofine_video_understanding_via_temporal_zoomin"},"authorids":{"value":["~Xiaoqian_Shen3","~Min-Hung_Chen2","~Yu-Chiang_Frank_Wang2","~Mohamed_Elhoseiny1","~Ryo_Hachiuma1"]},"authors":{"value":["Xiaoqian Shen","Min-Hung Chen","Yu-Chiang Frank Wang","Mohamed Elhoseiny","Ryo Hachiuma"]}},"version":2},{"content":{"venue":{"value":"CVPR 2026"},"abstract":{"value":"Pose-guided video generation refers to controlling the motion of subjects in generated video through a sequence of poses. It enables precise control over subject motion and has important applications in animation. However, current pose-guided video generation methods are limited to accepting only human poses as input, thus generalizing poorly to pose of other subjects. To address this issue, we propose PoseAnything, a general pose-guided video generation framework capable of handling both human and non-human characters, supporting arbitrary skeletal inputs. To enhance consistency preservation during motion, we introduce Part-aware Temporal Coherence Module, which divides the subject into different parts, establishes part correspondences, and computes cross-attention between corresponding parts across frames to achieve fine-grained part-level consistency. Additionally, we propose Subject and Camera Motion Decoupled CFG, a novel guidance strategy that, for the first time, enables independent camera movement control in pose-guided video generation, by separately injecting subject and camera motion control information into the positive and negative anchors of CFG. Furthermore, we present XPose, a high-quality public dataset containing 50,000 non-human pose-video pairs, along with an automated pipeline for annotation and filtering. Extensive experiments demonstrate that PoseAnything significantly outperforms state-of-the-art methods in both effectiveness and generalization."},"_bibtex":{"value":"@inproceedings{\nwang2026poseanything,\ntitle={PoseAnything: General Pose-guided Video Generation with Part-aware Temporal Coherence},\nauthor={Ruiyan Wang and Teng Hu and Kaihui Huang and Zihan Su and Ran Yi and Lizhuang Ma},\nbooktitle={Conference on Computer Vision and Pattern Recognition 2026},\nyear={2026},\nurl={https://openreview.net/forum?id=6LopOvSAQH}\n}"},"title":{"value":"PoseAnything: General Pose-guided Video Generation with Part-aware Temporal Coherence"},"pdf":{"value":"https://openaccess.thecvf.com/content/CVPR2026/papers/Wang_PoseAnything_General_Pose-guided_Video_Generation_with_Part-aware_Temporal_Coherence_CVPR_2026_paper.pdf"},"venueid":{"value":"thecvf.com/CVPR/2026/Conference"},"paperhash":{"value":"wang|poseanything_general_poseguided_video_generation_with_partaware_temporal_coherence"},"authorids":{"value":["~Ruiyan_Wang5","~Teng_Hu1","~Kaihui_Huang2","~Zihan_Su4","~Ran_Yi1","~Lizhuang_Ma1"]},"authors":{"value":["Ruiyan Wang","Teng Hu","Kaihui Huang","Zihan Su","Ran Yi","Lizhuang Ma"]}},"tmdate":1789656609013,"pdate":1789656438045,"tcdate":1765221302879,"writers":["thecvf.com/CVPR/2026/Conference","thecvf.com/CVPR/2026/Conference/Submission35428/Authors"],"signatures":["thecvf.com/CVPR/2026/Conference/Submission35428/Authors"],"forum":"6LopOvSAQH","license":"CC BY 4.0","number":35428,"cdate":1765221302879,"readers":["everyone"],"invitations":["thecvf.com/CVPR/2026/Conference/-/Submission","thecvf.com/CVPR/2026/Conference/Submission35428/-/Full_Submission","thecvf.com/CVPR/2026/Conference/-/Post_Submission","thecvf.com/CVPR/2026/Conference/Submission35428/-/Supplementary_Material","thecvf.com/CVPR/2026/Conference/-/Edit","thecvf.com/CVPR/2026/Conference/-/Compute_Flag"],"mdate":1789656609013,"odate":1789656438045,"domain":"thecvf.com/CVPR/2026/Conference","id":"6LopOvSAQH","version":2},{"content":{"TLDR":{"value":"We introduce a new task of recognizing chiral (temporally opposite) actions; we propose a self-supervised recipe to adapt image models to obtain compact time-sensitive video descriptors."},"venue":{"value":"NeurIPS 2025 poster"},"keywords":{"value":["Time-sensitive video representations","Time chirality","Action recognition"]},"supplementary_material":{"value":"/attachment/e7a0751eb0b1ae40fa2e7728b1560cae6deb0497.zip"},"primary_area":{"value":"applications"},"abstract":{"value":"Our objective is to develop compact video representations that are sensitive to visual change over time. To measure such time-sensitivity, we introduce a new task: chiral action recognition, where one needs to distinguish between a pair of temporally opposite actions, such as “opening vs. closing a door\", “approaching vs. moving away from something\", “folding vs. unfolding paper\", etc. Such actions (i) occur frequently in everyday life, (ii) require understanding of simple visual change over time (in object state, size, spatial position, count . . . ), and (iii) are known to be poorly represented by many video embeddings. Our goal is to build time aware video representations which offer linear separability between these chiral pairs. To that end, we propose a self-supervised adaptation recipe to inject time-sensitivity into a sequence of frozen image features. Our model is based on an auto-encoder with a latent space with inductive bias inspired by perceptual straightening. We show that this results in a compact but time-sensitive video representation for the proposed task across three datasets: Something-Something, EPIC-Kitchens, and Charade. Our method (i) outperforms much larger video models pre-trained on large-scale video datasets, and (ii) leads to an improvement in classification performance on standard benchmarks when combined with these existing models."},"_bibtex":{"value":"@inproceedings{\nbagad2025chirality,\ntitle={Chirality in Action: Time-Aware Video Representation Learning by Latent Straightening},\nauthor={Piyush Nitin Bagad and Andrew Zisserman},\nbooktitle={The Thirty-ninth Annual Conference on Neural Information Processing Systems},\nyear={2025},\nurl={https://openreview.net/forum?id=hchAR53gA0}\n}"},"title":{"value":"Chirality in Action: Time-Aware Video Representation Learning by Latent Straightening"},"pdf":{"value":"/pdf/5479e4c819d0187c36b0a6766d144285973c2d50.pdf"},"venueid":{"value":"NeurIPS.cc/2025/Conference"},"paperhash":{"value":"bagad|chirality_in_action_timeaware_video_representation_learning_by_latent_straightening"},"authorids":{"value":["~Piyush_Nitin_Bagad1","~Andrew_Zisserman1"]},"authors":{"value":["Piyush Nitin Bagad","Andrew Zisserman"]}},"tmdate":1783627305617,"pdate":1758217347036,"tcdate":1746978862318,"writers":["NeurIPS.cc/2025/Conference","NeurIPS.cc/2025/Conference/Submission22294/Authors"],"signatures":["NeurIPS.cc/2025/Conference/Submission22294/Authors"],"forum":"hchAR53gA0","license":"CC BY 4.0","number":22294,"cdate":1746978862318,"readers":["everyone"],"invitations":["NeurIPS.cc/2025/Conference/-/Submission","NeurIPS.cc/2025/Conference/-/Post_Submission","NeurIPS.cc/2025/Conference/Submission22294/-/Full_Submission","NeurIPS.cc/2025/Conference/Submission22294/-/Supplementary_Material","NeurIPS.cc/2025/Conference/-/Edit","NeurIPS.cc/2025/Conference/Submission22294/-/Camera_Ready_Revision"],"mdate":1783627305617,"odate":1761705025767,"domain":"NeurIPS.cc/2025/Conference","id":"hchAR53gA0","version":2},{"content":{"venue":{"value":"CoRR 2025"},"pdf":{"value":"http://arxiv.org/pdf/2509.08502v2"},"venueid":{"value":"dblp.org/journals/CORR/2025"},"paperhash":{"value":"bagad|chirality_in_action_timeaware_video_representation_learning_by_latent_straightening"},"authorids":{"value":["~Piyush_Bagad1","~Andrew_Zisserman1"]},"html":{"value":"https://doi.org/10.48550/arXiv.2509.08502"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2509-08502,\n  publtype={informal},\n  author={Piyush Bagad and Andrew Zisserman},\n  title={Chirality in Action: Time-Aware Video Representation Learning by Latent Straightening},\n  year={2025},\n  month={September},\n  cdate={1756684800000},\n  journal={CoRR},\n  volume={abs/2509.08502},\n  url={https://doi.org/10.48550/arXiv.2509.08502}\n}\n"},"abstract":{"value":"Our objective is to develop compact video representations that are sensitive to visual change over time. To measure such time-sensitivity, we introduce a new task: chiral action recognition, where one needs to distinguish between a pair of temporally opposite actions, such as \"opening vs. closing a door\", \"approaching vs. moving away from something\", \"folding vs. unfolding paper\", etc. Such actions (i) occur frequently in everyday life, (ii) require understanding of simple visual change over time (in object state, size, spatial position, count . . . ), and (iii) are known to be poorly represented by many video embeddings. Our goal is to build time aware video representations which offer linear separability between these chiral pairs. To that end, we propose a self-supervised adaptation recipe to inject time-sensitivity into a sequence of frozen image features. Our model is based on an auto-encoder with a latent space with inductive bias inspired by perceptual straightening. We show that this results in a compact but time-sensitive video representation for the proposed task across three datasets: Something-Something, EPIC-Kitchens, and Charade. Our method (i) outperforms much larger video models pre-trained on large-scale video datasets, and (ii) leads to an improvement in classification performance on standard benchmarks when combined with these existing models."},"title":{"value":"Chirality in Action: Time-Aware Video Representation Learning by Latent Straightening"},"authors":{"value":["Piyush Bagad","Andrew Zisserman"]}},"tmdate":1768557494437,"pdate":1735689600000,"externalIds":["dblp:journals/corr/abs-2509-08502"],"tcdate":1762691782613,"writers":["~"],"signatures":["~Piyush_Nitin_Bagad1"],"forum":"sngFioZAf4","license":"CC BY-SA 4.0","number":682337,"cdate":1756684800000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1768557494437,"domain":"DBLP.org","id":"sngFioZAf4","version":2},{"content":{"summary":{"value":"This work focuses on video undersdtanding. It analyzed current drawback for Vision-Language Model on video training and proposed to extend Vision-Language Model to support video with pooling operations. Experiments prove effectiveness on several video datasets."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"More studied are recommended to futher varified the proposed strategy. Please refer to the questions on the Weakness section."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. This work proposes to extend current Vision-Language Model to support video, which is important and promising.\n2. Extensive experiments on video-based benchmarks prove the effectiveness.\n3. The overall method is simple and easy to follow."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"It is interesting to find training with pooling can help video training. However, there could be some concerns regarding this paradigm that need to be solved. \n\n1. Whether this phenomenon is caused by the low-quality video training data? Because most of the visual tokens in video dataset could be redundant for simple video caption. A suggestion is to use high-quality data for model training, like LLaVA-Video [A] data. \n\n2. Whether the pooling strategy harms the performance on image? It's good to find the authors counduct extensive experiments for video. However, the performance on image-based settings is still not clear. Considering add image-based experiments to make it more clear.\n\n3. What's the baseline used in the experiments? Considering add experiments over the baseline to make the improvement more clear.\n\n4. LLaVA-Next-Video 34b in Table 4 should also be added in Table 2 and 3.\n\n[A] Video Instruction Tuning With Synthetic Data, 2024"}},"nonreaders":[],"tmdate":1731428035349,"tcdate":1730228794996,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission4096/Reviewer_vomH"],"signatures":["ICLR.cc/2025/Conference/Submission4096/Reviewer_vomH"],"forum":"Rs8fLyaOer","number":2,"license":"CC BY 4.0","cdate":1730228794996,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission4096/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428035349,"domain":"ICLR.cc/2025/Conference","replyto":"Rs8fLyaOer","id":"82kGThbNWw","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video Understanding;Parameter-efficient;Pooling"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Vision-language pre-training has significantly elevated performance across a wide range of image-language applications. Yet, the pre-training process for video-related tasks demands exceptionally large computational and data resources, which hinders the progress of video-language models. This paper investigates a straightforward, highly efficient, and resource-light approach to adapting an existing image-language pre-trained model for dense video understanding. Our preliminary experiments reveal that directly fine-tuning pre-trained image-language models with multiple frames as inputs on video datasets leads to performance saturation or even a drop. Our further investigation reveals that it is largely attributed to the bias of learned high-norm visual features.  Motivated by this finding, we propose a simple but effective pooling strategy to smooth the feature distribution along the temporal dimension and thus reduce the dominant impacts from the extreme features. The new model is termed Pooling LLaVA, or PLLaVA in short.  PLLaVA achieves new state-of-the-art performance on modern benchmark datasets for both video question-answer and captioning tasks. Notably, on the recent popular Video ChatGPT benchmark, PLLaVA achieves a score of 3.25 out of 5 on average of five evaluated dimensions. On the latest multi-choice benchmark MVBench, PLLaVA achieves 58.1\\% accuracy on average across 20 sub-tasks, 14.5\\% higher than GPT4V (IG-VLM)."},"_bibtex":{"value":"@misc{\nxu2025pllava,\ntitle={{PLL}a{VA}: Parameter-efficient {LL}a{VA} Extension from Image to Video Understanding},\nauthor={Lin Xu and Yilin Zhao and Daquan Zhou and Zhijie Lin and See-Kiong Ng and Jiashi Feng},\nyear={2025},\nurl={https://openreview.net/forum?id=Rs8fLyaOer}\n}"},"title":{"value":"PLLaVA: Parameter-efficient LLaVA Extension from Image to Video Understanding"},"pdf":{"value":"/pdf/e0c9f7447c0f6a644f1d455ff1811a92b895eefb.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"xu|pllava_parameterefficient_llava_extension_from_image_to_video_understanding"},"authorids":{"value":["~Lin_Xu3","~Yilin_Zhao4","~Daquan_Zhou1","~Zhijie_Lin1","~See-Kiong_Ng1","~Jiashi_Feng1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Lin Xu","Yilin Zhao","Daquan Zhou","Zhijie Lin","See-Kiong Ng","Jiashi Feng"]}},"version":2},{"content":{"venue":{"value":"ICLR 2027 Conference Submission"},"TLDR":{"value":"We identify an entropy shortcut in LLM correctness probes that causes failures on low-entropy incorrect and high-entropy correct responses, and propose probes that improve robustness across entropy bands."},"keywords":{"value":["Uncertainty Estimation","Abstention","Large Language Model","Multimodal Language Models","Confidence Estimation","LLM Probe","Hallucination Detection","Semantic Entropy"]},"primary_area":{"value":"alignment, fairness, safety, privacy, and societal considerations"},"abstract":{"value":"Large language models (LLMs) achieve strong performance across textual, visual, and other multimodal tasks, yet they frequently hallucinate, producing outputs that are incorrect and unsupported. Preventing such failures is difficult, hence recent works have increasingly focused on detecting them. Two promising approaches that do not modify the model parameters are: (a) sampling the model at high temperature and measuring the entropy of the sampled responses, and (b) training a linear probe on hidden states to estimate the probability of a correct response. Entropy-based detection rests on the hypothesis that sampling entropy and correctness are inversely associated. In this work, we show that a sizable fraction of mispredictions exhibit low sampling entropy. More importantly, standard linear correctness probes fail to distinguish these low-entropy mispredictions from high-entropy correct predictions. We trace this failure to an entropy shortcut- the probes exploit the entropy–correctness association rather than learning a correctness signal that remains stable across entropy regimes. To address this, we propose linear detectors that predict correctness with reduced entropy shortcut. Experiments with four LLMs and five multimodal LLMs across eight datasets, including safety-critical domains, establish the entropy shortcut as a systematic failure mode for LLM probes and demonstrate substantially more robust correctness detection across both entropy-aligned and entropy-misaligned cases."},"_bibtex":{"value":"@inproceedings{\nanonymous2026entropy,\ntitle={Entropy Shortcut in {LLM} Correctness Probes},\nauthor={Anonymous},\nbooktitle={Submitted to The Fifteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=kgDnO2WYhp},\nnote={under review}\n}"},"title":{"value":"Entropy Shortcut in LLM Correctness Probes"},"pdf":{"value":"/pdf/f6a3919f6bf06a8c0bef5210fb9dc86c174dfd6a.pdf"},"venueid":{"value":"ICLR.cc/2027/Conference/Submission"}},"tmdate":1791219684333,"tcdate":1787295458366,"writers":["ICLR.cc/2027/Conference","ICLR.cc/2027/Conference/Submission2509/Authors"],"signatures":["ICLR.cc/2027/Conference/Submission2509/Authors"],"forum":"kgDnO2WYhp","license":"CC BY 4.0","number":2509,"cdate":1787295458366,"readers":["everyone"],"invitations":["ICLR.cc/2027/Conference/-/Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Bidding","ICLR.cc/2027/Conference/Submission2509/-/Full_Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Reviewing","ICLR.cc/2027/Conference/-/Edit"],"mdate":1791219684333,"odate":1791031328286,"domain":"ICLR.cc/2027/Conference","id":"kgDnO2WYhp","version":2},{"content":{"summary":{"value":"This paper addresses limitations in text-driven video editing using generative diffusion models, particularly the challenge of nuanced editing when constrained by pre-trained word embeddings. The authors propose a concept-augmented video editing approach that allows for flexible and stable video generation through the use of abstract conceptual pairs. The key components include concept-augmented textual inversion, which provides plug-and-play guidance for stable diffusion to better capture target attributes, and a dual prior supervision mechanism that enhances video stability and fidelity. By integrating these elements, the approach overcomes the disruptions caused by direct keyword alterations and allows for more nuanced, stylized editing. Experiments demonstrate that this method outperforms existing state-of-the-art techniques in generating stable and lifelike videos."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Please refer to the weakness section."},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The paper is easy to follow.\n2. The clarity of the writing.\n3. SCAM loss and TCAM loss sound novel.\n4. Convincing metric for qualitative evaluation: Masked Peek-Signal-Noise Ratio."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Diffusion based video editing has developed quickly and Tune-A-Video and Fatezero are not the strong baselines at this point. Plus, how is this paper different from DreamVideo [1] and MotionBooth [2] paper? I think these works could serve as better baselines?\n2. In L36, Zeroscope is not a video editing method but an open-sourced video diffusion model.\n3. Is there a reason why image diffusion (stable diffusion) is used as a backbone? I don't see reasons that the method can't be applied to video diffusion models and compared to video diffusion based methods.\n4. L242-245, can the authors show results?\n5. CLP consistency fails to reflect temporal consistency. Using other metrics like DINOv2 (like the ones from VBench[3]) or LIPIPS would be better.\n6. In terms of temporal consistency, the resulting videos of MotionDirector look better.\n7. The concept augmented textual inversion lacks technical contribution.\n\n[1] DreamVideo: Composing Your Dream Videos with Customized Subject and Motion, CVPR 2024\n\n[2] MotionBooth: Motion-Aware Customized Text-to-Video Generation, Arxiv 2024\n\n[3] VBench: Comprehensive Benchmark Suite for Video Generative Models, CVPR 2024"}},"nonreaders":[],"tmdate":1731427306434,"tcdate":1729513463029,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission716/Reviewer_ZGWC"],"signatures":["ICLR.cc/2025/Conference/Submission716/Reviewer_ZGWC"],"forum":"CxS8mlkOH7","number":1,"license":"CC BY 4.0","cdate":1729513463029,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission716/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427306434,"domain":"ICLR.cc/2025/Conference","replyto":"CxS8mlkOH7","id":"V5xDus4j4g","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Text-Guided Video Editing"]},"primary_area":{"value":"generative models"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Text-driven video editing utilizing generative diffusion models has garnered significant attention due to their potential applications. However, existing approaches are constrained by the limited word embeddings provided in pre-training, which hinders nuanced editing targeting open concepts with specific attributes. Directly altering the keywords in target prompts often results in unintended disruptions to the attention mechanisms. To achieve more flexible editing easily, this work proposes an improved concept-augmented video editing approach that generates diverse and stable target videos flexibly by devising abstract conceptual pairs. Specifically, the framework involves concept-augmented textual inversion and a dual prior supervision mechanism. The former enables plug-and-play guidance of stable diffusion for video editing, effectively capturing target attributes for more stylized results. The dual prior supervision mechanism significantly enhances video stability and fidelity. Comprehensive evaluations demonstrate that our approach generates more stable and lifelike videos, outperforming state-of-the-art methods. The anonymous code is available at \\href{https://anonymous.4open.science/w/STIVE-PAGE-B4D4/}{https://anonymous.4open.science/w/STIVE-PAGE-B4D4/}."},"_bibtex":{"value":"@misc{\nguo2025shaping,\ntitle={Shaping a Stabilized Video by Mitigating Unintended Changes for Concept-Augmented Video Editing},\nauthor={Mingce Guo and Jingxuan He and Shengeng Tang and Zhangye Wang and Lechao Cheng},\nyear={2025},\nurl={https://openreview.net/forum?id=CxS8mlkOH7}\n}"},"title":{"value":"Shaping a Stabilized Video by Mitigating Unintended Changes for Concept-Augmented Video Editing"},"pdf":{"value":"/pdf/6ee805d6f2ff14368b91a3b977f139f02d6d57a4.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"guo|shaping_a_stabilized_video_by_mitigating_unintended_changes_for_conceptaugmented_video_editing"},"authorids":{"value":["~Mingce_Guo1","~Jingxuan_He2","~Shengeng_Tang1","~Zhangye_Wang1","~Lechao_Cheng2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Mingce Guo","Jingxuan He","Shengeng Tang","Zhangye Wang","Lechao Cheng"]}},"version":2},{"content":{"summary":{"value":"The authors provide an interpretability analysis of the most-used SoTA sign language embedding representations with regard to how different phonological parameters are learned. In order to do this, they constract minimal pairs, within which only one parameter changes. Then, they extract feature vectors from the model’s penultimate layer based on the hypothesis that if a model is sensitive to phonological change, the cosine similarity between two minimally different signs should be significantly lower than two instances of the same sign (produced by different signers).\n\nThrough a thorough statistical analysis, they find out that (a) training on non-sign language action data cannot capture SL phonological sensitivity, (b) that the two types of architectures (pixel vs. poses) exhibit distinct biases in what aspects of phonology they are sensitive to and (c) model sensitivity aligns with human perception whereas (d) domain shifts may significantly affect the model representations. Finally they measure how phonological sensitivity compares to human perception and they provide a dendroid  visualization of the hierarchical representations learned by the models."},"ethics_flag":{"value":1},"reasons_to_reject":{"value":"* lack of reproducibility, test-set not publicly available (or at least no info about that)"},"review":{"value":"Overarching comments:\n\n * The authors could also use the term “interpretability” to describe their work, as is it is highly relevant to what they are doing.\n * The authors would be really encourage to make their minimal pairs and calculation scripts publicly available\n * Can the authors specify if and how their minimal pairs could be used in a way that would improve the performance of training those representations in the future?\n\n\nSpecific comments:\n\n * l. 62: The use of minimal pairs in related work deserves a separate longer reference, as a practice, and the relevant citation(s). \n * l. 66: “[automatically-scored reference-based] accuracy”\n * Section 2: please clarify that you only focus to the manual phonological parameters. Signs contain also non-manual elements which are also considered phonological parameters.\n * similar works regarding the handling of face and mouthings as signing linguistic elements by SL models has been investigated with ablation in https://doi.org/10.1145/3742886.3756718 and https://scholar.google.gr/scholar?oi=bibs&cluster=8939081797178628299 . Please consider citing \n * 92: Can the authors clarify, how exactly the Deaf author define whether a pair is “minimal”. Possibly it is helpful to say, that a minimal pair is where one phonological parameter changes whereas the other are the same. \n* The first two conclusions in Section 5.2 are actually one, so they could be united and shorted. (training on random weights vs. action data vs. ASL). \n\n\nMinor comments:\n\n * l. 24: “However…”: this sentence needs syntactic improvement (conjugation issue)\n * l. 46: “handshapes[, and how this] compares” \n * l. 119: please concatenate two parentheses (i.e. vision versus pose-based; Desai et al., 2023)  using sth like \\cite[i.e.vision versus pose-based][]{desaietal2023}\n * l. 124: unbalanced parentheses here too\n * The authors should try to increase the font size of Table 1. \n * Not sure about this, but please check whether Table 3 is supposed to float to the top of the page, as many conference templates require so."},"confidence":{"value":3},"reasons_to_accept":{"value":"* very innovative analysis that has been long missing from sign language recognition and translation research. \n * given the severe stagnation of the evaluation practices of new models, the results of this analysis consist a valuable contribution \n * expected to have impact to the understanding of these models\n * the paper is well-structured and well-written, and the methods and experiments seem scientific sound and well performed."},"rating":{"value":8},"questions_to_authors":{"value":"See above"},"title":{"value":"Well written paper shading some light on the underlying representations of commonly used SL embeddings"}},"parentInvitations":"colmweb.org/COLM/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1787537390662,"tcdate":1778251107228,"writers":["colmweb.org/COLM/2026/Conference","colmweb.org/COLM/2026/Conference/Submission2298/Reviewer_xqw7"],"signatures":["colmweb.org/COLM/2026/Conference/Submission2298/Reviewer_xqw7"],"forum":"8MoSlQbdS8","number":3,"license":"CC BY 4.0","cdate":1778251107228,"readers":["everyone"],"invitations":["colmweb.org/COLM/2026/Conference/Submission2298/-/Official_Review","colmweb.org/COLM/2026/Conference/-/Edit"],"mdate":1787537390662,"domain":"colmweb.org/COLM/2026/Conference","replyto":"8MoSlQbdS8","id":"8ApDBLW6eJ","forumContent":{"TLDR":{"value":"We use minimal pairs to evaluate the phonological sensitivity of sign language recognition models."},"venue":{"value":"COLM 2026"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the COLM Code of Ethics on https://colmweb.org/CoE.html"},"LLM_usage":{"value":"In accordance with the Policy on the use of Large Language Models in the Call for Papers https://colmweb.org/cfp.html, I certify that my submission discloses any substantive use of LLMs in the research and paper-writing process."},"keywords":{"value":["sign language","phonology","interpretability"]},"abstract":{"value":"Sign languages are compositional systems where meaning arises by combining sublexical phonological parameters, such as handshape, location, and movement. While deep learning models for Sign Language Recognition (SLR) have achieved increased performance on translation benchmarks, it remains unclear whether these models distinguish abstract phonological features or merely rely on low-level statistical correlations. This work evaluates the phonological perception of SLR models trained on American Sign Language (ASL) by probing phonological sensitivity using minimal pairs and evaluating representational alignment with human behavioral data. Our results reveal that SLR models exhibit emergent phonological sensitivity, but with clear architectural trade-offs: pose-based models are sensitive to handshape contrasts, while pixel-based models better capture location changes. Furthermore, pose-based models learn latent representations that correlate with human perceptual similarity judgments (r≈0.49). These findings suggest that while SLR models exhibit emergent phonology, current training paradigms are insufficient to scale them beyond their architectural inductive biases."},"_bibtex":{"value":"@inproceedings{\nyin2026phonological,\ntitle={Phonological Perception of Sign Language Models},\nauthor={Kayo Yin and Jessica Carter and Alex Xijie Lu and Annemarie Kocab},\nbooktitle={Third Conference on Language Modeling},\nyear={2026},\nurl={https://openreview.net/forum?id=8MoSlQbdS8}\n}"},"title":{"value":"Phonological Perception of Sign Language Models"},"pdf":{"value":"/pdf/319cca5e40aad8f57bdcc95f6eb7e06b20525d32.pdf"},"venueid":{"value":"colmweb.org/COLM/2026/Conference"},"paperhash":{"value":"yin|phonological_perception_of_sign_language_models"},"authorids":{"value":["~Kayo_Yin1","~Jessica_Carter1","~Alex_Xijie_Lu1","~Annemarie_Kocab1"]},"author_guide":{"value":"I certify that this submission complies with the submission instructions as described on https://colmweb.org/AuthorGuide.html"},"authors":{"value":["Kayo Yin","Jessica Carter","Alex Xijie Lu","Annemarie Kocab"]}},"version":2},{"content":{"TLDR":{"value":"SSVL shortens long-horizon value propagation directly in primitive action-value learning using single-discount shortcut backups and selective reward compression, while retaining a flat goal-conditioned policy at deployment."},"venue":{"value":"ICLR 2027 Conference Submission"},"keywords":{"value":["Offline Reinforcement Learning","Goal-Conditioned Reinforcement Learning","Long-Horizon Reinforcement Learning","Temporal Abstraction","Value Learning","Implicit Q-Learning"]},"primary_area":{"value":"reinforcement learning"},"abstract":{"value":"One source of difficulty in long-horizon offline goal-conditioned reinforcement learning (GCRL) is that sparse rewards must propagate through many Bellman updates before primitive actions become well separated in value. We introduce Selective Shortcut Value Learning (SSVL), which shortens this propagation directly in primitive action-value learning using stopped trajectory segments and a single-discount shortcut continuation. A frozen shortcut-aligned reference selectively compresses segment rewards, while an auxiliary Shortcut Value Consistency (SVC) loss helps the separately learned IQL state value preserve shortcut relations; both are used only during training, and deployment remains a single primitive goal-conditioned policy. We analyze how shorter propagation strengthens progress-state separation and separately characterize the logged-continuation conditions under which shortcut regression targets preserve primitive-action preference. Across diverse state-based and visual OGBench tasks, SSVL is competitive with strong offline GCRL methods, including hierarchical approaches."},"_bibtex":{"value":"@inproceedings{\nanonymous2026selective,\ntitle={Selective Shortcut Value Learning for Long-Horizon Offline Goal-Conditioned {RL}},\nauthor={Anonymous},\nbooktitle={Submitted to The Fifteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=kfaXAaXdWH},\nnote={under review}\n}"},"title":{"value":"Selective Shortcut Value Learning for Long-Horizon Offline Goal-Conditioned RL"},"pdf":{"value":"/pdf/a7a8c8d1e5a75e7e8117a9aa20a27b4c676b3ddd.pdf"},"venueid":{"value":"ICLR.cc/2027/Conference/Submission"}},"tmdate":1791229945363,"tcdate":1789505886385,"writers":["ICLR.cc/2027/Conference","ICLR.cc/2027/Conference/Submission22335/Authors"],"signatures":["ICLR.cc/2027/Conference/Submission22335/Authors"],"forum":"kfaXAaXdWH","license":"CC BY 4.0","number":22335,"cdate":1789505886385,"readers":["everyone"],"invitations":["ICLR.cc/2027/Conference/-/Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Bidding","ICLR.cc/2027/Conference/-/Submission_Change_Before_Reviewing","ICLR.cc/2027/Conference/-/Edit"],"mdate":1791229945363,"odate":1791031328286,"domain":"ICLR.cc/2027/Conference","id":"kfaXAaXdWH","version":2},{"content":{"venue":{"value":"NeurIPS 2025 poster"},"keywords":{"value":["egocentric vision","video-language pretraining","multi-task learning","3D-aware video representation"]},"supplementary_material":{"value":"/attachment/51d585636157a26856e5c4258dab621fd2cbfc61.zip"},"primary_area":{"value":"applications"},"abstract":{"value":"Egocentric video-language pretraining has significantly advanced video representation learning. Humans perceive and interact with a fully 3D world, developing spatial awareness that extends beyond text-based understanding. However, most previous works learn from 1D text or 2D visual cues, such as bounding boxes, which inherently lack 3D understanding.  To bridge this gap, we introduce EgoDTM, an Egocentric Depth- and Text-aware Model, jointly trained through large-scale 3D-aware video pretraining and video-text contrastive learning. EgoDTM incorporates a lightweight 3D-aware decoder to efficiently learn 3D-awareness from pseudo depth maps generated by depth estimation models. To further facilitate 3D-aware video pretraining, we enrich the original brief captions with hand-object visual cues by organically combining several foundation models. Extensive experiments demonstrate EgoDTM's superior performance across diverse downstream tasks, highlighting its superior 3D-aware visual understanding. Code: \\url{https://github.com/xuboshen/EgoDTM}."},"_bibtex":{"value":"@inproceedings{\nxu2025egodtm,\ntitle={Ego{DTM}: Towards 3D-Aware Egocentric Video-Language Pretraining},\nauthor={Boshen Xu and Yuting Mei and liu xinbi and Sipeng Zheng and Qin Jin},\nbooktitle={The Thirty-ninth Annual Conference on Neural Information Processing Systems},\nyear={2025},\nurl={https://openreview.net/forum?id=r2GebY4MnU}\n}"},"title":{"value":"EgoDTM: Towards 3D-Aware Egocentric Video-Language Pretraining"},"pdf":{"value":"/pdf/be67ccce51579fa153f0fc8f563e204f372d6d4d.pdf"},"venueid":{"value":"NeurIPS.cc/2025/Conference"},"paperhash":{"value":"xu|egodtm_towards_3daware_egocentric_videolanguage_pretraining"},"authorids":{"value":["~Boshen_Xu1","~Yuting_Mei2","~liu_xinbi1","~Sipeng_Zheng1","~Qin_Jin1"]},"authors":{"value":["Boshen Xu","Yuting Mei","liu xinbi","Sipeng Zheng","Qin Jin"]}},"tmdate":1783627355510,"pdate":1758217090423,"tcdate":1746883499382,"writers":["NeurIPS.cc/2025/Conference","NeurIPS.cc/2025/Conference/Submission16359/Authors"],"signatures":["NeurIPS.cc/2025/Conference/Submission16359/Authors"],"forum":"r2GebY4MnU","license":"CC BY 4.0","number":16359,"cdate":1746883499382,"readers":["everyone"],"invitations":["NeurIPS.cc/2025/Conference/-/Submission","NeurIPS.cc/2025/Conference/-/Post_Submission","NeurIPS.cc/2025/Conference/Submission16359/-/Full_Submission","NeurIPS.cc/2025/Conference/Submission16359/-/Supplementary_Material","NeurIPS.cc/2025/Conference/-/Edit","NeurIPS.cc/2025/Conference/Submission16359/-/Camera_Ready_Revision"],"mdate":1783627355510,"odate":1761704931849,"domain":"NeurIPS.cc/2025/Conference","id":"r2GebY4MnU","version":2},{"content":{"summary":{"value":"This paper introduces a novel model, MindGrapher, designed to reconstruct video stimuli from fMRI data. The model operates in two stages: (1) reconstructing the first frame of the video from fMRI, and (2) completing the full video guided by fMRI using the initial reconstructed frame. Additionally, the paper propose a new evaluation metric, dynamic content fidelity (DCF), to assess the temporal dynamics of the reconstructed videos. The proposed model outperforms previous state-of-the-art methods on both quantitative (video-based and frame-based) and qualitative (user study) evaluation metrics. Extensive ablation studies further demonstrate the effectiveness of each module within the model."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"The questions have all been listed in the \"Weaknesses\" section."},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":1},"contribution":{"value":2},"strengths":{"value":"1. The authors take a reasonable approach to address the problem by breaking down video reconstruction into two parts: (1) using established image reconstruction algorithms to decode the first frame, and (2) employing an image-to-video model to complete the subsequent frames based on the first frame.\n2. Comprehensive experiments, including a detailed ablation study, are conducted to thoroughly evaluate and analyze the effectiveness of the proposed method."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Writing weaknesses\n\n(1) In Section 3.2 (Line 267), the authors utilize the image-to-video model AnimateDiff to complete the remaining video frames based on the reconstructed first frame. It would be beneficial to introduce the AnimateDiff model in the Related Works section (Video Generation with Diffusion Models) and clarify its role in the proposed method.\n\n(2) A key challenge in video reconstruction, compared to image reconstruction, lies in learning and decoding dynamic information. Therefore, the authors should emphasize in Section 3.2 how fine-grained integration of dynamics is achieved in the video reconstruction process (Line 267). For instance, in Line 277, the authors should specify which model the self-attention mechanism belongs to—whether it is from Versatile Diffusion or AnimateDiff.\n\n(3) The visualization method described in Section 6 (Line 511) is unclear, which may hinder understanding. Please clarify the specific steps of the visualization process. For instance, when projecting onto a cortical flatmap, is it the weights of a certain model layer, as in Mind-Video, or the activation values being projected? If it is the activation values, how are their dimensions aligned with the number of voxels?\n\n2. Model and experimental design weaknesses\n\n(1) In Line 230 and Equation (2), the authors state that 'minimizing the similarity with non-corresponding frames across the entire dataset.' This implies that negative samples are not selected from each training batch but instead include all samples in the dataset, excluding the target sample. This undoubtedly broadens the range of negative samples, potentially introducing 'false negatives' that could impact model training[1]. I am curious as to why the authors' model can still converge under these circumstances.\n\n(2) The authors propose utilizing text annotated with Video-LLaVA as dynamic information for videos on line 246. This approach is highly questionable, as the annotations provided by Video-LLaVA primarily encompass semantic information. The authors are requested to elucidate the rationale and justification for their choice.\n\n3. Issues existing in the conclusion\n\n（1）In Section 3.3, the authors claim to introduce a dynamic-aware evaluation metric; however, the description starting from Line 290 indicates that this metric assesses the cosine similarity between the ground truth and the reconstructed results based on Video-LlaVA text annotations. This is, in fact, a semantic evaluation metric and fails to accurately measure dynamic content. To evaluate the similarity of dynamic information between two videos, I recommend using optical flow-related metrics, such as Average End-point Error (AE) or End-point Error (EPE).\n\n（2）In Line 468, the authors discuss the impact of generative models on performance. While their premise is sound, there are some issues with the experimental design. Upon reviewing the Mind-Video paper, I noted that it employs the Stable Diffusion V1.5 model, whereas this work utilizes Versatile Diffusion for decoding the first frame of the video. To determine whether the improvements in evaluation metrics are due to the proposed method or the generative model, I recommend that the authors replace Versatile Diffusion with Stable Diffusion V1.5 and observe how the evaluation metrics change.\n\n(3) In Line 523, the authors state, 'After the injection, we observe increased activation in higher cognitive networks, while activity in the visual cortex decreases.' However, as shown in Figure 6(b), the overall shading on the graph becomes darker after the injection, indicating a decrease in activation according to the legend. This result contradicts the authors' report, and I request an explanation for this discrepancy.\n\n\n\n[1] Xu H, Xie S, Tan X E, et al. Demystifying clip data[J]. arXiv preprint arXiv:2309.16671, 2023."}},"nonreaders":[],"tmdate":1731427864761,"tcdate":1729588377680,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission3510/Reviewer_BW5R"],"signatures":["ICLR.cc/2025/Conference/Submission3510/Reviewer_BW5R"],"forum":"JfKF7Pdigi","number":1,"license":"CC BY 4.0","cdate":1729588377680,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission3510/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427864761,"domain":"ICLR.cc/2025/Conference","replyto":"JfKF7Pdigi","id":"p7pvnFn4ko","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["fMRI; fMRI-to-Video; Dynamic-Aware; Video Reconstruction from Brain Activities;"]},"primary_area":{"value":"applications to neuroscience & cognitive science"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Existing methods for fMRI-to-video reconstruction typically focus on accurately reconstructing visual content ($i.e.$, appearance), neglecting dynamic event information. However, as highlighted in cognitive neurology, these key dynamic events significantly influence brain signal changes during video perception. In this article, we introduce Mindgrapher, a two-stream framework designed to address this gap by enhancing the reconstruction of dynamic-aware videos from fMRI data. Mindgrapher comprises $i)$ a visual content reconstruction stream, that improves the accuracy of the reconstructed visual content from sparsely distributed fMRI data through a temporal dynamics enrichment approach and multi-moment multimodal contrastive learning; $ii)$ a dynamics injection stream, that firstly crafts dynamic-aware fMRI embeddings and then integrates them into the reconstruction process via a fine-grained approach, thereby producing videos that effectively perceive dynamic events. Moreover, to address the lack of suitable metrics for evaluating dynamic event information, we introduce a new evaluation metric named dynamic content fidelity (DCF), which measures how accurately dynamic events within the video are reconstructed. Upon evaluation with a publicly available fMRI dataset, Mindgrapher outperforms the state-of-the-arts on all metrics, $i.e.$, semantic classification accuracy, structural similarity index, and DCF. The reconstructed video results are available on our web page. Code shall be released."},"_bibtex":{"value":"@misc{\nquan2024mindgrapher,\ntitle={MindGrapher: Dynamic-Aware f{MRI}-to-Video Reconstruction},\nauthor={Ruijie Quan and Wensong Song and Liulei Li and Wenguan Wang and Yi Yang},\nyear={2024},\nurl={https://openreview.net/forum?id=JfKF7Pdigi}\n}"},"title":{"value":"MindGrapher: Dynamic-Aware fMRI-to-Video Reconstruction"},"pdf":{"value":"/pdf/4ddf08c68ba61444352b3a6245d4f2283624e38a.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"quan|mindgrapher_dynamicaware_fmritovideo_reconstruction"},"authorids":{"value":["~Ruijie_Quan1","~Wensong_Song1","~Liulei_Li1","~Wenguan_Wang4","~Yi_Yang4"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Ruijie Quan","Wensong Song","Liulei Li","Wenguan Wang","Yi Yang"]}},"version":2},{"content":{"summary":{"value":"The authors evaluate Kimi-K2.5 on the CT-RATE benchmark for two primary tasks: radiology report generation and multiple-choice visual question answering (VQA). Since the model is designed for video rather than 3D volumes, CT scans are resampled and converted into 256-slice axial video sequences using the RAVE framework. Results indicate that while CT-CHAT exhibits superior alignment with the specific reporting style of the dataset, Kimi-K2.5 matches or exceeds its performance on certain macro-classification metrics ($F1$ of $0.178$ vs $0.169$). The model also demonstrates a meaningful zero-shot VQA accuracy of $77.4\\%$, compared to the specialized reference of $91.2\\%$"},"review":{"value":"This paper provides a high-quality assessment of the current boundaries between generalist and specialist AI in medical imaging by leveraging video-based architectures to process volumetric data. The methodology is original and technically sound, utilizing the 1T-parameter Kimi-K2.5 model to achieve impressive zero-shot results that include temporal grounding, where the model links specific findings to time points in the slice sequence. While there remains a notable performance gap in VQA accuracy and reporting style compared to domain-specific models like CT-CHAT, the findings are significant for demonstrating that general-purpose pretraining on natural video contains transferable knowledge for complex 3D medical reasoning."},"strengths":{"value":"The study successfully proves that general-purpose pretraining on natural video and text contains transferable knowledge for medical 3D reasoning. The model's ability to maintain competitive clinical macro-precision and F1 scores in a zero-shot setting is particularly impressive and suggests that foundation models could serve as strong base-learners for medical applications."},"weaknesses":{"value":"The primary weakness is the inherent limitation of \"video-style\" interpretation, which may miss subtle voxel-level details required for highly specialized diagnoses compared to native 3D architectures. Additionally, the paper does not explore if fine-tuning Kimi-K2.5 would yield a \"best-of-both-worlds\" result, leaving the ceiling of this approach unknown"},"confidence":{"value":4},"rating":{"value":4},"justification_of_rating":{"value":"The paper provides valuable insights into the transferability of general AI to medical domains. The findings are relevant to the MIDL community as they challenge the necessity of ground-up domain training for all tasks, though the superiority of specialized models like CT-CHAT for specific reporting styles remains clear"},"title":{"value":"Review"}},"parentInvitations":"MIDL.io/2026/Short_Papers/-/Official_Review","nonreaders":[],"tmdate":1778421297587,"tcdate":1777754169377,"writers":["MIDL.io/2026/Short_Papers","MIDL.io/2026/Short_Papers/Submission88/Reviewer_XVc7"],"signatures":["MIDL.io/2026/Short_Papers/Submission88/Reviewer_XVc7"],"forum":"j6t3IKOfC2","number":1,"license":"CC BY 4.0","cdate":1777754169377,"readers":["everyone"],"invitations":["MIDL.io/2026/Short_Papers/Submission88/-/Official_Review","MIDL.io/2026/Short_Papers/-/Edit"],"mdate":1778421297587,"domain":"MIDL.io/2026/Short_Papers","replyto":"j6t3IKOfC2","id":"ODsfnjBhiY","forumContent":{"venue":{"value":"MIDL 2026 - Short Papers Poster"},"keywords":{"value":["CT","foundation models","report generation","visual question answering"]},"read_cfp_and_author_instructions":{"value":"Yes"},"originality_policy":{"value":"Yes"},"abstract":{"value":"General multimodal foundation models have shown strong performance in natural image and video understanding, but their ability to reason over 3D medical imaging remains underexplored. In this work, we investigate how well Kimi-K2.5, a novel general multimodal model, transfers to volumetric CT interpretation when chest CT scans are converted into video inputs. We evaluate on the CT-RATE validation set, a large-scale benchmark for 3D chest CT understanding, using two tasks: radiology report generation and multiple-choice visual question answering. As a domain-specialized reference, we compare against CT-CHAT. We find that CT-CHAT produces reports that align better with the reporting style reflected in the CT-RATE dataset, whereas both models achieve comparable performance on selected clinically relevant report evaluation metrics. For multiple-choice visual question answering, Kimi-K2.5 achieves meaningful zero-shot performance, although it does not match the performance of CT-CHAT. Overall, our results suggest that general multimodal pretraining can transfer to 3D CT reasoning."},"_bibtex":{"value":"@inproceedings{\nbuess2026how,\ntitle={How Well Does a General Multimodal Foundation Model Understand 3D {CT} Scans?},\nauthor={Lukas Buess and Franziska Weber and Andreas Maier},\nbooktitle={Medical Imaging with Deep Learning - Short Papers},\nyear={2026},\nurl={https://openreview.net/forum?id=j6t3IKOfC2}\n}"},"title":{"value":"How Well Does a General Multimodal Foundation Model Understand 3D CT Scans?"},"pdf":{"value":"/pdf/8c52525ea7ec470a84ef8b6095edafb9c8f72efd.pdf"},"visa":{"value":"No"},"single_blind_notice":{"value":"Yes"},"venueid":{"value":"MIDL.io/2026/Short_Papers"},"paperhash":{"value":"buess|how_well_does_a_general_multimodal_foundation_model_understand_3d_ct_scans"},"authorids":{"value":["~Lukas_Buess1","~Franziska_Weber2","~Andreas_Maier1"]},"registration":{"value":"Yes"},"authors":{"value":["Lukas Buess","Franziska Weber","Andreas Maier"]},"llm_policy_acknowledgment":{"value":"Yes"}},"version":2},{"content":{"venue":{"value":"CVPR 2023"},"pdf":{"value":"https://ieeexplore.ieee.org/iel7/10203037/10203050/10203697.pdf"},"venueid":{"value":"dblp.org/conf/CVPR/2023"},"paperhash":{"value":"wang|compressionaware_video_superresolution"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Yingwei_Wang:","~Takashi_Isobe1","~Xu_Jia1","https://dblp.org/search/pid/api?q=author:Xin_Tao_0001:","https://dblp.org/search/pid/api?q=author:Huchuan_Lu:","https://dblp.org/search/pid/api?q=author:Yu-Wing_Tai:"]},"html":{"value":"https://doi.org/10.1109/CVPR52729.2023.00200"},"_bibtex":{"value":"@inproceedings{DBLP:conf/cvpr/WangIJTLT23,\n  author={Yingwei Wang and Takashi Isobe and Xu Jia and Xin Tao and Huchuan Lu and Yu-Wing Tai},\n  title={Compression-Aware Video Super-Resolution},\n  year={2023},\n  cdate={1672531200000},\n  pages={2012-2021},\n  url={https://doi.org/10.1109/CVPR52729.2023.00200},\n  booktitle={CVPR},\n  crossref={conf/cvpr/2023}\n}\n"},"abstract":{"value":"Videos stored on mobile devices or delivered on the Internet are usually in compressed format and are of various unknown compression parameters, but most video super-resolution (VSR) methods often assume ideal inputs resulting in large performance gap between experimental settings and real-world applications. In spite of a few pioneering works being proposed recently to super-resolve the compressed videos, they are not specially designed to deal with videos of various levels of compression. In this paper, we propose a novel and practical compression-aware video super-resolution model, which could adapt its video enhancement process to the estimated compression level. A compression encoder is designed to model compression levels of input frames, and a base VSR model is then conditioned on the implicitly computed representation by inserting compression-aware modules. In addition, we propose to further strengthen the VSR model by taking full advantage of meta data that is embedded naturally in compressed video streams in the procedure of information fusion. Extensive experiments are conducted to demonstrate the effectiveness and efficiency of the proposed method on compressed VSR benchmarks. The codes will be available at https://github.com/aprBlue/CAVSR"},"title":{"value":"Compression-Aware Video Super-Resolution"},"authors":{"value":["Yingwei Wang","Takashi Isobe","Xu Jia","Xin Tao","Huchuan Lu","Yu-Wing Tai"]}},"tmdate":1731475658831,"pdate":1672531200000,"tcdate":1715742652965,"writers":["~"],"signatures":["~Takashi_Isobe1"],"forum":"DfcHb9rtiA","license":"CC BY-SA 4.0","number":12442,"cdate":1672531200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1731475658831,"domain":"DBLP.org","id":"DfcHb9rtiA","version":2},{"content":{"venue":{"value":"WACV 2025"},"pdf":{"value":"https://ieeexplore.ieee.org/iel8/10943266/10943193/10944101.pdf"},"venueid":{"value":"dblp.org/conf/WACV/2025"},"paperhash":{"value":"cao|towards_generalized_face_antispoofing_from_a_frequency_shortcut_view"},"authorids":{"value":["~Junyi_Cao1","https://dblp.org/search/pid/api?q=author:Chao_Ma_0004:"]},"html":{"value":"https://doi.org/10.1109/WACV61041.2025.00107"},"_bibtex":{"value":"@inproceedings{DBLP:conf/wacv/Cao025,\n  author={Junyi Cao and Chao Ma},\n  title={Towards Generalized Face Anti-Spoofing from a Frequency Shortcut View},\n  year={2025},\n  cdate={1735689600000},\n  pages={1005-1015},\n  url={https://doi.org/10.1109/WACV61041.2025.00107},\n  booktitle={WACV},\n  crossref={conf/wacv/2025}\n}\n"},"abstract":{"value":"The generalization capability of a face anti-spoofing (FAS) model is critical to its practicality in the real world. Recent studies have theoretically and empirically uncovered that neural networks tend to exploit easy-to-learn frequency sets for decisions. These simplicity-biased representations, depending on what best simplifies the training objective, may hamper generalization. This paper thus focuses on mitigating the frequency shortcut learning of prior FAS models for improved generalization. Specifically, we introduce a frequency-aware autoencoder to retain more frequency details in intermediate features via reconstruction, facilitating comprehensive judgment of FAS. Based on the encoder output, we propose a dynamic frequency masking mechanism to select and suppress the probable shortcut bands during training, enabling broader horizons on under-explored frequencies. Moreover, we employ a style inhibited modulation to weaken stylized information in frequency space to reduce the reliance on spurious style features. Experiment results on generalized FAS benchmarks verify the superiority of our framework over existing methods. Our code has been integrated into this project: https://github.com/VISION-SJTU/UniDefense."},"title":{"value":"Towards Generalized Face Anti-Spoofing from a Frequency Shortcut View"},"authors":{"value":["Junyi Cao","Chao Ma"]}},"tmdate":1747192375118,"pdate":1735689600000,"tcdate":1747192373386,"writers":["~"],"signatures":["~Junyi_Cao1"],"forum":"Bll0V4RCpm","license":"CC BY-SA 4.0","number":452872,"cdate":1735689600000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1747192375118,"domain":"DBLP.org","id":"Bll0V4RCpm","version":2},{"content":{"venue":{"value":"CoRR 2025"},"pdf":{"value":"https://arxiv.org/pdf/2506.04142v1"},"venueid":{"value":"dblp.org/journals/CORR/2025"},"paperhash":{"value":"zhu|establishing_trustworthy_llm_evaluation_via_shortcut_neuron_analysis"},"authorids":{"value":["","https://dblp.org/search/pid/api?q=author:Shangqing_Tu:","https://dblp.org/search/pid/api?q=author:Zhuoran_Jin:","~Lei_Hou2","https://dblp.org/search/pid/api?q=author:Juanzi_Li:","https://dblp.org/search/pid/api?q=author:Jun_Zhao_0001:"]},"html":{"value":"https://doi.org/10.48550/arXiv.2506.04142"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2506-04142,\n  publtype={informal},\n  author={Kejian Zhu and Shangqing Tu and Zhuoran Jin and Lei Hou and Juanzi Li and Jun Zhao},\n  title={Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis},\n  year={2025},\n  month={June},\n  cdate={1748736000000},\n  journal={CoRR},\n  volume={abs/2506.04142},\n  url={https://doi.org/10.48550/arXiv.2506.04142}\n}\n"},"abstract":{"value":"The development of large language models (LLMs) depends on trustworthy evaluation. However, most current evaluations rely on public benchmarks, which are prone to data contamination issues that significantly compromise fairness. Previous researches have focused on constructing dynamic benchmarks to address contamination. However, continuously building new benchmarks is costly and cyclical. In this work, we aim to tackle contamination by analyzing the mechanisms of contaminated models themselves. Through our experiments, we discover that the overestimation of contaminated models is likely due to parameters acquiring shortcut solutions in training. We further propose a novel method for identifying shortcut neurons through comparative and causal analysis. Building on this, we introduce an evaluation method called shortcut neuron patching to suppress shortcut neurons. Experiments validate the effectiveness of our approach in mitigating contamination. Additionally, our evaluation results exhibit a strong linear correlation with MixEval, a recently released trustworthy benchmark, achieving a Spearman coefficient ($ρ$) exceeding 0.95. This high correlation indicates that our method closely reveals true capabilities of the models and is trustworthy. We conduct further experiments to demonstrate the generalizability of our method across various benchmarks and hyperparameter settings. Code: https://github.com/GaryStack/Trustworthy-Evaluation"},"title":{"value":"Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis"},"authors":{"value":["Kejian Zhu","Shangqing Tu","Zhuoran Jin","Lei Hou","Juanzi Li","Jun Zhao"]}},"tmdate":1769654174937,"pdate":1767139200000,"externalIds":["dblp:journals/corr/abs-2506-04142"],"tcdate":1769062037909,"writers":["~"],"signatures":["~Lei_Hou2"],"forum":"eD1X3HwkkV","license":"CC BY-SA 4.0","number":787038,"cdate":1748736000000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1769654174937,"domain":"DBLP.org","id":"eD1X3HwkkV","version":2},{"content":{"venue":{"value":"ACL (1) 2025"},"pdf":{"value":"https://aclanthology.org/2025.acl-long.192.pdf"},"venueid":{"value":"dblp.org/conf/ACL/2025"},"paperhash":{"value":"zhu|establishing_trustworthy_llm_evaluation_via_shortcut_neuron_analysis"},"authorids":{"value":["~Kejian_Zhu1","https://dblp.org/search/pid/api?q=author:Shangqing_Tu:","https://dblp.org/search/pid/api?q=author:Zhuoran_Jin:","~Lei_Hou2","https://dblp.org/search/pid/api?q=author:Juanzi_Li:","https://dblp.org/search/pid/api?q=author:Jun_Zhao_0001:"]},"html":{"value":"https://aclanthology.org/2025.acl-long.192/"},"_bibtex":{"value":"@inproceedings{DBLP:conf/acl/ZhuTJHL025,\n  author={Kejian Zhu and Shangqing Tu and Zhuoran Jin and Lei Hou and Juanzi Li and Jun Zhao},\n  title={Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis},\n  year={2025},\n  cdate={1735689600000},\n  pages={3809-3822},\n  url={https://aclanthology.org/2025.acl-long.192/},\n  booktitle={ACL (1)},\n  crossref={conf/acl/2025-1}\n}\n"},"abstract":{"value":"The development of large language models (LLMs) depends on **trustworthy evaluation**. However, most current evaluations rely on public benchmarks, which are prone to data contamination issues that significantly compromise fairness. Previous researches have focused on constructing dynamic benchmarks to address contamination. However, continuously building new benchmarks is costly and cyclical.In this work, we aim to tackle contamination by analyzing the mechanisms of contaminated models themselves. Through our experiments, we discover that the overestimation of contaminated models is likely due to parameters acquiring shortcut solutions in training. We further propose a novel method for identifying shortcut neurons through **comparative and causal analysis**.Building on this, we introduce an evaluation method called **shortcut neuron patching** to suppress shortcut neurons. Experiments validate the effectiveness of our approach in mitigating contamination. Additionally, our evaluation results exhibit a strong linear correlation with MixEval, a recently released trustworthy benchmark, achieving a Spearman coefficient (𝜌) exceeding 0.95. This high correlation indicates that our method closely reveals true capabilities of the models and is trustworthy. We conduct further experiments to demonstrate the generalizability of our method across various benchmarks and hyperparameter settings. **Code**: https://github.com/GaryStack/Trustworthy-Evaluation."},"title":{"value":"Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis"},"authors":{"value":["Kejian Zhu","Shangqing Tu","Zhuoran Jin","Lei Hou","Juanzi Li","Jun Zhao"]}},"tmdate":1769654090592,"pdate":1767139200000,"externalIds":["dblp:conf/acl/ZhuTJHL025"],"tcdate":1769062034903,"writers":["~"],"signatures":["~Lei_Hou2"],"forum":"gRce78Frhc","license":"CC BY-SA 4.0","number":787017,"cdate":1735689600000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1769654090592,"domain":"DBLP.org","id":"gRce78Frhc","version":2},{"content":{"venue":{"value":"ICLR 2027 Conference Submission"},"keywords":{"value":["AI Security","Backdoor Defense","Backdoor Detection","Model Repair"]},"primary_area":{"value":"alignment, fairness, safety, privacy, and societal considerations"},"abstract":{"value":"Backdoor attacks pose a severe security threat to deep neural networks (DNNs) by implanting hidden behaviors activated by specific triggers. This threat is particularly concerning for downstream users acquiring models from untrusted sources, for whom reliable deployment entails assessing model integrity and, if compromised, localizing suspicious target classes and repairing the model. Backdoor learning introduces an additional trigger-to-target association beyond the intended task mapping, forming a shortcut that can leave class-specific irregularities in classifier-head weights and prediction responses, as well as directional changes in feature space. Building on these observations, we propose Shortcut-aware Class-wise Assessment and Revision (SCAR), a unified post-training framework for backdoor assessment and model recovery. SCAR first combines classifier-weight anomalies with perturbation-induced prediction preferences to assess model integrity and localize suspicious target classes. For each localized class, it then estimates a shortcut-related feature direction from class-specific weights and feature-wise variability, and selectively revises the corresponding classifier-head parameters along this direction to suppress backdoor behavior while preserving normal predictive utility. Extensive experiments across diverse datasets, attacks, and architectures demonstrate the effectiveness and generality of SCAR for both backdoor assessment and model recovery. Code will be released upon acceptance."},"_bibtex":{"value":"@inproceedings{\nanonymous2026scar,\ntitle={{SCAR}: Class-Wise Assessment and Shortcut-Guided Model Revision for Backdoor Defense},\nauthor={Anonymous},\nbooktitle={Submitted to The Fifteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=Eh9NQzBMRp},\nnote={under review}\n}"},"title":{"value":"SCAR: Class-Wise Assessment and Shortcut-Guided Model Revision for Backdoor Defense"},"pdf":{"value":"/pdf/dabebd50b535849c16a0264549846bc331b942d2.pdf"},"venueid":{"value":"ICLR.cc/2027/Conference/Submission"}},"tmdate":1791229480913,"tcdate":1789452629289,"writers":["ICLR.cc/2027/Conference","ICLR.cc/2027/Conference/Submission19904/Authors"],"signatures":["ICLR.cc/2027/Conference/Submission19904/Authors"],"forum":"Eh9NQzBMRp","license":"CC BY 4.0","number":19904,"cdate":1789452629289,"readers":["everyone"],"invitations":["ICLR.cc/2027/Conference/-/Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Bidding","ICLR.cc/2027/Conference/Submission19904/-/Full_Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Reviewing","ICLR.cc/2027/Conference/-/Edit"],"mdate":1791229480913,"odate":1791031328286,"domain":"ICLR.cc/2027/Conference","id":"Eh9NQzBMRp","version":2},{"content":{"summary":{"value":"The paper presents a benchmark dataset for the evaluation of video foundation models. It is focused on video classification (instead of video -to-text, video-question-answering, etc.). All the source videos in the dataset are selected from existing public datasets. The paper reported results for a number of existing video foundation models."},"soundness":{"value":2},"confidence":{"value":5},"questions":{"value":"Are video-language models regarded as video foundation models?\n\nCan the benchmark be used to evaluate video-language models such as GPT-4V, video-llava, etc?\n\nDoes the benchmark cover surveillance monitoring scenarios where there may be multiple people in the scene and each person performs a different action?\n\nDoes the benchmark cover the capability to evaluate how well a person performs certain actions?"},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":1},"strengths":{"value":"The dataset is useful for researchers to evaluate their video encoders."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The paper did not evaluate any of the large multimodality models such as GPT-4V, Gemini, video-llava, etc. It seems that the paper regards \"video foundation models\" as \"video encoders\", and the benchmark can only be used to evaluate video encoders.\n \nThe paper focuses on classification tasks. Modern video understanding tasks have shifted from classification to more fine-grained understanding such as (dense) text description, video question answering, etc. \n\nOnly a few (8 vidTab, 16 vidEB) frames are selected from each video. The benchmark won't be able to evaluate long video understanding capabilities, and it won't be able to capture the speed of motion."}},"nonreaders":[],"tmdate":1731427792155,"tcdate":1730004088227,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission7392/Reviewer_jNwF"],"signatures":["ICLR.cc/2025/Conference/Submission7392/Reviewer_jNwF"],"forum":"wMRFTQwp1d","number":3,"license":"CC BY 4.0","cdate":1730004088227,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission7392/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427792155,"domain":"ICLR.cc/2025/Conference","replyto":"wMRFTQwp1d","id":"3qBe2NS7lk","forumContent":{"TLDR":{"value":"A vision-centric evaluation method for video foundation models that is comprehensive, challenging, indicative, and low-cost."},"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video Understanding","Video Foundation Model","Benchmark"]},"primary_area":{"value":"datasets and benchmarks"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"With the accumulation of high-quality data and advancements in visual pretraining paradigms, recent Video Foundation Models (VFMs) have made significant progress, demonstrating remarkable performance on popular video understanding benchmarks. However, conventional benchmarks (e.g. Kinetics) and evaluation protocols are limited by their relatively poor diversity, high evaluation costs, and saturated performance metrics. In this work, we introduce a comprehensive benchmark suite to address these issues, namely **VideoEval**. We establish the **Vid**eo **T**ask **A**daption **B**enchmark (VidTAB) and the **Vid**eo **E**mbedding **B**enchmark (VidEB) from two perspectives: evaluating the task adaptability of VFMs under few-shot conditions and assessing their feature embedding's direct applicability to downstream tasks. With VideoEval, we conduct a large-scale study of 20 popular open-source vision foundation models. Our study reveals some insightful findings, 1) overall, current VFMs exhibit weak generalization across diverse tasks, 2) increasing video data, whether labeled or in video-text pairs, does not necessarily improve task performance, 3) the effectiveness of some pre-training paradigms may not be fully validated in previous benchmarks, and 4) combining different pre-training paradigms can help develop models with better generalization capabilities. We believe this study serves as a important complement to the current evaluation methods for VFMs and offers valuable insights for future research directions."},"_bibtex":{"value":"@misc{\nli2024videoeval,\ntitle={VideoEval: Comprehensive Benchmark Suite for Low-Cost Evaluation of Video Foundation Model},\nauthor={Xinhao Li and Zhenpeng Huang and Jing Wang and Kunchang Li and Limin Wang},\nyear={2024},\nurl={https://openreview.net/forum?id=wMRFTQwp1d}\n}"},"title":{"value":"VideoEval: Comprehensive Benchmark Suite for Low-Cost Evaluation of Video Foundation Model"},"pdf":{"value":"/pdf/652a4eb9e5126e3b38cd7820b52f1dacc0558c7a.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"li|videoeval_comprehensive_benchmark_suite_for_lowcost_evaluation_of_video_foundation_model"},"authorids":{"value":["~Xinhao_Li1","~Zhenpeng_Huang1","~Jing_Wang48","~Kunchang_Li1","~Limin_Wang1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Xinhao Li","Zhenpeng Huang","Jing Wang","Kunchang Li","Limin Wang"]}},"version":2},{"content":{"venue":{"value":"ICCV 2023"},"pdf":{"value":"https://ieeexplore.ieee.org/iel7/10376473/10376477/10378379.pdf"},"venueid":{"value":"dblp.org/conf/ICCV/2023"},"paperhash":{"value":"yoo|video_object_segmentationaware_video_frame_interpolation"},"authorids":{"value":["~Jun-Sang_Yoo2","~Hongjae_Lee1","~Seung-Won_Jung2"]},"html":{"value":"https://doi.org/10.1109/ICCV51070.2023.01132"},"_bibtex":{"value":"@inproceedings{DBLP:conf/iccv/0002LJ23,\n  author={Jun-Sang Yoo and Hongjae Lee and Seung-Won Jung},\n  title={Video Object Segmentation-aware Video Frame Interpolation},\n  year={2023},\n  cdate={1672531200000},\n  pages={12288-12299},\n  url={https://doi.org/10.1109/ICCV51070.2023.01132},\n  booktitle={ICCV},\n  crossref={conf/iccv/2023}\n}\n"},"abstract":{"value":"Video frame interpolation (VFI) is a very active re-search topic due to its broad applicability to many applications, including video enhancement, video encoding, and slow-motion effects. VFI methods have been advanced by improving the overall image quality for challenging sequences containing occlusions, large motion, and dynamic texture. This mainstream research direction neglects that foreground and background regions have different importance in perceptual image quality. Moreover, accurate synthesis of moving objects can be of utmost importance in computer vision applications. In this paper, we propose a video object segmentation (VOS)-aware training framework called VOS-VFI that allows VFI models to interpolate frames with more precise object boundaries. Specifically, we exploit VOS as an auxiliary task to help train VFI models by providing additional loss functions, including segmentation loss and bi-directional consistency loss. From extensive experiments, we demonstrate that VOS-VFI can boost the performance of existing VFI models by rendering clear object boundaries. Moreover, VOS-VFI displays its effectiveness on multiple benchmarks for different applications, including video object segmentation, object pose estimation, and visual tracking. The code is available at https://github.com/junsang7777/VOS-VFI"},"title":{"value":"Video Object Segmentation-aware Video Frame Interpolation"},"authors":{"value":["Jun-Sang Yoo","Hongjae Lee","Seung-Won Jung"]}},"tmdate":1747101431905,"pdate":1672531200000,"tcdate":1729052610394,"writers":["~"],"signatures":["~Hongjae_Lee1"],"forum":"pAf8sClSre","license":"CC BY-SA 4.0","number":152820,"cdate":1672531200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1747101431905,"domain":"DBLP.org","id":"pAf8sClSre","version":2},{"content":{"venue":{"value":"Signal Image Video Process. 2026"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/s11760-025-05090-8.pdf"},"venueid":{"value":"dblp.org/journals/SIVP/2026"},"paperhash":{"value":"kumari|a_hybrid_contextaware_video_violence_detection_framework_using_hierarchical_spatiotemporal_and_semantic_modeling"},"authorids":{"value":["",""]},"html":{"value":"https://doi.org/10.1007/s11760-025-05090-8"},"_bibtex":{"value":"@article{DBLP:journals/sivp/KumariK26,\n  author={Shobha Kumari and Vijay Kumar},\n  title={A hybrid context-aware video violence detection framework using hierarchical spatiotemporal and semantic modeling},\n  year={2026},\n  month={January},\n  cdate={1767225600000},\n  journal={Signal Image Video Process.},\n  volume={20},\n  number={1},\n  pages={12},\n  url={https://doi.org/10.1007/s11760-025-05090-8}\n}\n"},"abstract":{"value":"Video violence detection poses significant challenges due to the complex integration of spatial, temporal, and contextual features. Conventional methods, such as 3D Convolutional Neural Networks, Long Short-Term Memory Models, and You Only Look Once, exhibit limited scene-level semantic understanding, high computational costs, and poor generalization across diverse environments. This work proposes a novel hybrid context-aware framework that combines a lightweight MobileNetV2 with a Bidirectional Long Short-Term Memory (BiLSTM) network for efficient spatiotemporal feature extraction. YOLOv8 is used for real-time object detection and Bootstrapping Language-Image Pre-training is utilized for generating the natural language captions to enhance high-level semantic understanding of scenes. The violence detection module is trained on the Real-Life Violence Situations (RLVS) dataset, achieving classification accuracy of 92.49%."},"title":{"value":"A hybrid context-aware video violence detection framework using hierarchical spatiotemporal and semantic modeling"},"authors":{"value":["Shobha Kumari","Vijay Kumar"]}},"tmdate":1784551617866,"pdate":1798675200000,"externalIds":["dblp:journals/sivp/KumariK26"],"tcdate":1772176556163,"writers":["~"],"signatures":["~Dr_Vijay_Kumar1"],"forum":"Nlkm4pikol","license":"CC BY-SA 4.0","number":831813,"cdate":1767225600000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1784551617866,"domain":"DBLP.org","id":"Nlkm4pikol","version":2},{"content":{"venue":{"value":"Signal, Image and Video Processing"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/s11760-025-05090-8.pdf"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"kumari|a_hybrid_contextaware_video_violence_detection_framework_using_hierarchical_spatiotemporal_and_semantic_modeling"},"html":{"value":"https://doi.org/10.1007/s11760-025-05090-8"},"abstract":{"value":"Video violence detection poses significant challenges due to the complex integration of spatial, temporal, and contextual features. Conventional methods, such as 3D Convolutional Neural Networks, Long Short-Term Memory Models, and You Only Look Once, exhibit limited scene-level semantic understanding, high computational costs, and poor generalization across diverse environments. This work proposes a novel hybrid context-aware framework that combines a lightweight MobileNetV2 with a Bidirectional Long Short-Term Memory (BiLSTM) network for efficient spatiotemporal feature extraction. YOLOv8 is used for real-time object detection and Bootstrapping Language-Image Pre-training is utilized for generating the natural language captions to enhance high-level semantic understanding of scenes. The violence detection module is trained on the Real-Life Violence Situations (RLVS) dataset, achieving classification accuracy of 92.49%."},"title":{"value":"A hybrid context-aware video violence detection framework using hierarchical spatiotemporal and semantic modeling"},"authors":{"value":[{"fullname":"Shobha Kumari","username":"~Shobha_Kumari2"},{"fullname":"Vijay Kumar","username":"https://orcid.org/orcid-search/search?searchQuery=Vijay%20Kumar"}]}},"tmdate":1784551586208,"pdate":1767225600000,"externalIds":["doi:10.1007/s11760-025-05090-8"],"tcdate":1783696192739,"writers":["~","OpenReview.net/Public_Article/ORCID.org","OpenReview.net/Support"],"signatures":["~Shobha_kumari1"],"forum":"ADklDj9VC0","license":"CC BY-SA 4.0","number":80374,"cdate":1768443682333,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/ORCID.org/-/Record","OpenReview.net/Public_Article/-/Edit","OpenReview.net/Public_Article/-/Author_Removal","OpenReview.net/Public_Article/-/Authorship_Claim"],"mdate":1784551586208,"domain":"OpenReview.net/Public_Article","id":"ADklDj9VC0","version":2},{"content":{"summary":{"value":"This paper tackles the task of 4D reconstruction from monocular video. It introduces a training approach for a multi-view video generative model using a synthetic dataset of multi-view videos. The architecture uses a 3D-aware denoising diffusion model previously applied to multi-view images and extends it to accommodate multi-view videos. The model is fine-tuned using 1,000 synthetic multi-view videos from the Objaverse dataset. Then, score-distillation sampling (SDS) is used to generate a dynamic radiance field. The evaluation on videos of synthetic object-centric scenes demonstrates a slight improvement in terms of CLIP and FVD metrics over the recent Consistent4D work on the task of novel view synthesis from monocular video. Although the qualitative results show minor enhancement over baselines, concerns remain about the generalization to real-world videos and the evaluation, especially regarding potential training data leakage and significance of improvement over Consistent4D. Addressing these issues would warrant the acceptance of the paper."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"The paper assumes access to a multi-view video dataset. The rationale behind the need for SDS when a multi-view video diffusion model is already available is unclear. Could you explain why not fit the dynamic NeRF directly on the generated multi-view videos?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"- The paper addresses the significant and timely issue of generating 4D content using diffusion models.\n- The architectural extension of the 3D-aware diffusion model and its fine-tuning are good contributions that would be useful to know for the community. \n- The technical contribution is highlighted by impressive generalization performance (assuming no train data leakage). \n- This also highlights the scalability potential of synthetic Objaverse dataset for fine-tuning video diffusion models to perform 4D generation.\n- Both qualitative and quantitative results indicate improvements over the baselines, albeit modest compared to Consistent4D."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The training requires a multi-view video dataset which is difficult to obtain.\n- The evaluation is limited to synthetic, object-centric toy scenes without backgrounds.\n- I haven’t found a description of the validation and test dataset for experiments in Section 4.1\n- It is not clear whether assets in test-videos are unseen during training of all models (as some of them are also trained on objaverse).\n- Evaluation in Sec 4.1 is limited to CLIP and FVD metrics. Since multi-view video datasets were used for training, one could evaluate models for the novel view synthesis task using standard metrics such as LPIPS/PSNR (taking best of 10 due probabilistic nature of the task).\n- Minor: Given the small improvement over Consistent4D, further evaluation of statistical significance is needed.\n- Minor: The paper would benefit from more precise writing; particularly, the training description in lines 176-182 needs more clarity on each variable and the noising process, and Equation 10 lacks clarity regarding sampled variables used in expectations. The method description is overly complex, and the language is difficult to follow, containing several unclear sentences."},"limitations":{"value":"Authors addressed the limitations."}},"nonreaders":[],"tmdate":1730879348088,"tcdate":1721748595674,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission9911/Reviewer_Y3vx"],"signatures":["NeurIPS.cc/2024/Conference/Submission9911/Reviewer_Y3vx"],"forum":"SFk7AMpyhx","number":3,"license":"CC BY 4.0","cdate":1721748595674,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission9911/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730879348088,"domain":"NeurIPS.cc/2024/Conference","replyto":"SFk7AMpyhx","id":"dcSLssYp9U","forumContent":{"TLDR":{"value":"we present a novel 4D generation pipeline, 4Diffusion, to create high-quality spatial-temporally consistent 4D content from a monocular video."},"venue":{"value":"NeurIPS 2024 poster"},"keywords":{"value":["Diffusion Model","4D Generation","NeRF"]},"supplementary_material":{"value":"/attachment/93d614b0ed825bfaa519b19418cdb02188b26762.zip"},"primary_area":{"value":"generative_models"},"abstract":{"value":"Current 4D generation methods have achieved noteworthy efficacy with the aid of advanced diffusion generative models. However, these methods lack multi-view spatial-temporal modeling and encounter challenges in integrating diverse prior knowledge from multiple diffusion models, resulting in inconsistent temporal appearance and flickers. In this paper, we propose a novel 4D generation pipeline, namely $\\textbf{4Diffusion}$, aimed at generating spatial-temporally consistent 4D content from a monocular video. We first design a unified diffusion model tailored for multi-view video generation by incorporating a learnable motion module into a frozen 3D-aware diffusion model to capture multi-view spatial-temporal correlations. After training on a curated dataset, our diffusion model acquires reasonable temporal consistency and inherently preserves the generalizability and spatial consistency of the 3D-aware diffusion model. Subsequently, we propose 4D-aware Score Distillation Sampling loss, which is based on our multi-view video diffusion model, to optimize 4D representation parameterized by dynamic NeRF. This aims to eliminate discrepancies arising from multiple diffusion models, allowing for generating spatial-temporally consistent 4D content. Moreover, we devise an anchor loss to enhance the appearance details and facilitate the learning of dynamic NeRF. Extensive qualitative and quantitative experiments demonstrate that our method achieves superior performance compared to previous methods."},"_bibtex":{"value":"@inproceedings{\nzhang2024diffusion,\ntitle={4Diffusion: Multi-view Video Diffusion Model for 4D Generation},\nauthor={Haiyu Zhang and Xinyuan Chen and Yaohui Wang and Xihui Liu and Yunhong Wang and Yu Qiao},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=SFk7AMpyhx}\n}"},"title":{"value":"4Diffusion: Multi-view Video Diffusion Model for 4D Generation"},"pdf":{"value":"/pdf/c82cad9ac993b7e7299eb3787716e0166d8f972c.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"zhang|4diffusion_multiview_video_diffusion_model_for_4d_generation"},"authorids":{"value":["~Haiyu_Zhang1","~Xinyuan_Chen1","~Yaohui_Wang1","~Xihui_Liu1","~Yunhong_Wang1","~Yu_Qiao1"]},"authors":{"value":["Haiyu Zhang","Xinyuan Chen","Yaohui Wang","Xihui Liu","Yunhong Wang","Yu Qiao"]}},"version":2},{"content":{"summary":{"value":"The authors of this paper propose a hypothesis -- namely, that the in-context learning capabilities of transformers are limited by their tendency to be deceived by shortcuts (as demonstrated in [1]) -- and test it by benchmarking a transformer they introduce to prevent shortcut learning (explicit model) against a standard autoregressive transformer (implicit). The explicit model is a slightly modified architecture in which the parameters that are optimized to perform next-token prediction are not able to attend to specific hidden states, whereas the implicit model can do so. The benchmark is performed over a variety of tasks in in-domain (ID) and out-of-domain (OOD) settings. The results show that the explicit model is not superior to the implicit one on ID data. They show that the bottleneck of the explicit model encodes useful representations of the learned task but the next-token prediction parameters are suboptimal at exploiting it. Additional experiments to investigate the interpretability and scaling properties of the explicit and implicit models are performed.\n\n[1] Large Language Models Can be Lazy Learners: Analyze Shortcuts in In-Context Learning, Tang et al, 2023"},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. Do you have results comparing the explicit model's performance to the implicit model on a task in which triggers have been deliberately injected? If you do, please report them here. If they show that the explicit model is robust to shortcuts, I would find the result compelling.\n2. Beyond what you argue in the paragraph starting on line 70, does your understanding of OOD generalization include robustness to variation of incidental features on ID data?"},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":2},"strengths":{"value":"* **Originality** This paper brings a new perspective to recent work on the problem of shortcut learning. Designing a network to prevent shortcut learning and applying it in the way the authors have done is novel. Framing the investigation in a parametric vs. non-parametric setting is insightful and highlights the operation of attention mechanisms as a particular feature of interest.\n* **Quality** This is a high-quality paper. The evaluation on at least five types of tasks is thorough. The authors' use of tasks with known latents indicates a high level of rigor. The authors state the general hypothesis under investigation and provide additional hypotheses (line 165) about which model -- implicit or explicit -- are best suited to the evaluation tasks. The additional experiments (starting line 294) are informative.\n* **Clarity** This paper is clear and well written. The setup of the problem and the motivation for the work is communicated concisely.\n* **Significance** The contribution of this paper is a novel perspective on a timely problem. Robustness of large transformers is an area of active investigation -- including, shortcut learning (see [1]) -- and this paper shows to an extent that when a network is constrained in a targeted way, generalization does not necessarily improve."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"While the overall quality of the paper is quite good, much of the quality can be said to be tactical (i.e. the execution) rather than strategic (i.e. the motivation and contribution). In tactical terms, the paper is beautiful. Strategically, it can be improved.\n\nWhile I acknowledge and appreciate the principled argument in favor of the explicit model in the context of parametric vs non-parametric learning (see p. on line 70, Section 3, and Figure 1), I believe the paper would be very much improved if it verified empirically that the explicit model is itself resistant to shortcut learning in the presence of **injected** shortcuts. Whereas [1] injects shortcuts (\"triggers\") into data in an ICL setting, the work presented here appears to use datasets without such injections. Creating, if necessary, a synthetic dataset for evaluating robustness to shortcuts, and empirically demonstrating that the explicit model is resistant to shortcuts would strengthen the paper by showing a method by which existing large-scale models might be made resistant to shortcuts. Further showing that the explicit model exhibited consistent resistance to shortcuts as the model size increases would be a substantial contribution. Finally, showing this would provide solid empirical ground on which the argument in p. on line 70 rests.\n\nI would prefer that the authors removed \"commonly\" from their conclusion (line 512). It's non-controversial that shortcuts are a likely explanation for well-demonstrated non-robustness of transformers to incidental variations of their inputs. In the context of the sentence on line 512, however, the use of \"commonly\" can be taken to imply that most people believe that shortcut learning is the reason for poor OOD generalization. If that particular view w.r.t. OOD generalization is commonly held, I would expect there to be be some literature -- particularly since [1] was published in 2023 -- claiming exactly that. If it exists, it should be cited in the introduction.\n\nNit: The classification tasks in the first row of Figure 4 are missing the label on the Y axis."}},"nonreaders":[],"tmdate":1732230348676,"tcdate":1730570748820,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission3860/Reviewer_uCxn"],"signatures":["ICLR.cc/2025/Conference/Submission3860/Reviewer_uCxn"],"forum":"TSlJ3ikcBZ","number":2,"license":"CC BY 4.0","cdate":1730570748820,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission3860/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732230348676,"domain":"ICLR.cc/2025/Conference","replyto":"TSlJ3ikcBZ","id":"2szEX9uCuV","forumContent":{"TLDR":{"value":"We test whether explicit latent-variable inference leads to better performance in in-context learning and identify the benefits and shortcomings of such methods."},"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["in-context learning","transformers","attention","latent variable","shortcuts"]},"supplementary_material":{"value":"/attachment/8c2a5092e3d71c14ab15b857468f12dfed70c230.zip"},"primary_area":{"value":"interpretability and explainable AI"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Large autoregressive models like Transformers can solve tasks through in-context learning (ICL) without learning new weights, suggesting avenues for efficiently solving new tasks. For many tasks, e.g., linear regression, the data factorizes: examples are independent given a task latent that generates the data, e.g., linear coefficients. While an optimal predictor leverages this factorization by inferring task latents, it is unclear if Transformers implicitly do so or if they instead exploit heuristics and statistical shortcuts enabled by attention layers. Both scenarios have inspired active ongoing work. In this paper, we systematically investigate the effect of explicitly inferring task latents. We minimally modify the Transformer architecture with a bottleneck designed to prevent shortcuts in favor of more structured solutions, and then compare performance against standard Transformers across various ICL tasks. Contrary to intuition and some recent works, we find little discernible difference between the two; biasing towards task-relevant latent variables does not lead to better out-of-distribution performance, in general. Curiously, we find that while the bottleneck effectively learns to extract latent task variables from context, downstream processing struggles to utilize them for robust prediction. Our study highlights the intrinsic limitations of Transformers in achieving structured ICL solutions that generalize, and shows that while inferring the right latents aids interpretability, it is not sufficient to alleviate this problem."},"_bibtex":{"value":"@misc{\nmittal2025does,\ntitle={Does learning the right latent variables necessarily improve in-context learning?},\nauthor={Sarthak Mittal and Eric Elmoznino and Leo Gagnon and Sangnie Bhardwaj and Guillaume Lajoie and Dhanya Sridhar},\nyear={2025},\nurl={https://openreview.net/forum?id=TSlJ3ikcBZ}\n}"},"title":{"value":"Does learning the right latent variables necessarily improve in-context learning?"},"pdf":{"value":"/pdf/7d96622853f6d91960f66f10620c5a4e9ef1e837.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"mittal|does_learning_the_right_latent_variables_necessarily_improve_incontext_learning"},"authorids":{"value":["~Sarthak_Mittal1","~Eric_Elmoznino1","~Leo_Gagnon1","~Sangnie_Bhardwaj1","~Guillaume_Lajoie1","~Dhanya_Sridhar2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Sarthak Mittal","Eric Elmoznino","Leo Gagnon","Sangnie Bhardwaj","Guillaume Lajoie","Dhanya Sridhar"]}},"version":2},{"content":{"summary":{"value":"This paper introduces Lumos-1, an autoregressive video generator based on an LLM architecture with minimal modifications, which aims to create a unified multimodal model. A key achievement is the demonstration of video generation using LLM architecture, which paves the way for a truly unified foundation model and eliminates the need for external text encoders. \n\nAdditionally, Lumos-1 incorporates two innovations:\n1.  It proposes MM-ROPE to address imbalanced frequency spectrums in 3D RoPE, enhancing spatiotemporal correlation modeling via distributed channel allocation and scaled 3D positions.\n2. Lumos-1 employs autoregressive discrete diffusion forcing (AR-DF) to mitigate frame-wise loss imbalance and spatial information redundancy in autoregressive video generation training.\n\nThe paper provides detailed descriptions of the methods employed, thorough experimental evaluations against robust benchmarks, and insightful ablation studies that effectively demonstrate the merits of the innovations. However, Lumos-1 is trained from scratch despite a high structural similarity with Llama, so the model needs to learn both language and vision simultaneously, which can introduce instability and inefficiency during training. \n\nIn summary, this is an innovative work with significant community impact, and I appreciate its architectural exploration; however, I remain concerned about its training efficiency and stability, and I am curious about its performance with LLM initialization."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. Since the model structure is highly similar to LLM, why not initialize the parameters with pretrained Llama?\n2. Based on the experimental results, the unified architecture improves semantic consistency significantly, yet the visual quality still lags behind some baseline methods. Have you attempted to explain this trade-off or conducted a deeper analysis of its causes?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":4},"strengths":{"value":"1. This work proposes Lumios, which demonstrates the effectiveness of video generation using LLM architecture, paving the way for a truly unified foundation model and eliminating the need for external text encoders. \n2. It proposes MM-ROPE to address imbalanced frequency spectrums in 3D RoPE, enhancing spatiotemporal correlation modeling via distributed channel allocation and scaled 3D positions.\n3. Lumos-1 employs autoregressive discrete diffusion forcing (AR-DF) to mitigate frame-wise loss imbalance and spatial information redundancy in autoregressive video generation training.\n4. Comprehensive experiments and complete ablation studies validate the advantage of MM-ROPE and AR-DF, and the effectiveness of LLM-based video generation."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Although structurally similar to Llama, Lumos-1 is trained from scratch, requiring simultaneous learning of language and vision, which may lead to training instability and inefficiency. The validation curve in Figure 7(a) supports concerns about the instability. Furthermore, as a foundation model, the full cost of the training, particularly in terms of time, is not disclosed, which is a cause for concern in terms of the inefficiency of the training.\n2. The paper claims minimal structural modification from LLM for a unified text-visual model, but it lacks a comprehensive structural comparison with other established unified generative models or LLM-based autoregressive models, making the claimed advantage less substantiated."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762916814483,"tcdate":1761925883603,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission3552/Reviewer_uiBr"],"signatures":["ICLR.cc/2026/Conference/Submission3552/Reviewer_uiBr"],"forum":"wWAxwSCKR2","number":2,"license":"CC BY 4.0","cdate":1761925883603,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission3552/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762916814483,"domain":"ICLR.cc/2026/Conference","replyto":"wWAxwSCKR2","id":"ARGDoKvqU7","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Video generation","autoregressive models","unified models"]},"primary_area":{"value":"generative models"},"abstract":{"value":"Autoregressive large language models (LLMs) have unified a vast range of language tasks, inspiring preliminary efforts in autoregressive (AR) video generation. Existing AR video generators either diverge from standard LLM architectures, depend on bulky external text encoders, or incur prohibitive latency due to next-token decoding. In this paper, we introduce Lumos-1, an LLM-based unified model for AR video generation with efficient discrete diffusion. Firstly, to fit videos with LLMs, we identify that 1D RoPE is ill-suited for visual spatiotemporal correlation modeling, and while demonstrated to be useful, naive 3D RoPE exhibits imbalanced frequency spectra. Therefore, we propose MM‑RoPE, which preserves the original textual RoPE while seamlessly accommodating video data with comprehensive frequency spectra and scaled 3D positions. Secondly, to fit the video data's nature and overcome the inefficiency of next-token decoding, we adopt a parallel and mask-based discrete diffusion with the intra-frame bidirectional and inter-frame causal attention masks. Based on this attention mask, we uncover the frame‑wise loss imbalance issue caused by spatial information redundancy and propose Autoregressive Discrete Diffusion Forcing, which introduces temporal tube masking during training with a compatible inference‑time masking policy to avoid quality degradation. Despite using only 48 GPUs for pre-training, limited data and a discrete tokenizer, Lumos-1 achieves results surpassing those of Show-o2 on GenEval, COSMOS-Video2World on VBench-I2V, and OpenSoraPlan on VBench-T2V."},"_bibtex":{"value":"@inproceedings{\nyuan2026lumos,\ntitle={Lumos-1: On Autoregressive Video Generation with Discrete Diffusion from a Unified Model Perspective},\nauthor={Hangjie Yuan and Weihua Chen and Jun CEN and Hu Yu and Jingyun Liang and Shuning Chang and Zhihui Lin and Tao Feng and Pengwei Liu and Jiazheng Xing and Hao Luo and Jiasheng Tang and Fan Wang and Yi Yang},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=wWAxwSCKR2}\n}"},"title":{"value":"Lumos-1: On Autoregressive Video Generation with Discrete Diffusion from a Unified Model Perspective"},"pdf":{"value":"/pdf/977b5e71a64f9e2dba23028bbf1fd70760a93c5f.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"yuan|lumos1_on_autoregressive_video_generation_with_discrete_diffusion_from_a_unified_model_perspective"},"authorids":{"value":["~Hangjie_Yuan1","~Weihua_Chen1","~Jun_CEN1","~Hu_Yu2","~Jingyun_Liang1","~Shuning_Chang1","~Zhihui_Lin2","~Tao_Feng3","~Pengwei_Liu1","~Jiazheng_Xing1","~Hao_Luo1","~Jiasheng_Tang1","~Fan_Wang6","~Yi_Yang22"]},"authors":{"value":["Hangjie Yuan","Weihua Chen","Jun CEN","Hu Yu","Jingyun Liang","Shuning Chang","Zhihui Lin","Tao Feng","Pengwei Liu","Jiazheng Xing","Hao Luo","Jiasheng Tang","Fan Wang","Yi Yang"]}},"version":2},{"content":{"summary":{"value":"The paper presents a novel framework designed to improve text-to-video (T2V) generation, especially for complex scenarios involving multiple objects and the composition of different objects.\nThe proposed VideoTetris introduces spatio-temporal compositional diffusion, a dynamic-aware data processing pipeline, and a consistency regularization method to enhance the quality and coherence of generated videos.\nExtensive experiments demonstrate the framework's superior performance in generating both short and long videos with complex and multi-object prompts."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"- Does the proposed framework have the scalability limitations in terms of the length and complexity of the generated videos?\n- How does the framework perform in real-world scenarios where the input text prompts may be less structured and more diverse?"},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"- I like the idea that the model first splits a story prompt into frame-wise short prompt with object composition. This approach supports sharing global information among long synthesic video (which is usually computationally expensive and hard to achieve).\n- The visual quality provided in the supplementary looks great (although the video number is limited)\n- The experiments include the quantitative comparison with the state-of-the-arts in the tasks of short and long video generation with multi-object prompts, and also ablation study, which can provide better understanding about the performance of the proposed VideoTetris."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- My major concern is that the model heavily relies on the performance of LLM. In the proposed method, LLM acts as a director to split an input story prompt into spatiotemporal prompts. I am wondering whether LLM is robust enough to generate temporal-consistent prompt with spatial-consistent composition.\n- Following the concern above, the supplementary only provides one example of long video generation. The authors should more detailed and more samples. For example, provide 100 story prompts and the corresponding frame-wise prompts with object composition. For what we have now, it is hard to fairly justify the model performance.\n- Another concern is the autoregressive strategy of long video synthesis. While the long video shows good consistency of motion and subject identity, I found the visual quality gradually decreases in the video. For example, in the sample \"fig5-ours.mp4\", the contour of leaves is visible at beginning; however, it gets more and more blurry until the end of the video. Is this also happens on other long synthesis videos?"},"limitations":{"value":"- Scalability: as mentioned in the weakness, the model highly relies on the performance of LLM which has unstable performance and sometimes has hallucination issue. Hence, it requires manual check and limits the scalability.\n- Limited length of long video generation: also mentioned in the weakness, autoregressive strategy of long video synthesis limits the length of generated video. Could the author provide more or even longer video to verify this limitation?"}},"nonreaders":[],"tmdate":1730879408354,"tcdate":1720912938006,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission10727/Reviewer_Qhp4"],"signatures":["NeurIPS.cc/2024/Conference/Submission10727/Reviewer_Qhp4"],"forum":"RPM7STrnVz","number":4,"license":"CC BY 4.0","cdate":1720912938006,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission10727/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730879408354,"domain":"NeurIPS.cc/2024/Conference","replyto":"RPM7STrnVz","id":"dRELS1BCEf","forumContent":{"venue":{"value":"NeurIPS 2024 poster"},"keywords":{"value":["Text-to-Video Generation","Video Diffusion Models"]},"supplementary_material":{"value":"/attachment/f72c15c52e21ddf70f17cb82f2289107037072b9.zip"},"primary_area":{"value":"generative_models"},"abstract":{"value":"Diffusion models have demonstrated great success in text-to-video (T2V) generation. However, existing methods may face challenges when handling complex (long) video generation scenarios that involve multiple objects or dynamic changes in object numbers. To address these limitations, we propose VideoTetris, a novel framework that enables compositional T2V generation. Specifically, we propose spatio-temporal compositional diffusion to precisely follow complex textual semantics by manipulating and composing the attention maps of denoising networks spatially and temporally. Moreover, we propose a new dynamic-aware data processing pipeline and a consistency regularization method to enhance the consistency of auto-regressive video generation. Extensive experiments demonstrate that our VideoTetris achieves impressive qualitative and quantitative results in compositional T2V generation. Code is available at: https://github.com/YangLing0818/VideoTetris"},"_bibtex":{"value":"@inproceedings{\ntian2024videotetris,\ntitle={VideoTetris: Towards Compositional Text-to-Video Generation},\nauthor={Ye Tian and Ling Yang and Haotian Yang and Yuan Gao and Yufan Deng and Xintao Wang and Zhaochen Yu and Xin Tao and Pengfei Wan and Di ZHANG and Bin CUI},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=RPM7STrnVz}\n}"},"title":{"value":"VideoTetris: Towards Compositional Text-to-Video Generation"},"pdf":{"value":"/pdf/bb68997ead13efc218660269a1fa4d189f0588a2.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"tian|videotetris_towards_compositional_texttovideo_generation"},"authorids":{"value":["~Ye_Tian15","~Ling_Yang1","~Haotian_Yang1","~Yuan_Gao32","~Yufan_Deng2","~Xintao_Wang1","~Zhaochen_Yu2","~Xin_Tao3","~Pengfei_Wan1","~Di_ZHANG3","~Bin_CUI2"]},"authors":{"value":["Ye Tian","Ling Yang","Haotian Yang","Yuan Gao","Yufan Deng","Xintao Wang","Zhaochen Yu","Xin Tao","Pengfei Wan","Di ZHANG","Bin CUI"]}},"version":2},{"content":{"summary":{"value":"1. The paper addresses a new video popularity prediction task of distinguishing between globally and locally popular content.\n2. The authors develop a Video Popularity Dataset (VPD)  of 17,000 YouTube videos, with titles, descriptions, and detailed metadata, along with the countries the video was trending in amongst 109 countries.\n3. The authors transform video content into textual representations using Vision-Language Models (VLMs) to bridge the modality gap between video data and LLMs. \n4. The authors test extensively on prompting paradigms for the video popularity prediction task, including Thinking, ICL, Hypothesis generation, and supervised signal into the prompt."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1. What steps were taken to filter the videos for this dataset? Languages, Tags\n2. How was the dataset collected\nSee weaknesses"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. The paper introduces a new dataset and task for local and global popularity prediction.\n2. The insights like reach vs popularity, ablation studies, are useful.\n3. The authors conduct human studies to verify video verbalization and hypothesis generation."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The frame-to-text method is not novel, there are multiple works [1,2,3,4] that use titles, channel, generated visual captions, ASR of youtube videos for tasks like Views, Likes/Views, replay predictions [2,3,4].\n2. There is no explanation of the 17k videos and metadata collection, one subsection on the tools and data sources is necessary.\n3. The dataset consists of only global and local hits, and not negative samples, looking at locally/globally popular videos from the examples (and https://kworb.net/youtube/trending_overall.html#google_vignette , since any other source was not mentioned in the paper) there seems to be a huge language bias (Hindi/Korean/.. languages being local hits)\n4. The human evaluation is not signifcant for both tasks,14,11 subjects passed the screening and they were only given 2 video-caption pairs, totalling 28, 22 interactions only.\n5. The technical contributions are severely limited, the authors only show combinations of known prompting strategies for evaluation.\n\n[1] https://aclanthology.org/2023.emnlp-main.608/\n[2] https://arxiv.org/abs/2309.00378\n[3] https://arxiv.org/abs/2405.00942\n[4] https://openreview.net/forum?id=TrKq4Wlwcz"}},"nonreaders":[],"tmdate":1733122919719,"tcdate":1730690506669,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission11188/Reviewer_HVCp"],"signatures":["ICLR.cc/2025/Conference/Submission11188/Reviewer_HVCp"],"forum":"iKsTtpzBtc","number":3,"license":"CC BY 4.0","cdate":1730690506669,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission11188/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1733122919719,"domain":"ICLR.cc/2025/Conference","replyto":"iKsTtpzBtc","id":"MXJ3ZSOwaX","forumContent":{"TLDR":{"value":"We use Large Language Models (LLMs) to predict video popularity by transforming multimodal content into text with Vision-Language Models (VLMs), achieving higher accuracy and interpretability than traditional machine learning methods."},"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Large Language Models (LLMs)","Vision-Language Models (VLMs)","Multimodal Textual Representations","Video Popularity Prediction"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Predicting video popularity is typically formalized as a supervised learning problem, where models classify videos as popular or unpopular. Traditional approaches rely heavily on meta-information and aggregated user engagement data, but video popularity is highly context-dependent, influenced by cultural, social, and temporal factors that these approaches fail to capture. We argue that Large Language Models (LLMs), with their deep contextual awareness, are well-suited to address these challenges. A key difficulty, however, lies in bridging the modality gap between pixel-based video data and token-based LLMs. To overcome this, we transform frame-level visual data into sequential text representations using Vision-Language Models (VLMs), enabling LLMs to process multimodal video content—titles, frame-based descriptions, and captions—and capture rich contextual information for more accurate predictions. Evaluating on a newly introduced dataset of 17,000 videos, we show that while a supervised neural network using content embeddings achieved 80% accuracy, our LLM-based method reached 82% without fine-tuning. A combined approach, integrating the neural network's predictions into the LLM, further improved accuracy to 85.5%. Additionally, the LLM generates interpretable hypotheses explaining its predictions based on theoretically sound attributes. Survey-based manual validations confirm the quality of these hypotheses and address concerns about hallucinations in the video-to-text conversion process. Our findings highlight that LLMs, equipped with textually transformed multimodal representations, offer a powerful, interpretable, and data-efficient solution to the context-dependent challenge of video popularity prediction."},"_bibtex":{"value":"@misc{\nkayal2025large,\ntitle={Large Language Models Are Natural Video Popularity Predictors},\nauthor={Pratik Kayal and Pascal Mettes and Nima Dehmamy and Minsu Park},\nyear={2025},\nurl={https://openreview.net/forum?id=iKsTtpzBtc}\n}"},"title":{"value":"Large Language Models Are Natural Video Popularity Predictors"},"pdf":{"value":"/pdf/a10ecc11aa3f6e85952b5fee0633e5822c8a7ee9.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"kayal|large_language_models_are_natural_video_popularity_predictors"},"authorids":{"value":["~Pratik_Kayal1","~Pascal_Mettes1","~Nima_Dehmamy1","~Minsu_Park1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Pratik Kayal","Pascal Mettes","Nima Dehmamy","Minsu Park"]}},"version":2},{"content":{"summary":{"value":"The contributions of this paper include three aspects. 1. It implemented an automatic annotation process, providing a high-quality, large-scale QA dataset for egocentric videos. 2. It constructed a challenging egocentric QA benchmark, consisting of 629 videos and 7026 questions and introduced a new metric to mitigate inevitable language biases in evaluated models. 3. It proposed a novel model structure, including the global glimpse step and fallback step. By fine-tuning, MM Ego was built and demonstrated excellent performance in egocentric video understanding."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"Please refer to \"Weaknesses\"."},"rating":{"value":6},"details_of_ethics_concerns":{"value":"None"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"- A new model structure is proposed, with the motivation of incrementally understanding videos: first, providing an overview of the entire video, then focusing on specific details, and keeping in mind particular questions.\n\n- This paper provides dialogue examples of MM Ego in real-world scenarios, offering a paradigm for understanding human instructions in real-world settings.\n\n- This work provides a large amount of Ego video data to the community through an automated annotation process and it is valuable to assess language biases in the LLM evaluation process."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The authors suggest that in terms of benchmarking, existing egocentric video benchmarks either focus on shorter videos, such as EgoSchema and QaEgo4D or on internet video content, such as video MME. This results in a significant gap in egocentric video understanding benchmarks, but the data for the proposed benchmark EgoMemoria still comes from Ego4D. I do not feel that EgoMemoria is significantly different from the benchmarks mentioned earlier, and the authors need to clarify this point further.\n\n- Is it reasonable to only use GPT-4o in the automated annotation process? When generating QA pairs, I think that videos are equally important as dense captions, which will further ensure the quality of the QA pairs. The authors need to provide ablation experiments to validate that the existing method is more reasonable and ensures higher annotation quality.\n\n- To ensure reproducibility, the authors need to provide the prompt templates for GPT-4o and ChatGPT used in the automated annotation and evaluation processes.\n\n- The benchmark proposed in this work includes 629 videos. The authors classified them by length but did not provide the distribution of videos for each length category. This is crucial for the comprehensiveness and robustness of model evaluation. The authors should further elaborate on the distribution of video lengths.\n\n- In Table 4, MM Ego's performance on Video-MME is not as impressive as LLaVA OV. The authors should provide a more comprehensive analysis to explain the reasons for this discrepancy. Otherwise, I might conclude that the construction method of MM Ego sacrifices its ability to understand general videos in favor of enhancing ego understanding.\n\n- Language bias is inevitable in LLM, but whether it will be influenced by random responses in ego videos also needs to be verified through an ablation study.\n\n- The Global Glimpse Step and Fallback Step described by the author would be more complete and convincing if they could be correlated with mechanisms related to cognitive neuroscience and brain science.\n\n- Figure 5 may be better represented using a word cloud.\n\n- Please note that it is not GPT4-o, but GPT-4o.\n\nIf the author can address the questions above, I will improve my rating."}},"nonreaders":[],"tmdate":1732367656862,"tcdate":1729527900090,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission962/Reviewer_g472"],"signatures":["ICLR.cc/2025/Conference/Submission962/Reviewer_g472"],"forum":"67sSPPAZiG","number":1,"license":"CC BY 4.0","cdate":1729527900090,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission962/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732367656862,"domain":"ICLR.cc/2025/Conference","replyto":"67sSPPAZiG","id":"JODElQbfMD","forumContent":{"venue":{"value":"ICLR 2025 Poster"},"TLDR":{"value":"MM-Ego, an egocentric multimodal LLM that shows powerful performance on egocentric video understanding with the contribution on: (i) Egocentric QA Data Engine (ii) Memory Pointer Prompting Mechanism (iii)  EgoMemoria benchmark."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["multimodal models"]},"primary_area":{"value":"foundation or frontier models, including LLMs"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"This research aims to comprehensively explore building a multimodal foundation model for egocentric video understanding.\nTo achieve this goal, we work on three fronts. \nFirst, as there is a lack of QA data for egocentric video understanding, we automatically generate 7M high-quality QA samples for egocentric videos ranging from 30 seconds to one hour long in Ego4D based on human-annotated data.\nThis is one of the largest egocentric QA datasets.\nSecond, we contribute a challenging egocentric QA benchmark with 629 videos and 7,026 questions to evaluate the models' ability in recognizing and memorizing visual details across videos of varying lengths. We introduce a new de-biasing evaluation method to help mitigate the unavoidable language bias present in the models being evaluated.\nThird, we propose a specialized multimodal architecture featuring a novel ``Memory Pointer Prompting\" mechanism. This design includes a global glimpse step to gain an overarching understanding of the entire video and identify key visual information, followed by a fallback step that utilizes the key visual information to generate responses. This enables the model to more effectively comprehend extended video content.\nWith the data, benchmark, and model, we build MM-Ego, an egocentric multimodal LLM that shows powerful performance on egocentric video understanding."},"_bibtex":{"value":"@inproceedings{\nye2025mmego,\ntitle={{MME}go: Towards Building Egocentric Multimodal {LLM}s for Video {QA}},\nauthor={Hanrong Ye and Haotian Zhang and Erik Daxberger and Lin Chen and Zongyu Lin and Yanghao Li and Bowen Zhang and Haoxuan You and Dan Xu and Zhe Gan and Jiasen Lu and Yinfei Yang},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=67sSPPAZiG}\n}"},"title":{"value":"MMEgo: Towards Building Egocentric Multimodal LLMs for Video QA"},"pdf":{"value":"/pdf/cceef0909c06cbce21c0d7c9990986913f0f3499.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"ye|mmego_towards_building_egocentric_multimodal_llms_for_video_qa"},"authorids":{"value":["~Hanrong_Ye1","~Haotian_Zhang3","~Erik_Daxberger1","~Lin_Chen11","~Zongyu_Lin1","~Yanghao_Li1","~Bowen_Zhang2","~Haoxuan_You1","~Dan_Xu4","~Zhe_Gan1","~Jiasen_Lu2","~Yinfei_Yang1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Hanrong Ye","Haotian Zhang","Erik Daxberger","Lin Chen","Zongyu Lin","Yanghao Li","Bowen Zhang","Haoxuan You","Dan Xu","Zhe Gan","Jiasen Lu","Yinfei Yang"]}},"version":2},{"content":{"summary":{"value":"The paper introduces QueryStream, a novel framework designed to enhance the efficiency and interactivity of streaming video understanding models by incorporating query-awareness into the core processing loop. This work addresses the limitations of existing approaches, which typically rely on a flawed, query-agnostic “change-is-important” principle that conflates raw visual dynamics with true semantic relevance, leading to computational waste and interaction errors in real-time online video scenarios."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"See weakness above, please."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. Query-Aware Differential Pruning (QDP): QDP is a novel token pruning mechanism that is unique because it employs a dual criterion that jointly assesses semantic relevance to the user's query and temporal novelty. Previous token pruning methods were often query-agnostic.\n2. Dynamically Smoothed History (DSH): Within QDP, temporal novelty is assessed not against the immediately preceding frame, but against a Dynamically Smoothed History (DSH). This adaptive historical context ensures the pruning is robust to transient noise and slow visual drifts.\n3. Relevance-Triggered Active Response (RTAR): RTAR is a novel, logic-driven dual-gated mechanism that schedules responses based on a specific confluence of two signals: high query relevance and significant information density. This departs from prior reactive (passive) models or proactive models relying on heavily trained, specialized modules.\n4. Lightweight and Training-Free: QueryStream is designed as a lightweight, training-free module that operates as an intelligent pre-processing gateway, allowing for seamless, zero-shot integration with existing off-the-shelf Video-LLMs."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The effectiveness of the core Query-Aware Differential Pruning (QDP) mechanism is \"fundamentally bound by the representational quality of the pre-trained OpenCLIP encoder\". The authors utilized OpenCLIP-ViT-L/14, which was chosen as a pragmatic trade-off between feature quality and real-time computational efficiency. However, the inherent constraints of this encoder in discerning \"fine-grained details or abstract relationships\" may challenge the \"pruning precision in semantically nuanced scenarios\". This reliance risks causing critical but subtle events to be erroneously pruned and missed.\n2. The logic-based gating mechanisms (QDP and RTAR) rely on a \"set of fixed hyperparameters\". These values were determined empirically via sweeps and grid searches on a held-out validation set to find an optimal balance. However, these static thresholds, while effective in the evaluated benchmarks, \"may not be universally optimal\" and could reduce the model's robustness across diverse video domains, streaming qualities, or different types of user queries.\n3. The crucial ablation study for the Relevance-Triggered Active Response (RTAR) policy (Table 5), which demonstrates its superiority in timeliness-aware scoring, was conducted using a simulated evaluation protocol. This simulation was necessary because the official real-time evaluation infrastructure for the OVO-Bench benchmark and the TimeChat-Online code were unavailable at the time of experiments. While the simulation aimed to be fair by adhering to temporal constraints during response generation, the initial identification of trigger points leveraged the full video stream."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762925778885,"tcdate":1761977327272,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission15489/Reviewer_vYGP"],"signatures":["ICLR.cc/2026/Conference/Submission15489/Reviewer_vYGP"],"forum":"738HjJEbml","number":3,"license":"CC BY 4.0","cdate":1761977327272,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission15489/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762925778885,"domain":"ICLR.cc/2026/Conference","replyto":"738HjJEbml","id":"1TO1oaa1iD","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Multimodal Large Language Model","Online Video Understanding","Streaming Video Understanding"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"The increasing demand for real-time interaction in online video scenarios necessitates a new class of efficient streaming video understanding models. However, existing approaches often rely on a query-agnostic ''change-is-important'' assumption, which conflates visual dynamics with semantic relevance, leading to computational redundancy and mistimed responses. To address this, we propose QueryStream, a novel framework that integrates query-awareness into the core of video processing and response scheduling. QueryStream features two synergistic components: (1) Query-Aware Differential Pruning (QDP), a policy that filters the token stream by jointly assessing semantic relevance to the query and temporal novelty against a dynamically smoothed history; and (2) Relevance-Triggered Active Response (RTAR), a dual-gated mechanism that schedules responses based on both high query relevance and significant information density. As a lightweight, training-free module, QueryStream achieves state-of-the-art performance on benchmarks such as StreamingBench and OVO-Bench under moderate pruning, and matches full-token baselines while pruning over 70\\% of visual tokens. Notably, our pruning mechanism generalizes to offline tasks, where it serves as a context-denoising module that benefits long-form video understanding. This work not only reveals the vast semantic redundancy in video streams relative to user intent but also establishes a promising, intent-driven direction for efficient and robust online video understanding. Code is available at: https://github.com/Zhangkr2003/QueryStream."},"_bibtex":{"value":"@inproceedings{\nzhang2026querystream,\ntitle={QueryStream: Advancing Streaming Video Understanding with Query-Aware Pruning and Proactive Response},\nauthor={Kairui Zhang and Zhenyu Yang and Bing Wang and Shengsheng Qian and Changsheng Xu},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=738HjJEbml}\n}"},"title":{"value":"QueryStream: Advancing Streaming Video Understanding with Query-Aware Pruning and Proactive Response"},"pdf":{"value":"/pdf/cd7eb509371a7f036f82fb9c86406725b54a1cbb.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"zhang|querystream_advancing_streaming_video_understanding_with_queryaware_pruning_and_proactive_response"},"authorids":{"value":["~Kairui_Zhang1","~Zhenyu_Yang6","~Bing_Wang20","~Shengsheng_Qian1","~Changsheng_Xu1"]},"authors":{"value":["Kairui Zhang","Zhenyu Yang","Bing Wang","Shengsheng Qian","Changsheng Xu"]}},"version":2},{"content":{"summary":{"value":"This paper introduces MemLens, a novel method for detecting data contamination in Large Language Models (LLMs) by analyzing the activation trajectories of numeric tokens. The authors posit that existing methods, which rely on lexical overlap or perplexity, are brittle and fail when benchmark data is implicitly contaminated (e.g., rephrased or translated). MemLens operates on the hypothesis that memorized (contaminated) samples induce \"shortcut\" reasoning paths, causing the model to \"lock onto\" an answer with high confidence in its early layers. In contrast, clean samples show more gradual evidence accumulation across the model's depth.\n\nThe method extracts layer-wise probabilities for digit tokens, along with metrics like entropy and maximum confidence. These trajectories are fed into a 1D CNN discriminator to classify samples as \"contaminated\" or \"clean\". The authors validate this approach across four LLMs , showing that MemLens is significantly more robust to rephrasing, perturbation, and translation than baseline methods. Crucially, they provide strong causal evidence using LoRA injection: fine-tuning a model on clean data causes MemLens to identify those samples as contaminated, demonstrating it captures a genuine signal of memorization. Finally, the method is applied to the MMLU-Math benchmark, suggesting that 94.9% of the dataset is \"seen-like\" and that model performance is significantly inflated on this subset."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. Would you hypothesize the current approach generalizes to non-numeric tasks?\n2. Do you have a guess for what is it, about the internal activation trajectory, that shed light on memorization?\n3. How would one practically deploy MemLens for a new model without an existing, labeled dataset of clean/contaminated samples?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"** 1. Novelty **\n\nThe paper core contribution is moving detection from surface-level statistics (overlap, perplexity) to internal activation dynamics. The \"shortcut hypothesis\" provides a clear intuition for why memorization should manifest differently in activation space. The strong empirical evidence shows there is reliable signal for memorization in activations space. The performance is maintained across rephrased, translated, and perturbed inputs.\n\n** 2. Strong Causal Validation **\n\nThe LoRA injection experiments (Section 4.3, Table 2) provide compelling causal evidence. The monotonic increase in detection rates with increasing LoRA rank (2.2% → 45.1%) and the controlled single-sample case study (Figure 4) strongly suggest the method captures genuine memorization signals rather than spurious correlations.\n\n** 3. Implications **\n\nThe ability to detect data contamination is a powerful tool for robust model evaluation and selection, benefiting both developers and end-users. Furthermore, identifying memorization is critical for mitigating legal and ethical risks, such as the potential for models to reproduce copyrighted or private data, which compromises claims of generalization and trustworthiness."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"**1. Limited Scope**\n\nThe method explicitly focuses on digit tokens (Eq. 7), severely limiting applicability. The generalizability claim is undermined by this narrow scope.\n\n*Recommendation:* show, even a small demonstration of success on non-numeric tasks to encourage further research along those lines.\n\n**2. Limited Mechanistic Insight**\n\nWhile the paper compellingly demonstrates that a reliable memorization signal exists in the activation space, it falls short of explaining the underlying mechanism. The \"shortcut hypothesis\"  is presented as intuition but isn't proven or deeply investigated. Instead, the method relies on training a black-box discriminator (a 1D CNN ) on a set of engineered features. This approach successfully detects the signal but does not explain the precise relationship between the activations and the memorization.\n\n*Recommendation:* A feature ablation study, as suggested, would be a valuable starting point to understand the relative importance of different trajectory features. A deeper, more localized analysis could also examine which specific model components (e.g., MLPs vs. attention layers, or even individual neurons or attention heads) are the primary carriers of this signal, moving toward finding a minimal, interpretable pattern.\n\n**3. Supervised Approach:**\n\nThe method requires training a supervised discriminator with an already existing set of seen and unseen examples. The paper's setup is excellent for evaluating MemLens, but it doesn't address the practical challenge of how a developer would bootstrap this detector for a new model without already knowing which samples are contaminated."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1764359168483,"tcdate":1761726257593,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission9758/Reviewer_VAtL"],"signatures":["ICLR.cc/2026/Conference/Submission9758/Reviewer_VAtL"],"forum":"YtOAjOsL4N","number":2,"license":"CC BY 4.0","cdate":1761726257593,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission9758/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1764359168483,"domain":"ICLR.cc/2026/Conference","replyto":"YtOAjOsL4N","id":"j3AWXCn4PN","forumContent":{"venue":{"value":"ICLR 2026 Conference Desk Rejected Submission"},"keywords":{"value":["LLM","Interpretability","Memorization"]},"primary_area":{"value":"interpretability and explainable AI"},"abstract":{"value":"Large language models (LLMs) are commonly evaluated on challenging benchmarks such as AIME and Math500, which are susceptible to contamination and risk at being memorized. Contamination can be explicit, where samples appear verbatim, or implicit, where samples are rephrased, perturbed, or translated but still memorized. Existing detecting baselines focuses on surface-level lexcial overlap and perplexity, performing well for explicit contamination but degrading significantly under implicit cases. We propose MemLens (An Activation Lens for Memorization Detection) to detect memorization by analyzing the probability trajectories of numeric tokens during generation. We observe that contaminated and clean samples exhibit distinct and well-separated reasoning trajectories. To further validate the observation, we inject carefully designed samples into the model through LoRA fine-tuning and observe the same trajectory patterns as in naturally contaminated data. These results provide strong evidence that MemLens captures genuine signals of memorization rather than spurious correlations."},"_bibtex":{"value":"@misc{\nanonymous2026memlens,\ntitle={MemLens: Uncovering Memorization in {LLM}s with Activation Trajectories},\nauthor={Anonymous},\nyear={2026},\nurl={https://openreview.net/forum?id=YtOAjOsL4N}\n}"},"title":{"value":"MemLens: Uncovering Memorization in LLMs with Activation Trajectories"},"pdf":{"value":"/pdf/ea3df2790265756b4f9309f42237da5a1be88d24.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Desk_Rejected_Submission"},"paperhash":{"value":"he|memlens_uncovering_memorization_in_llms_with_activation_trajectories"},"authorids":{"value":["~Zirui_He1","~Haiyan_Zhao3","~Ali_Payani1","~Mengnan_Du1"]},"authors":{"value":["Zirui He","Haiyan Zhao","Ali Payani","Mengnan Du"]}},"version":2},{"content":{"summary":{"value":"This paper proposed a motion-coherent human video generation model, called MoSA. Formally, MoSA decouples the structure and appearance generation through a 3D structure generation and an appearance generation branch. The structure generation branch comprises a text-to-motion Transformer, structure-gen blocks, and Human-Aware Dynamic Control (HADC) blocks, while the appearance generation is based on a pretrained video generation model. Mask loss and tracking loss are additionally incorporated to improve the performance. Detailed experiments show that the proposed method enjoys visually coherent results. Moreover, this paper proposed a new human video dataset, called MoVid, including 30K human motion videos exhibiting diverse action categories and complex motions."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"I recommend that the authors further compare their method to the SOTA human animation methods, such as Uni3C, to confirm their performance. Morevoer, an interesting discussion should be included for the unified training of both text-to-motion and controllable video generation."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. This paper proposed an interesting idea for text-to-human generation, unifying both 3D text-to-motion and controllable video generation.   \n2. The writing quality of this paper is good, while qualitative and quantitative experiments are sufficient to cover SOTA video generation and user studies.\n3. The proposed dataset, MoVid, seems to be useful to the community."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. As detailed in Appendix.D, both the video generation model and text-to-motion transformer are frozen during the training. It would be helpful and interesting to investigate whether the text-to-motion model can also be enhanced with this end-to-end training strategy.\n\n2. Although this paper included detailed experiments, some results are still lacking. For example, a qualitative ablation study for HADC and its visual outcomes can help the readers to understand the effect of HADC better.\n\n3. For the comparison of human animation methods, the authors only compare to Animate Anyone, Champ, and HumanVid with inferior video backbones. More in-depth discussion and comparison are required, such as the Wan-based Uni3C (text-to-smpl can be used to simulate the Uni3C's input conditions as discussed in their paper)."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762920992429,"tcdate":1761727412116,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission9376/Reviewer_4dkv"],"signatures":["ICLR.cc/2026/Conference/Submission9376/Reviewer_4dkv"],"forum":"JKIa2rTxdB","number":1,"license":"CC BY 4.0","cdate":1761727412116,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission9376/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762920992429,"domain":"ICLR.cc/2026/Conference","replyto":"JKIa2rTxdB","id":"sgBqBZ3ITD","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"TLDR":{"value":"We propose a structure-appearance decoupling framework with human-aware dynamic control, dense tracking constraints, 3D contact constraints,and a large-scale human video dataset to generate motion-coherent human videos from given prompts."},"keywords":{"value":["Human Video Generation","Coherent Video Generation","Human Video Dataset"]},"supplementary_material":{"value":"/attachment/633a933b4e96aaef8edba288e42f07388a8a0005.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Existing video generation models predominantly emphasize appearance fidelity while exhibiting limited ability to synthesize complex human motions, such as whole-body movements, long-range dynamics, and fine-grained human–environment interactions. This often leads to unrealistic or physically implausible movements with inadequate structural coherence. To conquer these challenges, we propose MoSA, which decouples the process of human video generation into two components, i.e., structure generation and appearance generation. MoSA first employs a 3D structure transformer to generate a human motion sequence from the text prompt. The remaining video appearance is then synthesized under the guidance of this structural sequence. We achieve fine-grained control over the sparse human structures by introducing Human-Aware Dynamic Control modules with a dense tracking constraint during training. The modeling of human–environment interactions is improved through the proposed contact constraint. Those two components work comprehensively to ensure the structural and appearance fidelity across the generated videos. This paper also contributes a large-scale human video dataset, which features more complex and diverse motions than existing human video datasets. We conduct comprehensive comparisons between MoSA and a variety of approaches, including general video generation models, human video generation models, and human animation models. Experiments demonstrate that MoSA substantially outperforms existing approaches across the majority of evaluation metrics."},"_bibtex":{"value":"@inproceedings{\nwang2026mosa,\ntitle={Mo{SA}: Motion-Coherent Human Video Generation via Structure-Appearance Decoupling},\nauthor={Haoyu Wang and Hao Tang and Donglin Di and Zhilu Zhang and Wangmeng Zuo and Feng Gao and Siwei Ma and Shiliang Zhang},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=JKIa2rTxdB}\n}"},"title":{"value":"MoSA: Motion-Coherent Human Video Generation via Structure-Appearance Decoupling"},"pdf":{"value":"/pdf/b834dbf8a9f4c9c235dd39a566520e3602a75a37.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"wang|mosa_motioncoherent_human_video_generation_via_structureappearance_decoupling"},"authorids":{"value":["~Haoyu_Wang13","~Hao_Tang6","~Donglin_Di1","~Zhilu_Zhang2","~Wangmeng_Zuo3","~Feng_Gao8","~Siwei_Ma1","~Shiliang_Zhang2"]},"authors":{"value":["Haoyu Wang","Hao Tang","Donglin Di","Zhilu Zhang","Wangmeng Zuo","Feng Gao","Siwei Ma","Shiliang Zhang"]}},"version":2},{"content":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Large Language Model","Reinforcement Learning","Self-improving","Self-evolution"]},"primary_area":{"value":"foundation or frontier models, including LLMs"},"abstract":{"value":"Reinforcement learning (RL) has demonstrated potential in enhancing the reasoning capabilities of large language models (LLMs), but such training typically demands substantial efforts in creating and annotating data. In this work, we explore improving LLMs through RL with minimal data. Our approach alternates between the LLM proposing a task and then attempting to solve it. To minimize data dependency, we introduce two novel mechanisms grounded in self-awareness: (1) self-aware difficulty prediction, where the model learns to assess task difficulty relative to its own abilities and prioritize challenging yet solvable tasks, and (2) self-aware limit breaking, where the model recognizes when a task is beyond its capability boundary and proactively requests external data to break through that limit. Extensive experiments on nine benchmarks showing a 53.8% relative improvement with less than 1.2% extra data demonstrate the efficacy of self-aware RL and underscore the promise of self-evolving agent training. Our code is available at https://anonymous.4open.science/r/SARL-B7FE/."},"_bibtex":{"value":"@misc{\nzhang2026selfaware,\ntitle={Self-Aware Reinforcement Learning for Improving {LLM}s with Minimal Data},\nauthor={Hangfan Zhang and Siyuan Xu and Zhimeng Guo and Huaisheng Zhu and Shicheng Liu and Xinrun Wang and Qiaosheng Zhang and Yang Chen and Peng Ye and LEI BAI and Shuyue Hu},\nyear={2026},\nurl={https://openreview.net/forum?id=k3ylkWMJAc}\n}"},"title":{"value":"Self-Aware Reinforcement Learning for Improving LLMs with Minimal Data"},"pdf":{"value":"/pdf/ea41327a5c14c4efc734dff643439272ae0657ed.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"zhang|selfaware_reinforcement_learning_for_improving_llms_with_minimal_data"},"authorids":{"value":["~Hangfan_Zhang1","~Siyuan_Xu4","~Zhimeng_Guo1","~Huaisheng_Zhu1","~Shicheng_Liu1","~Xinrun_Wang1","~Qiaosheng_Zhang2","~Yang_Chen12","~Peng_Ye4","~LEI_BAI1","~Shuyue_Hu1"]},"authors":{"value":["Hangfan Zhang","Siyuan Xu","Zhimeng Guo","Huaisheng Zhu","Shicheng Liu","Xinrun Wang","Qiaosheng Zhang","Yang Chen","Peng Ye","LEI BAI","Shuyue Hu"]}},"tmdate":1770805021244,"tcdate":1758242327788,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission14713/Authors"],"signatures":["ICLR.cc/2026/Conference/Submission14713/Authors"],"forum":"k3ylkWMJAc","license":"CC BY 4.0","number":14713,"cdate":1758242327788,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/-/Submission","ICLR.cc/2026/Conference/-/Post_Submission","ICLR.cc/2026/Conference/Submission14713/-/Full_Submission","ICLR.cc/2026/Conference/-/Edit"],"mdate":1770805021244,"odate":1759896705795,"domain":"ICLR.cc/2026/Conference","id":"k3ylkWMJAc","version":2},{"content":{"venue":{"value":"CDC 2018"},"pdf":{"value":"https://ieeexplore.ieee.org/iel7/8592870/8618647/08618951.pdf"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"sarkar|minimal_realization_problems_for_jump_linear_systems"},"html":{"value":"https://doi.org/10.1109/CDC.2018.8618951"},"_bibtex":{"value":"@inproceedings{DBLP:conf/cdc/SarkarRD18,\n  author={Tuhin Sarkar and Mardavij Roozbehani and Munther A. Dahleh},\n  title={Minimal Realization Problems for Jump Linear Systems},\n  year={2018},\n  cdate={1514764800000},\n  pages={5670-5675},\n  url={https://doi.org/10.1109/CDC.2018.8618951},\n  booktitle={CDC},\n  crossref={conf/cdc/2018}\n}\n"},"abstract":{"value":"This paper addresses two fundamental problems in the context of jump linear systems (JLS). The first problem is concerned with characterizing the minimal state space dimension solely from input-output pairs and without any knowledge of the number of mode switches. The second problem is concerned with characterizing the number of discrete modes of the JLS. For the first problem, we develop a linear system theory based approach and construct an appropriate Hankel–like matrix. The rank of this matrix gives us the state space dimension. For the second problem we show that minimal number of modes corresponds to the minimal rank of a positive semi–definite matrix obtained via a non–convex formulation."},"title":{"value":"Minimal Realization Problems for Jump Linear Systems"},"authors":{"value":[{"fullname":"Tuhin Sarkar","username":""},{"fullname":"Mardavij Roozbehani","username":"~Mardavij_Roozbehani1"},{"fullname":"Munther A. Dahleh","username":""}]}},"tmdate":1778183883284,"pdate":1546214400000,"externalIds":["dblp:conf/cdc/SarkarRD18"],"tcdate":1778183878833,"writers":["~","OpenReview.net/Public_Article/DBLP.org","OpenReview.net/Support"],"signatures":["~Mardavij_Roozbehani1"],"forum":"pJDUR0CjWR","license":"CC BY-SA 4.0","number":13311,"cdate":1514764800000,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/DBLP.org/-/Record","OpenReview.net/Public_Article/-/Edit"],"mdate":1778183883284,"domain":"OpenReview.net/Public_Article","id":"pJDUR0CjWR","version":2},{"content":{"venue":{"value":"GECCO (Companion) 2018"},"venueid":{"value":"dblp.org/conf/GECCO/2018"},"paperhash":{"value":"parque|towards_bundling_minimal_trees_in_polygonal_maps"},"authorids":{"value":["~Victor_Parque1","https://dblp.org/search/pid/api?q=author:Tomoyuki_Miyashita:"]},"html":{"value":"https://doi.org/10.1145/3205651.3208316"},"_bibtex":{"value":"@inproceedings{DBLP:conf/gecco/ParqueM18a,\n  author={Victor Parque and Tomoyuki Miyashita},\n  title={Towards bundling minimal trees in polygonal maps},\n  year={2018},\n  cdate={1514764800000},\n  pages={1813-1820},\n  url={https://doi.org/10.1145/3205651.3208316},\n  booktitle={GECCO (Companion)},\n  crossref={conf/gecco/2018c}\n}\n"},"abstract":{"value":"Minimal trees in polygonal maps aim at minimizing the connectivity in a network while avoiding obstacle collision. Being closely related to the Steiner Tree Problem, yet with a different scope, minimal trees aim at connecting origin-destination pairs, given in a bipartite network, to allow the joint transport of information, goods, resources and people. In this paper, we propose a method to tackle the bundling problem of minimal trees in modular bipartite networks by using a two-layer optimization based on Differential Evolution with a convex representation of coordinates. Our computational experiments in polygonal domains considering both convex and non-convex geometry show the feasibility and the efficiency of the proposed approach."},"title":{"value":"Towards bundling minimal trees in polygonal maps"},"authors":{"value":["Victor Parque","Tomoyuki Miyashita"]}},"tmdate":1738825631391,"pdate":1514764800000,"tcdate":1738825557529,"writers":["~"],"signatures":["~Victor_Parque1"],"forum":"IA0bla0tY0","license":"CC BY-SA 4.0","number":303670,"cdate":1514764800000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1738825631391,"domain":"DBLP.org","id":"IA0bla0tY0","version":2},{"content":{"venue":{"value":"CVPR 2025"},"pdf":{"value":"https://openaccess.thecvf.com/content/CVPR2025/papers/Jeong_Track4Gen_Teaching_Video_Diffusion_Models_to_Track_Points_Improves_Video_CVPR_2025_paper.pdf"},"venueid":{"value":"dblp.org/conf/CVPR/2025"},"paperhash":{"value":"jeong|track4gen_teaching_video_diffusion_models_to_track_points_improves_video_generation"},"authorids":{"value":["~Hyeonho_Jeong1","~Chun-Hao_P._Huang1","~Jong_Chul_Ye1","~Niloy_J._Mitra1","https://dblp.org/search/pid/api?q=author:Duygu_Ceylan:"]},"html":{"value":"https://openaccess.thecvf.com/content/CVPR2025/html/Jeong_Track4Gen_Teaching_Video_Diffusion_Models_to_Track_Points_Improves_Video_CVPR_2025_paper.html"},"_bibtex":{"value":"@inproceedings{DBLP:conf/cvpr/JeongHYMC25,\n  author={Hyeonho Jeong and Chun-Hao P. Huang and Jong Chul Ye and Niloy J. Mitra and Duygu Ceylan},\n  title={Track4Gen: Teaching Video Diffusion Models to Track Points Improves Video Generation},\n  year={2025},\n  cdate={1735689600000},\n  pages={7276-7287},\n  url={https://openaccess.thecvf.com/content/CVPR2025/html/Jeong_Track4Gen_Teaching_Video_Diffusion_Models_to_Track_Points_Improves_Video_CVPR_2025_paper.html},\n  booktitle={CVPR},\n  crossref={conf/cvpr/2025}\n}\n"},"abstract":{"value":"While recent foundational video generators produce visually rich output, they still struggle with appearance drift, where objects gradually degrade or change inconsistently across frames, breaking visual coherence. We hypothesize that this is because there is no explicit supervision in terms of spatial tracking at the feature level. We propose Track4Gen, a spatially aware video generator that combines video diffusion loss with point tracking across frames, providing enhanced spatial supervision on the diffusion features. Track4Gen merges the video generation and point tracking tasks into a single network by making minimal changes to existing video generation architectures. Using Stable Video Diffusion as a backbone, Track4Gen demonstrates that it is possible to unify video generation and point tracking, which are typically handled as separate tasks. Our extensive evaluations show that Track4Gen effectively reduces appearance drift, resulting in temporally stable and visually coherent video generation. Project page: https://hyeonho99.github.io/track4gen/"},"title":{"value":"Track4Gen: Teaching Video Diffusion Models to Track Points Improves Video Generation"},"authors":{"value":["Hyeonho Jeong","Chun-Hao P. Huang","Jong Chul Ye","Niloy J. Mitra","Duygu Ceylan"]}},"tmdate":1772729918628,"pdate":1735689600000,"externalIds":["dblp:conf/cvpr/JeongHYMC25"],"tcdate":1754178814461,"writers":["~"],"signatures":["~Hyeonho_Jeong1"],"forum":"sbwXEUtRpw","license":"CC BY-SA 4.0","number":611965,"cdate":1735689600000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1772729918628,"domain":"DBLP.org","id":"sbwXEUtRpw","version":2},{"content":{"venue":{"value":"CoRR 2024"},"pdf":{"value":"http://arxiv.org/pdf/2412.06016v3"},"venueid":{"value":"dblp.org/journals/CORR/2024"},"paperhash":{"value":"jeong|track4gen_teaching_video_diffusion_models_to_track_points_improves_video_generation"},"authorids":{"value":["~Hyeonho_Jeong1","https://dblp.org/search/pid/api?q=author:Chun-Hao_Paul_Huang:","~Jong_Chul_Ye1","~Niloy_J._Mitra1","https://dblp.org/search/pid/api?q=author:Duygu_Ceylan:"]},"html":{"value":"https://doi.org/10.48550/arXiv.2412.06016"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2412-06016,\n  publtype={informal},\n  author={Hyeonho Jeong and Chun-Hao Paul Huang and Jong Chul Ye and Niloy J. Mitra and Duygu Ceylan},\n  title={Track4Gen: Teaching Video Diffusion Models to Track Points Improves Video Generation},\n  year={2024},\n  cdate={1704067200000},\n  journal={CoRR},\n  volume={abs/2412.06016},\n  url={https://doi.org/10.48550/arXiv.2412.06016}\n}\n"},"abstract":{"value":"While recent foundational video generators produce visually rich output, they still struggle with appearance drift, where objects gradually degrade or change inconsistently across frames, breaking visual coherence. We hypothesize that this is because there is no explicit supervision in terms of spatial tracking at the feature level. We propose Track4Gen, a spatially aware video generator that combines video diffusion loss with point tracking across frames, providing enhanced spatial supervision on the diffusion features. Track4Gen merges the video generation and point tracking tasks into a single network by making minimal changes to existing video generation architectures. Using Stable Video Diffusion as a backbone, Track4Gen demonstrates that it is possible to unify video generation and point tracking, which are typically handled as separate tasks. Our extensive evaluations show that Track4Gen effectively reduces appearance drift, resulting in temporally stable and visually coherent video generation. Project page: hyeonho99.github.io/track4gen"},"title":{"value":"Track4Gen: Teaching Video Diffusion Models to Track Points Improves Video Generation"},"authors":{"value":["Hyeonho Jeong","Chun-Hao Paul Huang","Jong Chul Ye","Niloy J. Mitra","Duygu Ceylan"]}},"tmdate":1772729918319,"pdate":1704067200000,"externalIds":["dblp:journals/corr/abs-2412-06016"],"tcdate":1754178814697,"writers":["~"],"signatures":["~Hyeonho_Jeong1"],"forum":"4x5aMXlQn3","license":"CC BY-SA 4.0","number":611966,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1772729918319,"domain":"DBLP.org","id":"4x5aMXlQn3","version":2},{"content":{"summary":{"value":"This paper presents several architecture modifications for video VAE typically used for latent video diffusion model training. The modifications are based on two observations: (a) a naive initialization from image VAE that uses the same latent dimension for video VAE is suboptimal, (b) and the per-frame reconstruction within the same block processed per each causal step (e.g., 4 frames) is not equal. Thus, the authors introduce an alternative initialization that uses half of the latent dimension for spatial compression and half of the latent dimension for temporal compression. They also propose several blocks to mitigate different per-frame compression in the same block of consecutive frames. Based on these modifications, the authors argue the proposed architecture outperform previous video VAEs and can improve the video generation performance."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"- Will the authors release the VAE checkpoint and code for reproducibility?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"- The paper is well-motivated, and the observations are interesting (although these might be quite trivial).\n- The proposed method seems to show clear improvement compared with existing video VAEs. \n- A new benchmark to evaluate video VAEs on high-resolution and dynamic video is interesting.\n- The paper tries to demonstrate the effectiveness of the proposed VAE in video generation, not limited to showing better video reconstruction."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Overall, I am positive about the paper, but in its current status, I cannot give it a very high rating because my contribution to the community is a bit marginal. Specifically, although the paper shows an improved performance compared with baselines under the same compression ratio, the proposed method does not compress the video more extensively (like more temporal compression than 4x). It can improve the video generation quality slightly (as shown by video generation results on Kinetics-600 and Skytimelapse), but the proposed video VAE still suffers from a fundamental problem of video generation models: handling memory or computation burdens. That being said, I think some observations and architectural modifications will be interesting to the community.\n- Due to the increased number of layers (e.g., 2D + 3D in KTC unit), I expect the encoding/decoding time would be slower than other autoencoder variants. The paper seems to compare the time only with CogX-VAE on Table 6, but I think this part should be analyzed more extensively. \n- The hyperparameter details for the diffusion model (using Latte) experiments are missing\n- Visualizing videos as a stack of image frames is really a bad direction to visualize the video results; could the authors provide exact video files in the supplementary material or anonymized site? I've observed so many cases that two of them are not corresponds very well to each other. \n- Also qualitative comparison on video reconstruction and generation is insufficient; could the authors add more examples?"}},"nonreaders":[],"tmdate":1731427508029,"tcdate":1730616351486,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission1903/Reviewer_RwZv"],"signatures":["ICLR.cc/2025/Conference/Submission1903/Reviewer_RwZv"],"forum":"e5288Iu4Zc","number":1,"license":"CC BY 4.0","cdate":1730616351486,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission1903/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427508029,"domain":"ICLR.cc/2025/Conference","replyto":"e5288Iu4Zc","id":"DjsZINkwMn","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video VAE","Variational Autoencoder"]},"primary_area":{"value":"generative models"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Variational Autoencoder (VAE) aims to compress pixel data into low-dimensional latent space, playing an important role in OpenAI's Sora and other latent video diffusion generation models. While most of existing video VAEs inflate a pretrained image VAE into the 3D causal structure for temporal-spatial compression, this paper presents two astonishing findings: (1) The initialization from a well-trained image VAE with the same latent dimensions suppresses the improvement of subsequent temporal compression capabilities. (2) The adoption of causal reasoning leads to unequal information interactions and unbalanced performance between frames. To alleviate these problems, we propose a keyframe-based temporal compression (KTC) architecture and a group causal convolution (GCConv) module to further improve video VAE (IV-VAE). Specifically, the KTC architecture divides the latent space into two branches, in which one half completely inherits the compression prior of keyframes from lower-dimension image VAEs while the other half involves temporal compression into the 3D group causal convolution, reducing temporal-spatial conflicts and accelerating the convergence speed of video VAE. The GCConv in above 3D half uses standard convolution within each frame group to ensure inter-frame equivalence, and employs causal logical padding between groups to maintain flexibility in processing variable frame video. Extensive experiments on five benchmarks demonstrate the SOTA video reconstruction and generation capabilities of the proposed IV-VAE. The source code and weights will be made available to the public."},"_bibtex":{"value":"@misc{\nwu2024improved,\ntitle={Improved Video {VAE} for Latent Video Diffusion Model},\nauthor={Pingyu Wu and Kai Zhu and Yu Liu and Liming Zhao and Wei Zhai and Yang Cao and Zheng-Jun Zha},\nyear={2024},\nurl={https://openreview.net/forum?id=e5288Iu4Zc}\n}"},"title":{"value":"Improved Video VAE for Latent Video Diffusion Model"},"pdf":{"value":"/pdf/ab776342023195907bcd087e1ac2d0423b4dea80.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"wu|improved_video_vae_for_latent_video_diffusion_model"},"authorids":{"value":["~Pingyu_Wu1","~Kai_Zhu4","~Yu_Liu23","~Liming_Zhao1","~Wei_Zhai1","~Yang_Cao5","~Zheng-Jun_Zha2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Pingyu Wu","Kai Zhu","Yu Liu","Liming Zhao","Wei Zhai","Yang Cao","Zheng-Jun Zha"]}},"version":2},{"content":{"summary":{"value":"This paper proposes an online retrieval-augmented dense video captioning approach that utilizes action phrases with text assistance. Additionally, the authors introduce image-based simulated video pretraining, which reduces reliance on extensive video datasets by using image-level text-paired data aligned with the online video captioning format. They conducted experiments on the ViTT, YouCook2, and ActivityNet benchmarks to validate its effectiveness."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. Could the authors provide a comparison of inference times in the tables, considering that both retrieval and streaming output are time-consuming processes, to demonstrate the model's performance in practical use cases?\n2. Could you further elaborate on the segmentation strategy, particularly regarding cases where the same sequence is divided into different segments?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. This work focuses on dense video captioning in video streaming setting. The problem to be studied is explained very clearly.\n2. The figures are informative and the tables are well-organized.\n3. The ablation experiments effectively validate the effectiveness of online action retrieval and integration and image-based simulated video pretraining.\n4. The authors conducted comparisons with existing dense video caption models across multiple benchmarks, validating the effectiveness of their proposed model."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The authors need to provide a clearer explanation of the segmentation strategy. If a sequence is divided into different segments, does that imply it represents two distinct actions? This seems rather counterintuitive.\n2. The authors should scale their validation methods to longer video sequences. The motivation of the paper is to address evolving contexts in video streams through an online strategy. Therefore, using longer video lengths, such as those found in ego datasets ego4d[1],egoSchema[2] ranging from dozens of minutes to hours, would better validate the effectiveness of their approach.\n\n[1] Ego4D: Around the World in 3,000 Hours of Egocentric Video. https://arxiv.org/pdf/2110.07058\n[2] EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding. https://egoschema.github.io"}},"nonreaders":[],"tmdate":1731427795467,"tcdate":1730293885863,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission3294/Reviewer_UpTg"],"signatures":["ICLR.cc/2025/Conference/Submission3294/Reviewer_UpTg"],"forum":"oO3oXJ19Pb","number":1,"license":"CC BY 4.0","cdate":1730293885863,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission3294/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427795467,"domain":"ICLR.cc/2025/Conference","replyto":"oO3oXJ19Pb","id":"gEPnIDtR0I","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Dense video captioning","Online dense video captioning"]},"supplementary_material":{"value":"/attachment/22be9efeac9199c2db56442e763f88fbc635769f.pdf"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Dense video captioning requires solving the challenging tasks of temporally localizing events and generating descriptive captions within long video sequences. Existing methods often struggle to capture the evolving context within video streams and to produce accurate temporal alignment. To address this, we propose an online retrieval-augmented approach that processes video segments incrementally while dynamically retrieving relevant action phrases from a pre-constructed action-text corpus. This enriches the contextual information for both the video representation and the subsequent text decoder, improving the caption generation. Additionally, we present image-based simulated video pretraining, which mitigates the reliance on extensive video datasets by using image-level text-paired data aligned with the online video captioning format. Our experiments on the ViTT, YouCook2, and ActivityNet benchmarks demonstrate that our model significantly outperforms both existing global and online methods, validating its effectiveness."},"_bibtex":{"value":"@misc{\nkim2024actions,\ntitle={Actions Inspire Every Moment: Online Action-Augmented Dense Video Captioning},\nauthor={Dahun Kim and AJ Piergiovanni and Anelia Angelova},\nyear={2024},\nurl={https://openreview.net/forum?id=oO3oXJ19Pb}\n}"},"title":{"value":"Actions Inspire Every Moment: Online Action-Augmented Dense Video Captioning"},"pdf":{"value":"/pdf/ced2509c52424b1ede3d852681c4a02933dc58a7.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"kim|actions_inspire_every_moment_online_actionaugmented_dense_video_captioning"},"authorids":{"value":["~Dahun_Kim1","~AJ_Piergiovanni1","~Anelia_Angelova1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Dahun Kim","AJ Piergiovanni","Anelia Angelova"]}},"version":2},{"content":{"venue":{"value":"CVPR 2023"},"pdf":{"value":"https://ieeexplore.ieee.org/iel7/10203037/10203050/10204030.pdf"},"venueid":{"value":"dblp.org/conf/CVPR/2023"},"paperhash":{"value":"wang|adapting_shortcut_with_normalizing_flow_an_efficient_tuning_framework_for_visual_recognition"},"authorids":{"value":["~Yaoming_Wang1","https://dblp.org/search/pid/api?q=author:Bowen_Shi:","https://dblp.org/search/pid/api?q=author:Xiaopeng_Zhang_0008:","https://dblp.org/search/pid/api?q=author:Jin_Li:","~Yuchen_Liu4","https://dblp.org/search/pid/api?q=author:Wenrui_Dai:","~Chenglin_Li2","https://dblp.org/search/pid/api?q=author:Hongkai_Xiong:","https://dblp.org/search/pid/api?q=author:Qi_Tian_0001:"]},"html":{"value":"https://doi.org/10.1109/CVPR52729.2023.01532"},"_bibtex":{"value":"@inproceedings{DBLP:conf/cvpr/WangSZLLDLXT23,\n  author={Yaoming Wang and Bowen Shi and Xiaopeng Zhang and Jin Li and Yuchen Liu and Wenrui Dai and Chenglin Li and Hongkai Xiong and Qi Tian},\n  title={Adapting Shortcut with Normalizing Flow: An Efficient Tuning Framework for Visual Recognition},\n  year={2023},\n  cdate={1672531200000},\n  pages={15965-15974},\n  url={https://doi.org/10.1109/CVPR52729.2023.01532},\n  booktitle={CVPR},\n  crossref={conf/cvpr/2023}\n}\n"},"abstract":{"value":"Pretraining followed by fine-tuning has proven to be effective in visual recognition tasks. However, fine-tuning all parameters can be computationally expensive, particularly for large-scale models. To mitigate the computational and storage demands, recent research has explored Parameter-Efficient Fine-Tuning (PEFT), which focuses on tuning a minimal number of parameters for efficient adaptation. Existing methods, however, fail to analyze the impact of the additional parameters on the model, resulting in an unclear and suboptimal tuning process. In this paper, we introduce a novel and effective PEFT paradigm, named SNF (Shortcut adaptation via Normalization Flow), which utilizes normalizing flows to adjust the shortcut layers. We highlight that layers without Lipschitz constraints can lead to error propagation when adapting to downstream datasets. Since modifying the over-parameterized residual connections in these layers is expensive, we focus on adjusting the cheap yet crucial shortcuts. Moreover, learning new information with few parameters in PEFT can be challenging, and information loss can result in label information degradation. To address this issue, we propose an information-preserving normalizing flow. Experimental results demonstrate the effectiveness of SNF. Specifically, with only 0.036M parameters, SNF surpasses previous approaches on both the FGVC and VTAB-1k benchmarks using ViT/B-16 as the backbone. The code is available at https://github.com/Wang-Yaoming/SNF"},"title":{"value":"Adapting Shortcut with Normalizing Flow: An Efficient Tuning Framework for Visual Recognition"},"authors":{"value":["Yaoming Wang","Bowen Shi","Xiaopeng Zhang","Jin Li","Yuchen Liu","Wenrui Dai","Chenglin Li","Hongkai Xiong","Qi Tian"]}},"tmdate":1762323036421,"pdate":1672531200000,"tcdate":1731503652347,"writers":["~"],"signatures":["~Chenglin_Li2"],"forum":"phV9MeM5vZ","license":"CC BY-SA 4.0","number":227415,"cdate":1672531200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1762323036421,"domain":"DBLP.org","id":"phV9MeM5vZ","version":2},{"content":{"summary":{"value":"This paper proposes a new video VAE for video generation by incorporating design principles from traditional video codec standards into the VAE architecture. It initializes the model using pretrained weights from an image VAE and employs Temporal Dynamic Difference Convolution (TDC) to capture temporal dynamics. Evaluations on several benchmarks demonstrate that the proposed method outperforms existing video VAEs."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"+ Reconstruction videos should be provided.\n+ A video model needs to be trained based on the proposed video VAE."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":3},"strengths":{"value":"+ Integrates traditional video codec design principles into the VAE framework.\n+ Carefully initializes the model with pretrained image VAE weights.\n+ Achieves superior performance over other video VAEs across multiple benchmarks."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The performance improvements on WebVid-10M and UCF-101 appear marginal compared to state-of-the-art video VAEs when FCR is 4×8×8, despite strong text reconstruction results. The reason for this inconsistency is not clearly explained.\n- The supplementary materials do not include video samples, making it difficult to assess temporal consistency.\n- Although the title claims the VAE is designed for latent video diffusion models, there is no experiment training a full video generation model based on the proposed VAE. Thus, it remains unclear how well the VAE performs in actual video generation tasks."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762915460682,"tcdate":1761727951440,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission162/Reviewer_zj9n"],"signatures":["ICLR.cc/2026/Conference/Submission162/Reviewer_zj9n"],"forum":"UBsmQXhXg8","number":1,"license":"CC BY 4.0","cdate":1761727951440,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission162/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762915460682,"domain":"ICLR.cc/2026/Conference","replyto":"UBsmQXhXg8","id":"XIjdYgw3Ix","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Video VAE"]},"primary_area":{"value":"generative models"},"abstract":{"value":"Video Variational Auto-Encoders (Video VAEs) compress video data from the highly redundant pixel space into a compact latent representation, playing an important role in state-of-the-art video generation models. However, existing methods typically learn inter-frame correlations implicitly, overlooking the potential of breaking down video compression into two separate parts: keyframe encoding and inter-frame dynamic encoding, which is a fundamental design of traditional video codecs. To address this, we incorporate traditional video codec standard design into the Video VAE and introduce VC-VAE, a model that explicitly separates keyframe and inter-frame dynamic compression. We start by establishing a high-fidelity static keyframe anchor through initialization from a powerful pre-trained image VAE. Then, to explicitly model dynamic relative to this anchor, we introduce the Temporal Dynamic Difference Convolution (TDC), an operator designed to learn sparse motion residuals from inter-frame differences while maintaining a separate pathway for static content. Qualitative and quantitative experiments show that our proposed VC-VAE significantly outperforms baseline models in reconstruction quality, dynamic modelling, and training efficiency."},"_bibtex":{"value":"@misc{\nge2025vcvae,\ntitle={{VC}-{VAE}: Enhancing Video {VAE} with Video Codec Standard for Latent Video Diffusion Model},\nauthor={Xinxu Ge and Shang Chai and Litong Gong and Zitong YU and Xin Liu and Tiezheng Ge},\nyear={2025},\nurl={https://openreview.net/forum?id=UBsmQXhXg8}\n}"},"title":{"value":"VC-VAE: Enhancing Video VAE with Video Codec Standard for Latent Video Diffusion Model"},"pdf":{"value":"/pdf/6bc69a0444086367135743c6092fe44ec0a7ab3f.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"ge|vcvae_enhancing_video_vae_with_video_codec_standard_for_latent_video_diffusion_model"},"authorids":{"value":["~Xinxu_Ge1","~Shang_Chai1","~Litong_Gong1","~Zitong_YU3","~Xin_Liu25","~Tiezheng_Ge3"]},"authors":{"value":["Xinxu Ge","Shang Chai","Litong Gong","Zitong YU","Xin Liu","Tiezheng Ge"]}},"version":2},{"content":{"summary":{"value":"This paper presents EPiC (Efficient and Precise Camera Control), a framework for efficient learning of camera motion control in video diffusion models (VDMs). Instead of relying on point-cloud rendering and camera trajectory estimation—which often introduce pixel-level misalignment—the authors propose a visibility-based masking strategy to construct well-aligned anchor videos directly from source videos. They further introduce a lightweight Anchor-ControlNet, accounting for less than 1% of the backbone parameters, which conditions video generation on anchor videos through visibility-aware output masking. EPiC enables efficient training (only 5K videos and 500 iterations) and achieves state-of-the-art camera control accuracy on RealEstate10K and MiraData, while also demonstrating strong zero-shot generalization from image-to-video (I2V) to video-to-video (V2V) tasks. Extensive quantitative, qualitative, and ablation studies validate the proposed approach’s efficiency, robustness, and alignment advantages over existing baselines such as CameraCtrl, AC3D, and ViewCrafter."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"See Weakness."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. Clear and interpretable design.\n2. The idea of ​​converting the problem into anchor-video construction is very interesting."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Converting camera-guided video generation into the task of supplementing an anchor video is an interesting idea. However, when applied to I2V, although the authors preserve foreground dynamics by masking foreground regions during guidance, the background is effectively forced to remain static. This imposes a limitation when treating video generation as a world model. Moreover, the approach depends to some extent on the performance of foreground extraction models (e.g., *GroundingDINO* used in the paper), which requires users to provide additional inputs.\n\n2. In the Introduction, the authors state that the Anchor-ControlNet is “injected into the first 25% of backbone layers”, while the original ControlNet is injected into 50% of layers. What specific design or empirical analysis motivated selecting 25%? Some works (e.g., *InstantStyle*) experimentally show that particular layers predominantly control generation style, and therefore only control those layers. Did the authors observe similar layer-wise effects? Please provide more design details and/or empirical evidence supporting the 25% choice.\n\n3. Minor :\n\n   a. Line 76: change “We propose” → “we propose” (lowercase “we”).  \n   b. Make figure references consistent throughout the manuscript; do not mix full form and abbreviation. For example, Line 367 uses “Figure 1” while Line 375 uses “Fig. 4”. Please standardize notation across the paper."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762924356316,"tcdate":1760798831631,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission13832/Reviewer_uFzJ"],"signatures":["ICLR.cc/2026/Conference/Submission13832/Reviewer_uFzJ"],"forum":"LP947kVbT4","number":2,"license":"CC BY 4.0","cdate":1760798831631,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission13832/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762924356316,"domain":"ICLR.cc/2026/Conference","replyto":"LP947kVbT4","id":"wLZjV10Jj6","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Video Generation","Camera Control","Efficiency"]},"supplementary_material":{"value":"/attachment/cbdbed4553c2ea98cfa0d3b6f49a22bd069b2004.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Controlling camera motion in video diffusion models is highly sought after for content creation, yet remains a significant challenge.\nRecent approaches often create anchor videos\n(i.e., rendered videos that approximate desired camera motions)\nto guide diffusion models as a structured prior,\nby rendering from estimated point clouds following camera trajectories. \nHowever, errors in point cloud and camera trajectory estimation often lead to inaccurate anchor videos during training. Furthermore, these inherent errors lead to higher training cost and inefficiency, since the model is forced to compensate for rendering misalignments.\nTo address these limitations, we introduce EPiC, an efficient and precise camera control learning framework\nthat constructs well-aligned training anchor videos\nwithout the need for camera pose or point cloud estimation.\nConcretely, we create highly precise anchor videos by masking source videos based on first-frame visibility.\nThis approach ensures strong alignment, eliminates the need for camera/point cloud estimation, and thus can be readily applied to any in-the-wild video\nto generate image-to-video (I2V) training pairs.\nFurthermore, we introduce Anchor-ControlNet, a lightweight conditioning module that integrates anchor video guidance in visible regions to pretrained video diffusion models, with less than 1\\% of backbone model parameters. By combining the proposed anchor video data and ControlNet module, EPiC achieves efficient training with substantially fewer parameters, training steps, and less data, without requiring modifications to the diffusion model backbone.\nAlthough being trained on masking-based anchor videos, our method generalizes robustly to anchor videos made with point clouds at test time, enabling precise 3D-informed camera control.\nEPiC achieves state-of-the-art performance on RealEstate10K and MiraData for I2V camera control task, demonstrating precise and robust camera control ability both quantitatively and qualitatively.\nNotably, EPiC also exhibits strong zero-shot generalization to video-to-video (V2V) scenarios. This is compelling as it is trained exclusively on I2V data, where anchor videos are derived with only source videos' first frame as visibility referencing."},"_bibtex":{"value":"@misc{\nwang2026epic,\ntitle={{EP}iC: Efficient Video Camera Control Learning with Precise Anchor-Video Guidance},\nauthor={Zun Wang and Jaemin Cho and Jialu Li and Han Lin and Jaehong Yoon and Yue Zhang and Mohit Bansal},\nyear={2026},\nurl={https://openreview.net/forum?id=LP947kVbT4}\n}"},"title":{"value":"EPiC: Efficient Video Camera Control Learning with Precise Anchor-Video Guidance"},"pdf":{"value":"/pdf/f74c4bfd853beb2025dedd7b9e3209359a974e71.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"wang|epic_efficient_video_camera_control_learning_with_precise_anchorvideo_guidance"},"authorids":{"value":["~Zun_Wang1","~Jaemin_Cho1","~Jialu_Li2","~Han_Lin1","~Jaehong_Yoon1","~Yue_Zhang14","~Mohit_Bansal2"]},"authors":{"value":["Zun Wang","Jaemin Cho","Jialu Li","Han Lin","Jaehong Yoon","Yue Zhang","Mohit Bansal"]}},"version":2},{"content":{"summary":{"value":"This paper examines video generation models' ability to learn physical laws from visual data, testing their generalization across in-distribution (ID), out-of-distribution (OOD), and combinatorial scenarios. Claimed key contributions include:\n1. Evaluation Framework: A structured assessment for how well models follow physical laws across different data distributions.\n2. Scaling Insights: While data scaling improves ID generalization, it has minimal effect on OOD.\n3. Prioritization in Generalization: Models prioritize attributes (color > size > velocity > shape) when referencing training data.\nThese findings could provide some insights for building robust world models in video generation."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"Have you compared your findings with a long list of previous works on examining OOD and IND generalization of neural network models in general? For example, this paper:\nVisual Representation Learning Does Not Generalize Strongly Within the Same Domain"},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"The paper presents limited originality, as it largely builds on existing approaches to video generation and world modeling, applying these to well-studied physical scenarios without introducing novel methodologies or frameworks. (see weaknesses)\n\nHowever, it demonstrates good quality in experimental design and data scaling, using systematic evaluations across in-distribution, out-of-distribution, and combinatorial scenarios, which provides valuable insights into model limitations. \n\nThe **clarity of the paper is strong**, with well-defined sections on methodology, findings, and practical insights, making the results and their implications easily accessible to the reader. In terms of **significance**, the contributions are moderate, as the findings—while informative—are more incremental than groundbreaking. They highlight known limitations in model generalization rather than offering transformative advances for video generation or physical law discovery. \n\nOverall, the paper is a solid study with strengths in clarity and quality but limited novelty and moderate significance in its contributions to the field."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The paper's main weaknesses revolve around its misalignment between its stated focus on physical law discovery and the actual content of its experiments:\n\n1. Although the paper claims to investigate the learning of physical laws, the experiments primarily assess model sensitivity to superficial attributes like color, size, and shape. This focus on object properties, rather than fundamental dynamics such as force or acceleration, detracts from the paper’s relevance to physical law discovery, as these experiments fail to represent deeper physical principles.\n\n2.  The findings are primarily reiterations of known phenomena—such as the ease of in-distribution (ID) generalization versus the challenges of out-of-distribution (OOD) generalization. These are well-established issues in machine learning, and the paper does not contribute fresh insights specific to physical law generalization.\n\n3. Most observations (e.g., the prioritization of color over shape) apply to general video model generalization rather than to any specific domain of physics. Consequently, the paper’s insights into generalization mechanisms are not uniquely tied to physical laws, making the connection to physical modeling somewhat tenuous.\n\n4.  The paper appears to position itself as a study on physical law discovery, yet the experimental design and findings align more closely with a standard exploration of OOD and ID generalization in video generation. This disconnect may suggest an attempt to frame familiar findings in a new light, without sufficiently novel contributions to the study of physical law generalization.\n\nOverall, the paper does not fully deliver on its premise of exploring physical laws, focusing instead on known generalization issues in video models without advancing the state of research in physics-informed AI."}},"nonreaders":[],"tmdate":1731427313439,"tcdate":1730012618314,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission759/Reviewer_r3mF"],"signatures":["ICLR.cc/2025/Conference/Submission759/Reviewer_r3mF"],"forum":"ZyLkNVHBZF","number":1,"license":"CC BY 4.0","cdate":1730012618314,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission759/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427313439,"domain":"ICLR.cc/2025/Conference","replyto":"ZyLkNVHBZF","id":"nJEETfVJZQ","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"TLDR":{"value":"we conduct systematic experiments to investigate \"how far is video generation model from world model\" from the physical law perspetive."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["video generation","diffusion model","world model"]},"primary_area":{"value":"generative models"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"OpenAI's Sora highlights the potential of video generation for developing world models that adhere to fundamental physical laws. \nHowever, the ability of video generation models to discover such laws purely from visual data without human priors can be questioned.\nA world model learning the true law should give predictions robust to nuances and correctly extrapolate on unseen scenarios.\nIn this work, we evaluate across three key scenarios: in-distribution, out-of-distribution, and combinatorial generalization.\nWe developed a 2D simulation testbed for object movement and collisions to generate videos deterministically governed by one or more classical mechanics laws.\nThis provides unlimited supply of data for large-scale experimentation, and enables quantitative evaluation for the law in generated videos. \nWe trained diffusion-based video generation models to predict object movements based on initial frames.\nOur scaling experiments show perfect generalization within the distribution, measurable scaling behavior for combinatorial generalization, but failure in out-of-distribution scenarios.\nFurther experiments reveal two key insights about the generalization mechanisms of these models: (1) the models fail to abstract general physical rules and instead exhibit ``case-based'' generalization behavior, \\textit{i.e.}, mimicking the closest training example; (2) when generalizing to new cases, models are observed to prioritize different factors when referencing training data: color $>$ size $>$ velocity $>$ shape.\nOur study suggests that scaling alone is insufficient for video generation models to uncover fundamental physical laws, despite its role in Sora's broader success."},"_bibtex":{"value":"@misc{\nkang2025how,\ntitle={How Far Is Video Generation from World Model: A Physical Law Perspective},\nauthor={Bingyi Kang and Yang Yue and Rui Lu and Zhijie Lin and Yang Zhao and Kaixin Wang and Gao Huang and Jiashi Feng},\nyear={2025},\nurl={https://openreview.net/forum?id=ZyLkNVHBZF}\n}"},"title":{"value":"How Far Is Video Generation from World Model: A Physical Law Perspective"},"pdf":{"value":"/pdf/61786c392aeb1743704ab9483f7f67d63ce837d2.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"kang|how_far_is_video_generation_from_world_model_a_physical_law_perspective"},"authorids":{"value":["~Bingyi_Kang1","~Yang_Yue1","~Rui_Lu2","~Zhijie_Lin1","~Yang_Zhao14","~Kaixin_Wang1","~Gao_Huang1","~Jiashi_Feng1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Bingyi Kang","Yang Yue","Rui Lu","Zhijie Lin","Yang Zhao","Kaixin Wang","Gao Huang","Jiashi Feng"]}},"version":2},{"content":{"summary":{"value":"This paper introduces a reinforcement RL based framework for equipping MLLMs with long-video temporal grounding skills. The authors propose two main technical innovations: (1) Token-aware KL Regularization to balance exploration (for timestamp tokens) and exploitation (for general language tokens) during RL fine-tuning, and (2) a Center Distance Reward (CenDist) to provide denser learning signals for temporal localization. Additionally, they construct a new dataset, SceneTG, with sharp scene boundaries and unambiguous queries to facilitate effective RL training. The model achieves state-of-the-art results on several temporal grounding benchmarks and demonstrates that improved temporal localization skills generalize to long-video question answering (QA) tasks."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"Check above."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"The RL approach to long video temporal grounding is a nice idea an approach, a natural extension to techniques used for short-term videos. The details about token-aware KL regularization (theoretically motivated and practically effective according to the ablation) for balancing exploration and exploitation is interesting. Another good technical contribution is the CenDist reward addresses the reward sparsity problem, which is particularly acute in long-video settings. Additionally, the authors also construct a dataset of 16k samples which seems to be well constructed for the task and gives a boost in performance. The evaluation and ablation section are extensive also."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"While the idea is interesting and full of experiments, I believe there are some major issues:\n\n1 - Overclaim: The authors claim to outperform previous works on multiple out of domain datasets, but according to table 2 they perform better only in 2: ReXTime (significant improvement) and ANet (very marginal, +0.2), while being outperformed in other 4 datasets. To me this means that the method is actually not that effective and improvement are only marginal despite the use of also a custom dataset (while this data is smaller is also more qualitative). \n\n2 -  General QA with full video input (Table 3), the method is still outperformed in 2 out of 4 benchmarks.\n\n3 - Ablation study lack completeness: For example experiments are made only on 30% of the data, is not a big issue but given that the improvements are marginal maybe those improvements might not hold in full data. Additionally, an ablation on alpha is important to understand the effect of CenDist.\n\nWhile the contributions are really interesting, the claims are not fully supported by the results. My initial evaluation will be 4, the paper has a lot of things to improve but I am willing to increase my score if authors address my concerns."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762925944799,"tcdate":1761626991866,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission15694/Reviewer_R286"],"signatures":["ICLR.cc/2026/Conference/Submission15694/Reviewer_R286"],"forum":"8H1HmGH8ua","number":2,"license":"CC BY 4.0","cdate":1761626991866,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission15694/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762925944799,"domain":"ICLR.cc/2026/Conference","replyto":"8H1HmGH8ua","id":"EVc4iYjE9x","forumContent":{"TLDR":{"value":"We present LongVTG-R1, the first RL-based framework for long-video temporal grounding, which leverages novel regularization, reward design, and dataset construction to achieve state-of-the-art performance and generalize to QA tasks."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["video understanding","temporal grounding","multimodal large language model (MLLM)"]},"supplementary_material":{"value":"/attachment/a7526f2447a3b01b5d091f2498a198a2aae0ee49.zip"},"primary_area":{"value":"foundation or frontier models, including LLMs"},"abstract":{"value":"We present the first Reinforcement Learning (RL)-based framework that equips Multimodal Large Language Models (MLLMs) with long-video temporal grounding skills, and demonstrate that this approach also generalizes to improve performance on general question-answering (QA) tasks. Unlike dominant supervised fine-tuning (SFT) methods, RL enables models to acquire temporal grounding abilities without risking catastrophic forgetting of their core understanding. However, adopting RL for long-video temporal grounding reveals a challenge in balancing exploitation of pre-trained knowledge with exploration of new localization skills. To address this, we propose Token-aware KL Regularization, which selectively relaxes the KL-divergence regularization on timestamp-related tokens to guide exploration. Moreover, effective optimization requires a learning signal that alleviates the sparsity of key events in long videos, for which we introduce a denser reward, the Center Distance Reward (CenDist). To further mitigate grounding ambiguity between language queries and visually similar content, and to facilitate effective RL training, we propose an automatic data construction method and construct a small but high-quality dataset, SceneTG. Our resulting model, QwenLongTG, delivers substantial improvements across three long-video temporal grounding datasets among efficiently fine-tuned MLLMs, and even approaches the performance of densely pre-trained or continually trained models. Beyond temporal grounding, we further verify its generalization to long-video QA: under a “Ground-then-Answer” strategy, QwenLongTG consistently enhances downstream QA performance, serving as an effective first-stage grounding module."},"_bibtex":{"value":"@misc{\nzhang2026longvtgr,\ntitle={Long{VTG}-R1: Reinforcement Learning for Robust Long-Video Temporal Grounding},\nauthor={Zheyu Aqa Zhang and Shixing Chen and Ziqi Pang and Xiang Hao and Kushan Thakkar and Yu-Xiong Wang},\nyear={2026},\nurl={https://openreview.net/forum?id=8H1HmGH8ua}\n}"},"title":{"value":"LongVTG-R1: Reinforcement Learning for Robust Long-Video Temporal Grounding"},"pdf":{"value":"/pdf/1a44d3a0d532c7e585afb40db0defe29d293c3cd.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"zhang|longvtgr1_reinforcement_learning_for_robust_longvideo_temporal_grounding"},"authorids":{"value":["~Zheyu_Aqa_Zhang1","~Shixing_Chen1","~Ziqi_Pang1","~Xiang_Hao1","~Kushan_Thakkar1","~Yu-Xiong_Wang1"]},"authors":{"value":["Zheyu Aqa Zhang","Shixing Chen","Ziqi Pang","Xiang Hao","Kushan Thakkar","Yu-Xiong Wang"]}},"version":2},{"content":{"venue":{"value":"ICLR 2027 Conference Submission"},"keywords":{"value":["Quantization-Aware Training","Straight-Through Estimator","Bundle Methods","Frank-Wolfe Algorithm"]},"primary_area":{"value":"optimization"},"abstract":{"value":"The Straight-Through Estimator (STE) is a widely used heuristic for Quantization-Aware Training (QAT), but its surrogate gradients can exhibit substantial mismatch with the underlying quantized objective, leading to noisy updates and parameter oscillations, particularly in ultra-low-bit regimes. We propose the Quantization-Aware Minimal-Norm Optimizer (Q-MINO), a temporal bundle method that combines gradient consensus, state-drift regularization, and an alignment constraint to construct stabilized, minimum-norm update directions from recent optimization states. Q-MINO solves the resulting constrained subproblem using a warm-started Frank--Wolfe procedure with a feasible fallback initialization. Theoretically, via a stochastic quasi-Lyapunov Kurdyka--\\L ojasiewicz (KL) framework, we show that Q-MINO achieves almost sure convergence to a stationary point. Moreover, we detail numerical experiments with Q-MINO at various quantizations."},"_bibtex":{"value":"@inproceedings{\nanonymous2026qmino,\ntitle={Q-{MINO}: A Minimal-Norm Method for Quantization-Aware Training},\nauthor={Anonymous},\nbooktitle={Submitted to The Fifteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=XPAC09cMfg},\nnote={under review}\n}"},"title":{"value":"Q-MINO: A Minimal-Norm Method for Quantization-Aware Training"},"pdf":{"value":"/pdf/5434d6c932ce1923d6f7af4e5f29a2216f9e092b.pdf"},"venueid":{"value":"ICLR.cc/2027/Conference/Submission"}},"tmdate":1791235389470,"tcdate":1789810399755,"writers":["ICLR.cc/2027/Conference","ICLR.cc/2027/Conference/Submission58811/Authors"],"signatures":["ICLR.cc/2027/Conference/Submission58811/Authors"],"forum":"XPAC09cMfg","license":"CC BY 4.0","number":58811,"cdate":1789810399755,"readers":["everyone"],"invitations":["ICLR.cc/2027/Conference/-/Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Bidding","ICLR.cc/2027/Conference/Submission58811/-/Full_Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Reviewing","ICLR.cc/2027/Conference/-/Edit"],"mdate":1791235389470,"odate":1791031328286,"domain":"ICLR.cc/2027/Conference","id":"XPAC09cMfg","version":2},{"content":{"summary":{"value":"This paper presents a text-to-video generation model and a large-scale dataset of text/video pairs. The video generator is initialized from ModelScope (which uses the LDM backbone under the hood) and fine-tines on the collected dataset with a novel \"swapped spatio-temporal attention\" module. Then, a GAN-based super-resolution model is trained on top of it. For the dataset, the authors collect HD youtube videos, split them into clips with PySceneDetect and annotate the clips with BLIP-2 using the middle frame. For the video generator, it is then conducts a rigorous evaluation on UCF101, MSR-VTT, WebVid-10M and also includes human evaluation results."},"presentation":{"value":"3 good"},"contribution":{"value":"2 fair"},"soundness":{"value":"2 fair"},"strengths":{"value":"- The model is quite cheap to train, though it fine-tunes from ModelScope (VideoFusion), which fine-tunes from LDM.\n- The proposed HD-VG-130M dataset is a solid contribution and should help the community develop better generators.\n- The evaluation is quite thorough and encompasses text-video alignment, visual quality and the human studies. The paper also makes a good effort comparing to closed-source models (Make-a-Video and ImagenVideo). And also includes the failure case study.\n- I appreciate that training costs have been reported, and also throughpout/memory requirements. I believe that reporting such information to be crucial for foundational DL models.\n- The paper is well written and provides rich information about both the dataset, the main model, and the training information."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- When comparing FVDs on UCF-101, the paper included many previous works — but only if their FVD score is higher: https://paperswithcode.com/sota/video-generation-on-ucf-101. In this way, it omitted Make-A-Video, VDIM, LVDM, etc. despite comparing it in other regards. I do not see how it could be justified.\n- The contributions of the proposed components (datasets, new attention mechanism, etc) are entangled between each other and the ablations are confusing. Table 2 shows the scores for different components on WebVid, but there is no information about the ablation experiments: have each model was trained anew from ModelScope? For how many steps? Was it trained on webvid or HD-VG + LAION? I do not intuitively understand how swapped spatio-temporal attention can outperform just the \"brute-force\" global attention, if the latter one is given enough compute/training data.\n- The dataset annotation with captions is fairly simple and cannot capture any dynamics. For example, if a person comes to a table and takes a cup, then the annotation would be something like \"a person walks towards a table\", because only the middle frame was used for static caption annotation."},"confidence":{"value":"4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work."},"questions":{"value":"- Do I get it right that the model takes ~2.5 days of training on 32 V100s? Also, am I understanding this correctly that in total the model has been trained for ~40k steps with the batch size of 128, which implies that it has seen only ~5M videos (+ images)?\n- Why does Table 4 omit the methods which have better FVDs? (https://paperswithcode.com/sota/video-generation-on-ucf-101)\n- I suspect problems in human studies. I do not understand how Imagen-Video or Make-a-Video can be so much worse in terms of the performance in Table 7. Could you please share the videos which were used for the human studies?\n- For ablations, is it possible to train a model completely from scratch (without initializing it from anything) on some small-scale dataset, like UCF-101?"},"rating":{"value":"5: marginally below the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636616928,"tcdate":1698789824822,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission5839/Reviewer_NUsL"],"signatures":["ICLR.cc/2024/Conference/Submission5839/Reviewer_NUsL"],"forum":"dUDwK38MVC","number":4,"license":"CC BY 4.0","cdate":1698789824822,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission5839/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636616928,"domain":"ICLR.cc/2024/Conference","replyto":"dUDwK38MVC","id":"V6AVweiy7W","forumContent":{"TLDR":{"value":"Open-domain text-to-video generation with high-definition (1376×768), widescreen (16:9), and watermark-free advantages."},"venue":{"value":"ICLR 2024 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video Generation","Diffusion Models"]},"supplementary_material":{"value":"/attachment/aaf3d9311b58f7ff53b7f0bea6b9f15158737327.zip"},"primary_area":{"value":"generative models"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"We present VideoFactory, an innovative framework for generating high-quality open-domain videos. VideoFactory excels in producing high-definition (1376$\\times$768), widescreen (16:9) videos without watermarks, creating an engaging user experience. Generating videos guided by text instructions poses significant challenges, such as modeling the complex relationship between space and time, and the lack of large-scale text-video paired data. Previous approaches extend pretrained text-to-image generation models by adding temporal 1D convolution/attention modules for video generation. However, these approaches overlook the importance of jointly modeling space and time, inevitably leading to temporal distortions and misalignment between texts and videos. In this paper, we propose a novel approach that strengthens the interaction between spatial and temporal perceptions. In particular, we utilize a swapped cross-attention mechanism in 3D windows that alternates the “query” role between spatial and temporal blocks, enabling mutual reinforcement for each other. To fully unlock model capabilities for high-quality video generation, we curate a large-scale video dataset called HD-VG-130M. This dataset comprises 130 million text-video pairs from the open-domain, ensuring high-definition, widescreen and watermark-free characters. Objective metrics and user studies demonstrate the superiority of our approach in terms of per-frame quality, temporal correlation, and text-video alignment, with clear margins."},"_bibtex":{"value":"@misc{\nwang2024videofactory,\ntitle={VideoFactory: Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation},\nauthor={Wenjing Wang and Huan Yang and Zixi Tuo and Huiguo He and Junchen Zhu and Jianlong Fu and Jiaying Liu},\nyear={2024},\nurl={https://openreview.net/forum?id=dUDwK38MVC}\n}"},"title":{"value":"VideoFactory: Swap Attention in Spatiotemporal Diffusions for Text-to-Video Generation"},"pdf":{"value":"/pdf/14ece4592731b2ab0642015eb1769affda66a4e8.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Withdrawn_Submission"},"paperhash":{"value":"wang|videofactory_swap_attention_in_spatiotemporal_diffusions_for_texttovideo_generation"},"authorids":{"value":["~Wenjing_Wang1","~Huan_Yang4","~Zixi_Tuo2","~Huiguo_He3","~Junchen_Zhu1","~Jianlong_Fu1","~Jiaying_Liu1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Wenjing Wang","Huan Yang","Zixi Tuo","Huiguo He","Junchen Zhu","Jianlong Fu","Jiaying Liu"]}},"version":2},{"content":{"summary":{"value":"This paper introduces AdaFlow, a training-free method designed to address memory constraints in long video editing. AdaFlow incorporates two key strategies: Adaptive Attention Slimming, which selectively reduces token use in self-attention to decrease memory requirements, and Adaptive Keyframe Selection, which optimizes frame selection for enhanced editing quality. According to the authors, these techniques enable AdaFlow to handle videos exceeding 1,000 frames on a single A800 GPU, reportedly achieving lengths ten times greater than prior methods. The paper also presents LongV-EVAL, a benchmark for assessing long video edits, where AdaFlow demonstrates potential advantages in both efficiency and quality over existing approaches."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"See weakness."},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"- The paper presents AdaFlow, a training-free approach that effectively addresses memory constraints, offering a feasible alternative for handling longer videos compared to traditional methods.\n\n- AdaFlow combines Adaptive Attention Slimming and Adaptive Keyframe Selection to optimize memory usage and frame selection, respectively. This combined approach not only reduces computational load by focusing on essential tokens in self-attention but also enhances editing quality by selecting keyframes that capture critical scene changes.\n\n- The introduction of LongV-EVAL provides the field with a dedicated benchmark for long video editing, complete with detailed annotations and varied scenarios, which can serve as a valuable tool for assessing future developments in this area.\n\n- Initial results on LongV-EVAL indicate that AdaFlow may outperform existing methods in both efficiency and quality, positioning it as a promising approach for text-driven long video editing."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. **Insufficient Evidence for Claimed Contributions:** Although the paper lists three main contributions, some are not substantiated in sufficient detail. For instance, the claim of “effective” memory optimization mainly references spatial memory savings, with no discussion of runtime performance or the extra computation that Adaptive Keyframe Selection  and Adaptive Attention Slimming might add. Additionally, the benchmark’s description in the paper is limited, with definitions for key evaluation metrics lacking clarity.\n\n2. **Limited and Incomplete Experiment Comparisons:** The experiments primarily compare AdaFlow to methods focused on consistency in short video editing, lacking comparisons to dedicated long video editing techniques. Furthermore, the ablation study is minimal, with few quantitative measures and no in-depth analysis of other components within the approach. Relying on visual clarity in a few cases does not provide sufficient evidence of AdaFlow’s overall performance.\n\n3. **Unclear Visualization and Algorithm Details:** The pipeline visualization, particularly for the AKS module, is difficult to interpret, and specific steps like the \"window_check\" are not well-explained, leaving some ambiguity regarding the AKS process and its impact on overall results.\n\n4. **Limitations of Feature-Matched Latent Propagation:** If the intended edits involve adding new content or altering major background elements, rather than just subtle changes (e.g., color or style shifts), the proposed feature-matched propagation approach may fail to preserve coherence, potentially limiting its applicability for more complex or structural video modifications."}},"nonreaders":[],"tmdate":1732527229389,"tcdate":1730363591452,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission5625/Reviewer_EZqB"],"signatures":["ICLR.cc/2025/Conference/Submission5625/Reviewer_EZqB"],"forum":"yP0iKsinmk","number":3,"license":"CC BY 4.0","cdate":1730363591452,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission5625/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732527229389,"domain":"ICLR.cc/2025/Conference","replyto":"yP0iKsinmk","id":"3ATOvJvqml","forumContent":{"TLDR":{"value":"AdaFlow enables efficient, high-quality editing of minute-long videos by adaptively selecting keyframes and pruning redundant tokens, achieving state-of-the-art results on a single GPU."},"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["video editing","diffusion model","keyframe selection","token slimming"]},"supplementary_material":{"value":"/attachment/fbc8c97a3722b16443b9d1c6d1be25d05bd94a92.zip"},"primary_area":{"value":"generative models"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Text-driven video editing is an emerging research hot spot in deep learning. Despite great progress, long video editing is still notoriously challenging mainly due to excessive memory overhead. To tackle this problem, recent efforts have simplified this task into a two-step process of keyframe translation and interpolation generation, enabling the editing of more frames. However, the token-wise keyframe translation still plagues the upper limit of video length. In this paper, we propose a novel and training-free approach towards efficient and effective long video editing, termed AdaFlow. We first reveal that not all tokens of video frames hold equal importance for keyframe-consistency editing, based on which we propose an Adaptive Attention Slimming scheme for AdaFlow to squeeze the $KV$ sequence of extended self-attention. This enhancement allows AdaFlow to increase the number of keyframes for translations by an order of magnitude. In addition, an Adaptive Keyframe Selection scheme is also equipped to select the representative frames for joint editing, further improving generation quality. With these innovative designs, AdaFlow achieves high-quality long video editing of minutes in one inference, i.e., more than 1$k$ frames on one A800 GPU, which is about ten times longer than the compared methods. To validate AdaFlow, we also build a new benchmark for long video editing with high-quality annotations, termed LongV-EVAL. The experimental results show that our AdaFlow can achieve obvious advantages in both the efficiency and quality of long video editing. Our code is anonymously released at https://anonymous.4open.science/r/AdaFlow-C28F."},"_bibtex":{"value":"@misc{\nzhang2025adaflow,\ntitle={AdaFlow: Efficient Long Video Editing via Adaptive Attention Slimming And Keyframe Selection},\nauthor={Shuheng Zhang and Yuqi Liu and Hongbo Zhou and Jun Peng and Yiyi Zhou and Xiaoshuai Sun and Rongrong Ji},\nyear={2025},\nurl={https://openreview.net/forum?id=yP0iKsinmk}\n}"},"title":{"value":"AdaFlow: Efficient Long Video Editing via Adaptive Attention Slimming And Keyframe Selection"},"pdf":{"value":"/pdf/584fe8d5f50c3800c934f792d0940fe5a4784734.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"zhang|adaflow_efficient_long_video_editing_via_adaptive_attention_slimming_and_keyframe_selection"},"authorids":{"value":["~Shuheng_Zhang1","~Yuqi_Liu7","~Hongbo_Zhou3","~Jun_Peng2","~Yiyi_Zhou1","~Xiaoshuai_Sun3","~Rongrong_Ji5"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Shuheng Zhang","Yuqi Liu","Hongbo Zhou","Jun Peng","Yiyi Zhou","Xiaoshuai Sun","Rongrong Ji"]}},"version":2},{"content":{"summary":{"value":"This paper introduces DFVEdit, a novel zero-shot video editing method specifically tailored for modern Video Diffusion Transformers (Video DiTs). The core contribution is a flow-transformation framework that operates directly on latents, bypassing the need for computationally expensive attention modification or model finetuning. By unifying the editing and sampling processes under a continuous flow perspective, the method proposes a Conditional Delta Flow Vector (CDFV) to estimate the transformation from the source to the target video. This approach, enhanced by Implicit Cross-Attention (ICA) guidance and Embedding Reinforcement (ER), achieves state-of-the-art results in fidelity and consistency while offering speed-up and memory reduction."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Please see weaknesses"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"-The primary strength lies in its conceptual shift away from attention engineering. Instead of manipulating the internal query/key/value matrices, the method cleverly reframes editing as a continuous flow transformation in the latent space. Conditional Delta Flow Vector (CDFV) as a theoretically-backed estimate of the \"delta\" between the source and target latents is an clever thing that directly helps the use of Video DiTs. \n\n- The paper is well-written and clearly structured. The motivation for the work is well established  in the first figure.\n\n- The authors provide strong motivation for their work, focusing on the critical need for an efficient editing solution for Video DiTs. The evaluation is thorough, using a well-chosen suite of metrics to measure distinct properties: CLIP-F for temporal consistency, E_warp for motion fidelity, M.PSNR and LPIPS for fidelity and background preservation, and CLIP-T for prompt alignment. This comprehensive quantitative and qualitative analysis, along with a user study, makes a convincing case for the method's superiority."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The reported CLIP-F score of 0.9924 (Table 1) is exceptionally high. While this is presented as a strength, a score this close to 1.0 could imply that the edited frames are almost identical to each other, suggesting the CLIP-F metric may not be sensitive enough to detect subtle, fine-grained changes or potential flickering. It seems unlikely for a video to be meaningfully edited and still retain this level of inter-frame similarity unless the video was already static. It would be beneficial for the authors to provide a more detailed, one-by-one results breakdown for the evaluation dataset, or at least discuss this high score and why it doesn't indicate a lack of meaningful editing.\n\n- The comparison setup could be strengthened. The paper's method (DFVEdit) is applied to Video DiT backbones (CogVideoX-5B, Wan2.1-14B). However, many of the baselines (e.g., FateZero, TokenFlow, VideoDirector) are evaluated on their original, often U-Net-based backbones like Stable Diffusion 1.5 (as noted in Appendix D.5 and Table T2). While the paper does test extending some baselines to CogVideoX (Fig 1b) to show they are computationally infeasible, the primary qualitative and quantitative comparisons are between methods running on different underlying generative models. Text-to-Video (T2V) models or methods specifically designed for Video DiTs would serve as a more direct and fair comparison group to truly isolate the contribution of the editing method (DFVEdit) versus the power of the backbone model (Video DiTs).\n\n- The paper does not clearly state the total size of the evaluation dataset used for the main quantitative results in Table 1. Appendix D.1 mentions using the public DAVIS 2017 dataset and Pexels videos. Appendix D.2 mentions \"10 DAVIS videos, 30 diverse prompts\" for the M.PSNR metric, and Appendix D.3 mentions \"80 video-prompt pairs\" for the user study. However, the total number of videos used to calculate CLIP-F, E_warp, LPIPS, and CLIP-T in Table 1 is not specified. Clarifying the scale of the automated evaluation would help in assessing the robustness of the reported quantitative claims."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762925393605,"tcdate":1761855429280,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission15070/Reviewer_KCoA"],"signatures":["ICLR.cc/2026/Conference/Submission15070/Reviewer_KCoA"],"forum":"MnQD69han5","number":3,"license":"CC BY 4.0","cdate":1761855429280,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission15070/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762925393605,"domain":"ICLR.cc/2026/Conference","replyto":"MnQD69han5","id":"tyiUt8BjrP","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["zeroshot","video editing","traning free","video transformer"]},"supplementary_material":{"value":"/attachment/2784d3f149701c4cd1ba06007ab98c63d582b7a9.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"The advent of Video Diffusion Transformers (Video DiTs) marks a milestone in video generation. However, directly applying existing video editing methods to Video DiTs often incurs substantial computational overhead, due to resource-intensive attention modification or fine-tuning. To alleviate this problem, we present DFVEdit, an efficient zero-shot video editing method tailored for Video DiTs. DFVEdit eliminates the need for both attention engineering and fine-tuning by directly operating on clean latents via flow transformation. To be more specific, we observe that editing and sampling can be unified under the continuous flow perspective. Building upon this foundation, we propose the Conditional Delta Flow Vector (CDFV) -- a theoretically unbiased estimation of DFV -- and integrate Implicit Cross Attention (ICA) guidance as well as Embedding Reinforcement (ER) to further enhance editing quality. DFVEdit excels in practical efficiency, offering at least 20x inference speed-up and 85% memory reduction on Video DiTs compared to attention-engineering-based editing methods. Extensive quantitative and qualitative experiments demonstrate that DFVEdit can be seamlessly applied to popular Video DiTs (\\emph{e.g.}, CogVideoX and Wan2.1), attaining state-of-the-art performance on structural fidelity, spatial-temporal consistency, and editing quality."},"_bibtex":{"value":"@misc{\ncai2026dfvedit,\ntitle={{DFVE}dit: Conditional Delta Flow Vector for Zero-shot Video Editing},\nauthor={Lingling Cai and Kang Zhao and Hangjie Yuan and Xiang Wang and Yingya Zhang and Kejie Huang},\nyear={2026},\nurl={https://openreview.net/forum?id=MnQD69han5}\n}"},"title":{"value":"DFVEdit: Conditional Delta Flow Vector for Zero-shot Video Editing"},"pdf":{"value":"/pdf/2f7e77ef973cd7047441766f1d096bd246bc013c.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"cai|dfvedit_conditional_delta_flow_vector_for_zeroshot_video_editing"},"authorids":{"value":["~Lingling_Cai2","~Kang_Zhao7","~Hangjie_Yuan1","~Xiang_Wang9","~Yingya_Zhang3","~Kejie_Huang1"]},"authors":{"value":["Lingling Cai","Kang Zhao","Hangjie Yuan","Xiang Wang","Yingya Zhang","Kejie Huang"]}},"version":2},{"content":{"summary":{"value":"The paper introduces a benchmark for the evaluation of video foundation models (VFM). The benchmark is carefully curated based on a selection of existing video datasets. It is split into two sets: (1) a task adaptation benchmark (Video-TAB) that tests model performance on a variety of video tasks after lightweight adapter-based fine-tuning, and (2) an embedding benchmark (VidEB) that tests the performance of video embeddings, extracted using the benchmarked foundation models, on retrieval-related video tasks. Based on these benchmarks, the paper provides a survey of 20 current VFMs and offers insights into their applicability to different tasks. Furthermore, it discusses the effect of training data and training recipes on model performance."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"* Will this benchmark be publicly available? If so, this would be a major contribution and should be mentioned in the paper.  \n* Tab. 4 motivates the choice to use attentive probe as the adapter method of choice for the evaluation by showing its good cost-performace tradeoff. However, results are only shown on V-JEPA-H. Do the findings hold true for other methods?  \n* As a reader unfamiliar with the attentive probe adapter, I was not able to understand how it works from just Fig. 4\\. Could authors provide literature references and maybe expand the explanation in Sec. 3.1?  \n* Sec. 3.1: What other tasks / datasets were considered that did not end up in the benchmark? What were the reasons for them being excluded?  \n* In Tab. 5 it would be interesting to see which models perform well at spatial vs. temporal tasks. Have you considered dividing the evaluation results into the categories from Tab. 3?  \n* Fig. 3 motivates the choice of averaging three few-shot settings for scoring by showing that few-to-medium shot settings better separate the model accuracies. However this is only shown on two tasks. Do the other tasks exhibit the same separation properties?  \n* Sec. 4.1: How are the 8 / 16 frames that are fed into the models sampled from the videos? Does the evaluation use multiple temporal crops?  \n* Sec. 4.1: How are image models applied to video?  \n* I do not understand conclusion (5) in Sec. 4.2. “The effectiveness of pre-training paradigms in scaling model size might not be adequately validated on popular action recognition benchmarks.” What benchmarks is scaling not validated on and what are the results? What does “adaptation performance” mean? Could the authors clarify the meaning of this statement and explain how the following discussion supports it?  \n* Some numbers in Sec. 4.2 do not match the table, e.g. “+3.6” in l 470, 8.0 in l. 471 and 37.7, 44.4 in l. 476\\.  \n* Sec. 4.3: How are videos ranked when using image models?  \n* l. 521: “This is also consistent with the performance differences observed between DINO and CLIP-style pre-training methods.” Which results does this statement refer to?  \n* Sec. 4/3, conclusion 4: “Labels bring new semantic information or disrupt existing finer-grained semantic information.” Could authors provide more context on the pre-training data used for the models discussed in this section? The conclusions are hard to understand without this context.  \n* l. 471: “previous conclusions” should be cited.  \n* l. 484: “Effective adaptation method for FMs is crucial.” This conclusion is not clear to me. Was the attentive probe not used for image-to-video methods? If not, which results are they being compared to?\n\nMinor points:\n\n* It’s often difficult to map author names of cited approaches to the actual names of the approaches. For citations that propose an approach with a well-known name, I’d suggest providing the names along with the citation, e.g. “V-JEPA (Bardes et al., 2023)”. That would also make it easier to map the methods mentioned in the related work section to Tab. 5\\.  \n* Citations should have parentheses around them, otherwise it looks like they are part of the sentence. Please use “\\\\citep” to achieve this.  \n* The name VideoEval is very generic and would apply to any video benchmark. I’d suggest choosing a more specific name, e.g. “VideoFMEval”  \n* Minor grammatical error in the title: Models should be plural, so the title should be “\\[...\\] evaluation of video foundation model**s**”  \n* The font size of Tab. 2 is a little small and the table is hard to read in print.  \n* The radar chart in Fig. 1 (top right) is very interesting but hard to read at this size. I’d suggest enlarging it or even moving it into its own figure.  \n* Please right-justify numeric columns in tables.  \n* l. 413: The table reference should be changed to Tab. 5\\.  \n* Could authors better explain the terms DSVs, CSVs and ISVs in Sec. 3.2?"},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":3},"strengths":{"value":"Originality:\n\n* The paper performs a large-scale evaluation of video foundation models on a new benchmark consisting of a diverse set of tasks. While this is not the only such survey, I am not aware of any studies that evaluate such a large amount of VFMs on such a large diversity of tasks, so I see the main novelty in the sheer breadth of the evaluation and the conclusions drawn from the state of the art in video foundation models.\n\nQuality:\n\n* The benchmark was carefully curated to contain high quality, challenging video datasets with good diversity.  \n* The paper evaluates 20 different models, including image-only baselines, image models adopted to video and video foundation models.  \n* The benchmark suite is comprehensive and diverse, consisting of 6 different domains and 12 different tasks and thus offers a good overview of VFM applicability to different tasks  \n* Design / scope choices made are clearly motivated and backed up with experiments, e.g. focus on few-shot setting, using attentive probe for model adaptation.\n\nClarity:\n\n* The paper is well-organized and well-written.\n\nSignificance:\n\n* Both adaptation and retrieval with video embeddings are evaluated  \n* The paper draws several interesting conclusions from their findings that point out weaknesses and directions for further research."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"It seems to me that there is a disconnect between the discussion of the results and the actual results. Conclusions drawn are not supported by the data while some clear trends in the data are not discussed. While I would not recommend this paper for acceptance in its current state, I would be in favor of accepting it if the weaknesses above can be addressed sufficiently.\n\n* The main weakness of the paper is that many of the conclusions it draws are not fully supported by the experimental results.  \n  * On VidTAB:  \n    * l. 421: “current vision FMs struggle to adapt to unseen video tasks”. What is the evidence for this statement? A comparison of VFM performance to SotA on these datasets would help support this.  \n    * l. 422: “VFMs outperform IFMs”: The best IFM beats 9 out of the 12 evaluated VFMs, so this statement doesn’t seem to be generally true, and also not in action and behavior tasks where some IFMs show stronger performance than most VFMs. So this statement should be weakened and results discussed in more nuance. This is actually an interesting negative result that warrants further discussion.  \n    * l. 431: “Similar findings are observed with ViCLIP-L, where post-pretraining on a large-scale video dataset improves Action-related tasks but diminishes performance in other domains (Science, Safety, Quality, Emotion)”: This statement is not supported by results in Tab. 5: VICLIP-L performance actually increases on science tasks, only slightly drops on one safety task while improving on the other, and increases on emotion.  \n  * On VidEB:  \n    * Conclusion (1): “The contrastive learning (CL) based approach consistently excels in embedding evaluation.”: This does not seem to be the case as two MVM methods perform better than the best CL method.  \n    * Conclusion (2): “The effectiveness of masked video modeling is closely tied to the targets it reconstructs or aligns with.”: This statement is hard to understand without context about the training data. What are the targets? How is “higher semantic richness” introduced?  \n* Results also suggest conclusions that are not discussed in the paper. These should be addressed and discussed in more detail:  \n  * VidTAB: A missing insight is that VFMs show no gain over IFMs on spatial tasks (Safety and Quality).  \n  * Image-to-video methods do better than VFMs on most tasks.  \n* From the evaluation results on both benchmarks it is not clear what the gap between VFMs and state of the art on each dataset is since results on specialist models on each task are not given.  \n* There are significant differences in results between the VidTAB and VidEB benchmarks that are not explained.  \n  * Why does stage 2 training improve UMT-L and InternVideo2 performance on VidTAB but degrades performance on VidEB?  \n  * Similarly, why does fine-tuning on K710 increase performance on VidTAB but hurts performance on VidEB?  \n* The related work section does not discuss image foundation models and image-to-video adapter approaches. It would be good to at least discuss those approaches that are evaluated. This would allow readers who are less familiar with the literature in this area to better interpret the results.  \n* Evaluation of task adaptation is limited to one method (attentive probe).  \n* Future research directions are not clearly pointed out, though some ideas are given in the results discussion. A separate section on this could help direct future research."}},"nonreaders":[],"tmdate":1731427792110,"tcdate":1729721213352,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission7392/Reviewer_MzWr"],"signatures":["ICLR.cc/2025/Conference/Submission7392/Reviewer_MzWr"],"forum":"wMRFTQwp1d","number":2,"license":"CC BY 4.0","cdate":1729721213352,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission7392/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427792110,"domain":"ICLR.cc/2025/Conference","replyto":"wMRFTQwp1d","id":"qygsyLNTss","forumContent":{"TLDR":{"value":"A vision-centric evaluation method for video foundation models that is comprehensive, challenging, indicative, and low-cost."},"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video Understanding","Video Foundation Model","Benchmark"]},"primary_area":{"value":"datasets and benchmarks"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"With the accumulation of high-quality data and advancements in visual pretraining paradigms, recent Video Foundation Models (VFMs) have made significant progress, demonstrating remarkable performance on popular video understanding benchmarks. However, conventional benchmarks (e.g. Kinetics) and evaluation protocols are limited by their relatively poor diversity, high evaluation costs, and saturated performance metrics. In this work, we introduce a comprehensive benchmark suite to address these issues, namely **VideoEval**. We establish the **Vid**eo **T**ask **A**daption **B**enchmark (VidTAB) and the **Vid**eo **E**mbedding **B**enchmark (VidEB) from two perspectives: evaluating the task adaptability of VFMs under few-shot conditions and assessing their feature embedding's direct applicability to downstream tasks. With VideoEval, we conduct a large-scale study of 20 popular open-source vision foundation models. Our study reveals some insightful findings, 1) overall, current VFMs exhibit weak generalization across diverse tasks, 2) increasing video data, whether labeled or in video-text pairs, does not necessarily improve task performance, 3) the effectiveness of some pre-training paradigms may not be fully validated in previous benchmarks, and 4) combining different pre-training paradigms can help develop models with better generalization capabilities. We believe this study serves as a important complement to the current evaluation methods for VFMs and offers valuable insights for future research directions."},"_bibtex":{"value":"@misc{\nli2024videoeval,\ntitle={VideoEval: Comprehensive Benchmark Suite for Low-Cost Evaluation of Video Foundation Model},\nauthor={Xinhao Li and Zhenpeng Huang and Jing Wang and Kunchang Li and Limin Wang},\nyear={2024},\nurl={https://openreview.net/forum?id=wMRFTQwp1d}\n}"},"title":{"value":"VideoEval: Comprehensive Benchmark Suite for Low-Cost Evaluation of Video Foundation Model"},"pdf":{"value":"/pdf/652a4eb9e5126e3b38cd7820b52f1dacc0558c7a.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"li|videoeval_comprehensive_benchmark_suite_for_lowcost_evaluation_of_video_foundation_model"},"authorids":{"value":["~Xinhao_Li1","~Zhenpeng_Huang1","~Jing_Wang48","~Kunchang_Li1","~Limin_Wang1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Xinhao Li","Zhenpeng Huang","Jing Wang","Kunchang Li","Limin Wang"]}},"version":2},{"content":{"summary":{"value":"The paper introduces GSCV, a novel method for compressing 3D Gaussian Splatting sequences using standard video codecs. Unlike previous approaches that rely on optimization-based compression (A-3DGS), GSCV focuses on compressing the trained GS data (I-3DGS).The core innovation is Inter-PLAS, an improved version of the Parallel Linear Assignment Sorting (PLAS) method and anchor-based refinement, which ensures that consecutive GS frames (I- and P-frames) produce similar images, enhancing inter-frame prediction in video codecs like HEVC and VVC. Experiments on MPEG datasets show that GSCV outperforms existing video- and point-cloud-based methods in both tracked and semi-tracked GS sequences at some bitrate region."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1. Generally, methods leveraging existing codecs, even if not achieving optimal overall performance, still significantly outperform training-based approaches in certain aspects, such as encoding time. This allows for some compromise in rate-distortion (RD) performance. However, PLAS-based methods still require thousands of iteration steps—does this undermine the practical value of such approaches? Furthermore, could some encoding time results be provided in the experiments?\n\n2. In the low-bitrate region of the rate-distortion curve (bartender, cinema), the performance of the GSCV method falls below that of the GPCC all-intra mode. Does this indicate that the proposed inter-frame model still has certain limitations?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. Effective Inter-Frame Compression.  Spatial-Stable Initialization (SSI) and anchor-based refinement reduce randomness in PLAS. These modules  improve frame similarity, leading to better inter-frame-prediction \n\n2. Compatibility with Standard Codecs. Compression techniques for pre-trained Gaussian Splatting models are practical in certain scenarios, while leveraging existing video encoders also offers better adaptability, making them applicable in environments lacking specialized acceleration hardware."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Limited Compression Efficiency on PLAS Images. PLAS-generated images are structurally different from natural images, limiting the full potential of standard video codecs.\n\n2. The experimental comparisons are insufficient. Although the technical paradigms are not entirely identical, there are already several studies on Gaussian Splatting sequences or 4D Gaussians. Conducting more comparisons on more dataset would more comprehensively reflect the performance of the proposed method in this paper. \n\nsome previous work:\n\n[1] 4D Gaussian Splatting for Real-Time Dynamic Scene Rendering\n\n[2] 4DGC: Rate-Aware 4D Gaussian Compression for Efficient Streamable Free-Viewpoint Video"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762915540237,"tcdate":1761721215281,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission529/Reviewer_qDJS"],"signatures":["ICLR.cc/2026/Conference/Submission529/Reviewer_qDJS"],"forum":"LafrTql67m","number":2,"license":"CC BY 4.0","cdate":1761721215281,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission529/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762915540237,"domain":"ICLR.cc/2026/Conference","replyto":"LafrTql67m","id":"Tx51tDMvLP","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Gaussian Splatting","Compression"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"This paper presents a novel effective Gaussian Splatting (GS) sequence Compression method that utilizes the Video codec (GSCV). Existing video-based GS sequence compression relies on the Parallel Linear Assignment Sorting (PLAS) to convert GS into smooth 2D maps. Using the vanilla PLAS, however, can generate images exhibiting weak inter-frame correlation, due to its stochastic nature. GSCV incorporates a simple yet efficient Inter-PLAS method to produce close images between the I- and P-frames of GS, enhancing the inter-frame performance of video codec greatly. GSCV also realizes a new pipeline based on the state-of-the-art video codecs with high bit-depth GS images, achieving higher compressibility while simultaneously raising the upper limit of the quality. Experiment results show that the proposed GSCV exhibits evidently improved performance over MPEG video and point cloud-based anchors in GS sequence compression."},"_bibtex":{"value":"@misc{\nyang2026gscv,\ntitle={{GSCV}: Compressing Gaussian Splatting Sequence with Video Codec},\nauthor={Qi Yang and Shuting Xia and Le Yang and Geert Van der Auwera and Zhu Li},\nyear={2026},\nurl={https://openreview.net/forum?id=LafrTql67m}\n}"},"title":{"value":"GSCV: Compressing Gaussian Splatting Sequence with Video Codec"},"pdf":{"value":"/pdf/c58f0e248b816a496c15d84dc6aafc53be2e5a9d.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"yang|gscv_compressing_gaussian_splatting_sequence_with_video_codec"},"authorids":{"value":["~Qi_Yang4","~Shuting_Xia2","~Le_Yang6","~Geert_Van_der_Auwera1","~Zhu_Li1"]},"authors":{"value":["Qi Yang","Shuting Xia","Le Yang","Geert Van der Auwera","Zhu Li"]}},"version":2},{"content":{"summary":{"value":"The paper introduces StreamingBench, a benchmark designed to evaluate the streaming video understanding abilities of Multimodal Large Language Models (MLLMs). Traditional MLLMs are effective in offline video comprehension but struggle with real-time, streaming scenarios that require instant processing, synchronizing visual and audio inputs, and understanding context over time. StreamingBench addresses this by presenting 900 videos across diverse real-world scenarios, structured into 18 tasks and 4,300 human-curated question-answer pairs. These tasks test MLLMs on real-time visual, omni-source, and contextual understanding, aiming to bridge the gap between MLLMs and human-level comprehension in streaming contexts. Testing 13 MLLMs, including state-of-the-art proprietary models, revealed significant limitations in current models, especially in omni-source and contextual tasks, suggesting that MLLMs need further development to match human performance in real-time understanding. In general, it is a solid paper and I would recommend acceptance to it."},"soundness":{"value":4},"confidence":{"value":5},"questions":{"value":"Please see the weaknesses."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. It is the first valid benchmark on streaming long videos. The questions are designed properly to reflect the information gained in a streaming long video, and highly resembles what human will ask when continuously watching a video.\n2. The evaluation and discussion are both very solid.\n3. Human performance is another plus."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The real-time understanding part is nice, but seems a little bit trivial. From all kinds of NIAH evaluations, all models can best answer questions near the ending part of the input, and questions related to \"current moments\" (which is actually the ending part of input as implemented) might not be so important. Would love to see the understanding on \"remembering earlier moments\" and the discrepancy from \"current moments\" for LMMs.\n\n2. The omni-source (visual+audio) part is good. However, how are LMMs without audio abilities evaluated? As these `audio'-related questions seem to be mostly about speeches, do authors plan to interleave text ASR into the model for evaluation? At present, sadly we only see a black-box Gemini-1.5-Pro (for which we do not know how they integrate audio and video) being evaluated with audio.\n\n3. A minor suggestion: the omnisource part of the benchmark is related to \"referring reasoning\" part of LongVideoBench, an interleaved benchmark for frames and ASR texts, which also needs to judge between concurrent video and audio information in a video. As some other long video benchmarks discussed in Tab 1, please also try to discuss it in the revised paper."}},"nonreaders":[],"tmdate":1731427999879,"tcdate":1730648327315,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission14232/Reviewer_wiNw"],"signatures":["ICLR.cc/2025/Conference/Submission14232/Reviewer_wiNw"],"forum":"qnAZqlMGTB","number":2,"license":"CC BY 4.0","cdate":1730648327315,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission14232/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427999879,"domain":"ICLR.cc/2025/Conference","replyto":"qnAZqlMGTB","id":"nu0p7u8bvt","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Benchmark","Streaming Video Understanding","Multimodal Large Language Models","Video Benchmark","Evaluation"]},"primary_area":{"value":"datasets and benchmarks"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"The rapid development of Multimodal Large Language Models (MLLMs) has expanded their capabilities from image comprehension to video understanding. However, most of these MLLMs focus primarily on ofﬂine video comprehension, necessitating extensive processing of all video frames before any queries can be made. This presents a signiﬁcant gap compared to the human ability to watch, listen, think, and respond to streaming inputs in real time, highlighting the limitations of current MLLMs. In this paper, we introduce StreamingBench, the ﬁrst comprehensive benchmark designed to evaluate the streaming video understanding capabilities of MLLMs. StreamingBench assesses three core aspects of streaming video understanding: (1) real-time visual understanding, (2) omni-source understanding and (3) contextual understanding. The benchmark consists of 18 tasks, featuring 900 videos and 4,500 human-curated QA pairs. Each video features ﬁve questions presented at different time points to simulate a continuous streaming scenario. We conduct experiments on StreamingBench with 15 open-source and proprietary MLLMs and ﬁnd that even the most advanced proprietary MLLMs like Gemini 1.5 Pro and GPT-4o perform signiﬁcantly below human-level streaming video understanding capabilities. We hope our work can facilitate further advancements for MLLMs, empowering them to approach human-level video comprehension and interaction in more realistic scenarios."},"_bibtex":{"value":"@misc{\nlin2025streamingbench,\ntitle={StreamingBench: Assessing the Gap for {MLLM}s to Achieve Streaming Video Understanding},\nauthor={Junming Lin and Zheng Fang and Zihao Wan and Fuwen Luo and Chi Chen and Peng Li and Yang Liu and Maosong Sun},\nyear={2025},\nurl={https://openreview.net/forum?id=qnAZqlMGTB}\n}"},"title":{"value":"StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding"},"pdf":{"value":"/pdf/510ba84e9bb94cbbd1bcb5060d0bf89b4703ddb7.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"lin|streamingbench_assessing_the_gap_for_mllms_to_achieve_streaming_video_understanding"},"authorids":{"value":["~Junming_Lin1","~Zheng_Fang14","~Zihao_Wan2","~Fuwen_Luo1","~Chi_Chen1","~Peng_Li2","~Yang_Liu19","~Maosong_Sun1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Junming Lin","Zheng Fang","Zihao Wan","Fuwen Luo","Chi Chen","Peng Li","Yang Liu","Maosong Sun"]}},"version":2},{"content":{"venue":{"value":"MICCAI (6) 2024"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-031-72089-5_14.pdf"},"venueid":{"value":"dblp.org/conf/MICCAI/2024"},"paperhash":{"value":"zhang|depthaware_endoscopic_video_inpainting"},"authorids":{"value":["~Francis_Xiatian_Zhang1","https://dblp.org/search/pid/api?q=author:Shuang_Chen_0010:","https://dblp.org/search/pid/api?q=author:Xianghua_Xie:","https://dblp.org/search/pid/api?q=author:Hubert_P._H._Shum:"]},"html":{"value":"https://doi.org/10.1007/978-3-031-72089-5_14"},"_bibtex":{"value":"@inproceedings{DBLP:conf/miccai/ZhangCXS24,\n  author={Francis Xiatian Zhang and Shuang Chen and Xianghua Xie and Hubert P. H. Shum},\n  title={Depth-Aware Endoscopic Video Inpainting},\n  year={2024},\n  cdate={1704067200000},\n  pages={143-153},\n  url={https://doi.org/10.1007/978-3-031-72089-5_14},\n  booktitle={MICCAI (6)},\n  crossref={conf/miccai/2024-6}\n}\n"},"abstract":{"value":"Video inpainting fills in corrupted video content with plausible replacements. While recent advances in endoscopic video inpainting have shown potential for enhancing the quality of endoscopic videos, they mainly repair 2D visual information without effectively preserving crucial 3D spatial details for clinical reference. Depth-aware inpainting methods attempt to preserve these details by incorporating depth information. Still, in endoscopic contexts, they face challenges including reliance on pre-acquired depth maps, less effective fusion designs, and ignorance of the fidelity of 3D spatial details. To address them, we introduce a novel Depth-aware Endoscopic Video Inpainting (DAEVI) framework. It features a Spatial-Temporal Guided Depth Estimation module for direct depth estimation from visual features, a Bi-Modal Paired Channel Fusion module for effective channel-by-channel fusion of visual and depth information, and a Depth Enhanced Discriminator to assess the fidelity of the RGB-D sequence comprised of the inpainted frames and estimated depth images. Experimental evaluations on established benchmarks demonstrate our framework’s superiority, achieving a 2% improvement in PSNR and a 6% reduction in MSE compared to state-of-the-art methods. Qualitative analyses further validate its enhanced ability to inpaint fine details, highlighting the benefits of integrating depth information into endoscopic inpainting."},"title":{"value":"Depth-Aware Endoscopic Video Inpainting"},"authors":{"value":["Francis Xiatian Zhang","Shuang Chen","Xianghua Xie","Hubert P. H. Shum"]}},"tmdate":1748196485236,"pdate":1704067200000,"tcdate":1748196473426,"writers":["~"],"signatures":["~Francis_Xiatian_Zhang1"],"forum":"lkyzy9r0BZ","license":"CC BY-SA 4.0","number":551048,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1748196485236,"domain":"DBLP.org","id":"lkyzy9r0BZ","version":2},{"content":{"summary":{"value":"This paper proposes a physics-inspired video spatiotemporal flow representation method called ReynoldsFlow, which directly derives spatiotemporal features with physical interpretability from video data. The method adopts an unsupervised and training-free design, constructing texture-preserving and dynamic-aware feature representations by integrating flow field magnitude with frame intensity. It performs exceptionally well in tasks such as small object detection and pose estimation."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"Please refer to the weaknesses."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. Lightweight plug-and-play utility: ReynoldsFlow directly computes from raw video frames without training process and large-scale annotated data, with computational cost comparable to traditional optical flow methods. Its three-channel representation (optical flow magnitude, complementary flow magnitude, frame intensity) can be seamlessly integrated into existing detection and pose estimation models, without task-specific preprocessing or model reconstruction, offering strong engineering practicality and compatibility.\n\n2. The paper breaks through the limitations of traditional optical flow that rely on the brightness constancy assumption, introducing fluid dynamics into video spatiotemporal feature modeling, and first proposes ReynoldsFlow which simultaneously includes traditional optical flow (divergence-free component) and complementary flow (curl-free component). This design covers information beyond motion, such as illumination changes and non-rigid deformations, from a principled perspective, solving the robustness issues of traditional optical flow in complex scenarios, and the physical modeling endows features with stronger interpretability."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The generalization capability of the task coverage is insufficient: the experimental validation focuses on two specific tasks, UAV target detection and golf swing pose estimation, and lacks performance evaluation on more general video understanding tasks (such as action recognition, video semantic segmentation). The current results are insufficient to fully demonstrate the generalization ability of ReynoldsFlow, especially its performance in scenarios with non-rigid targets and dense motion trajectories is still unclear. I believe that the dataset's performance on more video tasks will help in verifying the model's capabilities.\n\n2. The data typesetting in Table 1 is chaotic and the data is missing, which affects the readability of the results."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762943239962,"tcdate":1761975348702,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission24909/Reviewer_hLGr"],"signatures":["ICLR.cc/2026/Conference/Submission24909/Reviewer_hLGr"],"forum":"Tf29oMgErW","number":2,"license":"CC BY 4.0","cdate":1761975348702,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission24909/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762943239962,"domain":"ICLR.cc/2026/Conference","replyto":"Tf29oMgErW","id":"yMWk5m33Pu","forumContent":{"TLDR":{"value":"We propose ReynoldsFlow, a physics-inspired spatiotemporal flow representation that is lightweight, interpretable, and robust to photometric and structural variations for efficient video representation learning."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Video Representation Learning","Spatiotemporal Modeling","Physics-Inspired Flow","Helmholtz Decomposition","Reynolds Transport Theorem"]},"primary_area":{"value":"unsupervised, self-supervised, semi-supervised, and supervised representation learning"},"abstract":{"value":"Representation learning for videos has largely relied on spatiotemporal modules embedded in deep architectures, which, while effective, often require heavy computation and heuristic design. Existing approaches, such as 3D convolutional modules or optical flow networks, may also overlook changes in illumination, scale variations, and structural deformations in video sequences. To address these challenges, we propose ReynoldsFlow, a physics-inspired flow representation that leverages the Helmholtz decomposition and the Reynolds transport theorem to derive principled spatiotemporal features directly from video data. Unlike classical optical flow, ReynoldsFlow captures both divergence-free and curl-free components under more general assumptions, enabling robustness to photometric variation while preserving intrinsic structure. Beyond its theoretical grounding, ReynoldsFlow remains lightweight and adaptable, combining frame intensity with flow magnitude to yield texture-preserving and dynamics-aware representations that substantially enhance tiny object detection. Experiments on benchmarks with various target scales demonstrate that ReynoldsFlow is consistently comparable to or outperforms existing flow-based features, while also improving interpretability and efficiency. These results position ReynoldsFlow as a compelling representation for video understanding and a strong foundation for downstream model learning. The code will be made publicly available."},"_bibtex":{"value":"@misc{\nchen2026reynoldsflow,\ntitle={ReynoldsFlow: Spatiotemporal Flow Representations for Video Learning},\nauthor={Yu-Hsi Chen and Chin-Tien Wu},\nyear={2026},\nurl={https://openreview.net/forum?id=Tf29oMgErW}\n}"},"title":{"value":"ReynoldsFlow: Spatiotemporal Flow Representations for Video Learning"},"pdf":{"value":"/pdf/6bb21571b1246cfd028e658bfa7f4293a26c8b43.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"chen|reynoldsflow_spatiotemporal_flow_representations_for_video_learning"},"authorids":{"value":["~Yu-Hsi_Chen2","~Chin-Tien_Wu1"]},"authors":{"value":["Yu-Hsi Chen","Chin-Tien Wu"]}},"version":2},{"content":{"summary":{"value":"The paper introduces VidEgoThink, a benchmark for evaluating egocentric video understanding capabilities of Multi-modal Large Language Models (MLLMs). It has four key tasks: video question-answering, hierarchy planning, visual grounding, and reward modeling. The authors develop an automatic data generation pipeline to generate the data using GPT-4o from Ego4D dataset. Human annotators are employed to filter the generated data. Experiments are conducted with various MLLMs, including API-based, open-source image-based, and video-based models, showing that current models are still lacking in egocentric video understanding."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"It would be helpful if the authors could clarify the human annotator filtering process and the explanation for weakness points 2,3,4."},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The benchmark is comprehensive. Four key tasks of video question-answering, hierarchy planning, visual grounding, and reward modeling are designed, and many models including open and closed source LLMs are evaluated.\n\n2. The experimental results reveal that even the most advanced MLLMs struggle with egocentric video understanding, providing valuable insights into areas for potential research and development."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The data generation pipeline heavily relies on GPT-4o. The authors claim to use manual inspection on the automatically generated data, however, no detail on the human filtering is provided. For example, how the inspection is done, what instructions are given to the annotators, how many percentages of data are filtered, where the annotators are hired, and at what salary.\n\n2. The benchmark is mainly generated using GPT-4o, and the evaluation is in an open-ended manner. The evaluator is also GPT-4o. This raises the concern that the evaluation is not fair to other models. The evaluation may be not entirely based on the correctness, but partially based on the (language style) similarity to GPT-4o.\n\n3. For the mid-to-low planning task, the authors converted the output style into a \"function-like\" format. Since the function here is purely verb and noun pairs, I think this is totally unnecessary. It can be also seen from the example figure of this task, where not all models can follow the instructions to generate the answers of the same format, which potentially influences the evaluation result. From the result, all models except GPT-4o get almost completely wrong outputs. I understand the author's claim that this is closer to the \"actions\" of embodied AI, however in this work it is a simple text re-writing, and I think it is more harmful to the evaluation compared with being \"closer to embodied AI\".\n\n4. For the object grounding task, the prompt given to the LLMs does not seem to contain enough information. The evaluation is on the last frame of the video, but the prompt just says \"Please give out the bounding box coordinates of the object.\" (L1231)\n\n5. The title of this paper is \"for embodied AI\". However, only real-world human data is used. While it is understandable that the design of tasks is inspired by embodied AI, I believe this claim is not appropriate. This benchmark is more on the egocentric video understanding side."}},"nonreaders":[],"tmdate":1731427552215,"tcdate":1729240018924,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission2168/Reviewer_ptWK"],"signatures":["ICLR.cc/2025/Conference/Submission2168/Reviewer_ptWK"],"forum":"Z5nqeTH24j","number":1,"license":"CC BY 4.0","cdate":1729240018924,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission2168/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427552215,"domain":"ICLR.cc/2025/Conference","replyto":"Z5nqeTH24j","id":"kIDT7ASWfx","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Multi-modal Large Language Models","Egocentric Video Understanding","Embodied AI","Benchmark"]},"primary_area":{"value":"datasets and benchmarks"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Recent advancements in Multi-modal Large Language Models (MLLMs) have opened new avenues for applications in Embodied AI.\nBuilding on previous work, EgoThink, we introduce VidEgoThink, a comprehensive benchmark for evaluating egocentric video understanding capabilities. To bridge the gap between MLLMs and low-level control in Embodied AI, we design four key interrelated tasks: video question-answering, hierarchy planning, visual grounding and reward modeling. To minimize manual annotation costs, we develop an automatic data generation pipeline based on the Ego4D dataset, leveraging the prior knowledge and multimodal capabilities of GPT-4o. Three human annotators then filter the generated data to ensure diversity and quality, resulting in the VidEgoThink benchmark. We conduct extensive experiments with three types of models: API-based MLLMs, open-source image-based MLLMs, and open-source video-based MLLMs. Experimental results indicate that all MLLMs, including GPT-4o, perform poorly across all tasks related to egocentric video understanding. These findings suggest that foundation models still require significant advancements to be effectively applied to first-person scenarios in Embodied AI. In conclusion, VidEgoThink reflects a research trend towards employing MLLMs for egocentric vision, akin to human capabilities, enabling active observation and interaction in the complex real-world environments."},"_bibtex":{"value":"@misc{\ncheng2024videgothink,\ntitle={VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied {AI}},\nauthor={Sijie Cheng and Kechen Fang and Yangyang Yu and Sicheng Zhou and Bohao Li and Ye Tian and Tingguang Li and Lei Han and Yang Liu},\nyear={2024},\nurl={https://openreview.net/forum?id=Z5nqeTH24j}\n}"},"title":{"value":"VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI"},"pdf":{"value":"/pdf/04b24f0a982536bd150df66a7ea0233cf41dc649.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"cheng|videgothink_assessing_egocentric_video_understanding_capabilities_for_embodied_ai"},"authorids":{"value":["~Sijie_Cheng1","~Kechen_Fang1","~Yangyang_Yu2","~Sicheng_Zhou2","~Bohao_Li1","~Ye_Tian1","~Tingguang_Li1","~Lei_Han1","~Yang_Liu19"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Sijie Cheng","Kechen Fang","Yangyang Yu","Sicheng Zhou","Bohao Li","Ye Tian","Tingguang Li","Lei Han","Yang Liu"]}},"version":2},{"content":{"venue":{"value":"IEEE Trans. Circuits Syst. Video Technol. 2024"},"pdf":{"value":"https://ieeexplore.ieee.org/iel7/76/10739868/10508586.pdf"},"venueid":{"value":"dblp.org/journals/TCSV/2024"},"paperhash":{"value":"fang|linguistic_hallucination_for_textbased_video_retrieval"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Sheng_Fang:","~Tiantian_Dang1","~Shuhui_Wang1","~Qingming_Huang1"]},"html":{"value":"https://doi.org/10.1109/TCSVT.2024.3393843"},"_bibtex":{"value":"@article{DBLP:journals/tcsv/FangDWH24,\n  author={Sheng Fang and Tiantian Dang and Shuhui Wang and Qingming Huang},\n  title={Linguistic Hallucination for Text-Based Video Retrieval},\n  year={2024},\n  month={October},\n  cdate={1727740800000},\n  journal={IEEE Trans. Circuits Syst. Video Technol.},\n  volume={34},\n  number={10},\n  pages={9692-9705},\n  url={https://doi.org/10.1109/TCSVT.2024.3393843}\n}\n"},"abstract":{"value":"Text-based video retrieval is a crucial technology for video and multimodal applications. Although in traditional Text-Video Retrieval caption-video pairs are supposed to be entirely relevant, there is still information missing in text when compared to the video content. In a specific application scenario of Text-Video Retrieval, where the given caption corresponds to only a segment of the target video, the challenge of aligning two modalities becomes particularly difficult. To address this issue, we introduce context information as an auxiliary to enrich text representation and enhance alignment. In this work, we propose an effective Linguistic Hallucination framework, which incorporates context captions during training and replaces them with hallucinated textual representations predicted from the source sentence at inference. Specific hallucination loss and consistency loss are designed to supervise the learning process. Besides, Curriculum Learning is introduced at both data-level and model-level, which makes the training procedure more stable and improves the retrieval performance simultaneously. Extensive comparison experiments and ablation studies on benchmark datasets demonstrate the effectiveness of our framework. Moreover, we also apply our proposed method to other cross-modal tasks and the promising experimental results prove its generalization ability. Our codes and datasets are available in https://github.com/silenceFS/Linguistic-Hallucination ."},"title":{"value":"Linguistic Hallucination for Text-Based Video Retrieval"},"authors":{"value":["Sheng Fang","Tiantian Dang","Shuhui Wang","Qingming Huang"]}},"tmdate":1767802540722,"pdate":1704067200000,"tcdate":1731503869517,"writers":["~"],"signatures":["~Qingming_Huang2"],"forum":"FdwtErKufN","license":"CC BY-SA 4.0","number":227578,"cdate":1727740800000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1767802540722,"domain":"DBLP.org","id":"FdwtErKufN","version":2},{"content":{"summary":{"value":"The paper introduces L-C4, a novel framework for language-based video colorization that guides the colorization process of monochrome videos using natural language descriptions, addressing the primary challenges faced in video colorization tasks: ambiguous color\nassignment, limited creative imagination, and vulnerable color consistency. The proposed framework builds upon pre-trained cross-modal generative models to utilize their strong language understanding capabilities. To ensure accurate instance-level colorization, the paper proposes the Cross-Modality Pre-Fusion (CMPF) module to enhance text embeddings with instance-aware capability. Deformable Attention (TDA) block and Cross-Clip Fusion (CCF) are designed to ensure temporal color consistency across long sequences. The paper is well-written with clear motivation and rationales. The experiments are thorough, demonstrating their effectiveness."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"- What's the generation process of color-related masks? What if the text prompts are complicated or intricate? One example:\n  \"The sunset painted the sky with hues of burnt umber, casting a surreal glow over the landscape.\" Could the model handle these cases successfully, and to what extent will the quality and type of text prompts influence its performance?\n\n- Temporally Deformable Attention (TDA) generates offset estimations for reference points. Would the performance get worse or better if we replace the offset estimation with an off-the-shelf model, such as an optical flow estimation model?\n\nThank the authors for the detailed explanation. My concerns have been fully addressed. Regarding the failure cases, I think that fine-grained color manipulation may be better controlled by other kinds of user interactions, such as reference images or strokes, given the high-level and semantic nature of text prompts. Overall, I think this is a paper that tackles text-guided video colorization problems with meaningful contributions and good demonstrations. I will raise my score."},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":3},"strengths":{"value":"- Each component is well-motivated, with clear explanations of how they address specific challenges in video colorization. The paper is well-written, with logical organization and flow that makes it easy to understand.\n\n- The proposed Cross-Modality Pre-Fusion (CMPF) module enhances text embeddings with instance-aware capability, ensuring that the objects are assigned with desired colors, enabling creative and precise colorization given user input.\n\n- The Temporally Deformable Attention (TDA) and Cross-Clip Fusion (CCF) effectively maintain color consistency across frames in a video. These components help prevent flickering and color shifts, leading to smoother and more visually appealing results.\n\n- The experiments are thourough and well-designed, demonstrating superiority over state-of-the-art approaches."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Regarding the generation of instance-aware embeddings, the paper does not fully explain or justify the creation of color-related masks. This might limit the model's applicability in complex cases, such as when dealing with intricate or lengthy textual descriptions.\n\n- The paper lacks discussions on limitations. \n\n- Some typos: Line 311, \"cross-flip\" should be \"cross-clip\"."}},"nonreaders":[],"tmdate":1733228766103,"tcdate":1730645210885,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission4813/Reviewer_RHjY"],"signatures":["ICLR.cc/2025/Conference/Submission4813/Reviewer_RHjY"],"forum":"lessla98Wp","number":2,"license":"CC BY 4.0","cdate":1730645210885,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission4813/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1733228766103,"domain":"ICLR.cc/2025/Conference","replyto":"lessla98Wp","id":"zpAY4SwcNC","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video Editing","Diffusion Model","Colorization"]},"supplementary_material":{"value":"/attachment/ea2320a353b8a201e3a7ad1aba8f4d9f19a83817.zip"},"primary_area":{"value":"generative models"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Automatic video colorization is inherently an ill-posed problem because each monochrome frame has multiple optional color candidates.\nPrevious exemplar-based video colorization methods restrict the user's imagination due to the elaborate retrieval process. Alternatively, conditional image colorization methods combined with post-processing algorithms still struggle to maintain temporal consistency. To address these issues, we present Language-based video Colorization for Creative and Consistent Colors (L-C4) to guide the colorization process using user-provided language descriptions. Our model is built upon a pre-trained cross-modality generative model, leveraging its comprehensive language understanding and robust color representation abilities. We introduce the cross-modality pre-fusion module to generate instance-aware text embeddings, enabling the application of creative colors. Additionally, we propose temporally deformable attention to prevent flickering or color shifts, and cross-clip fusion to maintain long-term color consistency.  Extensive experimental results demonstrate that L-C4 outperforms relevant methods, achieving semantically accurate colors, unrestricted creative correspondence, and temporally robust consistency."},"_bibtex":{"value":"@misc{\nchang2025lc,\ntitle={L-C4: Language-Based Video Colorization for Creative and Consistent Colors},\nauthor={Zheng Chang and Shuchen Weng and Huan Ouyang and Yu Li and Si Li and Boxin Shi},\nyear={2025},\nurl={https://openreview.net/forum?id=lessla98Wp}\n}"},"title":{"value":"L-C4: Language-Based Video Colorization for Creative and Consistent Colors"},"pdf":{"value":"/pdf/cf748b9abc4c76995735764dab844c96b52d12d4.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"chang|lc4_languagebased_video_colorization_for_creative_and_consistent_colors"},"authorids":{"value":["~Zheng_Chang2","~Shuchen_Weng1","~Huan_Ouyang2","~Yu_Li4","~Si_Li5","~Boxin_Shi3"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Zheng Chang","Shuchen Weng","Huan Ouyang","Yu Li","Si Li","Boxin Shi"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a novel two-step pre-training process for ego-exo video alignment, aimed at improving multi-view video understanding. The approach starts with traditional ego-exo video pair pre-training, followed by an extension to grouped ego-exo videos alignment, allowing the model to capture denser cross-view relationships. The model employs contrastive losses at each layer as auxiliary supervision to enhance training efficiency. The method is evaluated on the Assembly101 dataset, demonstrating its effectiveness in improving downstream tasks like temporal action segmentation and action anticipation."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"In Figure 1 (a), the top labels all describe \"the ego video in scene\" for four items. Should two of these items be labeled as \"exo\" videos instead of \"ego\"?\n\n\nSection 3.4 discusses the auxiliary loss but does not explain how the values of α and β are selected for each layer during the first and second steps of training. Can you provide more detail on how the weight 0.2, noted in Section 4.2, was chosen?\n\n\nIn formula (1), what is the meaning of the two logarithmic terms? Specifically, what distinguishes the first log from the second in terms of its contribution to the loss?"},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"Originality: The paper presents a novel extension of ego-exo video alignment by moving from individual pair alignments to group-based alignments. This shift is supposed to capture richer inter-view relationships, which is a creative and meaningful advance over existing approaches. \n\n\nQuality: The experimental results are supported by ablation studies, showing improvements on certain evaluation metrics over the chosen base model, but not the state-of-art for the same task, such as ASQuery. \n\n\nClarity: The method is clearly described, with well-organized sections explaining the motivation, approach, and results. Figures, such as the visualizations of video alignment processes, effectively illustrate key concepts.\n\n\nSignificance: The contributions have the potential to improve video understanding, particularly for tasks involving multi-view data."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Data Dependency: The reliance on synchronized ego-exo video pairs imposes a strong constraint on the data, limiting its applicability to datasets where such synchronization is not available.\n\n\nMixed Results in Multi-View: The performance of the two-view setting ((d) Base + Step1 + Step2) does not consistently outperform the single-view setting in key metrics (e.g., F1 and Edit scores in Table 1). This contradicts the expectation of multi-view alignment improving overall video understanding. Additionally, it is reasonable to expect that extending the training for more epochs could improve performance in some cases. To ensure a fair comparison, (b) and (c) should be trained for the same total number of epochs or training time as (d). This would demonstrate whether the improvement observed in (d) is due to the proposed method, rather than simply the result of extended training in earlier steps.\n\n\nBaseline Comparison: The paper compares the proposed method against C2F-TCN, a model not specifically designed for view-invariant features learning. To fully assess the effectiveness of the approach, comparisons with state-of-the-art models in the same domain would provide a stronger baseline, such as the methods listed in section \"1. INTRODUCTION\"  or a method named ASQuery, which reports better F1@10 performance on Assembly101.\n\n\nLimited Dataset Coverage: The evaluation is restricted to the limited validation split of Assembly101 dataset, given the model itself trained with the Assembly101 dataset. Including results on additional public datasets like CMU-MMAC, H2O, or Ego-Exo4D would help generalize the findings and demonstrate the robustness of the method across diverse settings."}},"nonreaders":[],"tmdate":1731428551322,"tcdate":1729639998237,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission5997/Reviewer_f9HT"],"signatures":["ICLR.cc/2025/Conference/Submission5997/Reviewer_f9HT"],"forum":"F0K0zxi62U","number":1,"license":"CC BY 4.0","cdate":1729639998237,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission5997/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428551322,"domain":"ICLR.cc/2025/Conference","replyto":"F0K0zxi62U","id":"o8Rd3jQag2","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["multiview video","view-invariant pretraining","view alignment","ego-exo pair alignment"]},"supplementary_material":{"value":"/attachment/0e02a68395c5d5f7c0c30f848f4b30b2db1e9fcd.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Aligning egocentric and exocentric videos facilitates the learning of view-invariant features, which significantly contributes to video understanding. While previous approaches have primarily focused on aligning individual ego-exo video pairs, our method extends this concept by aligning groups of synchronized egocentric and exocentric videos.  This strategy enables the model to capture more comprehensive cross-view relationships across densely captured viewpoints, enhancing its capacity for robust multi-view understanding.\nTherefore, we develop a pipeline based on contrastive learning for \\textbf{E}gocentric-exocentric \\textbf{V}ideo \\textbf{G}roups \\textbf{A}lignment \\textbf{P}re-training (EVGAP). \nOur method introduces several key innovations: 1) a novel video pre-training paradigm that extends alignment from ego-exo video pairs to ego-exo video group alignments; 2) an innovative two-step training process that leverages the abundant ego-exo video pair data to support the learning of ego-exo video group alignments, transitioning from sparse to dense viewpoints; and 3) the application of auxiliary losses to progressively align videos from different perspectives.\nExtensive ablations illustrate the effectiveness of our approach in single-view and multi-view downstream tasks. We also find that our approach facilitates the tasks inluding novel views. The codes will be available upon acceptance."},"_bibtex":{"value":"@misc{\nwang2024evgap,\ntitle={{EVGAP}: Egocentric-Exocentric Video Groups Alignment Pre-training},\nauthor={Peiyao Wang and Haibin Ling},\nyear={2024},\nurl={https://openreview.net/forum?id=F0K0zxi62U}\n}"},"title":{"value":"EVGAP: Egocentric-Exocentric Video Groups Alignment Pre-training"},"pdf":{"value":"/pdf/9dbc7419b37fbcd3e8d90f9d4c03bc38178865b7.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"wang|evgap_egocentricexocentric_video_groups_alignment_pretraining"},"authorids":{"value":["~Peiyao_Wang2","~Haibin_Ling1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Peiyao Wang","Haibin Ling"]}},"version":2},{"content":{"summary":{"value":"The paper proposes Video-ToC, a video reasoning framework designed to enhance the performance and interpretability of Video Large Language Models (Video LLMs) by introducing a tree-guided visual cue localization mechanism and a dynamic, reasoning-demand-based reward for reinforcement learning. The approach includes a novel rationale annotation pipeline to generate progressive tree-structured reasoning samples (Video-ToC-SFT-1k) and an RL training set with automatically estimated reasoning demand (Video-ToC-RL-2k). The model is evaluated on six video understanding and one hallucination benchmark, showing improvements over previous state-of-the-art methods in both overall performance and hallucination mitigation."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"1. What happens if you scale the training data beyond 3k samples — for example, to 5k, 10k, or 20k examples? Does the improvement persist or saturate? \n2. Could you provide a more detailed qualitative analysis of failure cases?\nSpecifically, which categories of reasoning remain problematic — e.g., temporal reasoning, causal inference, spatial relations, or fine-grained attribute localization? Understanding what fails could clarify where the tree-guided cue localization or reasoning-demand reward are still insufficient.\n3. Could you show how sensitive the model performance is to multiply GRPO advantage  with reasoning-demand coefficient? \n4. Could add clip-selection ablations if possible, such as (i) random clips; (ii) top-K by caption similarity to the question; (iii) oracle (human/GT) when available. If varing tree depth and branching influence the model performance?"},"rating":{"value":6},"details_of_ethics_concerns":{"value":"No ethics review needed."},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"**1. Innovation of Tree of Cues:**\nThe paper introduces a Tree of Cues framework that models reasoning as a progressive traversal from coarse to fine video segments, offering a structured, interpretable process for visual cue localization. This tree-based formulation, clearly depicted in the pipeline diagram (Figure 2) and illustrated through supplementary examples (Figure 5), enhances both interpretability and the alignment between visual perception and reasoning steps.\n\n**2. Comprehensive Experimental Evaluation**: \nThe paper presents a comprehensive experimental evaluation, covering a broad set of benchmarks (Tables 1–6) and showing consistent, non-trivial performance improvements over strong baselines in both video reasoning and hallucination mitigation tasks. Ablation studies (Table 4 for tree vs. single-cue structures; Table 5 for reward design) and detailed visual analyses (Figures 3, 4, 5, 10) further clarify where and how the proposed components contribute to these gains."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"**1. Marginal Reward Novelty:**\nThe proposed Reasoning-Demand Reward adaptively adjusts incentives according to the estimated complexity of each query. However, it functions as a heuristic reweighting of RL rewards derived from an external MLLM, without theoretical grounding or sufficient empirical justification. The design closely resembles prior difficulty-aware or confidence-weighted RL schemes.\n\n**2. Tiny Scale of Dataset:**\nThe “Video-ToC-SFT-1k” and “Video-ToC-RL-2k” are extremely small (3k total training samples), such a limited scale is insufficient to convincingly demonstrate generalizable reasoning improvements in video LLMs. The observed performance boost may stem primarily from strong prompt supervision or curated data quality rather than genuine advancement in reasoning capability. Scaling and data-size control experiments are required to validate whether the framework truly improves reasoning or merely benefits from targeted supervision.\n\n**3. “Tree” Advantage Remains Ambiguous:**\nThe Single-Cue vs. Tree-of-Cue ablation (Table 4) shows gains, but it still entangles captioner quality, LLM summarization, and clip selection. There is no random clips or manually verified clips (if possible) to compare with, nor an ablation of backtracking depth.\n\n**4. Missing Ablation of Advantage Calculation:**\nThe paper introduces a modified advantage computation that multiplies normalized group advantages by the reasoning-demand coefficient (Eq. 5), but there is no ablation or sensitivity analysis isolating this design choice.\n\n**5. Trade-offs between Gains and Compute Costs:**\nRL with reasoning-demand rewards and tree-structured preprocessing seems computationally expensive."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762919599236,"tcdate":1760973427690,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission7483/Reviewer_49M5"],"signatures":["ICLR.cc/2026/Conference/Submission7483/Reviewer_49M5"],"forum":"sHOpjCILyz","number":1,"license":"CC BY 4.0","cdate":1760973427690,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission7483/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762919599236,"domain":"ICLR.cc/2026/Conference","replyto":"sHOpjCILyz","id":"R9T6IaGkCa","forumContent":{"TLDR":{"value":"This paper proposes Video-ToC, a video reasoning framework that employs a tree-of-cue reasoning pattern and a reasoning-demand reward mechanism to enhance Video LLMs, achieving superior performance on video understanding and hallucination benchmarks."},"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["video understanding","multimodal large language model reasoning"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Existing Video Large Language Models (Video LLMs) struggle with complex video understanding, exhibiting limited reasoning capabilities and potential hallucinations. In particular, these methods tend to perform reasoning solely relying on the pretrained inherent reasoning rationales whilst lacking perception-aware adaptation to the input video content. To address this, we propose \\textbf{Video-ToC}, a novel video reasoning framework that enhances video understanding through tree-of-cue reasoning. Specifically, our approach introduces three key innovations: (1) A tree-guided visual cue localization mechanism, which endows the model with enhanced fine-grained perceptual capabilities through structured reasoning patterns; (2) A reasoning-demand reward mechanism, which dynamically adjusts the reward value for reinforcement learning (RL) based on the estimation of reasoning demands, enabling on-demand incentives for more effective reasoning strategies; and (3) An automated annotation pipeline that constructs the Video-ToC-SFT-1k and Video-ToC-RL-2k datasets for supervised fine-tuning (SFT) and RL training, respectively. Extensive evaluations on six video understanding benchmarks and a video hallucination benchmark demonstrate the superiority of Video-ToC over baselines and recent methods. All code, models, and datasets will be released."},"_bibtex":{"value":"@misc{\ntan2025videotoc,\ntitle={Video-ToC: Video Tree-of-Cue Reasoning},\nauthor={Qizhong Tan and Zhuotao Tian and Guangming Lu and Jun Yu and Wenjie Pei},\nyear={2025},\nurl={https://openreview.net/forum?id=sHOpjCILyz}\n}"},"title":{"value":"Video-ToC: Video Tree-of-Cue Reasoning"},"pdf":{"value":"/pdf/e83a4455e536c5f24d49f0a1308feb26b90eccbf.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"tan|videotoc_video_treeofcue_reasoning"},"authorids":{"value":["~Qizhong_Tan1","~Zhuotao_Tian1","~Guangming_Lu2","~Jun_Yu1","~Wenjie_Pei1"]},"authors":{"value":["Qizhong Tan","Zhuotao Tian","Guangming Lu","Jun Yu","Wenjie Pei"]}},"version":2},{"content":{"venue":{"value":"CoRR 2025"},"pdf":{"value":"https://arxiv.org/pdf/2506.02350v1"},"venueid":{"value":"dblp.org/journals/CORR/2025"},"paperhash":{"value":"wan|truth_over_tricks_measuring_and_mitigating_shortcut_learning_in_misinformation_detection"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Herun_Wan:","https://dblp.org/search/pid/api?q=author:Jiaying_Wu:","~Minnan_Luo1","~Zhi_Zeng4","https://dblp.org/search/pid/api?q=author:Zhixiong_Su:"]},"html":{"value":"https://doi.org/10.48550/arXiv.2506.02350"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2506-02350,\n  publtype={informal},\n  author={Herun Wan and Jiaying Wu and Minnan Luo and Zhi Zeng and Zhixiong Su},\n  title={Truth over Tricks: Measuring and Mitigating Shortcut Learning in Misinformation Detection},\n  year={2025},\n  month={June},\n  cdate={1748736000000},\n  journal={CoRR},\n  volume={abs/2506.02350},\n  url={https://doi.org/10.48550/arXiv.2506.02350}\n}\n"},"abstract":{"value":"Misinformation detection models often rely on superficial cues (i.e., \\emph{shortcuts}) that correlate with misinformation in training data but fail to generalize to the diverse and evolving nature of real-world misinformation. This issue is exacerbated by large language models (LLMs), which can easily generate convincing misinformation through simple prompts. We introduce TruthOverTricks, a unified evaluation paradigm for measuring shortcut learning in misinformation detection. TruthOverTricks categorizes shortcut behaviors into intrinsic shortcut induction and extrinsic shortcut injection, and evaluates seven representative detectors across 14 popular benchmarks, along with two new factual misinformation datasets, NQ-Misinfo and Streaming-Misinfo. Empirical results reveal that existing detectors suffer severe performance degradation when exposed to both naturally occurring and adversarially crafted shortcuts. To address this, we propose SMF, an LLM-augmented data augmentation framework that mitigates shortcut reliance through paraphrasing, factual summarization, and sentiment normalization. SMF consistently enhances robustness across 16 benchmarks, encouraging models to rely on deeper semantic understanding rather than shortcut cues. To promote the development of misinformation detectors, we have published the resources publicly at https://github.com/whr000001/TruthOverTricks."},"title":{"value":"Truth over Tricks: Measuring and Mitigating Shortcut Learning in Misinformation Detection"},"authors":{"value":["Herun Wan","Jiaying Wu","Minnan Luo","Zhi Zeng","Zhixiong Su"]}},"tmdate":1769184877363,"pdate":1767139200000,"externalIds":["dblp:journals/corr/abs-2506-02350"],"tcdate":1769164093136,"writers":["~"],"signatures":["~Minnan_Luo1"],"forum":"XLcs7dB99d","license":"CC BY-SA 4.0","number":791509,"cdate":1748736000000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1769184877363,"domain":"DBLP.org","id":"XLcs7dB99d","version":2},{"content":{"venue":{"value":"ICML 2026 regular"},"keywords":{"value":["learn physics from video","ODE","identifiability","representation learning"]},"supplementary_material":{"value":"/attachment/83b33cd5f2700d971184b1e05aaefc366333a7c3.zip"},"_bibtex":{"value":"@inproceedings{\nwang2026physics,\ntitle={Physics from Video: Identifiability of Time-Invariant Second-Order {ODE}s under Minimal Trajectory Conditions},\nauthor={Yuanyuan Wang and Wenjie Wang and Kun Zhang and Mingming Gong},\nbooktitle={Forty-third International Conference on Machine Learning},\nyear={2026},\nurl={https://openreview.net/forum?id=xrXxqLadvS}\n}"},"title":{"value":"Physics from Video: Identifiability of Time-Invariant Second-Order ODEs under Minimal Trajectory Conditions"},"paperhash":{"value":"wang|physics_from_video_identifiability_of_timeinvariant_secondorder_odes_under_minimal_trajectory_conditions"},"originally_submitted_PDF":{"value":"/pdf/9698e05a0199baa9187423112c9a6b4ac4ef30a6.pdf"},"TLDR":{"value":"Learning Physics from Video with Identifiability Guarantees"},"primary_area":{"value":"general_machine_learning->representation_learning"},"abstract":{"value":"Bridging the gap between visual realism and physical understanding is a core challenge for video-based world models. We study the structural identifiability of continuous-time physical laws from raw pixels, focusing on whether an encoder-only pipeline can uniquely recover the parameters of second-order linear ODEs. We prove that a level-set slope-coverage condition ensures the learned latent space is locally affine to the true physical state, enabling exact parameter recovery. Our theory provides the first characterization of minimal data requirements across damping regimes, establishing that underdamped systems are identifiable from a single video clip, whereas other regimes require three diverse trajectories. We further introduce a variance-floor regularizer to stabilize the decoder-free objective and prevent latent collapse. Validated on synthetic and real-world data, our approach demonstrates that interpretable physical constants can be reliably estimated from video without the need for compute-intensive pixel reconstruction, ensuring both physical correctness and transparency. Code is available at https://github.com/wenjiewang3/PhysicsFromVideo."},"link_to_code":{"value":"https://github.com/wenjiewang3/PhysicsFromVideo"},"pdf":{"value":"/pdf/696d2ba363c457efbe0895606b533c29b6671b87.pdf"},"lay_summary":{"value":"Many real-world systems, like swinging pendulums and vibrating beams, follow simple physical laws. Being able to recover these laws automatically from a video would give us a contactless tool for science and engineering, requiring no sensors and no manual labeling. The catch is that cameras record pixels rather than physical states. The same motion can produce very different videos depending on lighting, background, and viewpoint, so it is not obvious whether a model trained on video has captured the underlying physics or just imitated the visual appearance.          \n\nWe answer this question for a basic class of physical systems: those described by second-order linear differential equations, which govern oscillations, damping, and ring-down behavior. We prove when a single video clip is enough to uniquely recover the physical parameters, and when multiple clips are needed. We also propose a simple training adjustment that prevents the model from collapsing to a trivial shortcut. Experiments on simulated systems and on real videos confirm the theory."},"venueid":{"value":"ICML.cc/2026/Conference"},"authorids":{"value":["~Yuanyuan_Wang5","~Wenjie_Wang3","~Kun_Zhang1","~Mingming_Gong1"]},"authors":{"value":["Yuanyuan Wang","Wenjie Wang","Kun Zhang","Mingming Gong"]}},"tmdate":1790067258046,"pdate":1777576140682,"tcdate":1767875356885,"writers":["ICML.cc/2026/Conference","ICML.cc/2026/Conference/Submission1039/Authors"],"signatures":["ICML.cc/2026/Conference/Submission1039/Authors"],"forum":"xrXxqLadvS","license":"CC BY 4.0","number":1039,"cdate":1767875356885,"readers":["everyone"],"invitations":["ICML.cc/2026/Conference/-/Submission","ICML.cc/2026/Conference/-/Post_Submission","ICML.cc/2026/Conference/Submission1039/-/Full_Submission","ICML.cc/2026/Conference/-/Edit","ICML.cc/2026/Conference/Submission1039/-/Camera_Ready_Revision"],"mdate":1790067258046,"odate":1782341894777,"domain":"ICML.cc/2026/Conference","id":"xrXxqLadvS","version":2},{"content":{"summary":{"value":"This paper introduces Property-Generated Solver (PGS) which provides minimal high-level debug feedback for LLM code refinement. The key insights are: 1) high-level property-based feedback is more effective than raw feedback containing the exact buggy outputs; 2) minimal feedback is more effective to reduce the cognitive load for an LLM. PGS follows these two principles to generate key properties a correct program should satisfy then generate test inputs for those properties and select the minimal ones. Across 4 code generation benchmarks with 3 LLMs, PGS show improved pass rates compared to other iterative refinement baselines."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. Similar to the discussion above about \"pass@1\", Figure 2 needs a more nuanced discussion as the two feedback methods include 2 generated programs.\n\n2. Please add error bars to the main results."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. This paper is well-written is easy to follow. The two principles: property-oriented and structurally minimal, are well presented with experiments to back up their advantages.\n\n2. The empirical results span multiple benchmark and model combinations, providing evidence that PGS is a general method for code refinement."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. While the evaluation metric is named \"pass@1\", there are some subtleties worth discussing. Iterative refinement methods generate multiple programs so calling it pass@1 is misleading. When comparing multiple methods, a proper measure to control for the amount of compute is necessary, for example, the number of tokens. For example, PGS generates 5 candidate properties and 64 new test inputs in each iteration step. Properly counting the number of tokens would provide a more complete compute vs performance tradeoffs. \n\n2. The benchmarks are small scale code generation. SWE-bench is used only to demonstrate the advantage of structurally minimal feedback. Including experiment results on SWE-bench for PGS would make the paper stronger. The three models used in the experiments include two that are released last year and I would not call them \"state-of-the-art\" (line 161). Incorporating more recent open-weight models such as Deepseek-v3 and Qwen3 family would strengthen the relevance of the work."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762920502151,"tcdate":1762739056769,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission8694/Reviewer_N9kz"],"signatures":["ICLR.cc/2026/Conference/Submission8694/Reviewer_N9kz"],"forum":"NuzRgYrBXo","number":4,"license":"CC BY 4.0","cdate":1762739056769,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission8694/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762920502151,"domain":"ICLR.cc/2026/Conference","replyto":"NuzRgYrBXo","id":"BnzuunqQx4","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Code Generation","Agent","LLM","Test-Driven Development"]},"supplementary_material":{"value":"/attachment/801de6f4120a05b77fd956bdcb69c53a8fd34d76.zip"},"primary_area":{"value":"foundation or frontier models, including LLMs"},"abstract":{"value":"Large Language Models (LLMs) excel at code generation, yet ensuring the functional correctness of their outputs remains a persistent challenge. \nWhile recent studies have applied Test-Driven Development (TDD) to refine code, these methods are fundamentally undermined by poor feedback quality, stemming from the scarcity of high-quality test cases and noisy signals from auto-generated ones.\nIn this work, we shift the focus from test quantity to feedback quality. \nWe introduce the Property-Generated Solver (PGS), a novel paradigm designed to generate highly effective feedback by adhering to two principles:\nit must be property-oriented, to provide semantic guidance beyond simple I/O mismatches, and structurally minimal, to reduce cognitive load and isolate the error's root cause.\nPGS operates by checking high-level program properties (e.g., a sorting function must produce a non-decreasing sequence) and then providing the simplest failing counterexample to the LLM.\nThis property-driven, minimal feedback steers LLMs toward more correct and generalizable solutions. \nAcross a diverse suite of programming benchmarks, PGS consistently demonstrates a superior corrective power, achieving a bug fix rate 1.4x-1.6x higher than the strongest debugging-based approaches and establishing a new state-of-the-art in automated code refinement.\nThe source code and data are available in the supplementary."},"_bibtex":{"value":"@misc{\nhe2026propertyoriented,\ntitle={Property-Oriented and Structurally Minimal Feedback for Effective {LLM}-based Code Refinement},\nauthor={Lehan He and Zeren Chen and Zhe Zhang and Xiang Gao and Jing Shao and Lu Sheng},\nyear={2026},\nurl={https://openreview.net/forum?id=NuzRgYrBXo}\n}"},"title":{"value":"Property-Oriented and Structurally Minimal Feedback for Effective LLM-based Code Refinement"},"pdf":{"value":"/pdf/118263969314daa579a63a2b484a53df949fb2a6.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"he|propertyoriented_and_structurally_minimal_feedback_for_effective_llmbased_code_refinement"},"authorids":{"value":["~Lehan_He1","~Zeren_Chen1","~Zhe_Zhang31","~Xiang_Gao9","~Jing_Shao3","~Lu_Sheng1"]},"authors":{"value":["Lehan He","Zeren Chen","Zhe Zhang","Xiang Gao","Jing Shao","Lu Sheng"]}},"version":2},{"content":{"venue":{"value":"NeurIPS 2025 poster"},"keywords":{"value":["Video Shadow Processing","Video Generation"]},"supplementary_material":{"value":"/attachment/529fc44bb31fa5b21e33d1c711424c0c49b92383.zip"},"primary_area":{"value":"deep_learning"},"abstract":{"value":"Conventional video outpainting methods primarily focus on maintaining coherent textures and visual consistency across frames.\nHowever, they often fail at handling dynamic scenes due to the complex motion of objects or camera movement, leading to temporal incoherence and visible flickering artifacts across frames. This is primarily because they lack instance-aware modeling to accurately separate and track individual object motions throughout the video. In this paper, we propose a novel video outpainting framework that explicitly takes shadow-object pairs into consideration to enhance the temporal and spatial consistency of instances, even when they are temporarily invisible. Specifically, we first track the shadow-object pairs across frames and predict the instances in the scene to unveil the spatial regions of invisible instances. Then, these prediction results are fed to guide the instance-aware optical flow completion to unveil the temporal motion of invisible instances. Next, these spatiotemporal guidances of instances are used to guide the video outpainting process. Finally, a video-aware discriminator is implemented to enhance alignment among dynamic shadows and the extended semantics in the scene. Comprehensive experiments underscore the superiority of our approach, outperforming existing state-of-the-art methods in widely recognized benchmarks."},"_bibtex":{"value":"@inproceedings{\nli2025dynamic,\ntitle={Dynamic Shadow Unveils Invisible Semantics for Video Outpainting},\nauthor={Ruilin Li and Hang Yu and Jiayan Qiu},\nbooktitle={The Thirty-ninth Annual Conference on Neural Information Processing Systems},\nyear={2025},\nurl={https://openreview.net/forum?id=irni27kAeP}\n}"},"title":{"value":"Dynamic Shadow Unveils Invisible Semantics for Video Outpainting"},"pdf":{"value":"/pdf/23f320173fb0cd8864e1f7ef7f33c8b9c9ddbc54.pdf"},"venueid":{"value":"NeurIPS.cc/2025/Conference"},"paperhash":{"value":"li|dynamic_shadow_unveils_invisible_semantics_for_video_outpainting"},"authorids":{"value":["~Ruilin_Li2","~Hang_Yu9","~Jiayan_Qiu1"]},"authors":{"value":["Ruilin Li","Hang Yu","Jiayan Qiu"]}},"tmdate":1783627326592,"pdate":1758217247205,"tcdate":1746946600242,"writers":["NeurIPS.cc/2025/Conference","NeurIPS.cc/2025/Conference/Submission19821/Authors"],"signatures":["NeurIPS.cc/2025/Conference/Submission19821/Authors"],"forum":"irni27kAeP","license":"CC BY 4.0","number":19821,"cdate":1746946600242,"readers":["everyone"],"invitations":["NeurIPS.cc/2025/Conference/-/Submission","NeurIPS.cc/2025/Conference/-/Post_Submission","NeurIPS.cc/2025/Conference/Submission19821/-/Full_Submission","NeurIPS.cc/2025/Conference/Submission19821/-/Supplementary_Material","NeurIPS.cc/2025/Conference/-/Edit","NeurIPS.cc/2025/Conference/Submission19821/-/Camera_Ready_Revision"],"mdate":1783627326592,"odate":1761704986453,"domain":"NeurIPS.cc/2025/Conference","id":"irni27kAeP","version":2},{"content":{"review":{"value":"The paper provides an insightful diagnostic account of challenge development. The strongest contribution is the demonstration that resampling to 1.5 mm caps even a perfect-label round trip at DSC 0.864 for a structure with median maximum inscribed radius of 1.82 mm. Correcting the spacing improves real hidden-test DSC from 0.651 to 0.789, illustrating how preprocessing can dominate architecture for thin anatomy. The video brightness-shortcut analysis is also useful and leads to targeted color augmentation. The report includes negative findings - misleading resampled-space validation, foreground-only validation bias, ineffective snapshot ensembling, unhelpful higher resolution, and the lack of temporal adjacency - that will be valuable to other participants.\nThe main weaknesses are limited experimental depth and incomplete modality coverage. Task 2 is omitted, and the final video distance metrics remain very large despite improved Dice. The authors attribute this partly to generalization from six labeled videos, but should provide case-level analysis of false positives, empty masks, and shortcut persistence after augmentation. The internal ablation table is sparse, and several changes are grouped together, preventing attribution of the gain to normalization, augmentation, boundary loss, patch size, or training length. The report also relies heavily on hidden-test iterations, which can become a form of adaptive tuning; the number of submissions and selection protocol should be discussed explicitly. Reproducibility would improve with released code, exact seeds, and a description of the robustness fallback behavior.\nDespite these limitations, the paper is clear, honest, and technically useful. I recommend acceptance as a challenge report, particularly because it documents preprocessing ceilings and failed approaches that are rarely"},"confidence":{"value":4},"rating":{"value":7},"title":{"value":"This concise challenge report analyzes mitral-valve segmentation for CT and surgical video. It identifies target spacing as a hard performance ceiling for thin CT structures and a brightness shortcut in the video data. The final hidden-test results are DSC/HD/ASD of 0.811/6.86/0.810 for CT and 0.775/361/233 for video."}},"parentInvitations":"MICCAI.org/2026/Workshop/MWM/-/Official_Review","nonreaders":[],"tmdate":1785779921778,"tcdate":1785770999235,"writers":["MICCAI.org/2026/Workshop/MWM","MICCAI.org/2026/Workshop/MWM/Submission113/Reviewer_1mHA"],"signatures":["MICCAI.org/2026/Workshop/MWM/Submission113/Reviewer_1mHA"],"forum":"SVX3aNHR3R","number":1,"license":"CC BY 4.0","cdate":1785770999235,"readers":["everyone"],"invitations":["MICCAI.org/2026/Workshop/MWM/Submission113/-/Official_Review","MICCAI.org/2026/Workshop/MWM/-/Edit"],"mdate":1785779921778,"domain":"MICCAI.org/2026/Workshop/MWM","replyto":"SVX3aNHR3R","id":"9fz38Cq3nx","forumContent":{"venue":{"value":"MWM2026"},"pdf":{"value":"/pdf/82e9cbe324606672de89c783d0affff62e6b49a3.pdf"},"keywords":{"value":["mitral valve","segmentation","nnU-Net","challenge","semi-supervised learning","cardiac CT","surgical video"]},"venueid":{"value":"MICCAI.org/2026/Workshop/MWM"},"paperhash":{"value":"nahata|mitral_valve_segmentation_in_cardiac_ct_and_surgical_video_mvaa_2026_challenge_report"},"authorids":{"value":["~Valmik_Nahata1"]},"abstract":{"value":"We present our solution to the MVAA 2026 challenge for mitral valve segmentation. Our final submission achieves a Task 1 (cardiac CT) hidden test score of DSC 0.811, HD 6.86 mm, and ASD 0.810 mm. For Task 3 (surgical video), we achieve DSC 0.775, HD 361 mm, and ASD 233 mm. We adopted a diagnostic methodology focused on identifying pipeline bottlenecks. We found that resampling target spacing imposes a hard ceiling on achievable accuracy for thin structures. A perfect prediction resampled to 1.5 mm and back scores only 0.864 DSC, whereas the mitral valve is a thin sheet (median maximum inscribed radius 1.82 mm). Correcting the spacing to the dataset median improved our real hidden-test DSC from 0.651 to 0.789. We also show that our Task 3 model exploited a brightness shortcut, which we mitigated using targeted augmentations."},"_bibtex":{"value":"@inproceedings{\nnahata2026mitral,\ntitle={Mitral Valve Segmentation in Cardiac {CT} and Surgical Video: {MVAA} 2026 Challenge Report},\nauthor={Valmik Nahata},\nbooktitle={The 1st MICCAI Workshop on Medical World Models},\nyear={2026},\nurl={https://openreview.net/forum?id=SVX3aNHR3R}\n}"},"title":{"value":"Mitral Valve Segmentation in Cardiac CT and Surgical Video: MVAA 2026 Challenge Report"},"authors":{"value":["Valmik Nahata"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a zero-shot framework for enhancing the realism of synthetic videos using a structure-aware denoising process built on top of a pre-trained video diffusion model. \nThe paper proposes a method for enhancing the realism of synthetic videos by integrating several existing techniques: Classifier-Free Guidance (CFG), latent inversion generation, ControlNet with structure-aware guidance, and the EDM scheduler. The authors apply these components to the video enhancement domain, aiming to improve visual fidelity and structural consistency in generated video frames."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. Can you provide a discussion on individual impact of CFG, latent inversion, ControlNet, and EDM on the final video quality?\n2. Have you considered evaluating different types of structure-aware guidance (e.g., pose, depth, edge maps)? A comparative analysis could help clarify the strengths and limitations of your approach.\n3. Is there any novel insight or adaptation in how these components are combined for video enhancement, beyond straightforward integration?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. Practical Engineering: The paper presents a well-integrated pipeline that combines several state-of-the-art techniques, demonstrating solid engineering and implementation.\n2. Clarity: The methodology is clearly described, and the paper is easy to follow for readers familiar with diffusion models and video synthesis.\n3. Significance: The task of synthetic video realism enhancement is important for applications in virtual production, simulation, and content creation."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Lack of Novelty: The core components, i.e. CFG, latent inversion, ControlNet, and EDM, are not new, and the paper does not offer significant innovation in how they are applied. The techniques used are well-established and widely applied in similar contexts, and the paper does not introduce novel algorithms or insights beyond their combination.\n2. Missing Ablation Study: There is no ablation analysis to isolate the contribution of each component, which makes it difficult to assess the effectiveness of the proposed pipeline.\n3. No Comparative Evaluation of Structure Guidance: The paper does not explore or compare different types of structure-aware guidance, which could have strengthened the evaluation and provided deeper insights."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762919760624,"tcdate":1761905941118,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission7702/Reviewer_JCw1"],"signatures":["ICLR.cc/2026/Conference/Submission7702/Reviewer_JCw1"],"forum":"4VzVWXUkhf","number":2,"license":"CC BY 4.0","cdate":1761905941118,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission7702/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762919760624,"domain":"ICLR.cc/2026/Conference","replyto":"4VzVWXUkhf","id":"3RVeVaxJwR","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Video Enhancement","Controllable Video Synthesis"]},"supplementary_material":{"value":"/attachment/28a729d15ec51947df7b96d15613c43b300b00cd.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"We propose an approach to enhancing synthetic video realism, which can re-render synthetic videos from a simulator in photorealistic fashion. \nOur realism enhancement approach is a zero-shot framework that focuses on preserving the multi-level structures from synthetic videos into the enhanced one in both spatial and temporal domains, built upon a diffusion video foundational model without further fine-tuning. Specifically, we incorporate an effective modification to have the generation/denoising process conditioned on estimated structure-aware information from the synthetic video, such as depth maps, semantic maps, and edge maps, by an auxiliary model, rather than extracting the information from a simulator. This guidance ensures that the enhanced videos are consistent with the original synthetic video at both the structural and semantic levels. Our approach is a simple yet general and powerful approach to enhancing synthetic video realism: we show that our approach outperforms existing  baselines in structural consistency with the original video while maintaining state-of-the-art photorealism quality in our experiments."},"_bibtex":{"value":"@misc{\nwang2026zeroshot,\ntitle={Zero-shot Synthetic Video Realism Enhancement via Structure-aware Denoising},\nauthor={Yifan Wang and Liya Ji and Zhanghan Ke and Harry Yang and Ser-Nam Lim and Qifeng Chen},\nyear={2026},\nurl={https://openreview.net/forum?id=4VzVWXUkhf}\n}"},"title":{"value":"Zero-shot Synthetic Video Realism Enhancement via Structure-aware Denoising"},"pdf":{"value":"/pdf/cbe9865a7ae1742bbe182f08d143fe138669bf35.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"wang|zeroshot_synthetic_video_realism_enhancement_via_structureaware_denoising"},"authorids":{"value":["~Yifan_Wang47","~Liya_Ji1","~Zhanghan_Ke1","~Harry_Yang2","~Ser-Nam_Lim3","~Qifeng_Chen1"]},"authors":{"value":["Yifan Wang","Liya Ji","Zhanghan Ke","Harry Yang","Ser-Nam Lim","Qifeng Chen"]}},"version":2},{"content":{"venue":{"value":"ICCV 2025"},"pdf":{"value":"https://ieeexplore.ieee.org/iel8/11443115/11443287/11445971.pdf"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"lee|cavis_contextaware_video_instance_segmentation"},"html":{"value":"https://doi.org/10.1109/ICCV51701.2025.00429"},"_bibtex":{"value":"@inproceedings{DBLP:conf/iccv/LeeSHCI25,\n  author={Seunghun Lee and Jiwan Seo and Kiljoon Han and Minwoo Choi and Sunghoon Im},\n  title={CAVIS: Context-Aware Video Instance Segmentation},\n  year={2025},\n  cdate={1735689600000},\n  pages={4507-4517},\n  url={https://doi.org/10.1109/ICCV51701.2025.00429},\n  booktitle={ICCV},\n  crossref={conf/iccv/2025}\n}\n"},"abstract":{"value":"In this paper, we introduce the Context-Aware Video Instance Segmentation (CAVIS), a novel framework designed to enhance instance association by integrating contextual information adjacent to each object. To efficiently extract and leverage this information, we propose the Context-Aware Instance Tracker (CAIT), which merges contextual data surrounding the instances with the core instance features to improve tracking accuracy. Additionally, we design the Prototypical Cross-frame Contrastive (PCC) loss, which ensures consistency in object-level features across frames, thereby significantly enhancing matching accuracy. CAVIS demonstrates superior performance over state-of-the-art methods on all benchmark datasets in video instance segmentation (VIS) and video panoptic segmentation (VPS). Notably, our method excels on the OVIS dataset, known for its particularly challenging videos. Project page: this https URL"},"title":{"value":"CAVIS: Context-Aware Video Instance Segmentation"},"authors":{"value":[{"fullname":"Seunghun Lee","username":""},{"fullname":"Jiwan Seo","username":""},{"fullname":"Kiljoon Han","username":""},{"fullname":"Minwoo Choi","username":"~Minwoo_Choi1"},{"fullname":"Sunghoon Im","username":""}]}},"tmdate":1780205402694,"pdate":1767139200000,"externalIds":["dblp:conf/iccv/LeeSHCI25"],"tcdate":1780205400779,"writers":["~","OpenReview.net/Public_Article/DBLP.org","OpenReview.net/Support"],"signatures":["~Minwoo_Choi1"],"forum":"qKoV15Ej5M","license":"CC BY-SA 4.0","number":41074,"cdate":1735689600000,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/DBLP.org/-/Record","OpenReview.net/Public_Article/-/Edit"],"mdate":1780205402694,"domain":"OpenReview.net/Public_Article","id":"qKoV15Ej5M","version":2},{"content":{"summary":{"value":"This study proposed CatVLM, a temporal boundary-aware Vision Language Model that learns to represent fine grained temporal dynamics in untrimmed cataract surgery videos. \nCatVLM facilitates moment level reasoning of 3 clinical tasks, namely Video Moment Retrieval, Video Captioning, and Counting. \nBased on task-related QA annotations and timestamp-integrated video clips, this study evaluated CatVLM on two publicly available datasets, setting good baselines and demonstrating that explicit modeling of temporal boundaries is crucial to medical video perception."},"justification_of_the_preliminary_rating":{"value":"This paper presents a new method of temporal reasoning in surgical videos that has a sound methodology and assessment. \n\nAlthough the performance, though not always high, is favourable in all tasks, the input is significant and has potential in future studies."},"confidentiality_llm_acknowledgment":{"value":"Yes"},"strengths":{"value":"The article introduces a new boundary sensitive VLM of cataract surgery videos, which fills a significant gap in the temporal reasoning. It is properly organized, well-written, scientifically grounded, and can be useful in the further development of medical video knowledge and research"},"weaknesses":{"value":"This paper indicates that there are limited performance improvements on a few tasks with the obvious overfitting on video captioning and moment retrieval. \n\nIt is also only evaluated using cataract datasets, which limits the ability to apply the results to other areas of surgery."},"confidence":{"value":5},"detailed_comments":{"value":"Nil"},"questions_to_address_in_the_rebuttal":{"value":"Nil"},"preliminary_rating":{"value":5}},"parentInvitations":"MIDL.io/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1771530028936,"tcdate":1768462007375,"writers":["MIDL.io/2026/Conference","MIDL.io/2026/Conference/Submission77/Reviewer_FoYa"],"signatures":["MIDL.io/2026/Conference/Submission77/Reviewer_FoYa"],"forum":"J7fsqSkqRU","number":3,"license":"CC BY 4.0","cdate":1768462007375,"readers":["everyone"],"invitations":["MIDL.io/2026/Conference/Submission77/-/Official_Review","MIDL.io/2026/Conference/-/Edit"],"mdate":1771530028936,"domain":"MIDL.io/2026/Conference","replyto":"J7fsqSkqRU","id":"g6zlHwolaj","forumContent":{"TLDR":{"value":"A Vision Language Model (VLM) for answering temporal queries in cataracts surgery"},"venue":{"value":"MIDL 2026 Poster"},"midl_latex_submission_checklist":{"value":["The paper compiles correctly using the pdflatex compiler.","Created a single midl26_NNN.zip file with midl26_NNN.tex, midl26_NNN.bib, all necessary figures and files.","The LaTeX file includes the correct header commands before \\title.","Hyperref is preloaded; do not modify it.","The times package is not used.","All co-authors are listed with correct firstname, lastname, and any suffix/prefix.","Math in the title and abstract uses valid LaTeX.","References are provided through the .bib file only.","Tables and figures stay within the page margins.","The zip archive contains all necessary figures and no unused files.","Special formatting from rebuttal has been removed.","All special characters use LaTeX commands.","Appendices and supplementary material are included in the same PDF after references.","The main paper does not exceed 12 pages."]},"keywords":{"value":["Cataract","Vision Language Models","Video Understanding"]},"read_cfp_and_author_instructions":{"value":"Yes"},"originality_policy":{"value":"Yes"},"abstract":{"value":"Recent studies have shown the effectiveness of Vision Language Models (VLMs) for understanding and analyzing videos in the medical domain and supporting various Question-Answer (QA) tasks. Yet, current VLMs fall short in addressing queries that require temporal reasoning—a critical capability for surgical video understanding. In this work, we introduce CatVLM, a boundary-aware VLM, designed to capture temporal dynamics in untrimmed cataract surgery videos. CatVLM is capable of performing three clinically relevant tasks that demand moment-level awareness: Video Moment Retrieval (VMR), Video Captioning (VC), and Counting. To facilitate the training of such a model, we generate a bank of QA annotations for each task and propose a method to integrate video clips with the timestamps they occur at. To the best of our knowledge, this work is one of the first approaches to explicitly incorporate temporal boundary awareness into VLMs for cataracts as well as the medical domain. We evaluate CatVLM on two public cataract surgery datasets, establishing new baselines across all three tasks. All the code, model checkpoints and annotations will be released post-review"},"_bibtex":{"value":"@inproceedings{\nparanjape2026catvlm,\ntitle={Cat{VLM}: Enhancing Temporal Understanding in Cataract Surgery Videos with Boundary-Aware {VLM}},\nauthor={Jay Nitin Paranjape and Nisarg A Shah and Nanthini Narayanan and Shameema Sikder and S. Swaroop Vedula and Vishal M. Patel},\nbooktitle={Medical Imaging with Deep Learning},\nyear={2026},\nurl={https://openreview.net/forum?id=J7fsqSkqRU}\n}"},"title":{"value":"CatVLM: Enhancing Temporal Understanding in Cataract Surgery Videos with Boundary-Aware VLM"},"latex_code":{"value":"/attachment/b48a04f4600ecfa5ae91468c96c4b1476fdce8d9.zip"},"secondary_subject_area":{"value":"Foundation Models"},"pdf":{"value":"/pdf/90bd5f872d4ba0fa4b970944aa6ed9349e75a8a0.pdf"},"copyright_form":{"value":"/attachment/b5aeaeebd7fd872a59bd8f66b2575fbfd1c51abe.pdf"},"visa":{"value":"Yes"},"single_blind_notice":{"value":"Yes"},"venueid":{"value":"MIDL.io/2026/Conference"},"paperhash":{"value":"paranjape|catvlm_enhancing_temporal_understanding_in_cataract_surgery_videos_with_boundaryaware_vlm"},"primary_subject_area":{"value":"Application: Ophthalmology"},"authorids":{"value":["~Jay_Nitin_Paranjape1","~Nisarg_A_Shah1","nnaraya4@alumni.jh.edu","~Shameema_Sikder1","~S._Swaroop_Vedula2","~Vishal_M._Patel1"]},"registration":{"value":"Yes"},"authors":{"value":["Jay Nitin Paranjape","Nisarg A Shah","Nanthini Narayanan","Shameema Sikder","S. Swaroop Vedula","Vishal M. Patel"]},"llm_policy_acknowledgment":{"value":"Yes"}},"version":2},{"content":{"summary":{"value":"To solve the issue of any-reference video generation task, this paper proposes a unified framework, MAGREF. The framework contains region-aware masking mechanism and subject disentanglement mechanism. Experiments show the scalable, controllable, and high-fidelity any-reference video synthesis results."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"See above weaknesses"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"-  The writing is easy to understand, and the painting is well-drawn.\n-  The proposed region-aware masking method preserves subject identity without backbone changes.\n-  Experimentally, the paper achieves best single-ID and multi-subject score."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- 1.\tIn this paper, multi-subjects are introduced into the videos through a blank canvas. Several subjects are directly added to the canvas with their pixels values. Although this function is useful, my major concerns are listed below:\n -      a)\tThe canvas size is limited. How many subjects can be placed on the canvas without harming the model’s generation ability?\n -      b)\tThe positions of these subjects are randomly shuffled during training, in my opinion, the locations of different subjects may contain implicit relationship, such as decide the distance of two subjects in the generated videos. However, this is not discussed in the paper and related ablations are not considered.\n -      c)\tSimilar to the absolute locations, the scale of each subject in the canvas is still missing. Because such model relies on the explicit vision cues to catch the details of the reference subjects, the scale is an important factor that needs to be considered.\n- 2. In subsection “Pixel-wise channel concatenation”, the composited image I_{comp} is encoded by VAE encoder and then concatenated with noised video latents along the channel dimension. I cannot find technically something new that differs from existing methods. Existing methods also apply the pixel-wise image/video with VAE encoder and concatenate the latent with noised video.\n- 3. The paper claims that the methods support arbitrary subject categories in Lines 098-103, but in their methods, it is unclear how the model support such ability.\n- 4. The qualitative comparison in Fig. 5 seems cannot demonstrate the superiority of the proposed methods, such as compared with close-source method Kling1.6 or the open-source method VACE. Besides, why does the multi-subject result of Skyreels method shows show a poor first frame?\n- 5. The experimental results lack more deep analysis to explain why the proposed methods outperform the previous methods in qualitative and quantitative results.\n-6. The failure cases and more analyses should be discussed."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762915743817,"tcdate":1761908264898,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission1347/Reviewer_fwvD"],"signatures":["ICLR.cc/2026/Conference/Submission1347/Reviewer_fwvD"],"forum":"Nbl43eAVaE","number":3,"license":"CC BY 4.0","cdate":1761908264898,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission1347/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762915743817,"domain":"ICLR.cc/2026/Conference","replyto":"Nbl43eAVaE","id":"VkFUwNvo6o","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Video generation; Diffusion models"]},"supplementary_material":{"value":"/attachment/535323e9c9270387046e58b902eb18b2d2b6c3fd.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"We tackle the task of any-reference video generation, which aims to synthesize videos conditioned on arbitrary types and combinations of reference subjects, together with textual prompts. This task faces persistent challenges, including identity inconsistency, entanglement among multiple reference subjects, and copy-paste artifacts. To address these issues, we introduce MAGREF, a unified and effective framework for any-reference video generation. Our approach incorporates masked guidance and a subject disentanglement mechanism, enabling flexible synthesis conditioned on diverse reference images and textual prompts. Specifically, masked guidance employs a region-aware masking mechanism combined with pixel-wise channel concatenation to preserve appearance features of multiple subjects along the channel dimension. This design preserves identity consistency and maintains the capabilities of the pre-trained backbone, without requiring any architectural changes. To mitigate subject confusion, we introduce a subject disentanglement mechanism which injects the semantic values of each subject derived from the text condition into its corresponding visual region. Additionally, we establish a four-stage data pipeline to construct diverse training pairs, effectively alleviating copy-paste artifacts. Extensive experiments on a comprehensive benchmark demonstrate that MAGREF consistently outperforms existing state-of-the-art approaches, paving the way for scalable, controllable, and high-fidelity any-reference video synthesis. The code and video demos are available in the supplementary materials."},"_bibtex":{"value":"@inproceedings{\ndeng2026magref,\ntitle={{MAGREF}: Masked Guidance for Any-Reference Video Generation with Subject Disentanglement},\nauthor={Yufan Deng and Yuanyang Yin and Xun Guo and Yizhi Wang and Jacob Zhiyuan Fang and Shenghai Yuan and Yiding Yang and Angtian Wang and Bo Liu and Haibin Huang and Chongyang Ma},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=Nbl43eAVaE}\n}"},"title":{"value":"MAGREF: Masked Guidance for Any-Reference Video Generation with Subject Disentanglement"},"pdf":{"value":"/pdf/935bdb178fe94c2d3605bb2a4651abec5aca1467.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"deng|magref_masked_guidance_for_anyreference_video_generation_with_subject_disentanglement"},"authorids":{"value":["~Yufan_Deng4","~Yuanyang_Yin1","~Xun_Guo2","~Yizhi_Wang2","~Jacob_Zhiyuan_Fang1","~Shenghai_Yuan2","~Yiding_Yang1","~Angtian_Wang2","~Bo_Liu16","~Haibin_Huang1","~Chongyang_Ma1"]},"authors":{"value":["Yufan Deng","Yuanyang Yin","Xun Guo","Yizhi Wang","Jacob Zhiyuan Fang","Shenghai Yuan","Yiding Yang","Angtian Wang","Bo Liu","Haibin Huang","Chongyang Ma"]}},"version":2},{"content":{"venue":{"value":"EMNLP 2024"},"pdf":{"value":"https://aclanthology.org/2024.emnlp-main.855.pdf"},"venueid":{"value":"dblp.org/conf/EMNLP/2024"},"paperhash":{"value":"tasawong|efficient_overshadowed_entity_disambiguation_by_mitigating_shortcut_learning"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Panuthep_Tasawong:","~Peerat_Limkonchotiwat1","~Potsawee_Manakul1","~Can_Udomcharoenchaikit1","~Ekapol_Chuangsuwanich1","https://dblp.org/search/pid/api?q=author:Sarana_Nutanong:"]},"html":{"value":"https://aclanthology.org/2024.emnlp-main.855"},"_bibtex":{"value":"@inproceedings{DBLP:conf/emnlp/TasawongLMUCN24,\n  author={Panuthep Tasawong and Peerat Limkonchotiwat and Potsawee Manakul and Can Udomcharoenchaikit and Ekapol Chuangsuwanich and Sarana Nutanong},\n  title={Efficient Overshadowed Entity Disambiguation by Mitigating Shortcut Learning},\n  year={2024},\n  cdate={1704067200000},\n  pages={15313-15321},\n  url={https://aclanthology.org/2024.emnlp-main.855},\n  booktitle={EMNLP},\n  crossref={conf/emnlp/2024}\n}\n"},"abstract":{"value":"Entity disambiguation (ED) is crucial in natural language processing (NLP) for tasks such as question-answering and information extraction. A major challenge in ED is handling overshadowed entities—uncommon entities sharing mention surfaces with common entities. The current approach to enhance performance on these entities involves reasoning over facts in a knowledge base (KB), increasing computational overhead during inference. We argue that the ED performance on overshadowed entities can be enhanced during training by addressing shortcut learning, which does not add computational overhead at inference. We propose a simple yet effective debiasing technique to prevent models from shortcut learning during training. Experiments on a range of ED datasets show that our method achieves state-of-the-art performance without compromising inference speed. Our findings suggest a new research direction for improving entity disambiguation via shortcut learning mitigation."},"title":{"value":"Efficient Overshadowed Entity Disambiguation by Mitigating Shortcut Learning"},"authors":{"value":["Panuthep Tasawong","Peerat Limkonchotiwat","Potsawee Manakul","Can Udomcharoenchaikit","Ekapol Chuangsuwanich","Sarana Nutanong"]}},"tmdate":1751467814849,"pdate":1704067200000,"tcdate":1738941449702,"writers":["~"],"signatures":["~Can_Udomcharoenchaikit1"],"forum":"w9GSsC9aiO","license":"CC BY-SA 4.0","number":311922,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1751467814849,"domain":"DBLP.org","id":"w9GSsC9aiO","version":2},{"content":{"venue":{"value":"CoRR 2025"},"pdf":{"value":"https://arxiv.org/pdf/2509.24997v1"},"venueid":{"value":"dblp.org/journals/CORR/2025"},"paperhash":{"value":"yin|panoworldx_generating_explorable_panoramic_worlds_via_sphereaware_video_diffusion"},"authorids":{"value":["","~Hao-Xiang_Guo1","","","","","","","",""]},"html":{"value":"https://doi.org/10.48550/arXiv.2509.24997"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2509-24997,\n  publtype={informal},\n  author={Yuyang Yin and Hao-Xiang Guo and Fangfu Liu and Mengyu Wang and Hanwen Liang and Eric Li and Yikai Wang and Xiaojie Jin and Yao Zhao and Yunchao Wei},\n  title={PanoWorld-X: Generating Explorable Panoramic Worlds via Sphere-Aware Video Diffusion},\n  year={2025},\n  month={September},\n  cdate={1756684800000},\n  journal={CoRR},\n  volume={abs/2509.24997},\n  url={https://doi.org/10.48550/arXiv.2509.24997}\n}\n"},"abstract":{"value":"Generating a complete and explorable 360-degree visual world enables a wide range of downstream applications. While prior works have advanced the field, they remain constrained by either narrow field-of-view limitations, which hinder the synthesis of continuous and holistic scenes, or insufficient camera controllability that restricts free exploration by users or autonomous agents. To address this, we propose PanoWorld-X, a novel framework for high-fidelity and controllable panoramic video generation with diverse camera trajectories. Specifically, we first construct a large-scale dataset of panoramic video-exploration route pairs by simulating camera trajectories in virtual 3D environments via Unreal Engine. As the spherical geometry of panoramic data misaligns with the inductive priors from conventional video diffusion, we then introduce a Sphere-Aware Diffusion Transformer architecture that reprojects equirectangular features onto the spherical surface to model geometric adjacency in latent space, significantly enhancing visual fidelity and spatiotemporal continuity. Extensive experiments demonstrate that our PanoWorld-X achieves superior performance in various aspects, including motion range, control precision, and visual quality, underscoring its potential for real-world applications."},"title":{"value":"PanoWorld-X: Generating Explorable Panoramic Worlds via Sphere-Aware Video Diffusion"},"authors":{"value":["Yuyang Yin","Hao-Xiang Guo","Fangfu Liu","Mengyu Wang","Hanwen Liang","Eric Li","Yikai Wang","Xiaojie Jin","Yao Zhao","Yunchao Wei"]}},"tmdate":1775011041476,"pdate":1767139200000,"externalIds":["dblp:journals/corr/abs-2509-24997"],"tcdate":1775011038437,"writers":["~"],"signatures":["~Hao-Xiang_Guo1"],"forum":"MAvrqIoa81","license":"CC BY-SA 4.0","number":861783,"cdate":1756684800000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1775011041476,"domain":"DBLP.org","id":"MAvrqIoa81","version":2},{"content":{"summary":{"value":"This paper proposes Video Policy, a framework that treats video generation as a proxy for robot policy learning. The central insight is that a video generative model, implicitly encodes policy information. A lightweight action decoder can then translate the video model’s latent dynamics into executable robot actions. The architecture jointly trains a video U-Net (adapted from Stable Video Diffusion, SVD) and an action U-Net that predicts robot end-effector actions conditioned on intermediate video features. The paper further highlights that action-free video data can enhance generalization to unseen tasks, suggesting that generative video modeling itself serves as a policy."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"See the weakness."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":2},"strengths":{"value":"- The paper is well-written. The introduction and method are easy to follow.\n\n- The experiments are very comprehensive to support the main claims in the paper.\n\n- The video prediction results look similar to the real world.\n\n- The results show that Video Policy is superior to baselines.\n\n- Using action-free data provides a potentially scalable way for data-driven robot learning."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The computation cost is high. The inference speed of the video model could slow down the policy rollout. Can the authors provide a comparison between Video Policy and other policy baselines? Can the authors propose some ways to accelerate the policy FPS?\n- It remains unclear if Video Policy still performs well in tasks with higher dynamics. The video model may show physics-inplausible results. It’s interesting if the authors can explore the behaviour of the policy and the video prediction model in more dynamic tasks.\n- Using more recent art in video models may improve the policy performance.\n- The idea of treating video prediction as a robot policy has already been proven working by some earlier work such as (Shuang et al,”Unified video action model.”). Can the authors provide some comparisons with these previous works."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762920945089,"tcdate":1761950348090,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission9310/Reviewer_hZje"],"signatures":["ICLR.cc/2026/Conference/Submission9310/Reviewer_hZje"],"forum":"cWczH8ontO","number":2,"license":"CC BY 4.0","cdate":1761950348090,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission9310/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762920945089,"domain":"ICLR.cc/2026/Conference","replyto":"cWczH8ontO","id":"I1lAnDeSYV","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Behavior Cloning","Video Generation","Representation Learning"]},"supplementary_material":{"value":"/attachment/524fa64b2ccd106dcea68fc0fcfa7b106bff0d0c.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Despite tremendous progress in dexterous manipulation, current visuomotor policies remain fundamentally limited by two challenges: they struggle to generalize under perceptual or behavioral distribution shifts, and their performance is constrained by the size of human demonstration data. In this paper, we use video generation as a proxy for robot policy learning to address both limitations simultaneously.\nWe propose Video Policy, a modular framework that combines video and action generation that can be trained end-to-end. Our results demonstrate that learning to generate videos of robot behavior allows for the extraction of policies with minimal demonstration data, significantly improving robustness and sample efficiency. Our method shows strong generalization to unseen objects, backgrounds, and tasks, both in simulation and the real world. We further highlight that task success is closely tied to the generated video, with action-free video data providing critical benefits for generalizing to novel tasks. By leveraging large-scale video generative models, we achieve superior performance compared to recent VLAs and video-action models, paving the way for more scalable and data-efficient robot policy learning."},"_bibtex":{"value":"@misc{\nliang2025video,\ntitle={Video Generators are Robot Policies},\nauthor={Junbang Liang and Pavel Tokmakov and Ruoshi Liu and Sruthi Sudhakar and Paarth Shah and Rares Andrei Ambrus and Carl Vondrick},\nyear={2025},\nurl={https://openreview.net/forum?id=cWczH8ontO}\n}"},"title":{"value":"Video Generators are Robot Policies"},"pdf":{"value":"/pdf/cdf5affcaceddafbd0c7100e9ca873e764edfd23.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"liang|video_generators_are_robot_policies"},"authorids":{"value":["~Junbang_Liang2","~Pavel_Tokmakov2","~Ruoshi_Liu2","~Sruthi_Sudhakar1","~Paarth_Shah2","~Rares_Andrei_Ambrus1","~Carl_Vondrick2"]},"authors":{"value":["Junbang Liang","Pavel Tokmakov","Ruoshi Liu","Sruthi Sudhakar","Paarth Shah","Rares Andrei Ambrus","Carl Vondrick"]}},"version":2},{"content":{"summary":{"value":"The paper introduces MagicTryOn, a diffusion transformer-based framework for garment-preserving video virtual try-on (VVT) that aims to improve both garment-detail fidelity and spatiotemporal consistency in generated try-on videos. The approach leverages a fine-grained garment-preservation strategy (decomposing cues into semantic, structure, and appearance), introduces a garment-aware spatiotemporal rotary position embedding (RoPE) for enhanced temporal consistency, and utilizes a mask-aware loss to enforce garment-region fidelity. Furthermore, it integrates a distribution-matching distillation to accelerate inference without compromising quality. The authors present quantitative and qualitative improvements over several state-of-the-art methods on public VVT benchmarks, supported by ablations and user studies."},"soundness":{"value":2},"confidence":{"value":5},"questions":{"value":"See weakness."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. Innovative Spatiotemporal Encoding: The extension of RoPE to “garment-aware spatiotemporal RoPE” is principled and directly addresses temporal instability; the subsuming of garment tokens into the full self-attention with grid extension is mathematically described and justified.\n2. Thorough Ablations: Table 3 and Figure 12 provide a detailed ablation of key architectural modules, with analyses that specify the performance and qualitative impact of removing each token stream or loss component, which aids scientific transparency.\n3. Clear writting."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Limited Novelty in Architectural Choices: While the garments’ semantic/structural/appearance decomposition and the cross-attention wiring are interesting, the overall method mainly combines existing mechanisms from prior VVT and diffusion transformer literature, such as patchification, CLIP feature usage, full self-attention, and cross-token fusion. As evident from the Related Work and as per foundational methods like ViViD, CatV2TON, or Hunyuan-DiT (all cited in the main text), many core strategies (semantic cross-attention, diffusion transformer backbone, pose-agnostic inputs, multi-modal fusion) are also used in recent related architectures.\n2. In addition, the semantic/structural/appearance sequence tokens are redundant and can be replaced by a limited number of keyframes, thereby further enhancing efficiency and reducing preprocessing requirements.\n3. Based on the video materials provided, the clothing replacements demonstrated in the paper appear limited—for instance, only showcasing exchanges between different T-shirts or between various pairs of jeans, while failing to present cross-category swaps such as skirts versus trousers. This limitation may be attributed to constraints inherent in the mask-based approach used."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762927234789,"tcdate":1761104441506,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission17298/Reviewer_xRnT"],"signatures":["ICLR.cc/2026/Conference/Submission17298/Reviewer_xRnT"],"forum":"JGiSS7hLRT","number":2,"license":"CC BY 4.0","cdate":1761104441506,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission17298/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762927234789,"domain":"ICLR.cc/2026/Conference","replyto":"JGiSS7hLRT","id":"4q4xKVr34i","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Video Virtual Try-On","Diffusion Model","Video Generation"]},"supplementary_material":{"value":"/attachment/11697514ce3d41656d51492563e88f52edb2457e.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Video Virtual Try-On (VVT) aims to synthesize garments that appear natural across consecutive video frames, capturing both their dynamics and interactions with human motion. Despite recent progress, existing VVT methods still suffer from inadequate garment fidelity and limited spatiotemporal consistency. The reasons are (i) under-exploitation of garment information, with limited garment cues being injected, resulting in weaker fine-detail fidelity, and (ii) the lack of spatiotemporal modeling, which hampers cross-frame identity consistency and causes temporal jitter and appearance drift. In this paper, we present MagicTryOn, a diffusion transformer–based framework for garment-preserving video virtual try-on. To preserve fine-grained garment details, we propose a fine-grained garment-preservation strategy that disentangles garment cues and injects these decomposed priors into the denoising process. To improve temporal garment consistency and suppress jitter, we introduce a garment-aware spatiotemporal rotary positional embedding (RoPE) that extends RoPE within full self-attention, using spatiotemporal relative positions to modulate garment tokens. We further impose a mask-aware loss during training to enhance fidelity within garment regions. Moreover, we adopt distribution-matching distillation to compress the sampling trajectory to four steps, enabling real-time inference without degrading garment fidelity. Extensive quantitative and qualitative experiments demonstrate that MagicTryOn outperforms existing methods, delivering superior garment-detail fidelity and temporal stability in unconstrained settings. Code will be made publicly available."},"_bibtex":{"value":"@misc{\nli2026magictryon,\ntitle={MagicTryOn: Harnessing Diffusion Transformer for Garment-Preserving Video Virtual Try-on},\nauthor={Guangyuan Li and Siming Zheng and Hao Zhang and Jinwei Chen and Junsheng Luan and Binkai Ou and Lei Zhao and Bo Li and Peng-Tao Jiang},\nyear={2026},\nurl={https://openreview.net/forum?id=JGiSS7hLRT}\n}"},"title":{"value":"MagicTryOn: Harnessing Diffusion Transformer for Garment-Preserving Video Virtual Try-on"},"pdf":{"value":"/pdf/5dbbd79b43084f5b68d8a72d876d9d5572ff0f22.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"li|magictryon_harnessing_diffusion_transformer_for_garmentpreserving_video_virtual_tryon"},"authorids":{"value":["~Guangyuan_Li1","~Siming_Zheng3","~Hao_Zhang52","~Jinwei_Chen3","~Junsheng_Luan1","~Binkai_Ou1","~Lei_Zhao3","~Bo_Li20","~Peng-Tao_Jiang1"]},"authors":{"value":["Guangyuan Li","Siming Zheng","Hao Zhang","Jinwei Chen","Junsheng Luan","Binkai Ou","Lei Zhao","Bo Li","Peng-Tao Jiang"]}},"version":2},{"content":{"summary":{"value":"This paper introduces MA-EgoQA, a benchmark dataset for question answering over long-duration, multi-agent egocentric videos. The benchmark is built on the EgoLife dataset, which features 266 hours of video from 6 agents interacting in a shared house over 7 days. The authors' core contribution is a set of 1.7k question-answer pairs specifically designed to be answerable only by aggregating information from multiple agents' video streams. The benchmark spans five challenging, multi-agent categories: Social Interaction, Task Coordination, Theory-of-Mind (ToM), Temporal Reasoning, and Environmental Interaction.\n\nThe paper uses a \"single-agent filtering\" step to remove any questions solvable by one agent's perspective alone. The authors also propose EgoMAS, a training-free baseline model that uses an event-based shared memory and agent-wise dynamic retrieval.\n\nExperimental results, testing over 16 baselines (including SOTA LLMs and Video LLMs), show that MA-EgoQA is challenging. The top model, Gemini-2.5-flash, achieves only 36.93% accuracy (vs. 20% random chance), and all models struggle with the Theory-of-Mind category."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"- More qualitative analysis \n\n- How do the VLM models handle long videos in the evaluation of Table 2? \n\n- What is the distribution of video lengths in the benchmark? How does this influence the performance?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"- The paper addresses a new and interesting topic of multi-agent egocentric videos, where video streams are captured continuously during operation.\n\n- The data is verified and selected by human annotators."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Not enough qualitative analysis. Can the author show and analyze more benchmark samples? One sample is not enough to help audience understand the scope and quality of the benchmark data.\n\n- The performance gain from adding video frames in the EgoMAS (Text+Video) variant over the EgoMAS (Text) variant is very modest (35.96% vs. 35.55%). This suggests either that the text is sufficient for most questions or that the model's method of incorporating video (sampling 8 frames) is not sophisticated enough to be impactful."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1764358099838,"tcdate":1761893350640,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission24786/Reviewer_oPM5"],"signatures":["ICLR.cc/2026/Conference/Submission24786/Reviewer_oPM5"],"forum":"wPr54hJYuF","number":1,"license":"CC BY 4.0","cdate":1761893350640,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission24786/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1764358099838,"domain":"ICLR.cc/2026/Conference","replyto":"wPr54hJYuF","id":"BoC54EjeA1","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Egocentric Video Understanding","Multi-Agent System"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"As embodied models become powerful, humans will collaborate with multiple embodied AI agents at their workplace or home in the future. To ensure better communication between human users and the multi-agent system, it is crucial to interpret incoming information from agents in parallel and refer to the appropriate context for each query. Existing challenges are to effectively compress and communicate high volumes of individual sensory inputs in the form of video and to correctly aggregate multiple egocentric videos to construct system-level memory. In this work, we first formally define a novel problem of understanding multiple long-horizon egocentric videos simultaneously collected from embodied agents. To facilitate research in this direction, we introduce MultiAgent-EgoQA (MA-EgoQA), a benchmark designed to systemically evaluate existing models in our scenario. MA-EgoQA provides 1.7k questions unique to multiple egocentric streams, spanning five categories: social interaction, task coordination, theory-of-mind, temporal reasoning, and environmental interaction. We further propose a simple baseline model for MA-EgoQA named EgoMAS, which leverages shared memory across embodied agents and agent-wise dynamic retrieval. Through comprehensive evaluation across diverse baselines and EgoMAS on MA-EgoQA, we find that current approaches are unable to effectively handle multiple egocentric streams, highlighting the need for future advances in this direction."},"_bibtex":{"value":"@misc{\nkim2026maegoqa,\ntitle={{MA}-Ego{QA}: Question Answering over Egocentric Videos  from Multiple Embodied Agents},\nauthor={Kangsan Kim and Yanlai Yang and Suji Kim and Woongyeong Yeo and Youngwan Lee and Sung Ju Hwang and Mengye Ren},\nyear={2026},\nurl={https://openreview.net/forum?id=wPr54hJYuF}\n}"},"title":{"value":"MA-EgoQA: Question Answering over Egocentric Videos  from Multiple Embodied Agents"},"pdf":{"value":"/pdf/3f47903c3ff0d002abd1d63319828a743214420a.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"kim|maegoqa_question_answering_over_egocentric_videos_from_multiple_embodied_agents"},"authorids":{"value":["~Kangsan_Kim1","~Yanlai_Yang1","~Suji_Kim2","~Woongyeong_Yeo1","~Youngwan_Lee1","~Sung_Ju_Hwang1","~Mengye_Ren1"]},"authors":{"value":["Kangsan Kim","Yanlai Yang","Suji Kim","Woongyeong Yeo","Youngwan Lee","Sung Ju Hwang","Mengye Ren"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a novel framework for identifying a minimal subset of entries in the adjacency matrix that uniquely distinguishes a directed acyclic graph (DAG) among all transitively-closed DAGs. The authors provide a provably optimal algorithm for computing this minimal set. They then leverage these insights to develop hierarchy-aware sampling, which allows training node embedding models more efficiently by focusing only on the essential graph information. Specifically, they prove that models with an inductive bias of transitivity, such as box embeddings, can learn faithful representations of hierarchies using substantially fewer training examples selected by their sampling approach. Experiments on synthetic and real-world graphs demonstrate that their method significantly reduces training data by up to 99% while maintaining or improving model performance and convergence rates compared to uniform negative sampling."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"- The framework of sidigraphs for distinguishing graphs with a given property seems quite general. Have the authors considered instantiating it for properties other than transitivity?\n- Do the authors have intuition or theory on why hierarchy-aware sampling provides a greater speed-up on some graphs than others (e.g. balanced trees vs Price graphs)?\n- Could hierarchy-aware sampling also improve the computational efficiency of non-embedding GNN approaches for learning hierarchical structures, since they also implicitly perform a form of negative sampling via contrastive estimation?"},"rating":{"value":7},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"- The theoretical grounding of the paper looks good to me, with key properties of the proposed framework and algorithms formally stated and proven. The experiments systematically evaluated the impact of algorithm design choices (e.g. positive/negative edge sets) and data characteristics (e.g. graph structures) on model performance.\n- Learning faithful graph representations with minimal data has important computational benefits and is conceptually valuable for understanding the essential structural information needed to distinguish different graphs. The reduction in training data achieved by the proposed sampling method is substantial. The theoretical framework and practical techniques in this paper could be built upon to develop further improved graph representation learning approaches."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- While the paper demonstrates the effectiveness of transitivity bias and hierarchy-aware negative sampling for DAGs, the authors could discuss whether these ideas extend to learning representations of other graph families characterized by different structural properties. Are there other useful inductive biases worth incorporating into node embeddings and corresponding \"structure-aware\" sampling strategies? \n- The experiments focus on evaluating hierarchy-aware sampling, but there is less empirical analysis of the FINDMINDISTINGUISHER algorithm itself, e.g. runtime complexity, actual sparsity of computed sidigraphs, etc. Knowing when the sidigraph is substantially smaller than the original graph is important for determining when hierarchy-aware sampling is computationally beneficial."},"limitations":{"value":"Yes, there is a section about the limitations of the work."}},"nonreaders":[],"tmdate":1730879683602,"tcdate":1720841647691,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission14075/Reviewer_gCXR"],"signatures":["NeurIPS.cc/2024/Conference/Submission14075/Reviewer_gCXR"],"forum":"HFS800reZK","number":3,"license":"CC BY 4.0","cdate":1720841647691,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission14075/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730879683602,"domain":"NeurIPS.cc/2024/Conference","replyto":"HFS800reZK","id":"f4nPfldfiG","forumContent":{"TLDR":{"value":"We provide the provably minimal set of entries from the adjacency matrix necessary to train representations for transitively-closed DAGs, provided our energy function has transitivity bias."},"venue":{"value":"NeurIPS 2024 poster"},"keywords":{"value":["graph embeddings","representation learning"]},"primary_area":{"value":"other"},"abstract":{"value":"When training node embedding models to represent large directed graphs (digraphs), it is impossible to observe all entries of the adjacency matrix during training. As a consequence most methods employ sampling. For very large digraphs, however, this means many (most) entries may be unobserved during training. In general, observing every entry would be necessary to uniquely identify a graph, however if we know the graph has a certain property some entries can be omitted - for example, only half the entries would be required for a symmetric graph. \nIn this work, we develop a novel framework to identify a subset of entries required to uniquely distinguish a graph among all transitively-closed DAGs. We give an explicit algorithm to compute the provably minimal set of entries, and demonstrate empirically that one can train node embedding models with greater efficiency and performance, provided the energy function has an appropriate inductive bias. We achieve robust performance on synthetic hierarchies and a larger real-world taxonomy, observing improved convergence rates in a resource-constrained setting while reducing the set of training examples by as much as 99%."},"_bibtex":{"value":"@inproceedings{\nrozonoyer2024learning,\ntitle={Learning Representations for Hierarchies with Minimal Support},\nauthor={Benjamin Rozonoyer and Michael Boratko and Dhruvesh Patel and Wenlong Zhao and Shib Sankar Dasgupta and Hung Le and Andrew McCallum},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=HFS800reZK}\n}"},"title":{"value":"Learning Representations for Hierarchies with Minimal Support"},"pdf":{"value":"/pdf/98c7ccf6ef86019ffc994aba434e5c6603739459.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"rozonoyer|learning_representations_for_hierarchies_with_minimal_support"},"authorids":{"value":["~Benjamin_Rozonoyer1","~Michael_Boratko1","~Dhruvesh_Patel1","~Wenlong_Zhao1","~Shib_Sankar_Dasgupta2","~Hung_Le4","~Andrew_McCallum1"]},"authors":{"value":["Benjamin Rozonoyer","Michael Boratko","Dhruvesh Patel","Wenlong Zhao","Shib Sankar Dasgupta","Hung Le","Andrew McCallum"]}},"version":2},{"content":{"summary":{"value":"* The paper develops a white box gradient-based image jailbreak method for multimodal fusion models. Prior work on gradient-based image jailbreaks has focused on VLMs due to the lack of open source fusion models, but this has recently changed with the release of Chameleon.  \n* The core challenge of doing this with fusion models is that gradients do not flow through to the input image due to a non-differentiable step in tokenization.   \n* The authors solve this problem by introducing a novel “tokenizer shortcut” technique, where they train a small MLP to approximate the image tokenizer in a differentiable way. The tokenizer is then replaced by this approximation during adversarial image optimization, allowing gradient-based optimization to succeed.  \n* Two versions of the tokenizer shortcut are developed, one mapping directly to embedding space and one producing a one-hot vocabulary encoding.  \n* A comprehensive set of experiments are conducted. Key findings:  \n  * Both shortcut methods produce images with high Attack Success Rate, but only the 1-hot shortcut images transfer to versions of the model that do not use the shortcut.  \n  * Circuit breakers substantially reduce ASR  \n  * Jailbreak images transfer easily across prompts but do not transfer across models  \n* The authors also conduct a series of ablations, including on response prefix, softmax temperature, and number of train prompts."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"* Do the authors have an explanation for why the embedding space shortcut attacks do not transfer to non-shortcut models while the 1-hot attacks do?"},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":4},"strengths":{"value":"* This is a novel method that solves the core challenge of creating gradient-based image jailbreaks for multimodal fusion models.  \n* Understanding the vulnerabilities in multimodal models is important for developing more robust systems, and gradient-based jailbreaking of fusion-based models has been under-explored.  \n* The authors use good baselines for their experiments (GCG and refusal direction attacks), and convincingly demonstrate the success of their method  \n* The experiments are thorough and informative, testing attack transfer as well as defence using circuit breakers. Several ablations are also performed."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"* The dataset used is quite small, with only 80 prompts in the test set for direct attacks and 20 in the test set for transfer attacks. The results would be more convincing if done on a larger dataset. In addition, only a single dataset is tested.  \n* The paper does not include any examples of jailbroken model responses - these are helpful for qualitative understanding of the attack.\n* With the exception of table 1, the results given are all for models using the tokenizer shortcut. It would be helpful to also include  the results when using the 1-hot jailbreak images on models without the shortcut in Tables 2 and 4."}},"nonreaders":[],"tmdate":1731427791330,"tcdate":1730697762628,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission10423/Reviewer_zzub"],"signatures":["ICLR.cc/2025/Conference/Submission10423/Reviewer_zzub"],"forum":"wNg0LibmQt","number":5,"license":"CC BY 4.0","cdate":1730697762628,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission10423/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427791330,"domain":"ICLR.cc/2025/Conference","replyto":"wNg0LibmQt","id":"dz56N96fCI","forumContent":{"TLDR":{"value":"We introduce the notion of a tokenizer shortcut that enables the first gradient-based image jailbreak attack against multimodal fusion models."},"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["jailbreak","adversarial examples","multimodal","language models"]},"primary_area":{"value":"alignment, fairness, safety, privacy, and societal considerations"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Augmenting language models with image inputs may enable more effective jailbreak attacks through continuous optimization, unlike text inputs that require discrete optimization. However, new *multimodal fusion models* tokenize all input modalities using non-differentiable functions, which hinders straightforward attacks. In this work, we introduce the notion of a *tokenizer shortcut* that approximates tokenization with a continuous function and enables continuous optimization. We use tokenizer shortcuts to create the first end-to-end gradient image attacks against multimodal fusion models. We evaluate our attacks on Chameleon models and obtain jailbreak images that elicit harmful information for 72.5% of prompts. Jailbreak images outperform text jailbreaks optimized with the same objective and require 3x lower compute budget to optimize 50x more input tokens. Finally, we find that representation engineering defenses, like Circuit Breakers, trained only on text attacks can effectively transfer to adversarial image inputs."},"_bibtex":{"value":"@misc{\nrando2025gradientbased,\ntitle={Gradient-based Jailbreak Images for Multimodal Fusion Models},\nauthor={Javier Rando and Hannah Korevaar and Erik Brinkman and Ivan Evtimov and Florian Tram{\\`e}r},\nyear={2025},\nurl={https://openreview.net/forum?id=wNg0LibmQt}\n}"},"title":{"value":"Gradient-based Jailbreak Images for Multimodal Fusion Models"},"pdf":{"value":"/pdf/f563e0ef8fb0d45d1ba1ab53bf4770ac8a21867e.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"rando|gradientbased_jailbreak_images_for_multimodal_fusion_models"},"authorids":{"value":["~Javier_Rando2","~Hannah_Korevaar1","~Erik_Brinkman1","~Ivan_Evtimov2","~Florian_Tramèr1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Javier Rando","Hannah Korevaar","Erik Brinkman","Ivan Evtimov","Florian Tramèr"]}},"version":2},{"content":{"venue":{"value":"CVPR 2026"},"abstract":{"value":"Tracking-any-point (TAP) answers query-conditioned correspondence but leaves the dense, all-pairs structure of a video implicit. We formulate All-Pairs Tracking (APT): given a video, predict dense displacement and visibility for every source-target frame pair, from which per-pixel trajectories can be read out. To this end, we propose PairFormer, a feed-forward transformer that addresses APT in a single pass. A spatio-temporal patch encoder computes temporally conditioned features for all frames. \\revaddtwo{CorrBank builds a learnable correlation bank for each frame pair and produces pairwise motion tokens.} A broadcast motion mixer aggregates trajectory-wise context and broadcasts it back to refine the pairwise motion tokens. A trajectory head first predicts coarse dense displacement, visibility, and confidence, and then refines them iteratively to form a coherent all-pairs trajectory field. To support APT at scale, we develop PAIRender, a data platform that synthesizes photo-realistic dynamic scenes with dense annotations. From PAIRender we derive a training set ($\\pi$-R10K) and a benchmark (APT-Bench) with an all-to-all evaluation protocol. Experiments show that PairFormer achieves strong performance on APT-Bench and competitive results on standard TAP benchmarks. Code and dataset will be released upon publication."},"_bibtex":{"value":"@inproceedings{\nwu2026matching,\ntitle={Matching Every Pair to Track Every Point: PairFormer for All-Pairs Tracking and Video Trajectory Fields},\nauthor={Guangyang Wu and Youran Ding and Xinyu Che and Benyuan Sun and Yi Yang and Xiaohong Liu},\nbooktitle={Conference on Computer Vision and Pattern Recognition 2026},\nyear={2026},\nurl={https://openreview.net/forum?id=3gbQM0Xuvb}\n}"},"title":{"value":"Matching Every Pair to Track Every Point: PairFormer for All-Pairs Tracking and Video Trajectory Fields"},"pdf":{"value":"https://openaccess.thecvf.com/content/CVPR2026/papers/Wu_Matching_Every_Pair_to_Track_Every_Point_PairFormer_for_All-Pairs_CVPR_2026_paper.pdf"},"venueid":{"value":"thecvf.com/CVPR/2026/Conference"},"paperhash":{"value":"wu|matching_every_pair_to_track_every_point_pairformer_for_allpairs_tracking_and_video_trajectory_fields"},"authorids":{"value":["~Guangyang_Wu1","~Youran_Ding1","~Xinyu_Che1","~Benyuan_Sun2","~Yi_Yang3","~Xiaohong_Liu2"]},"authors":{"value":["Guangyang Wu","Youran Ding","Xinyu Che","Benyuan Sun","Yi Yang","Xiaohong Liu"]}},"tmdate":1789656688518,"pdate":1789656438589,"tcdate":1765222480046,"writers":["thecvf.com/CVPR/2026/Conference","thecvf.com/CVPR/2026/Conference/Submission40408/Authors"],"signatures":["thecvf.com/CVPR/2026/Conference/Submission40408/Authors"],"forum":"3gbQM0Xuvb","license":"CC BY 4.0","number":40408,"cdate":1765222480046,"readers":["everyone"],"invitations":["thecvf.com/CVPR/2026/Conference/-/Submission","thecvf.com/CVPR/2026/Conference/Submission40408/-/Full_Submission","thecvf.com/CVPR/2026/Conference/-/Post_Submission","thecvf.com/CVPR/2026/Conference/Submission40408/-/Supplementary_Material","thecvf.com/CVPR/2026/Conference/-/Edit","thecvf.com/CVPR/2026/Conference/-/Compute_Flag"],"mdate":1789656688518,"odate":1789656438589,"domain":"thecvf.com/CVPR/2026/Conference","id":"3gbQM0Xuvb","version":2},{"content":{"venue":{"value":"NeurIPS 2023 poster"},"keywords":{"value":["shortcut learning","spurious correlations","perfect stable feature","perception tasks","implicit bias in optimization","improving inductive biases"]},"_bibtex":{"value":"@inproceedings{\npuli2023dont,\ntitle={Don{\\textquoteright}t blame Dataset Shift! Shortcut Learning due to Gradients and Cross Entropy},\nauthor={Aahlad Manas Puli and Lily H Zhang and Yoav Wald and Rajesh Ranganath},\nbooktitle={Thirty-seventh Conference on Neural Information Processing Systems},\nyear={2023},\nurl={https://openreview.net/forum?id=zyZkaqNnpa}\n}"},"title":{"value":"Don’t blame Dataset Shift! Shortcut Learning due to Gradients and Cross Entropy"},"paperhash":{"value":"puli|dont_blame_dataset_shift_shortcut_learning_due_to_gradients_and_cross_entropy"},"TLDR":{"value":"Implicit biases toward maximizing margins induce shortcut learning in ERM even in tasks with perfect stable features, controlling margins mitigates shortcuts"},"abstract":{"value":"Common explanations for shortcut learning assume that the shortcut improves prediction only under the training distribution. Thus, models trained in the typical way by minimizing log-loss using gradient descent, which we call default-ERM, should utilize the shortcut. However, even when the stable feature determines the label in the training distribution and the shortcut does not provide any additional information, like in perception tasks, default-ERM exhibits shortcut learning. Why are such solutions preferred when the loss can be driven to zero when using the stable feature alone? By studying a linear perception task, we show that default-ERM’s preference for maximizing the margin, even without overparameterization, leads to models that depend more on the shortcut than the stable feature. This insight suggests that default-ERM’s implicit inductive bias towards max-margin may be unsuitable for perception tasks. Instead, we consider inductive biases toward uniform margins. We show that uniform margins guarantee sole dependence on the perfect stable feature in the linear perception task and suggest alternative loss functions, termed margin control (MARG-CTRL), that encourage uniform-margin solutions. MARG-CTRL techniques mitigate shortcut learning on a variety of vision and language tasks, showing that changing inductive biases can remove the need for complicated shortcut-mitigating methods in perception tasks."},"pdf":{"value":"/pdf/956a72600022df2e6738151ad2ff6a638cc32154.pdf"},"venueid":{"value":"NeurIPS.cc/2023/Conference"},"authorids":{"value":["~Aahlad_Manas_Puli1","~Lily_H_Zhang1","~Yoav_Wald1","~Rajesh_Ranganath2"]},"authors":{"value":["Aahlad Manas Puli","Lily H Zhang","Yoav Wald","Rajesh Ranganath"]}},"tmdate":1698949793550,"pdate":1695326179171,"tcdate":1683835167534,"writers":["NeurIPS.cc/2023/Conference","NeurIPS.cc/2023/Conference/Submission15594/Authors"],"signatures":["NeurIPS.cc/2023/Conference/Submission15594/Authors"],"forum":"zyZkaqNnpa","number":15594,"cdate":1683835167534,"mdate":1698949793550,"readers":["everyone"],"invitations":["NeurIPS.cc/2023/Conference/-/Submission","NeurIPS.cc/2023/Conference/-/Post_Submission","NeurIPS.cc/2023/Conference/Submission15594/-/Revision","NeurIPS.cc/2023/Conference/Submission15594/-/Supplementary_Material_Revision","NeurIPS.cc/2023/Conference/-/Edit","NeurIPS.cc/2023/Conference/Submission15594/-/Camera_Ready_Revision"],"odate":1698949793534,"domain":"NeurIPS.cc/2023/Conference","id":"zyZkaqNnpa","version":2},{"content":{"summary":{"value":"This paper proposes Mani-WM, a trajectory-conditioned interactive video generation model for robot manipulation tasks. It proposes a novel trajectory-conditioning mechanism that ensures the generated video is aligned with the given trajectory. Experiment results show that Mani-WM outperforms all the comparing baseline methods and is more preferable in human evaluations, and the trained model supports interactive instructions and model-based planning."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"See the weakness."},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"Strongness:\n\n1.\tThe video generation results are good, especially on high-resolution videos. \n\n2.\tThe writing is clear, especially the method section. \n\n3.\tThe interactive application and the model-based planning demos are good to show the application of this paper."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Weakness:\n\n1.\tTo my understanding, the trajectory is to ensure the end-effector in the generated video follows the given trajectory. However, the authors didn’t explicitly evaluate this. Instead, they mainly evaluate the overall generated video quality. This is strange and unfair. So I think: 1) if the authors want to evaluate the quality of the generated videos, they should compare Mani-WM to other video generation models with enhancement techniques, such as video generation with VLM/human feedback, or self-conditioning techniques for diffusion models. Comparing Mani-WM to current baselines is unfair because they are not originally designed for extra trajectory conditions. 2) The author should explicitly evaluate how well the end-effector in the generated videos follow the given trajectory condition. \n\n2.\tThe authors didn’t explore adding explicit loss to ensure that the generated video meets the input trajectory conditions. For example, the author may consider [1]. Currently, the trajectory conditions are only implicitly injected to the model by training with trajectory conditions. The author may also need to show that videos generated by other generative models satisfy the trajectory condition (even if they do not explicitly use trajectories as input)?\n\n3.\tThe trajectory condition data acquisition is a limitation of this work. The authors didn’t show how to extract the trajectories of the end-effector from video datasets. This could contain several practical problems: 1) how to deal with the occlusion problem in datasets (the end-effector may be obscured by the objects)? This may cause the trajectory elimination issue, and I think this is very common for long-horizon videos, which is the main type of data that this article mainly promotes; 3) How to deal with stochastic environments dynamics? For example, consider the opening round of billiards. Even if you use the same trajectory every time you tee off, the other balls may not go the same way every time.\n\n4.\tUsing the end-effector trajectory as the general representation for the condition is not a universal idea for diverse robot video datasets, such as dexterous hand manipulations (where more than one point on the robot are useful) and robot unscrewing the cap (where the end-effector is not moving but rotating). This conflicts with the author's claim that “training a trajectory-to-video model only requires trajectory-video pairs, which is very scalable”.\n\n[1] Geng, Daniel, and Andrew Owens. \"Motion guidance: Diffusion-based image editing with differentiable motion estimators.\" arXiv preprint arXiv:2401.18085 (2024)."}},"nonreaders":[],"tmdate":1731428558126,"tcdate":1730619350468,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission6028/Reviewer_AVzL"],"signatures":["ICLR.cc/2025/Conference/Submission6028/Reviewer_AVzL"],"forum":"aVyJwS1fqQ","number":2,"license":"CC BY 4.0","cdate":1730619350468,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission6028/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428558126,"domain":"ICLR.cc/2025/Conference","replyto":"aVyJwS1fqQ","id":"GcJabguTpD","forumContent":{"TLDR":{"value":"We develop a novel method, Mani-WM, which leverages the power of generative models to generate realistic videos of a robot executing a given action trajectory, starting from an initial given frame."},"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["World Model","Video Generation","Robot Manipulation"]},"supplementary_material":{"value":"/attachment/548f40315f123100a0837c5ca8fd4f031ab6a611.zip"},"primary_area":{"value":"applications to robotics, autonomy, planning"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Scalable robot learning in the real world is limited by the cost and safety issues of real robots. In addition, rolling out robot trajectories in the real world can be time-consuming and labor-intensive. In this paper, we propose to learn an interactive world model for robot manipulation as an alternative. We present a novel method, Mani-WM, which leverages the power of generative models to generate realistic videos of a robot arm executing a given action trajectory, starting from an initial given frame. Mani-WM employs a novel frame-level conditioning technique to ensure precise alignment between actions and video frames and leverages a diffusion transformer for high-quality video generation. To validate the effectiveness of Mani-WM, we perform extensive experiments on four challenging real-robot datasets. Results show that Mani-WM outperforms all the comparing baseline methods and is more preferable in human evaluations. We further showcase the flexible action controllability of Mani-WM by controlling the virtual robots in datasets with trajectories 1) predicted by an autonomous policy and 2) collected by a keyboard or VR controller. Finally, we combine Mani-WM with model-based planning to showcase its usefulness on real-robot manipulation tasks. We hope that Mani-WM can serve as an effective and scalable approach to enhance robot learning in the real world. To promote research on manipulation world models, we opensource the code at https://anonymous.4open.science/r/Mani-WM."},"_bibtex":{"value":"@misc{\nzhu2024maniwm,\ntitle={Mani-{WM}: An Interactive World Model for Real-Robot Manipulation},\nauthor={Fangqi Zhu and Hongtao Wu and Song Guo and Yuxiao Liu and Chilam Cheang and Tao Kong},\nyear={2024},\nurl={https://openreview.net/forum?id=aVyJwS1fqQ}\n}"},"title":{"value":"Mani-WM: An Interactive World Model for Real-Robot Manipulation"},"pdf":{"value":"/pdf/9ec54f61ff8914ca5464dc5fc4c8454985e042fe.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"zhu|maniwm_an_interactive_world_model_for_realrobot_manipulation"},"authorids":{"value":["~Fangqi_Zhu1","~Hongtao_Wu2","~Song_Guo5","~Yuxiao_Liu3","~Chilam_Cheang1","~Tao_Kong3"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Fangqi Zhu","Hongtao Wu","Song Guo","Yuxiao Liu","Chilam Cheang","Tao Kong"]}},"version":2},{"content":{"correctness":{"value":"Yes"},"summary_and_contributions":{"value":"This paper proposes LVD-2M, a large-scale long-take video dataset with temporally dense captions. The proposed dataset features 1) long-form videos with high image quality and large motions, 2) temporally dense captions with rich information and high video-text alignment, and 3) varieties of video duration, caption length, and video categories. The authors fine-tuned an LM-based T2V model and a diffusion-based I2V model using LVD-2M to verify its effectiveness. Both qualitative and quantitative results show the proposed dataset could lead to a higher generation quality for both T2V and I2V models."},"confidence":{"value":5},"documentation":{"value":"Yes"},"rating":{"value":7},"title":{"value":"A good paper"},"ethics":{"value":"No"},"clarity":{"value":"Yes"},"review":{"value":"Pros:  \n\n1. Long-form video generation models face a shortage of high-quality video-text pairs. The proposed large-scale, high-quality dataset is valuable to the community.\n\n2. Technical routines are clear and intuitive. Each step is clearly elucidated and I believe these steps are reasonable and indispensable.\n\n3. The experiments are extensive, including both qualitative and quantitative results. The quality of both the proposed dataset and models fine-tuned by it are evaluated to make the technical contribution more robust.\n\nCons:\n\n1. For the methodology, I believe some ablations are needed for the expert models adopted for the pipelines. Since there are many alternatives, the rationale for selecting these models needs some discussion.  \n\n2. For the experiments, I am curious about the baseline selections. Why do not select some diffusion-based T2V model for evaluation? Does that mean the proposed method is only available for LM-based T2V models? More explanations are needed.\n\n3. Though the authors conduct some human studies to evaluate the quality of the dataset and models, I think some objective metrics are also required. For example, can the proposed pipeline be applied to some video caption benchmarks with long duration? Or if Clipsim or GPT score can be used to evaluate the caption quality? For the T2V and I2V models, can some video generation benchmarks such as VBench or EvalCrafter be applicable?  \n\nOverall I think the paper is good and the proposed dataset is valuable for the community. Though some concerns are addressed, I think they can be solved via some revisions. Therefore, I give my original rating as 7. I hope the authors can address these problems above during the rebuttal period."},"strengths":{"value":"1. Long-form video generation models face a shortage of high-quality video-text pairs. The proposed large-scale, high-quality dataset is valuable to the community.\n\n2. Technical routines are clear and intuitive. Each step is clearly elucidated and I believe these steps are reasonable and indispensable.\n\n3. The experiments are extensive, including both qualitative and quantitative results. The quality of both the proposed dataset and models fine-tuned by it are evaluated to make the technical contribution more robust."},"flag_for_ethics_review":{"value":"2: No, there are no or only very minor ethics concerns"},"relation_to_prior_work":{"value":"Yes"},"opportunities_for_improvement":{"value":"1. More ablations.  \n\n2. Using some diffusion-based T2V models as baselines.\n\n3. More objective evaluation."},"additional_feedback":{"value":"N/A"},"limitations":{"value":"The authors have adequately address the limitations and potential negative societal impact of their work yet it won't influence my positive feedback toward this paper."}},"nonreaders":[],"tmdate":1731500705709,"tcdate":1721976226420,"writers":["NeurIPS.cc/2024/Datasets_and_Benchmarks_Track","NeurIPS.cc/2024/Datasets_and_Benchmarks_Track/Submission1047/Reviewer_ouB8"],"signatures":["NeurIPS.cc/2024/Datasets_and_Benchmarks_Track/Submission1047/Reviewer_ouB8"],"forum":"H5bUdfM55S","number":4,"license":"CC BY 4.0","cdate":1721976226420,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Datasets_and_Benchmarks_Track/Submission1047/-/Official_Review","NeurIPS.cc/2024/Datasets_and_Benchmarks_Track/-/Edit"],"mdate":1731500705709,"domain":"NeurIPS.cc/2024/Datasets_and_Benchmarks_Track","replyto":"H5bUdfM55S","id":"H7PMYKLSSu","forumContent":{"venue":{"value":"NeurIPS 2024 Track Datasets and Benchmarks Poster"},"pdf":{"value":"/pdf/4dfd3d826849d2a9c45bb2e392d1ca998990dac8.pdf"},"keywords":{"value":["long video generation"]},"supplementary_material":{"value":"/attachment/f57dac0ce5fcfc1329de76f19efc005b506e7603.zip"},"venueid":{"value":"NeurIPS.cc/2024/Datasets_and_Benchmarks_Track"},"paperhash":{"value":"xiong|lvd2m_a_longtake_video_dataset_with_temporally_dense_captions"},"authorids":{"value":["~Tianwei_Xiong1","~Yuqing_Wang4","~Daquan_Zhou1","~Zhijie_Lin1","~Jiashi_Feng1","~Xihui_Liu1"]},"abstract":{"value":"The efficacy of video generation models heavily depends on the quality of their training datasets. Most previous video generation models are trained on short video clips, while recently there has been increasing interest in training long video generation models directly on longer videos. However, the lack of such high-quality long videos impedes the advancement long video generation. To promote research in long video generation, we desire a new dataset with four key features essential for training long video generation models: (1) long videos covering at least 10 seconds, (2) long-take videos without cuts, (3) large motion and diverse contents, and (4) temporally dense captions. To achieve this, we introduce a new pipeline for filtering high-quality long-take videos and generating temporally dense captions. Specifically, we define a set of metrics to quantitatively assess video quality including scene cuts, dynamic degrees, and semantic-level scores, enabling us to filter high-quality long-take videos from a large amount of source videos. Subsequently, we develop a hierarchical video captioning pipeline to annotate long videos with temporally-dense captions. With this pipeline, we curate the first long-take video dataset, LVD-2M, comprising 2 million long-take videos, each covering more than 10 seconds and annotated with temporally dense captions. We further validate the effectiveness of LVD-2M by fine-tuning video generation models to generate long videos with dynamic motions. We believe it will significantly contribute to future research in long video generation."},"_bibtex":{"value":"@inproceedings{\nxiong2024lvdm,\ntitle={{LVD}-2M: A Long-take Video Dataset with Temporally Dense Captions},\nauthor={Tianwei Xiong and Yuqing Wang and Daquan Zhou and Zhijie Lin and Jiashi Feng and Xihui Liu},\nbooktitle={The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track},\nyear={2024},\nurl={https://openreview.net/forum?id=H5bUdfM55S}\n}"},"title":{"value":"LVD-2M: A Long-take Video Dataset with Temporally Dense Captions"},"authors":{"value":["Tianwei Xiong","Yuqing Wang","Daquan Zhou","Zhijie Lin","Jiashi Feng","Xihui Liu"]}},"version":2},{"content":{"summary":{"value":"introduce Dreamix, a video editing framework that combines low-resolution spatio-temporal information from the original video with high-resolution synthesized information aligned with a text prompt. The framework maintains fidelity to the original video by using a preliminary stage of finetuning the model on the original video. The authors propose a mixed finetuning approach that improves motion editability by training the model on both the original video and its individual frames. Additionally, Dreamix can be used to animate still images and perform subject-driven video generation. The authors provide extensive experiments to showcase the capabilities of Dreamix."},"presentation":{"value":"3 good"},"contribution":{"value":"3 good"},"soundness":{"value":"3 good"},"strengths":{"value":"1. the proposed framework can be applied top multi-tasks like Video Editing, Image-driven Videos and Subject-driven Video Generation.\n\n2. the Mixed Video-Image Finetuning is sound and with reasonable performance \n\n3. The paper presents extensive experiments that demonstrate the ability of Dreamix."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1.  although the Mixed Video-Image Finetuning strategy is sound, the overall technical novelty is kind limited.  it is an application of VDMs with sophisticated strategy during finetuing. Also, it would be interesting to see how much would the base VDM would effect the finetuing results.\n\n2. although the paper claims high fidelity and quality for video editing, the resolutions are still low and with blurred details.  It looks like the input video/image provides layout information and the details are generated by VDMs. \n\n3. for Image-driven Videos, the method can not maintain the consistent between the input and first frame, and the quality is worse than more recent work like animatefiff.\n\n4. for Subject-driven Video Generation, the appearance feature and details can not be maintained, like the   toy in figure 2 and bear in figure 6."},"confidence":{"value":"3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked."},"questions":{"value":"Please check weakness"},"rating":{"value":"5: marginally below the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636049177,"tcdate":1698703629843,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission1227/Reviewer_cmbb"],"signatures":["ICLR.cc/2024/Conference/Submission1227/Reviewer_cmbb"],"forum":"2vAhX71UCL","number":2,"license":"CC BY 4.0","cdate":1698703629843,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission1227/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636049177,"domain":"ICLR.cc/2024/Conference","replyto":"2vAhX71UCL","id":"VZE0X3pldS","forumContent":{"venue":{"value":"ICLR 2024 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video Diffusion","Video Motion Editing"]},"supplementary_material":{"value":"/attachment/42a86714507cbe50b9d004fa457d752f84311cdf.zip"},"primary_area":{"value":"generative models"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Text-driven image and video diffusion models have recently achieved unprecedented generation realism. While diffusion models have been successfully applied for image editing, none can edit motion in video. We present the first diffusion-based method that is able to perform text-based motion and appearance editing of general, real-world videos. Our approach uses a video diffusion model to combine, at inference time, the low-resolution spatio-temporal information from the original video with new, high resolution information that it synthesized to align with the guiding text prompt. As maintaining high-fidelity to the original video requires retaining some of its high-resolution information, we add a preliminary stage of finetuning the model on the original video, significantly boosting fidelity. We propose to improve motion editability by using a mixed objective that jointly finetunes with full temporal attention and with temporal attention masking. We extend our method for animating images, bringing them to life by adding motion to existing or new objects, and camera movements. Extensive experiments showcase our method's remarkable ability to edit motion in videos."},"_bibtex":{"value":"@misc{\nmolad2024dreamix,\ntitle={Dreamix: Video Diffusion Models are General Video Editors},\nauthor={Eyal Molad and Eliahu Horwitz and Dani Valevski and Alex Rav-Acha and Yossi Matias and Yael Pritch and Yaniv Leviathan and Yedid Hoshen},\nyear={2024},\nurl={https://openreview.net/forum?id=2vAhX71UCL}\n}"},"title":{"value":"Dreamix: Video Diffusion Models are General Video Editors"},"pdf":{"value":"/pdf/268c76123e429eace4355f9d1a3ec0331d4a1a0e.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Withdrawn_Submission"},"paperhash":{"value":"molad|dreamix_video_diffusion_models_are_general_video_editors"},"authorids":{"value":["~Eyal_Molad1","~Eliahu_Horwitz1","~Dani_Valevski1","~Alex_Rav-Acha1","~Yossi_Matias2","~Yael_Pritch1","~Yaniv_Leviathan1","~Yedid_Hoshen3"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Eyal Molad","Eliahu Horwitz","Dani Valevski","Alex Rav-Acha","Yossi Matias","Yael Pritch","Yaniv Leviathan","Yedid Hoshen"]}},"version":2},{"content":{"summary":{"value":"The paper identifies two important shortcuts in the public AMPLIFAI dataset. Acquisition parameters reveal source cohort information, while missing lesion masks and zero lesion diameter strongly identify benign cases. A metadata only probe therefore achieves unexpectedly high validation performance without using voxel intensities.\n\nThe authors propose cohort grouped masked only validation and a clinically motivated model using peri lesional normalization, structured regional pooling, dense feature supervision, and phase difference channels. The method is then evaluated on the private challenge set."},"reproducibility":{"value":4},"confidence":{"value":4},"rating":{"value":4},"technical_soundness":{"value":4},"title":{"value":"A strong benchmark audit with a clinically motivated method. The shortcut analysis is convincing, although some validation claims should be stated more cautiously given the approximate grouping and substantial training variability."},"concerns":{"value":"- The acquisition grouping is only approximate for the largest cohort, so acquisition related leakage may not be completely removed.\n- Training variability is large, while several reported differences between model variants are relatively small. This limits the strength of the ablation conclusions.\n- Masked only evaluation leaves very few benign cases, making the ordinal evaluation unstable for the lower LI RADS categories."},"relevance":{"value":4},"external_resources_disclosure":{"value":"Adequately disclosed"},"review":{"value":"This is a clear and relevant paper with a useful analysis of shortcut learning in the AMPLIFAI benchmark. The most important contribution is showing that acquisition information and mask availability can produce very strong public validation scores without requiring meaningful image interpretation. The proposed validation protocol is therefore valuable, and the LI RADS motivated model design is reasonable and well explained.\n\nThe main limitation is that the evidence for the proposed model is less strong than the evidence for the benchmark shortcuts. One dominant acquisition signature still appears across folds, and the authors report large run to run variation, making small differences between model variants difficult to interpret.\n\nOverall, I find the benchmark analysis original and practically useful, and the paper is suitable for the workshop with minor revisions."},"strengths":{"value":"- The shortcut analysis is convincing and directly demonstrated using image independent probes.\n- The proposed validation protocol addresses an important issue in challenge model selection.\n- The model components are well connected to LI RADS imaging criteria rather than being purely architectural choices.\n- The paper includes useful error analysis and discusses negative results and limitations openly."},"presentation_recommendation":{"value":"Recommend for presentation"},"required_revisions":{"value":"- Please clarify the inconsistency between the five prior scenarios and the later reference to four targets.\n- Please also soften the claim that the proposed protocol fully removes acquisition related leakage, since the dominant acquisition signature is split across folds.\n- The limitations caused by training variability should be stated more explicitly when interpreting differences between variants B and C."}},"parentInvitations":"MICCAI.org/2026/Workshop/AMPLIFAI/-/Official_Review","nonreaders":[],"tmdate":1790106112422,"tcdate":1787372526279,"writers":["MICCAI.org/2026/Workshop/AMPLIFAI","MICCAI.org/2026/Workshop/AMPLIFAI/Submission4/Reviewer_awQa"],"signatures":["MICCAI.org/2026/Workshop/AMPLIFAI/Submission4/Reviewer_awQa"],"forum":"5DZQnMbNkq","number":1,"license":"CC BY 4.0","cdate":1787372526279,"readers":["everyone"],"invitations":["MICCAI.org/2026/Workshop/AMPLIFAI/Submission4/-/Official_Review","MICCAI.org/2026/Workshop/AMPLIFAI/-/Edit"],"mdate":1790106112422,"domain":"MICCAI.org/2026/Workshop/AMPLIFAI","replyto":"5DZQnMbNkq","id":"dByx90gtk2","forumContent":{"venue":{"value":"MICCAI 2026 Workshop AMPLIFAI Submission"},"keywords":{"value":["LI-RADS","Multi-phase CT","Ordinal classification","Shortcut learning","Challenge validation"]},"title":{"value":"Shortcut-Aware Validation for LI-RADS Categorisation on Multi-Phase CT"},"corresponding_author_name":{"value":"Subhamoy Mandal"},"paperhash":{"value":"vardhan|shortcutaware_validation_for_lirads_categorisation_on_multiphase_ct","readers":["everyone"]},"challenge_submission_date":{"value":"2026-08-10"},"TLDR":{"value":"Two non-transferable shortcuts inflate public AMPLIFAI cross-validation; a cohort-grouped, masked-only protocol removes them and reveals a 3D CNN baseline with zero ordinal agreement."},"abstract":{"value":"LI-RADS categorises focal liver lesions in at-risk patients on an ordinal scale from definitely benign to definitely hepatocellular carcinoma, and the assigned category informs clinical management. The AMPLIFAI challenge benchmarks automating this on multi-phase CT, scoring a composite of adjusted quadratic weighted kappa (QWK) and special-category recognition. Cross-validation scores on its public corpus\nfar exceed private-test performance—a gap usually blamed on distribution shift. We show that two non-transferable shortcuts account for\nmuch of it. First, slice thickness identifies the source cohort, whose label mixes are near-disjoint—the 0.8 mm stratum is 52% special-category, the 2.5 mm stratum 79% LR-5—so acquisition alone recognises the special categories well above chance, but carries no ordinal signal. Second, 70 lesion-free examinations are scored as LR-1 and 69 ship no segmentation, so lesion diameter—zero exactly when the mask is absent—identifies the anchoring benign class: that scalar alone reaches QWK 0.888, and 0.871 composite with acquisition added—within 0.04 of a full 3D CNN on the same folds, without reading a voxel. Every private case supplies a mask, so none of this transfers. Our cohort-grouped, masked-only protocol addresses both and discriminates where the public one cannot: a baseline scoring QWK 0.943 publicly has no ordinal agreement at all, QWK 0.000. Under it, a method built from the LI-RADS definitions—peri-lesional relative normalisation, region-structured pooling, dense feature supervision and phase-difference channels—scores 0.41 on 174 private cases from three unseen centres, where error analysis localises the residual deficit to small-volume LR-TIV read as LR-M. Both diagnostics are inexpensive and worth running on any harmonised benchmark."},"challenge_submission_id":{"value":"884586"},"team_name":{"value":"MedVisionAI"},"external_resources":{"value":"None were used in the submitted method: no external datasets, pretrained models, foundation models or pseudo-labels. All weights were trained from random initialisation on the AMPLIFAI public release alone.\n\nFor completeness, during development we evaluated the publicly available MCT-LTDiag dataset (Harvard Dataverse, doi:10.7910/DVN/S3RW15) as a domain-shift stress test. It degraded performance and was rejected; it contributed no data, weights or labels to the submitted method. This negative result is reported in Section 6 of the manuscript."},"pdf":{"value":"/pdf/42d602800c95a3abe628a3ea596c72d1f4f65a78.pdf","readers":["everyone"]},"venueid":{"value":"MICCAI.org/2026/Workshop/AMPLIFAI/Submission"},"authorids":{"readers":["everyone"],"value":["~Harsh_Vardhan5","~Sakshi_Suryawanshi1","~Subhamoy_Mandal1"]},"revision_response":{"value":"We thank the Reviewers and the Program Chairs for their constructive feedback. We have addressed all required revisions in the camera-ready manuscript as follows:\n\n1. Prior-Shift Scenario Construction & Reproducibility (Sects. 3.4 & 4.1): We clarified the prior-shift formulation and corrected the scenario count explanation. We explicitly specify the 5 evaluated scenarios—the empirical masked-only prior (needing no shift, alpha=0) plus four alternative targets (surveillance-like, flat-ordinal, HCC-heavy, few-special)—and detail the exact grid search over alpha in {0, 0.25, 0.5, 0.75, 1.0} combined with target distributions and group biases.\n\n2. Metadata Scalar Probe Benchmark (Sect. 4.2 & Table 1): We carried the zero-pixel metadata probe through the same nested decision rule (reaching 0.380) and 2.5 mm cohort transfer test (reaching 0.053 ± 0.000). This confirms, on identical criteria, that the image models' margin over legitimate scalar signals is real and large.\n\n3. Acquisition Leakage Qualification (Abstract, Sect. 2.3, & Conclusion): We softened the claim regarding acquisition shortcut removal. We explicitly note that because the dominant 0.8 x 0.78 mm acquisition signature covers 52% of the corpus and is randomly split across folds, cohort-grouping only approximately constrains acquisition-related leakage within this stratum and does not fully remove it.\n\n4. Training Variability & Model Selection (Sect. 4.2): We explicitly state the limitations imposed by training variability (up to a 0.20 reproducibility floor on single runs) and clarify that small single-run gaps (such as the 0.02 difference between variants B and C) cannot be trusted on single runs alone, justifying our selection based on nested rules and transfer test stability.","readers":["everyone"]},"authors":{"readers":["everyone"],"value":["Harsh Vardhan","Sakshi Suryawanshi","Subhamoy Mandal"]}},"version":2},{"content":{"venue":{"value":"ICLR 2027 Conference Submission"},"TLDR":{"value":"This paper proposes Alignment-aWare Distillation for improving vision-language distillation under image-text misalignment."},"keywords":{"value":["Knowledge Distillation","Vision-Language Models"]},"primary_area":{"value":"unsupervised, self-supervised, semi-supervised, and supervised representation learning"},"abstract":{"value":"Large-scale vision-language pretraining strongly benefits from data filtering, as poorly aligned image-text pairs provide noisy contrastive supervision and are often better discarded. We show that this intuition does not fully carry over to vision-language distillation. During distillation, weakly aligned pairs may indeed be harmful to the image-text contrastive objective but still provide valuable teacher supervision through feature, relation, or representation distillation. Rather than removing such samples, we propose to adapt how they are used. Specifically, we introduce an alignment-aware soft-gating mechanism that uses the teacher's image-text similarity to modulate, on a per-sample basis, alignment-sensitive objectives while preserving knowledge distillation. Well-aligned pairs receive stronger contrastive supervision, while weakly aligned pairs rely more heavily on the teacher signal. The gate is computed online from teacher embeddings already available for distillation, adding negligible overhead. Across CC12M and DataComp-M datasets, our approach consistently improves CLIP-KD and remains complementary to adaptive data selection with ACID. We further show that alignment weighting improves another distillation recipe, in TIPSv2, demonstrating that its benefits extend across different distillation settings. Our results suggest that, for vision-language distillation, noisy pairs are better reweighted than discarded. Code and pretrained weights will be open-sourced."},"_bibtex":{"value":"@inproceedings{\nanonymous2026dont,\ntitle={Don't Drop the Weak Pairs: Alignement-Aware Distillation of Vision-Language Models},\nauthor={Anonymous},\nbooktitle={Submitted to The Fifteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=mRI6ZexkJz},\nnote={under review}\n}"},"title":{"value":"Don't Drop the Weak Pairs: Alignement-Aware Distillation of Vision-Language Models"},"pdf":{"value":"/pdf/ef1447699d429ec6f534c14bddb89d26bf24ee14.pdf"},"venueid":{"value":"ICLR.cc/2027/Conference/Submission"}},"tmdate":1791230332310,"tcdate":1789554003995,"writers":["ICLR.cc/2027/Conference","ICLR.cc/2027/Conference/Submission24696/Authors"],"signatures":["ICLR.cc/2027/Conference/Submission24696/Authors"],"forum":"mRI6ZexkJz","license":"CC BY 4.0","number":24696,"cdate":1789554003995,"readers":["everyone"],"invitations":["ICLR.cc/2027/Conference/-/Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Bidding","ICLR.cc/2027/Conference/Submission24696/-/Full_Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Reviewing","ICLR.cc/2027/Conference/-/Edit"],"mdate":1791230332310,"odate":1791031328286,"domain":"ICLR.cc/2027/Conference","id":"mRI6ZexkJz","version":2},{"content":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Segmemt Anything Model","Efficient Deep Learning","Model Acceleration"]},"supplementary_material":{"value":"/attachment/b3481d522f61dec0cb42900d807bcdca9c57c111.zip"},"primary_area":{"value":"infrastructure, software libraries, hardware, systems, etc."},"abstract":{"value":"Segment Anything Model 2 (SAM2) shows excellent performance in video object segmentation tasks; however, the heavy computational burden hinders its application in real-time video processing.\nAlthough there have been efforts to improve the efficiency of SAM2, most of them focus on retraining a lightweight backbone, with little exploration into post-training acceleration.\nIn this paper, we observe that SAM2 exhibits sparse perception pattern as biological vision, which provides opportunities for eliminating redundant computation and acceleration:\ni) In mask decoder, the attention primarily focuses on the foreground objects, whereas the image encoder in the earlier stage exhibits a broad attention span, which results in unnecessary computation to background regions.\nii) In memory bank, only a small subset of tokens in each frame contribute significantly to memory attention, and the salient regions exhibit temporal consistency, making full-token computation redundant.\nWith these insights, we propose Efficient-SAM2, which promotes SAM2 to adaptively focus on object regions while eliminating task-irrelevant computations, thereby significantly improving inference efficiency.\nSpecifically, for image encoder, we propose object-aware Sparse Window Routing (SWR), a window-level computation allocation mechanism that leverages the consistency and saliency cues from the previous-frame decoder to route background regions into a lightweight shortcut branch.\nMoreover, for memory attention, we propose object-aware Sparse Memory Retrieval (SMR), which allows only the salient memory tokens in each frame to participate in computation, \nwith the saliency pattern reused from their first recollection.\nWith negligible additional parameters and minimal training overhead, Efficient-SAM2 delivers 1.68$\\times$ speedup on SAM2.1-L model with only 1.0\\% accuracy drop on SA-V test set, where SWR and SMR provide 1.83$\\times$ and 1.78$\\times$ speedups, respectively."},"_bibtex":{"value":"@inproceedings{\nzhang2026efficientsam,\ntitle={Efficient-{SAM}2: Accelerating {SAM}2 with Object-Aware Visual Encoding and Memory Retrieval},\nauthor={Jing Zhang and Zhikai Li and Xuewen Liu and Qingyi Gu},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=HcivRSezJp}\n}"},"title":{"value":"Efficient-SAM2: Accelerating SAM2 with Object-Aware Visual Encoding and Memory Retrieval"},"pdf":{"value":"/pdf/b1801056e42ab7eb9562f53bd0d351cf557070d0.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"zhang|efficientsam2_accelerating_sam2_with_objectaware_visual_encoding_and_memory_retrieval"},"authorids":{"value":["~Jing_Zhang58","~Zhikai_Li1","~Xuewen_Liu1","~Qingyi_Gu1"]},"authors":{"value":["Jing Zhang","Zhikai Li","Xuewen Liu","Qingyi Gu"]}},"tmdate":1775876960763,"pdate":1769435789961,"tcdate":1757817453483,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission4947/Authors"],"signatures":["ICLR.cc/2026/Conference/Submission4947/Authors"],"forum":"HcivRSezJp","license":"CC BY 4.0","number":4947,"cdate":1757817453483,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/-/Submission","ICLR.cc/2026/Conference/-/Post_Submission","ICLR.cc/2026/Conference/Submission4947/-/Full_Submission","ICLR.cc/2026/Conference/Submission4947/-/Rebuttal_Revision","ICLR.cc/2026/Conference/-/Edit","ICLR.cc/2026/Conference/Submission4947/-/Camera_Ready_Revision"],"mdate":1775876960763,"odate":1759896705795,"domain":"ICLR.cc/2026/Conference","id":"HcivRSezJp","version":2},{"content":{"summary":{"value":"This paper presents an approach to improve the explainability of video prediction models. In particular, the authors introduce two metrics, namely temporal importance scores (TIS) and TIS-aware importance map, to measure the importance of certain regions within certain frames with respect to the given video prediction task. Experiments on a synthetic video dataset and a real-world action classification dataset are conducted to verify the effectiveness of the proposed method."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"I wrote the questions within the weakness part."},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":1},"contribution":{"value":2},"strengths":{"value":"- The problem of improving the interpretability of deep neural networks for video prediction is important.\n\n&nbsp;\n\n- Although some technical details are missing, the proposed method is straightforward."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The proposed method seems largely ad-hoc. \n\n  1) In Eq. (3), it is unclear what variable the maximization is over in the notation $\\max(\\theta)$. My understanding is that rather than enumerating all possible window sizes w, the authors use Eq. (3) to determine the proper value of w. However, $\\theta$ is a function of w. If you maximize over w, I do not know what the computational benefit would be. More specifically, can you explain the computational complexity of the computation of TIS score?\n\n  2) In Eq. (6), what is the definition of the lowest baseline bound \\phi_0? It is introduced without proper explanation. \n\n  3) In Eq. (7), what is the \"vecsort” operator? Is it differentiable? If not, how are you going to minimize the loss?\n\n&nbsp;\n\n- The evaluation of the methods for the action classification task with real-world video datasets is not convincing enough. \n  1) It is quite subjective to define the so-called signature frames. Does the dataset provide the signature frames of individual videos, or do you collect them? What are the statistics of such signature frames, e.g., spatial/temporal proportion over all frames? Given a new video prediction task, how do we determine them without or with less human annotation?\n\n  2) The proposed “Temporal Pointing Game” metric does not consider the spatial extent within the signature frames. Even if there is only one pixel unmasked within the signature frames, the metric would be perfect.\n\n&nbsp;\n\n- The writing of the paper is quite confusing.\n\n  1) In the introduction section, the description of the motivating example is confusing. For example, what is exactly the prediction task? Without a proper understanding of the problem setup, it is impossible to understand the meaning of the experimental results. Unfortunately, such detailed descriptions are delayed to the experiment section.\n\n  2) The so-called STEP method is mentioned in the introduction section to motivate the problem. However, it was not introduced until the related work section.\n\n  3) Although interpretability has been mentioned many times, the paper only considers the important spatiotemporal regions as a way of interpretation. I am not sure how much this would add up to the interpretability of the black-box deep models. In other words, what is the exact expectation for improving interpretability? Are the proposed quantified measurements really what people care about when they use video prediction models?"}},"nonreaders":[],"tmdate":1731427863260,"tcdate":1730675089047,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission3509/Reviewer_7DWU"],"signatures":["ICLR.cc/2025/Conference/Submission3509/Reviewer_7DWU"],"forum":"TEjXRrhqtJ","number":3,"license":"CC BY 4.0","cdate":1730675089047,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission3509/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427863260,"domain":"ICLR.cc/2025/Conference","replyto":"TEjXRrhqtJ","id":"OT4VYbRf3A","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"TLDR":{"value":"This paper proposes TIEM, a novel video interpretation method that enhances explainability through dual perturbation that evaluates temporal importance across frames and generates spatio-temporal masks explicitly using this importance."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["XAI","visual explanation","extremal mask","dual perturbation","video prediction"]},"supplementary_material":{"value":"/attachment/4d590237bd717ac596a69e5687de1fab7bf231be.zip"},"primary_area":{"value":"interpretability and explainable AI"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Explaining video data predictions is challenging due to the complex spatio-temporal information in videos. In particular, the existing perturbation-based methods for video interpretation often fail to consider different temporal contexts, making them ineffective for dynamic videos where the important regions change rapidly or appear ephemerally across frames. To address this, we propose a novel video interpretation method, time importance score-aware extremal perturbation masks (TIEM), that enhances explainability by focusing on temporal dynamics in videos. TIEM exploits a dual perturbation process: first, it evaluates temporal importance across frames via temporal perturbation and then generates spatio-temporal extremal perturbation masks using the temporal importance explicitly. Our experimental results demonstrate that TIEM resolves the key challenges of the existing methods, providing more precise explanations across the time domain in synthetic white-box models and black-box models for real-world videos."},"_bibtex":{"value":"@misc{\nje-gal2024tiem,\ntitle={{TIEM}: Enhancing Explanation of Video Prediction via Temporal Dynamics-Focused Dual Perturbation},\nauthor={Hong Je-Gal and Hyun-Suk Lee},\nyear={2024},\nurl={https://openreview.net/forum?id=TEjXRrhqtJ}\n}"},"title":{"value":"TIEM: Enhancing Explanation of Video Prediction via Temporal Dynamics-Focused Dual Perturbation"},"pdf":{"value":"/pdf/ba9ded8ac2b175ad9e9b09d701a83d91e7c69278.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"jegal|tiem_enhancing_explanation_of_video_prediction_via_temporal_dynamicsfocused_dual_perturbation"},"authorids":{"value":["~Hong_Je-Gal1","~Hyun-Suk_Lee1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Hong Je-Gal","Hyun-Suk Lee"]}},"version":2},{"content":{"summary":{"value":"Whereas prior video editing methods often require inefficient video inversion or the design and training of complex control modules, this paper proposes an in-context approach based on a recently introduced DiT-based video generation model. The method concatenates multiple conditional signals and performs full attention over them, enabling a single unified model to handle diverse conditional editing tasks. To further support such a unified model, the authors introduce condition bias and task-aware RoPE, techniques that help distinguish different conditioning signals and improve performance. They also propose a sequential training scheme to make unified training feasible."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"- In L473–474, what does “temporal modeling” refer to? Task-aware RoPE seems to improve reference alignment (as Table 6 suggests), whereas “temporal modeling” sounds more like a measure of video quality, making its use here somewhat unclear.\n- In L102, what is meant by “index collisions”?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"- **Unified model for diverse video editing tasks:** The paper models various video editing tasks under a single formulation and demonstrates that learning a unified video editing model is feasible. The two new techniques introduced specifically for unified training are effective and constitute an additional strength.\n- **Timely extension to video:** As development is progressing from image models toward video models, the proposed work arrives at an opportune time.\n- **Clear and intuitive writing:** The paper is easy to follow, and the motivations behind the methods are clear and intuitive."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- **Slightly limited novelty:** Although the need to distinguish task-specific conditional signals in unified training is well-motivated and the proposed techniques appear appropriate, the overall design feels like a natural extension from a single signal to multiple signals. Moreover, the proposed methods do not seem tightly tailored to video editing, which makes the contribution feel somewhat less novel.\n\n- **Lack of detailed explanation for method and experiments:**\n    The sequential training scheme appears crucial to performance based on the presented results, yet it is not discussed in the method section. In addition, the paper does not specify which TTV base model is used during fine-tuning.    \n    Furthermore, in Table 6, the performance differences between D2–D4 and the baseline (D1) appear marginal depending on the metric, making it difficult to appreciate the improvements. Additional qualitative comparisons would help clarify the effectiveness of the two proposed components."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762922905365,"tcdate":1761746650916,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission11891/Reviewer_4wrd"],"signatures":["ICLR.cc/2026/Conference/Submission11891/Reviewer_4wrd"],"forum":"Vb4nE3WWf5","number":1,"license":"CC BY 4.0","cdate":1761746650916,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission11891/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762922905365,"domain":"ICLR.cc/2026/Conference","replyto":"Vb4nE3WWf5","id":"9BOZUP445l","forumContent":{"TLDR":{"value":"a parameter-efficient and unified framework for video editing tasks"},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["video editing; video generation; diffusion models"]},"primary_area":{"value":"generative models"},"abstract":{"value":"Recent advances in text-to-video generation have sparked interest in generative video editing tasks. Previous methods often rely on task-specific architectures (e.g., additional adapter modules) or dedicated customizations (e.g., DDIM inversion), which limit the integration of versatile editing conditions and the unification of various editing tasks. In this paper, we introduce UNified In-Context Video Editing (UNIC), a simple yet effective framework that unifies diverse video editing tasks within a single model in an in-context manner. To achieve this unification, we represent the inputs of various video editing tasks as three types of tokens: the source video tokens, the noisy video latent, and the multi-modal conditioning tokens that vary according to the specific editing task. Based on this formulation, our key insight is to integrate these three types into a single consecutive token sequence and jointly model them using the native attention operations of DiT, thereby eliminating the need for task-specific adapter designs. Nevertheless, direct task unification under this framework is challenging, leading to severe token collisions and task confusion due to the varying video lengths and diverse condition modalities across tasks. To address these, we introduce task-aware RoPE to facilitate consistent temporal positional encoding, and condition bias that enables the model to clearly differentiate different editing tasks. This allows our approach to adaptively perform different video editing tasks by referring the source video and varying condition tokens \"in context\", and support flexible task composition. To validate our method, we construct a unified video editing benchmark containing six representative video editing tasks. Results demonstrate that our unified approach achieves comparable performance with task specialists and exhibits emergent task composition abilities."},"_bibtex":{"value":"@inproceedings{\nye2026unified,\ntitle={Unified In-Context Video Editing},\nauthor={Zixuan Ye and Xuanhua He and Quande Liu and Qiulin Wang and Xintao Wang and Pengfei Wan and Di ZHANG and Kun Gai and Qifeng Chen and Wenhan Luo},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=Vb4nE3WWf5}\n}"},"title":{"value":"Unified In-Context Video Editing"},"pdf":{"value":"/pdf/b037cc220716c285a2c27112cbc052e5b1e7ef62.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"ye|unified_incontext_video_editing"},"authorids":{"value":["~Zixuan_Ye1","~Xuanhua_He1","~Quande_Liu1","~Qiulin_Wang1","~Xintao_Wang1","~Pengfei_Wan1","~Di_ZHANG3","~Kun_Gai1","~Qifeng_Chen1","~Wenhan_Luo1"]},"authors":{"value":["Zixuan Ye","Xuanhua He","Quande Liu","Qiulin Wang","Xintao Wang","Pengfei Wan","Di ZHANG","Kun Gai","Qifeng Chen","Wenhan Luo"]}},"version":2},{"content":{"summary":{"value":"This paper present a visual-centric video humor understanding benchmark, and give a comprehensive evaluation of many sota MLLMs."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"1. Although existing research points are repetitive, authors can still leverage their existing data for in-depth studies, such as model training, ablation experiments involving more modalities, and thorough analysis of human cognitive patterns regarding humor.\n2. Finally, **what are your thoughts and explanations regarding the overlap with HumorQA in FunQA**, and how do you plan to enhance  and reconstruct the value of your paper?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":1},"strengths":{"value":"1. The paper collects humor videos from both User-Generated Funny Videos and films, introduces new humor video sources.\n2. Compare to existing video humor benchmarks, this paper evaluates a new generation of MLLM models (newer, lager, wider), demonstrating improvements in model capabilities and introducing new challenges in humor comprehension."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. **Novelty**: The humor videoQA data type is already included in MVBench [1], Table 1 Action - Unexpected Action (What unexpected event contributes to the humor in the video?). The authors cited Mvbench in this paper but they didn't aware that their topic is already covered in it.\n2. **Repetitive work**: Upon more precise tracing of the MVbench source, this paper's overall conceptual framework **nearly entirely overlaps with the HumorQA subset within FunQA [2]**.  This raises concerns about duplicate publication.\nAfter a more precise comparison (FunQA vs v-HUB):  \ni) **Tasks**: Counter-intuitive timestamp - Caption matching, Title generation - Caption matching, Counter-intuitiveness reasoning - Humor explanation, FunQA-MCQA & Dialog subset - Open-ended QA.   \nii) **Datasize**: 1,769' avg 7s vs. 960'.   \niii) **Anno**: both annotate caption, description, explanation by human annotators.  \niv) **Common result**: The models heavily relies on text cues and weak visual reasoning in humor comprehension.   \nv) **Eval metrics**: BLEURT, GPT4 vs. BERTScore, METEOR\n\nThis is more like a **coincidental repetition of research topics** (humor videoQA). However, give the existing weakness, even this paper introduces new benchmark and evaluation, **the omissions in its literature review are significant**, leading the authors to overestimate the novelty of their work and limiting the paper's potential for further advancement in MLLM humor comprehension.\n\n[1] Li, K, et al. \"Mvbench: A comprehensive multi-modal video understanding benchmark.\" Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2024.  \n[2] Xie, B., et al. \"Funqa: Towards surprising video comprehension.\" European Conference on Computer Vision. Cham: Springer Nature Switzerland, 2024."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762917564382,"tcdate":1761817287230,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission4776/Reviewer_VXrF"],"signatures":["ICLR.cc/2026/Conference/Submission4776/Reviewer_VXrF"],"forum":"ltUQwfiFiC","number":3,"license":"CC BY 4.0","cdate":1761817287230,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission4776/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762917564382,"domain":"ICLR.cc/2026/Conference","replyto":"ltUQwfiFiC","id":"kPJTPyTt35","forumContent":{"TLDR":{"value":"challenging video LLMs on their capacity for visual-centric humor understanding"},"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["humor understanding","video LLMs"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"AI models capable of comprehending humor hold real-world promise—for example, enhancing engagement in human-machine interactions. To gauge and diagnose the capacity of multimodal large language models (MLLMs) for humor understanding, we introduce v-HUB, a novel visual-centric video humor understanding benchmark. v-HUB comprises a curated collection of minimally verbal short videos, sourced from classic silent films and online resources, and reflecting real-world scenarios where humor can be appreciated purely through visual cues. Each video clip is paired with rich annotations, including captions, descriptions, and explanations, supporting evaluation tasks like caption matching and humor explanation. To broaden its applicability, we further construct an open-ended video QA task, making it readily integrable into existing video understanding benchmarks. We evaluate a diverse set of MLLMs, from specialized Video-LLMs to versatile OmniLLMs that can process audio, covering both open-source and proprietary domains. The experimental results expose the difficulties MLLMs face in comprehending humor from visual cues alone. For example, all models exhibit a marked performance drop on caption matching when moving from text-based to video-based evaluation (without audio). Our findings also demonstrate that incorporating audio helps with video humor understanding, highlighting the informativeness of sound and the\npromise of integrating richer modalities for complex video understanding tasks."},"_bibtex":{"value":"@misc{\nshi2025vhub,\ntitle={v-{HUB}: A Visual-Centric Humor Understanding Benchmark for Video {LLM}s},\nauthor={Zhengpeng Shi and Hengli Li and Yanpeng Zhao and Jianqun Zhou and Yuxuan Wang and Qinrong Cui and Wei Bi and Song-Chun Zhu and Bo Zhao and Zilong Zheng},\nyear={2025},\nurl={https://openreview.net/forum?id=ltUQwfiFiC}\n}"},"title":{"value":"v-HUB: A Visual-Centric Humor Understanding Benchmark for Video LLMs"},"pdf":{"value":"/pdf/5384ea1089ae176962acee35a760e029e65d32dd.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"shi|vhub_a_visualcentric_humor_understanding_benchmark_for_video_llms"},"authorids":{"value":["~Zhengpeng_Shi2","~Hengli_Li1","~Yanpeng_Zhao1","~Jianqun_Zhou1","~Yuxuan_Wang4","~Qinrong_Cui1","~Wei_Bi1","~Song-Chun_Zhu1","~Bo_Zhao4","~Zilong_Zheng1"]},"authors":{"value":["Zhengpeng Shi","Hengli Li","Yanpeng Zhao","Jianqun Zhou","Yuxuan Wang","Qinrong Cui","Wei Bi","Song-Chun Zhu","Bo Zhao","Zilong Zheng"]}},"version":2},{"content":{"summary":{"value":"In order to address the current limitation of relying on linguistic reasoning of videos, authors propose: (1) 2-staged training method with 2 datasets. Training method performs a cold start with 3 tasks, and GRPO-based training with 2 tasks (2) an inference method aware of key elements and events by performing both dense and sparse sampling. This enables visual reasoning and authors demonstrated the effectiveness with 7 benchmarks."},"soundness":{"value":2},"confidence":{"value":5},"questions":{"value":"(major)\n1. Notation\n- The model is not generating the instruction tokens. So shouldn’t it be $X_{instruct}$ rather than $X_{instruct, <i}$ in Eq. 4?\n2. Sec 3.2\n- How does using trainable parameters <|event_start|> and <|event_end|> help to generate stable predictions? Explanation of the effects with trainable parameters and justification is required.\n- How big is the performance gap between adopting pre-trained models and post-trained SFT models?\n- Exactly which bias the post-trained SFT model possesses? I believe there is a way to estimate this bias. For example, if the authors are indicating the length of LLM responses, numbers of generated tokens or words can be compared to verify this.\n3. Experiments and analysis\n- Are there additional impacts of the proposed training and inference on long videos? In other words, any more findings specifically on LongVB, LVBench?\n- How can the authors argue that the proposed method is “stable” and “efficient” other than the number of the trainset? I believe a graph of the training curve [3] or the mean and standard deviation across multiple runs [4] should be reported for stability. This paper lacks validation of training efficiency. Also, FLOPs or latency comparison for regular inference vs proposed inference looks necessary. The proposed inference requires additional forwards of external modules (e.g., text encoder, visual encoders, additional generation calls), it does not look efficient.\n(minor)\n- Is Eq. 5 necessary in this presentation even though we have Eq. 2 and 3?\n\nReferences\n\n[1] Liang et al. Exploring format consistency for instruction tuning. TMLR 2024.\n\n[2] Li et al. Vidhalluc: evaluating temporal hallucinations in multimodal large language models for video understanding. CVPR 2025.\n\n[3] Lin et al. Rho-1: not all tokens are what you need. NeurIPS 2024.\n\n[4] Leng et al. Mitigating object hallucinations in large vision-language models through visual contrastive decoding. CVPR 2024."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":1},"contribution":{"value":2},"strengths":{"value":"1. Conceptual contribution: The main motivation of the paper of mitigating the sole reliance of linguistic reasoning is very solid. I also believe that relying on one modality is not a good choice, the community should focus on developing reasoning on both modalities.\n2. The proposed methods achieve benchmark improvements using only 8k training samples, where existing methods require a larger scale of dataset (165k, 260k pairs for Video-R1)."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Writing and readability\n- The logic flow is unclear, especially in the intro. e.g., it is unclear what “However, …” in line 74 wants to elaborate, referring to a Table located in the appendix in the introduction distracts the readability (line 83).\n- A brief explanation of Cold Start would enhance the readability.\n- Multiple notation inconsistency. For example, “MLLM” is declared more than one time and usage is inconsistent (line 243).\n- Section 3.1 looks redundant. It only has 3 sentences, and Eq.1 and 5 are identical, only substituted $q$ by $m$. Moving Section 3.1 to appendix and including more details on methods and experimental results would improve the paper.\n2. Data curation\n- Curating a dataset with a unified format is concerning. [1]. Also, why should we use specific prefixes for each task?\n- I am not against using proprietary models for automated annotations, but I believe there should be some careful filtering or verification process for key element labeling in the main script [2].\n- Adding actual system prompt and query at Figure 1 will be much more helpful for readers to understand the tasks.\n- Would be better if there’s a figure for data curation of 5k video pairs. Also, descriptions on how to generate the data are insufficient.\n3. GRPO design\n- More details of IoU and Format Reward should be elaborated. Is FormatReward = 1 if the model includes both \\<event_start\\> and \\<event_end\\> and 0 otherwise?\n- Do we need any normalizations for IoU rewards? Don’t we need to balance three rewards with hyperparameters?\n4. Experiments and analyses\n- No qualitative results are reported.\n- The experiments are only performed with Qwen2-VL-7B. We are curious how well it performs on other base MLLMs as well.\n\nI will increase the score once current concerns are addressed."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762915861307,"tcdate":1760765115421,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission1701/Reviewer_H3Vp"],"signatures":["ICLR.cc/2026/Conference/Submission1701/Reviewer_H3Vp"],"forum":"b3mk8XrH9N","number":1,"license":"CC BY 4.0","cdate":1760765115421,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission1701/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762915861307,"domain":"ICLR.cc/2026/Conference","replyto":"b3mk8XrH9N","id":"PxY8xsNcT8","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Multimodal large language model","video understanding"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Multimodal Large Language Models (MLLMs) have made great progress in video understanding tasks. However, when it comes to understanding complex or lengthy videos, MLLMs tend to overlook details or produce hallucinations. To alleviate these issues, recent work has attempted to leverage reinforcement learning (RL) to boost models' deep linguistic reasoning of complex videos. But these methods have two main problems: First, the RL framework they used has unstable training, high training costs, and is difficult to train satisfactory video reasoning models; Second, the linguistic reasoning process is difficult to guarantee the reliability of visual information. To alleviate these problems, we propose to use multimodal elements for reasoning, and we design a novel framework to build and enhance versatile video reasoning capabilities on MLLMs. We carefully design a multi-task cold start and multi-task reinforcement learning to improve the model's visual perception and proficiency in multiple capabilities. In the inference phase, we leverage multimodal reasoning and dynamic sampling to further improve the performance. We verified the efficiency of the framework on a base MLLM (Qwen2-VL-7B-Base). Through cold-start with 3k data and reinforcement learning training with 5k data, combined with inference design, our final model significantly outperforms the base model on seven public video benchmarks, even surpassing and approaching the state-of-the-art Instruct Models such as Qwen2.5-VL-7B-Instruct."},"_bibtex":{"value":"@misc{\nwang2025reinforcement,\ntitle={Reinforcement Learning for Versatile Video Reasoning Capabilities in Base Multimodal {LLM}s},\nauthor={Xiaodong Wang and Zhirong Wu and Langling Huang and Yuxi Zheng and Jinfa Huang and Peixi Peng},\nyear={2025},\nurl={https://openreview.net/forum?id=b3mk8XrH9N}\n}"},"title":{"value":"Reinforcement Learning for Versatile Video Reasoning Capabilities in Base Multimodal LLMs"},"pdf":{"value":"/pdf/7ad844176f120f86c26e7109a86cba19fc6f5b46.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"wang|reinforcement_learning_for_versatile_video_reasoning_capabilities_in_base_multimodal_llms"},"authorids":{"value":["~Xiaodong_Wang6","~Zhirong_Wu5","~Langling_Huang1","~Yuxi_Zheng3","~Jinfa_Huang2","~Peixi_Peng2"]},"authors":{"value":["Xiaodong Wang","Zhirong Wu","Langling Huang","Yuxi Zheng","Jinfa Huang","Peixi Peng"]}},"version":2},{"content":{"venue":{"value":"ISCC 2009"},"pdf":{"value":"https://ieeexplore.ieee.org/iel5/5179146/5202209/05202393.pdf"},"venueid":{"value":"dblp.org/conf/ISCC/2009"},"paperhash":{"value":"kondrad|media_aware_fec_for_ccalable_video_coding_transmission"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Lukasz_Kondrad:","https://dblp.org/search/pid/api?q=author:Imed_Bouazizi:","~Moncef_Gabbouj1"]},"html":{"value":"https://doi.org/10.1109/ISCC.2009.5202393"},"_bibtex":{"value":"@inproceedings{DBLP:conf/iscc/KondradBG09,\n  author={Lukasz Kondrad and Imed Bouazizi and Moncef Gabbouj},\n  title={Media aware FEC for Ccalable Video Coding transmission},\n  year={2009},\n  cdate={1230768000000},\n  pages={7-12},\n  url={https://doi.org/10.1109/ISCC.2009.5202393},\n  booktitle={ISCC},\n  crossref={conf/iscc/2009}\n}\n"},"abstract":{"value":"Scalable video coding (SVC) is an extension to H.264/AVC that enables encoding a video sequence once and using it subsequently at multiple operations points by different devices and applications. Forward Error Correction techniques may then be applied to the encoded video in order to enhance the robustness against transmission errors. Media aware FEC code construction caters for scalable video streams by adjusting the protection level for each video layer. In this paper, we discuss different approaches for media aware Forward Error Correction using Scalable Video Coded video that is transmitted over error prone channels. We describe two recent solutions for Unequal Error Protection and then evaluate those through multiple simulations."},"title":{"value":"Media aware FEC for Ccalable Video Coding transmission"},"authors":{"value":["Lukasz Kondrad","Imed Bouazizi","Moncef Gabbouj"]}},"tmdate":1741252405290,"pdate":1230768000000,"tcdate":1741251983239,"writers":["~"],"signatures":["~Moncef_Gabbouj1"],"forum":"FrrbkevwZV","license":"CC BY-SA 4.0","number":361012,"cdate":1230768000000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1741252405290,"domain":"DBLP.org","id":"FrrbkevwZV","version":2},{"content":{"venue":{"value":"ISCC 2009"},"pdf":{"value":"https://ieeexplore.ieee.org/iel5/5179146/5202209/05202393.pdf"},"venueid":{"value":"dblp.org/conf/ISCC/2009"},"paperhash":{"value":"kondrad|media_aware_fec_for_ccalable_video_coding_transmission"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Lukasz_Kondrad:","https://dblp.org/search/pid/api?q=author:Imed_Bouazizi:","~Moncef_Gabbouj1"]},"html":{"value":"https://doi.org/10.1109/ISCC.2009.5202393"},"_bibtex":{"value":"@inproceedings{DBLP:conf/iscc/KondradBG09,\n  author={Lukasz Kondrad and Imed Bouazizi and Moncef Gabbouj},\n  title={Media aware FEC for Ccalable Video Coding transmission},\n  year={2009},\n  cdate={1230768000000},\n  pages={7-12},\n  url={https://doi.org/10.1109/ISCC.2009.5202393},\n  booktitle={ISCC},\n  crossref={conf/iscc/2009}\n}\n"},"abstract":{"value":"Scalable video coding (SVC) is an extension to H.264/AVC that enables encoding a video sequence once and using it subsequently at multiple operations points by different devices and applications. Forward Error Correction techniques may then be applied to the encoded video in order to enhance the robustness against transmission errors. Media aware FEC code construction caters for scalable video streams by adjusting the protection level for each video layer. In this paper, we discuss different approaches for media aware Forward Error Correction using Scalable Video Coded video that is transmitted over error prone channels. We describe two recent solutions for Unequal Error Protection and then evaluate those through multiple simulations."},"title":{"value":"Media aware FEC for Ccalable Video Coding transmission"},"authors":{"value":["Lukasz Kondrad","Imed Bouazizi","Moncef Gabbouj"]}},"tmdate":1741252254327,"pdate":1230768000000,"tcdate":1741251884070,"writers":["~"],"signatures":["~Moncef_Gabbouj1"],"forum":"kmi0SxazvU","license":"CC BY-SA 4.0","number":360785,"cdate":1230768000000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1741252254327,"domain":"DBLP.org","id":"kmi0SxazvU","version":2},{"content":{"venue":{"value":"ISCC 2009"},"pdf":{"value":"https://ieeexplore.ieee.org/iel5/5179146/5202209/05202393.pdf"},"venueid":{"value":"dblp.org/conf/ISCC/2009"},"paperhash":{"value":"kondrad|media_aware_fec_for_ccalable_video_coding_transmission"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Lukasz_Kondrad:","https://dblp.org/search/pid/api?q=author:Imed_Bouazizi:","~Moncef_Gabbouj1"]},"html":{"value":"https://doi.org/10.1109/ISCC.2009.5202393"},"_bibtex":{"value":"@inproceedings{DBLP:conf/iscc/KondradBG09,\n  author={Lukasz Kondrad and Imed Bouazizi and Moncef Gabbouj},\n  title={Media aware FEC for Ccalable Video Coding transmission},\n  year={2009},\n  cdate={1230768000000},\n  pages={7-12},\n  url={https://doi.org/10.1109/ISCC.2009.5202393},\n  booktitle={ISCC},\n  crossref={conf/iscc/2009}\n}\n"},"abstract":{"value":"Scalable video coding (SVC) is an extension to H.264/AVC that enables encoding a video sequence once and using it subsequently at multiple operations points by different devices and applications. Forward Error Correction techniques may then be applied to the encoded video in order to enhance the robustness against transmission errors. Media aware FEC code construction caters for scalable video streams by adjusting the protection level for each video layer. In this paper, we discuss different approaches for media aware Forward Error Correction using Scalable Video Coded video that is transmitted over error prone channels. We describe two recent solutions for Unequal Error Protection and then evaluate those through multiple simulations."},"title":{"value":"Media aware FEC for Ccalable Video Coding transmission"},"authors":{"value":["Lukasz Kondrad","Imed Bouazizi","Moncef Gabbouj"]}},"tmdate":1741252061201,"pdate":1230768000000,"tcdate":1741251743379,"writers":["~"],"signatures":["~Moncef_Gabbouj1"],"forum":"oO3ObHjR64","license":"CC BY-SA 4.0","number":360462,"cdate":1230768000000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1741252061201,"domain":"DBLP.org","id":"oO3ObHjR64","version":2},{"content":{"summary":{"value":"The submission is motivated by the looming threat of adversarial image patterns for mission-critical applications using vision language models, or specifically, video large language models. The authors focus specifically on adversarial object hallucination (AOH) since it may affect downstream applications, leveraging known feature priors from  intermediate feature representations of the multi-modal fusion blocks in Vid-LLMs. The authors first collect a multi-source benchmark dataset that aims to represent (clean, target, target mask) triplets without having to add target objects manually. The authors propose to leverage existing datasets, including ROVI, Video-Sham, and HQVI, alongside a custom driving scenario dataset. For ROVI, the authors propose to use reverse inpainting to remove subjects from the video samples, effectively making them \"clean\" videos. These are later manually evaluated to filter for visual artifacts and realism. The target videos are fed into an annotation pipeline to generate text descriptions of the targets, create VQA pairs with existing LLMs, and validated against a handful of frames with ChatGPT. The threat model can be summarized as white box with bounded $l_\\inf$-norm such that the adversary may observe the target Vid-LLM. The authors propose to use PGD to optimize a video noise mask, such that the perturbed video semantic features are closer to those of the target video in the Vid-LLM connector. The attack is tested on a variety of Vid-LLM models from different architectures and model families, including InternVL and LLaVA, alongside ablations to check performance for random noise, mask-assisted random-noise (only noise in the target spatial dimensions). The authors analyze the attack success rate for each, check the cross-scale transferability, and check the high-level receptive field with GradCAM."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"- I would be interested to see the accuracy and scoring results split among dataset type, since the current results seem to take the aggregate of all video types (non-driving video + driving video).\n- Since the authors propose a new video dataset, they should carefully consider the distribution rights associated with the videos and clarify if they will be shared or not. If the dataset is distributed to the wider community, it should be properly annotated to clarify that some annotations were machine-generated."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":2},"strengths":{"value":"- The writing quality is excellent and conveys the technical points well. The notation appears to be free of major issues and is consistent with familiar convention in the field. Figures and tables are rendered properly and there were no issues viewing them at standard zoom levels.\n- The submission investigates the important concern of adversarial samples in Video-Language models, particularly the ability for those patterns to force hallucination via the high-level semantic alignment of a noise mask with a target video.\n- The authors highlight an interesting artifact of the AOH attack, which is the ability to generate successful attacks using smaller Vid-LLMs for larger models.\n- Through GradCAM visualizations, it is shown that the Vid-LLM attention may not necessarily be focused on the region of forced hallucination. Hence it may be that the semantic features (textures or shapes) of the target are embedded in other spatial regions of the frames."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- For some results it is difficult to check the statistical significance. In Table 1 for example, I suspect the false negatives/positives are highest for the smaller models, so the proposed attack accuracy may be within margin of error from the clean accuracy. The results for 7B and 8B models are more encouraging, although without an explicit measure such as standard deviation, it is difficult to make claims on models, for example InternVL2.5 which had 0.381 clean vs. 0.465 AOH. The same issue applies for the results in Section 5.3.\n- The dataset size is described in terms of amount of clean/target video pairs, but it can be difficult to gauge the scale since video datasets may have diverse clip lengths and resolutions. On L189 the authors should clarify the total runtime (minutes/hours) of the dataset and summarize their resolution, it appears to always be 480p but it wasn't clear if this was only HQVI. For VQA annotation, it should be clarified how much text was generated in number of tokens.\n- Overall the submission offers some interesting first looks at a few different aspects of video hallucination (success rate, receptive field, and transferability) but each aspect is not allocated much depth to have definitive takeaways. Specifically:\n    - The attack success rates show AOH has potential, but the analysis currently lacks statistical rigor. It would also help if the authors offered preliminary results for transferability to black-box systems, since that is a natural consequence of the use-cases described in the introduction, and preliminary results offer a first look that may not be available yet in the broader field.\n    - There is evidence of cross-scale transferability using smaller models, but ultimately it is still not well understood why the transferability occurs, so it is difficult for any action to be taken. The authors should investigate this more deeply since it seems the most interesting thread for a contribution.\n    - The Vid-LLM high-level receptive field does not necessarily attend to the hallucinated object, but it is also not well explored beyond the surface level evidence. I would expect those high-level semantic patterns to be embedded throughout the spatial dimensions since the loss gradient does not force those concepts to appear in the target region. Some deeper investigation of this behavior would strengthen the contribution. If the authors choose to stay with the white-box setting, then going deeper in this direction would probably offer more takeaways to the community."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762931565030,"tcdate":1761784117155,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission19729/Reviewer_UR3H"],"signatures":["ICLR.cc/2026/Conference/Submission19729/Reviewer_UR3H"],"forum":"cdhk58Z761","number":3,"license":"CC BY 4.0","cdate":1761784117155,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission19729/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762931565030,"domain":"ICLR.cc/2026/Conference","replyto":"cdhk58Z761","id":"Bx0vl2lZZ9","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Video Large Language Models","Adversarial Attack","Object Hallucination"]},"primary_area":{"value":"foundation or frontier models, including LLMs"},"abstract":{"value":"Video Large Language Models (Vid-LLMs) have rapidly advanced video understanding, yet their robustness against semantic adversarial manipulation, especially object hallucination, remains largely unexplored. We introduce Adversarial Object Hallucination (AOH), a novel attack that compels Vid-LLMs to ``see\" non-existent objects in videos by injecting visually imperceptible perturbations. Unlike prior attacks limited to inputs or outputs of videos, AOH directly manipulates intermediate connector features, aligning them with representations from a target video to induce controllable hallucinations. To systematically assess this threat, we curate a benchmark of 535 clean/target video pairs with high-quality VQA annotations. Extensive experiments show that AOH poses a severe threat to state-of-the-art Vid-LLMs, achieving highly effective attacks with alarming \\emph{cross-scale transferability}: adversarial examples optimized on smaller models transfer even more strongly to larger counterparts of the same architecture, amplifying attack impact while reducing adversarial cost. Further analyses reveal that perturbations encode semantic object contours, while Grad-CAM highlights their covert influence. These findings expose a severe and previously overlooked vulnerability in Vid-LLMs, raising urgent concerns about their secure deployment and providing a foundation for future adversarial research in video-language modeling."},"_bibtex":{"value":"@misc{\nzhang2025adversarial,\ntitle={Adversarial Object Hallucination Attacks in Video-Language Models via Intermediate Feature Alignment},\nauthor={Lu Zhang and Liang Zeng},\nyear={2025},\nurl={https://openreview.net/forum?id=cdhk58Z761}\n}"},"title":{"value":"Adversarial Object Hallucination Attacks in Video-Language Models via Intermediate Feature Alignment"},"pdf":{"value":"/pdf/504513cce9144d84fafcf5759b866e7e55223945.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"zhang|adversarial_object_hallucination_attacks_in_videolanguage_models_via_intermediate_feature_alignment"},"authorids":{"value":["~Lu_Zhang19","~Liang_Zeng1"]},"authors":{"value":["Lu Zhang","Liang Zeng"]}},"version":2},{"content":{"summary":{"value":"This paper proposes an efficient text-based video-editing method by adapting Instruct Pix2Pix from image to video editing, eliminating the need for additional training. To enable long term video editing, a Long Video Sampling Strategy is proposed to maintain long video consistency. Experimental comparisons with other methods demonstrate the advantages of the proposed approach."},"presentation":{"value":"2 fair"},"contribution":{"value":"2 fair"},"soundness":{"value":"3 good"},"strengths":{"value":"1. The paper is well-written and easy to follow\n2. The paper proposed a very intersting idea for universal one-model-all-video transfer idea for vid-to-vid transfer\n3. It proposed a novel synthetic dataset fo vid-to-vid transfer task.\n4. The experiments were well conducted, showcasing detailed numerical indicators and a user study to evaluate the video editing method."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The proposed method in this paper lacks significant innovation. The majority of the content is derived from Instruct Pix2Pix by adapting image editing to video editing without much improvement. \n2. The proposed sampling method to maintain long video consistency is a variation of inpainting sampling methods, which is also widely used in image/video generation tasks. The experimental section provides detailed numerical metrics for various evaluation indicators and user study. However, the compared baseline lacks strength in video editing tasks, there already are some video editing models based on image diffusion models, such as Pix2Video[1],  Render A Video [2], TokenFlow[3], which also do not require fine-tuning on a single video. The proposed method in this paper did not compare itself with those models mentioned. \n3. From the generated video results in the provided supplementary, it seems that the proposed method does not achieve superior results.\n\n [1] Pix2Video: Video Editing using Image Diffusion\n\n [2] Rerender A Video: Zero-Shot Text-Guided Video-to-Video Translation\n\n [3] TokenFlow: Consistent Diffusion Features for Consistent Video Editing"},"confidence":{"value":"4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work."},"questions":{"value":"1. In Section 3.1, the authors mentioned that the temporal attention layers are also replaced to adapt to video editing tasks. However, in the showcased results, most videos are style transferred frame by frame. Is there any attempt to simultaneously change the style of the video and modify the motion of the video, such as transforming a person walking towards the left to walking towards the right?\n2. In Section 3.2, how is the success rate calculated? Are there any other methods to improve the success rate?\n3. Does using video diffusion model based methods have any advantages over using optical flow in image diffusion models in video editing tasks?"},"rating":{"value":"5: marginally below the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636013654,"tcdate":1698861795440,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission873/Reviewer_99Pg"],"signatures":["ICLR.cc/2024/Conference/Submission873/Reviewer_99Pg"],"forum":"IoKRezZMxF","number":3,"license":"CC BY 4.0","cdate":1698861795440,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission873/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636013654,"domain":"ICLR.cc/2024/Conference","replyto":"IoKRezZMxF","id":"YrDOuQb3W7","forumContent":{"venue":{"value":"ICLR 2024 poster"},"TLDR":{"value":"We've developed a synthetic dataset to train a text-based video editing model, eliminating the need for per-video fine-tuning, and introduced a method for seamless long video editing."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Computer Vision","Video Editing","Diffusion Model"]},"supplementary_material":{"value":"/attachment/9cab55bbe2e33a78087b54a1213ffdd7a40cefee.zip"},"primary_area":{"value":"generative models"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"We introduce a novel and efficient approach for text-based video-to-video editing that eliminates the need for resource-intensive per-video-per-model finetuning. At the core of our approach is a synthetic paired video dataset tailored for video-to-video transfer tasks. Inspired by Instruct Pix2Pix's image transfer via editing instruction, we adapt this paradigm to the video domain. Extending the Prompt-to-Prompt to videos, we efficiently generate paired samples, each with an input video and its edited counterpart. Alongside this, we introduce the Long Video Sampling Correction during sampling, ensuring consistent long videos across batches. Our method surpasses current methods like Tune-A-Video, heralding substantial progress in text-based video-to-video editing and suggesting exciting avenues for further exploration and deployment."},"_bibtex":{"value":"@inproceedings{\ncheng2024consistent,\ntitle={Consistent Video-to-Video Transfer Using Synthetic Dataset},\nauthor={Jiaxin Cheng and Tianjun Xiao and Tong He},\nbooktitle={The Twelfth International Conference on Learning Representations},\nyear={2024},\nurl={https://openreview.net/forum?id=IoKRezZMxF}\n}"},"title":{"value":"Consistent Video-to-Video Transfer Using Synthetic Dataset"},"pdf":{"value":"/pdf/e91267e5675cf150b685ada8cd0303644c45a25f.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference"},"paperhash":{"value":"cheng|consistent_videotovideo_transfer_using_synthetic_dataset"},"authorids":{"value":["~Jiaxin_Cheng1","~Tianjun_Xiao1","~Tong_He5"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Jiaxin Cheng","Tianjun Xiao","Tong He"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a straightforward yet effective method for adapting a pre-trained video diffusion model into an action-conditioned world model. The core contributions are two-fold: 1) \"Video Diffusion Causalization,\" which enforces temporal causality through a causal masked attention layer and a temporal convolution layer with a novel \"extrapolative weight transfer\" technique, and 2) \"Causal Action Injection,\" which integrates action-conditioning via causal guidance with dropouts. The authors conduct extensive experiments across diverse domains, including robotic manipulation (RT-1), game simulation (CS:GO), and real-world navigation (NWM), demonstrating that their proposed model achieves higher fidelity in video prediction compared to existing world models."},"soundness":{"value":4},"confidence":{"value":4},"questions":{"value":"- Figure 6: This figure presents a crucial comparison of auto-regressive generation capabilities. However, the motivation for why this specific comparison is necessary and what insights it provides is detailed in Appendix C.6 but absent from the main text. Briefly explaining this context in the main paper would significantly improve the figure's clarity and impact."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"- Simplicity and Effectiveness: The primary strength of this work lies in its simple and elegant approach. The proposed modifications to transform a standard video diffusion model into a world model are minimal, yet they prove to be highly effective across multiple challenging domains.\n- Broad Applicability: The method's effectiveness is convincingly demonstrated on a variety of environments, from simulated robotics and games to real-world navigation. This suggests the proposed techniques are general and widely applicable.\n- High-Fidelity Generation: The paper provides strong quantitative and qualitative results, showing that the model can generate high-fidelity, fine-grained, and action-conditioned video sequences that are visually superior to prior work."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Insufficient Motivation for Extrapolative Weight Transfer: While the paper introduces \"extrapolative weight transfer\" as a key component, its theoretical and empirical motivation remains underdeveloped. The appendix (A.1) presents a simple counter-example involving a linearly increasing kernel, but this theoretical argument is not supported by empirical evidence. To strengthen this claim, the authors should provide experiments that verify the existence and prevalence of such temporal patterns in the learned kernels of baseline models, thereby justifying the need for their proposed solution.\n\n- Lack of Experiments Validating Utility as a World Model: The paper claims to build a \"world model,\" but the experiments primarily focus on video prediction fidelity (e.g., FVD, FID, SSIM). A true world model should be useful for downstream tasks like planning and control. The evaluation is missing crucial experiments, such as: (1) Reinforcement Learning Tasks: Can an RL agent be successfully trained using the learned model as the environment simulator? Demonstrating successful policy learning would provide compelling evidence of the model's utility and temporal consistency. If this is not feasible, a discussion of the model's limitations for such tasks is warranted. (2) Quantitative Interactivity Metrics: The current evaluation lacks metrics that specifically measure the model's fine-grained controllability and causal soundness in response to actions.\n\n- Ambiguity in the Role of Large-Scale Pre-training: The experimental setup creates a potential ambiguity. The proposed model leverages an internet-scale pre-trained video model, whereas the robotics and gaming experiments are conducted on domain-specific datasets. It is unclear which components of the model's success are attributable to the novel architectural changes versus the powerful, pre-trained backbone. An ablation study clarifying the benefits of internet-scale pre-training for domain-specific tasks like RT-1 or CS:GO would be highly valuable.\n\n- Incomplete Reporting of Metrics: For the real-world navigation (NWM) experiments, key fidelity metrics such as FVD, FID, and SSIM appear to be missing from the results tables. Including these would allow for a more direct and fair comparison with other models and experiments within the paper."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762917616235,"tcdate":1761886183963,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission4851/Reviewer_Yc6M"],"signatures":["ICLR.cc/2026/Conference/Submission4851/Reviewer_Yc6M"],"forum":"pFyzqbUiF9","number":1,"license":"CC BY 4.0","cdate":1761886183963,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission4851/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762917616235,"domain":"ICLR.cc/2026/Conference","replyto":"pFyzqbUiF9","id":"rWI5oAFcjH","forumContent":{"TLDR":{"value":"We propose Vid2World,  a general approach for leveraging and transferring pre-trained video diffusion models into interactive world models."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["World Models","Video Diffusion Models"]},"primary_area":{"value":"foundation or frontier models, including LLMs"},"abstract":{"value":"World models, which predict future transitions from past observation and action sequences, have shown great promise for improving data efficiency in sequential decision-making. However, existing world models often require extensive domain-specific training and still produce low-fidelity, coarse predictions, limiting their usefulness in complex environments. In contrast, video diffusion models trained on large-scale internet data have demonstrated impressive capabilities in generating high-quality videos that capture diverse real-world dynamics. In this work, we present _Vid2World_, a general approach for leveraging and transferring pre-trained video diffusion models into interactive world models. To bridge the gap, Vid2World systematically explores _video diffusion causalization_, reshaping both the architecture and training objective of pre-trained models to enable autoregressive generation. Additionally, it incorporates a _causal action guidance_ mechanism to enhance action controllability in the resulting interactive world models. Extensive experiments across multiple domains, including robot manipulation, 3D game simulation, and open-world navigation, demonstrate that our method offers a scalable and effective pathway for repurposing highly capable video diffusion models into interactive world models."},"_bibtex":{"value":"@inproceedings{\nhuang2026vidworld,\ntitle={Vid2World: Crafting Video Diffusion Models to Interactive World Models},\nauthor={Siqiao Huang and Jialong Wu and Qixing Zhou and Shangchen Miao and Mingsheng Long},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=pFyzqbUiF9}\n}"},"title":{"value":"Vid2World: Crafting Video Diffusion Models to Interactive World Models"},"pdf":{"value":"/pdf/bb7427c741f3d9f284a8429b4fc25037b01b4e8a.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"huang|vid2world_crafting_video_diffusion_models_to_interactive_world_models"},"authorids":{"value":["~Siqiao_Huang1","~Jialong_Wu1","~Qixing_Zhou1","~Shangchen_Miao1","~Mingsheng_Long5"]},"authors":{"value":["Siqiao Huang","Jialong Wu","Qixing Zhou","Shangchen Miao","Mingsheng Long"]}},"version":2},{"content":{"summary":{"value":"It presents the DTGVD model for video dialog response generation. It combines the strengths of video content understanding and conversation history analysis through dual temporal grounding. The model predicts relevant temporal regions in videos and dialogs, filters content, and uses contrastive learning to align video and dialog dynamics. Experiments on AVSD@DSTC datasets show improved performance."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. Since the information provided by the dialog history is from the video, why not directly find the important frames in videos and learn QA from the video and question? The dialog has noises, so why the dialog history is needed to solve this QA?\n2. In Figure 2, in the \"contrastive-enhanced answer generation\" module, why does the \"exclude\" branch link to the generation decoder?\n3. Since you select the top k relevant question-answer pair to help answer the final question. Why do you select top-k from the IoU other than directly searching the relevant pair from the question semantics?"},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"1. It introduces an innovative approach with the Dual Temporal Grounding-enhanced Video Dialog (DTGVD) model, which leverages dual temporal dynamics inherent in both video sequences and dialog histories.\n2. This model employs a temporal grounding module to explicitly model the attention shift of each dialog turn over the video, generating temporal masks to filter out irrelevant video frames and dialog history."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The motivation is not convincing enough. Why the temporal information in the dialog history is needed? Existing works can also model the video temporal structure and answer the question.\n\n2. the baseline models are too old and they are not the latest SOTA. There are some new works for example:\n[1] HEAR: Hearing Enhanced Audio Response for Video-grounded Dialogue\n[2] M2K-VDG: Model-Adaptive Multimodal Knowledge Anchor Enhanced Video-grounded Dialogue Generation\n\nThe grounded QA methods should also be compared. Since it solves the same questions essentially, like works\n[3] Can I Trust Your Answer? Visually Grounded Video Question Answering"}},"nonreaders":[],"tmdate":1731427477929,"tcdate":1730705631125,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission1727/Reviewer_npFV"],"signatures":["ICLR.cc/2025/Conference/Submission1727/Reviewer_npFV"],"forum":"rtUjj03qZv","number":3,"license":"CC BY 4.0","cdate":1730705631125,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission1727/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427477929,"domain":"ICLR.cc/2025/Conference","replyto":"rtUjj03qZv","id":"bNj3AhM7Lz","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["video dialog; multi-modal understanding; video grounding"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"In the realm of video dialog response generation, the understanding of video content and the temporal nuances of conversation history are paramount. While a segment of current research leans heavily on large-scale pretrained visual-language models and often overlooks temporal dynamics, another delves deep into spatial-temporal relationships within videos but demands intricate object trajectory pre-extractions and sidelines dialog temporal dynamics. \nThis paper introduces the Dual Temporal Grounding-enhanced Video Dialog model (DTGVD), strategically designed to merge the strengths of both dominant approaches.\nIt emphasizes dual temporal relationships by predicting dialog turn-specific temporal regions, filtering video content accordingly, and grounding responses in both video and dialog contexts. \nOne standout feature of DTGVD is its heightened attention to chronological interplay. By recognizing and acting upon the dependencies between different dialog turns, it captures more nuanced conversational dynamics. \nTo further bolster the alignment between video and dialog temporal dynamics, we've implemented a list-wise contrastive learning strategy. Within this framework, accurately grounded turn-clip pairings are designated as positive samples, while less precise pairings are categorized as negative. This refined classification is then funneled into our holistic end-to-end response generation mechanism. Evaluations using AVSD@DSTC-7 and AVSD@DSTC-8 datasets underscore the superiority of our methodology."},"_bibtex":{"value":"@misc{\nqin2024grounding,\ntitle={Grounding is All You Need? Dual Temporal Grounding for Video Dialog},\nauthor={You Qin and Wei Ji and Xinze Lan and Hao Fei and Xun Yang and Dan Guo and Roger Zimmermann and Lizi Liao},\nyear={2024},\nurl={https://openreview.net/forum?id=rtUjj03qZv}\n}"},"title":{"value":"Grounding is All You Need? Dual Temporal Grounding for Video Dialog"},"pdf":{"value":"/pdf/47362fac4c1848def6ca94f44620518960f2d0d5.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"qin|grounding_is_all_you_need_dual_temporal_grounding_for_video_dialog"},"authorids":{"value":["~You_Qin2","~Wei_Ji1","~Xinze_Lan1","~Hao_Fei1","~Xun_Yang1","~Dan_Guo1","~Roger_Zimmermann1","~Lizi_Liao1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["You Qin","Wei Ji","Xinze Lan","Hao Fei","Xun Yang","Dan Guo","Roger Zimmermann","Lizi Liao"]}},"version":2},{"content":{"summary":{"value":"The authors propose a framework to speed up video generation using Video Diffusion Transformers by optimizing attention computation and reducing sampling steps. A repetitive attention tile pattern in 3D attention maps is identified which allows for sparse attention that lowers complexity. The framework uses a three-stage training pipeline: multi-step consistency distillation to reduce sampling steps, a layer-wise search for optimal sparse attention masks, and knowledge distillation to retain performance. This approach claims to achieve up to a 7.8× speedup in video generation with minimal quality loss."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"What happens if the stages are switched, i.e., first obtain T_{sparse}, then T_{MCM}​ from T_{sparse}, and finally apply the knowledge distillation step?\n\nTable 4 needs additional quantitative metrics like aesthetic quality, subject consistency, imaging quality, and FVD to provide a complete understanding of the effect of parallelization.\n\nWhen comparing speed-up performance for parallelization, are the baseline models also trained with parallelization (Table 4)?\n\nHow does the proposed model achieve a lower FVD (Table 5) than the main base model, given that the proposed model is ultimately a distilled version of the main model?\n\nHow is the claim (lines 424 to 430) that model performance is within 1% of the base model accurate? It is evident that the numbers for imaging quality and subject class are significantly lower than those of the base model.\n\nAblation studies in Table 6 show that only MLCD can speed up the process by 5 to 8 times compared to the base model without significantly compromising quality. What is the justification, then, for the need for sparse attention maps on top of that?\n\nIt seems the main contribution is the sparse attention part. However, some doubts remain. Therefore, I can increase my rating if my questions and concerns in the weakness section and questions section are answered satisfactorily."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"The paper is well-written.\n\nThe computational complexity of video diffusion models presents a significant challenge, and the authors effectively highlight this issue and provide a good motivation for addressing it.\n\nTo tackle this, the solution provided by the authors of using a sparse attention map is interesting. Although thinking in this direction is not new, the way the authors motivate the solution and compute the attention maps is scientifically sound and has some novelty.\n\nThe computational speed-up achieved by the method looks impressive."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"In the video generation literature, there are models that generate frames sequentially or follow an auto-regressive approach [1,2]. These models may be less computationally expensive than those using full 3D attention heads, yet there is no empirical or theoretical comparison with such models in the paper.\n\nThere should be an ablation study with the separate effects of sparse attention (without the MLCD) to understand each component in more detail.\n\nThe sampling distillation stage (Stage 1) is not really new, either technically or conceptually. There has been a line of work that provides a similar methodology [3,4], etc. It is not clear how different the proposed distillation is from the existing literature. The same can be said for the knowledge distillation in the final stage (Stage 3).\n\nThe paper has only two qualitative video generation results (or at least what I have found), of which only four frames are shown. There should be a lot more generated videos shown side by side to compare the method qualitatively.\n\n[1] Diffusion forcing: Next-token prediction meets full-sequence diffusion. Chen et al. 2024.\n\n[2] Diffusion models are real-time game engines. Valevski et al. 2024.\n\n[3] MLCM: Multistep Consistency Distillation of Latent Diffusion Model. Xie et al. 2024.\n\n[4] SCott: Accelerating Diffusion Models with Stochastic Consistency Distillation. Liu et al. 2024."}},"nonreaders":[],"tmdate":1733872945563,"tcdate":1730210948438,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission2845/Reviewer_tExx"],"signatures":["ICLR.cc/2025/Conference/Submission2845/Reviewer_tExx"],"forum":"2ezRxhlAxJ","number":1,"license":"CC BY 4.0","cdate":1730210948438,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission2845/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1733872945563,"domain":"ICLR.cc/2025/Conference","replyto":"2ezRxhlAxJ","id":"DvoaBb58Gi","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Efficient inference","video generation","diffusion","Transformer"]},"primary_area":{"value":"infrastructure, software libraries, hardware, systems, etc."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Despite the promise of synthesizing high-fidelity videos, Diffusion Transformers (DiTs) with 3D full attention suffer from expensive inference due to the complexity of attention computation and numerous sampling steps. For example, the popular Open-Sora-Plan model consumes more than 9 minutes for generating a single video of 29 frames. This paper addresses the inefficiency issue from two aspects: 1) Prune the 3D full attention based on the redundancy within video data; We identify a prevalent tile-style repetitive pattern in the 3D attention maps for video data, and advocate a new family of sparse 3D attention that holds a linear complexity w.r.t. the number of video frames. 2) Shorten the sampling process based on multi-step consistency distillation; We split the entire sampling trajectory into several segments and perform consistency distillation within each one to activate few-step generation capacities. We further devise a three-stage training pipeline to conjoin the low-complexity attention and few-step generation capacities. Notably, with 0.1% pretraining data, we turn the Open-Sora-Plan-1.2 model into an efficient one that is 7.4x −7.8x faster for 29 and 93 frames 720p video generation with a marginal performance trade-off in VBench. In addition, we demonstrate that our approach is amenable to distributed inference, achieving an additional 3.91x speedup when running on 4 GPUs with sequence parallelism."},"_bibtex":{"value":"@misc{\nding2025efficientvdit,\ntitle={Efficient-vDiT: Efficient Video Diffusion Transformers With Attention Tile},\nauthor={Hangliang Ding and Dacheng Li and Runlong Su and Zhijie Deng and Ion Stoica and Hao Zhang},\nyear={2025},\nurl={https://openreview.net/forum?id=2ezRxhlAxJ}\n}"},"title":{"value":"Efficient-vDiT: Efficient Video Diffusion Transformers With Attention Tile"},"pdf":{"value":"/pdf/1009bb8176c4e1e6924213856a051436e5042cb8.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"ding|efficientvdit_efficient_video_diffusion_transformers_with_attention_tile"},"authorids":{"value":["~Hangliang_Ding1","~Dacheng_Li1","~Runlong_Su1","~Zhijie_Deng1","~Ion_Stoica1","~Hao_Zhang2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Hangliang Ding","Dacheng Li","Runlong Su","Zhijie Deng","Ion Stoica","Hao Zhang"]}},"version":2},{"content":{"summary":{"value":"This paper presents LVCap-Eval, a comprehensive benchmark for evaluating long-form video captioning in MLLMs. To promote model improvement, the authors build a 7,000-video training corpus using an automated hybrid annotation pipeline. Extensive experiments on 18 state-of-the-art models reveal that closed-source systems like Gemini-2.5-Pro significantly outperform open-source ones, whose performance declines with increasing video length. Further analyses show that the main limitation of current MLLMs lies in high-level narrative reasoning rather than visual perception, and that simple context-aware or memory-based mechanisms can effectively mitigate long-range dependency challenges."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"See weakness."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"The paper makes a contribution by introducing a long-form video captioning benchmark that jointly evaluates narrative coherence and factual accuracy, offering both methodological novelty and strong empirical validation. The analysis and experiments provide valuable insights and resources for advancing multimodal large language models in extended-duration video understanding."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. It is unclear whether the scene-level evaluation verifies the chronological order of events.\n\n2. Although the appendix explains the frame extraction strategy, it does not specify the resolution used. As far as I know, different models adopt different default resolutions.\n\n3. What is the impact of using models with different input sizes?\n\n4. How is the accuracy of the training corpus verified?\n\n6. How does caption length vary with video duration?\n\n7. The paper’s presentation, especially the figures, needs improvement. For example, the text in Figure 2 is too dense, Figure 3’s font is too small, and the row heights in Table 2 look inconsistent.\n\n8. The scene-level evaluation method appears almost identical to VDC, weakening the novelty and contribution.\n\n9. When comparing GPT-4-mini with other GPT variants, does it cause hallucinations or information leakage? How is the reliability of training data ensured?\n\n10. The paper only tests one model configuration for input frame sampling; conclusions may lack generality.\n\n11. Sampling up to 256 frames for a 20-minute video seems clearly unreasonable (even with the 128-frame “optimal” claim), implying about one frame every five seconds. If sampling by FPS rather than uniform intervals, what would the results look like?\n\n13. The paper lacks necessary statistical reporting, such as average caption length."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762917516815,"tcdate":1761617274665,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission4694/Reviewer_6RWE"],"signatures":["ICLR.cc/2026/Conference/Submission4694/Reviewer_6RWE"],"forum":"jMKQFEuSpi","number":3,"license":"CC BY 4.0","cdate":1761617274665,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission4694/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762917516815,"domain":"ICLR.cc/2026/Conference","replyto":"jMKQFEuSpi","id":"C8ol0mpdYi","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Caption","VLLMs","Benchmark"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Generating coherent and factually grounded captions for long-form videos is a critical yet underexplored challenge for multimodal large language models (MLLMs).\nExisting benchmarks, which predominantly feature short clips, are insufficient for evaluating a model's ability to capture narrative structure and fine-grained details over extended durations.\nTo address this gap, we introduce LVCap-Eval, a benchmark for long-form video captioning. \nLVCap-Eval comprises 200 videos from six diverse domains, ranging from 2 to 20 minutes, and features a dual-dimension evaluation protocol that assesses both scene-level narrative coherence and event-level factual accuracy.\nTo facilitate model improvement, we also provide a pipeline for generating a training corpus, demonstrating that fine-tuning with as few as 7,000 samples yields substantial gains.\nOur evaluation of existing MLLMs on this benchmark reveals a significant performance disparity:\nwhile leading closed-source models (e.g., Gemini-2.5-Pro) perform robustly across various video durations, their open-source counterparts degrade sharply as video length increases. \nFinally, our analysis of these model failures highlights potential directions for improving the long-video comprehension of MLLMs."},"_bibtex":{"value":"@misc{\nwang2025lvcapeval,\ntitle={{LVC}ap-Eval: Towards Holistic Long Video Caption Evaluation for Multimodal {LLM}s},\nauthor={Yanghai Wang and Zhe Cao and Yuanxing Zhang and Yifan Yao and Jiahao Wang and Liang Gong and Zhiyu Pan and Zijie Zhang and Zhidong Gan and Minxin Dai and Yonghong Lin and Shihao Li and Qianqian Xie and Xintao Wang and Jiaheng Liu and Zhaoxiang Zhang},\nyear={2025},\nurl={https://openreview.net/forum?id=jMKQFEuSpi}\n}"},"title":{"value":"LVCap-Eval: Towards Holistic Long Video Caption Evaluation for Multimodal LLMs"},"pdf":{"value":"/pdf/a12a1ba67120f20637b47231e5f00ec03ca0b262.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"wang|lvcapeval_towards_holistic_long_video_caption_evaluation_for_multimodal_llms"},"authorids":{"value":["~Yanghai_Wang1","~Zhe_Cao6","~Yuanxing_Zhang3","~Yifan_Yao4","~Jiahao_Wang36","~Liang_Gong1","~Zhiyu_Pan4","~Zijie_Zhang4","~Zhidong_Gan1","~Minxin_Dai1","~Yonghong_Lin2","~Shihao_Li5","~Qianqian_Xie2","~Xintao_Wang1","~Jiaheng_Liu1","~Zhaoxiang_Zhang3"]},"authors":{"value":["Yanghai Wang","Zhe Cao","Yuanxing Zhang","Yifan Yao","Jiahao Wang","Liang Gong","Zhiyu Pan","Zijie Zhang","Zhidong Gan","Minxin Dai","Yonghong Lin","Shihao Li","Qianqian Xie","Xintao Wang","Jiaheng Liu","Zhaoxiang Zhang"]}},"version":2},{"content":{"venue":{"value":"CoRR 2024"},"pdf":{"value":"http://arxiv.org/pdf/2410.00849v1"},"venueid":{"value":"dblp.org/journals/CORR/2024"},"paperhash":{"value":"rajendran|energyqualityaware_variable_framerate_paretofront_for_adaptive_video_streaming"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Prajit_T._Rajendran:","https://dblp.org/search/pid/api?q=author:Samira_Afzal:","https://dblp.org/search/pid/api?q=author:Vignesh_V._Menon:","~Christian_Timmerer1"]},"html":{"value":"https://doi.org/10.48550/arXiv.2410.00849"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2410-00849,\n  publtype={informal},\n  author={Prajit T. Rajendran and Samira Afzal and Vignesh V. Menon and Christian Timmerer},\n  title={Energy-Quality-aware Variable Framerate Pareto-Front for Adaptive Video Streaming},\n  year={2024},\n  cdate={1704067200000},\n  journal={CoRR},\n  volume={abs/2410.00849},\n  url={https://doi.org/10.48550/arXiv.2410.00849}\n}\n"},"abstract":{"value":"Optimizing framerate for a given bitrate-spatial resolution pair in adaptive video streaming is essential to maintain perceptual quality while considering decoding complexity. Low framerates at low bitrates reduce compression artifacts and decrease decoding energy. We propose a novel method, Decoding-complexity aware Framerate Prediction (DECODRA), which employs a Variable Framerate Pareto-front approach to predict an optimized framerate that minimizes decoding energy under quality degradation constraints. DECODRA dynamically adjusts the framerate based on current bitrate and spatial resolution, balancing trade-offs between framerate, perceptual quality, and decoding complexity. Extensive experimentation with the Inter-4K dataset demonstrates DECODRA's effectiveness, yielding an average decoding energy reduction of up to 13.45%, with minimal VMAF reduction of 0.33 points at a low-quality degradation threshold, compared to the default 60 fps encoding. Even at an aggressive threshold, DECODRA achieves significant energy savings of 13.45% while only reducing VMAF by 2.11 points. In this way, DECODRA extends mobile device battery life and reduces the energy footprint of streaming services by providing a more energy-efficient video streaming pipeline."},"title":{"value":"Energy-Quality-aware Variable Framerate Pareto-Front for Adaptive Video Streaming"},"authors":{"value":["Prajit T. Rajendran","Samira Afzal","Vignesh V. Menon","Christian Timmerer"]}},"tmdate":1741250341845,"pdate":1704067200000,"tcdate":1741250322709,"writers":["~"],"signatures":["~Christian_Timmerer1"],"forum":"EBUPxobURw","license":"CC BY-SA 4.0","number":359785,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1741250341845,"domain":"DBLP.org","id":"EBUPxobURw","version":2},{"content":{"venue":{"value":"VCIP 2024"},"pdf":{"value":"https://ieeexplore.ieee.org/iel8/10849624/10849781/10849829.pdf"},"venueid":{"value":"dblp.org/conf/VCIP/2024"},"paperhash":{"value":"rajendran|energyqualityaware_variable_framerate_paretofront_for_adaptive_video_streaming"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Prajit_T._Rajendran:","https://dblp.org/search/pid/api?q=author:Samira_Afzal:","https://dblp.org/search/pid/api?q=author:Vignesh_V._Menon:","~Christian_Timmerer1"]},"html":{"value":"https://doi.org/10.1109/VCIP63160.2024.10849829"},"_bibtex":{"value":"@inproceedings{DBLP:conf/vcip/RajendranAMT24,\n  author={Prajit T. Rajendran and Samira Afzal and Vignesh V. Menon and Christian Timmerer},\n  title={Energy-Quality-aware Variable Framerate Pareto-Front for Adaptive Video Streaming},\n  year={2024},\n  cdate={1704067200000},\n  pages={1-5},\n  url={https://doi.org/10.1109/VCIP63160.2024.10849829},\n  booktitle={VCIP},\n  crossref={conf/vcip/2024}\n}\n"},"abstract":{"value":"Optimizing framerate for a given bitrate-spatial resolution pair in adaptive video streaming is essential to maintain perceptual quality while considering decoding complexity. Low framerates at low bitrates reduce compression artifacts and decrease decoding energy. We propose a novel method, Decoding-complexity aware Framerate Prediction (DECODRA), which employs a Variable Framerate Pareto-front approach to predict an optimized framerate that minimizes decoding energy under quality degradation constraints. DECODRA dynamically adjusts the framerate based on current bitrate and spatial resolution, balancing trade-offs between framerate, perceptual quality, and decoding complexity. Extensive experimentation with the Inter-4K dataset demonstrates DECODRA’s effectiveness, yielding an average decoding energy reduction of up to 13.45 %, with minimal VMAF reduction of 0.33 points at a low-quality degradation threshold, compared to the default 60 fps encoding. Even at an aggressive threshold, DECODRA achieves significant energy savings of 13.45 % while only reducing VMAF by 2.11 points. In this way, DECODRA extends mobile device battery life and reduces the energy footprint of streaming services by providing a more energy-efficient video streaming pipeline."},"title":{"value":"Energy-Quality-aware Variable Framerate Pareto-Front for Adaptive Video Streaming"},"authors":{"value":["Prajit T. Rajendran","Samira Afzal","Vignesh V. Menon","Christian Timmerer"]}},"tmdate":1741250339972,"pdate":1704067200000,"tcdate":1741250322603,"writers":["~"],"signatures":["~Christian_Timmerer1"],"forum":"ThDWlE82d1","license":"CC BY-SA 4.0","number":359777,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1741250339972,"domain":"DBLP.org","id":"ThDWlE82d1","version":2},{"content":{"TLDR":{"value":"We study the timescale of data-reuse in a simple model of shortcut learning with linear networks using non-rigorous statistical physics tools."},"venue":{"value":"ICLR 2027 Conference Submission"},"keywords":{"value":["Shortcut Learning","Offline Learning","Dynamical Mean Field Theory","Statistical Physics of Learning"]},"primary_area":{"value":"learning theory"},"abstract":{"value":"Many dynamical theories of shortcut learning use fresh samples or population gradients, whereas most practical training repeatedly uses a finite dataset. We study full-batch gradient descent with square loss in a synthetic linear model of spurious correlations. Our analysis uses a non-rigorous dynamical mean-field calculation from statistical physics. The resulting equations separate shortcut and core learning from fitting directions specific to the training dataset.\nA perturbative analysis predicts a reuse crossover on a timescale proportional to $\\sqrt{P/D}$, where $P$ is dataset size and $D$ is input dimension. Without regularisation, the interpolation peak at $P=D$ can degrade test performance without impairing core or shortcut acquisition. For $P<D$, overlap between the features can amplify the shortcut coefficient. Ridge suppresses this peak but can increase the ratio of shortcut to core coefficients.\nSimulations validate the theory, while image benchmarks provide qualitative evidence beyond the solvable setting."},"_bibtex":{"value":"@inproceedings{\nanonymous2026the,\ntitle={The Role of Data-Reuse on Shortcut Learning with Linear Networks},\nauthor={Anonymous},\nbooktitle={Submitted to The Fifteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=7Lje1KCKrE},\nnote={under review}\n}"},"title":{"value":"The Role of Data-Reuse on Shortcut Learning with Linear Networks"},"pdf":{"value":"/pdf/03b306373b374635330d0c59f631f275b826c4a8.pdf"},"venueid":{"value":"ICLR.cc/2027/Conference/Submission"}},"tmdate":1791219520069,"tcdate":1787221976109,"writers":["ICLR.cc/2027/Conference","ICLR.cc/2027/Conference/Submission1619/Authors"],"signatures":["ICLR.cc/2027/Conference/Submission1619/Authors"],"forum":"7Lje1KCKrE","license":"CC BY 4.0","number":1619,"cdate":1787221976109,"readers":["everyone"],"invitations":["ICLR.cc/2027/Conference/-/Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Bidding","ICLR.cc/2027/Conference/Submission1619/-/Full_Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Reviewing","ICLR.cc/2027/Conference/-/Edit"],"mdate":1791219520069,"odate":1791031328286,"domain":"ICLR.cc/2027/Conference","id":"7Lje1KCKrE","version":2},{"content":{"summary":{"value":"In this paper, the authors propose the Video Q-Former model, a large multimodal language model for video understanding that leverages an attentive module and spatio-temporal query transformers. Compared to previous video understanding methods, it adaptively extracts spatio-temporal features from videos while considering spatial information within frames and temporal dynamics between frames."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"See weakness."},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":1},"contribution":{"value":2},"strengths":{"value":"Video Q-Former adaptively extracts spatiotemporal features from videos while considering spatial information within frames and temporal dynamics between frames. The Spatiotemporal Q-Former integrates three video experts, namely SP-FFN, T-FFN, and SM-FFN, to simultaneously extract semantically aligned spatial and temporal video representations, achieving better performance."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The authors claim that their method achieves state-of-the-art performance. However, the comparison methods lack recent works, such as VideoChat2, LLaVA-OneVision, MiniCPMv, and PLLaVA. The related work section also lacks recent works.\n\n- There is a lack of novelty. The proposed Attentive Pooling Module has been widely used in previous works, such as [1].\n\n[1] Temporal-attentive Covariance Pooling Networks for Video Recognition. NeurIPS 2021.\n\n- It is recommended that the authors visualize the results of the Attentive Pooling Module to better demonstrate the effectiveness of this module.\n\n- Video-ChatGPT and Video-LLaMA are the main comparison methods, as the Video Q-Former can be considered a combination of the two. However, in the experiments, the authors did not clarify whether all models were trained on the same dataset when comparing Video-LLaMA and Video-ChatGPT. This raises confusion as to whether the performance advantage of Video Q-Former comes from the model architecture or the dataset.\n\n- The results in Table 6 show that the improvement brought by SP Q-Former is quite limited. For example, in terms of the CIDEr metric, it only improved by 0.37. The ablation experiment in Table 8 also shows that compared to \"Spatial only,\" the improvement of \"Video Q-Former\" is also quite limited, improving by only 0.07 on the CIDEr metric."}},"nonreaders":[],"tmdate":1731428918806,"tcdate":1730288349409,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission11192/Reviewer_NyYZ"],"signatures":["ICLR.cc/2025/Conference/Submission11192/Reviewer_NyYZ"],"forum":"R6sIi9Kbxv","number":2,"license":"CC BY 4.0","cdate":1730288349409,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission11192/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428918806,"domain":"ICLR.cc/2025/Conference","replyto":"R6sIi9Kbxv","id":"jpoLZTct9c","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Multimodal Large Language Model","Vision-Language Pretraining","Video Understanding"]},"primary_area":{"value":"foundation or frontier models, including LLMs"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Large language models (LLMs) have made remarkable strides in natural language processing tasks. However, effectively processing and understanding visual information remains a challenge for these models. To address this, multimodal large language models have been proposed, which integrate pre-trained visual encoders with LLMs. Although existing image-based approaches have shown success in aligning visual and textual modalities, extending these advancements to videos is challenging due to the richer visual and temporal information they contain. Current methods, including Video-ChatGPT and Video-LLaMA, have limitations in capturing inter-frame relationships and providing sufficient semantic context. To overcome these challenges, we propose Video Q-Former, a model that adaptively extracts spatiotemporal features from videos with a spatio-temporal querying transformer, enhancing the LLM’s comprehension of visual-language alignment. Extensive experiments demonstrate that our model achieves state-of-the-art performance across various datasets in zero-shot video question answering tasks."},"_bibtex":{"value":"@misc{\nnie2025video,\ntitle={Video Q-Former: Multimodal Large Language Model with Spatio-Temporal Querying Transformer Towards Video Understanding},\nauthor={Yuxiang Nie and Han Wang and Yanjie Wang and Can Huang and Liang Lin and Guanbin Li},\nyear={2025},\nurl={https://openreview.net/forum?id=R6sIi9Kbxv}\n}"},"title":{"value":"Video Q-Former: Multimodal Large Language Model with Spatio-Temporal Querying Transformer Towards Video Understanding"},"pdf":{"value":"/pdf/6a4b2bd8b1e48662f75e7fca3b2b64f4848d6d91.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"nie|video_qformer_multimodal_large_language_model_with_spatiotemporal_querying_transformer_towards_video_understanding"},"authorids":{"value":["~Yuxiang_Nie2","~Han_Wang18","~Yanjie_Wang2","~Can_Huang1","~Liang_Lin1","~Guanbin_Li2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Yuxiang Nie","Han Wang","Yanjie Wang","Can Huang","Liang Lin","Guanbin Li"]}},"version":2},{"content":{"summary":{"value":"The paper introduces a novel self-training approach, Video Self-Training with augmented Reasoning (Video-STaR), to improve the performance of Large Multi-modal Models (LMMs) in video instruction tuning. By leveraging any labeled video dataset, Video-STaR enhances video question answering and adapts LMMs to new downstream tasks. The method cycles between generating answers, verifying them using the video labels, and fine-tuning the model. Experimental results show significant improvements in video understanding tasks, such as a 6.1% increase in Video QA performance and enhanced action quality assessment accuracy."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1. **Clarification on Dataset Labeling**: Can the authors clarify how they handle datasets where video labels are ambiguous or not well-defined? It would be helpful to understand how such cases are managed during the label verification step to avoid generating incorrect instructions.\n\n2. **Self-Training Cycles**: Could the authors provide more details on how they determine when the self-training process has plateaued? Is there a specific performance threshold or metric used to stop the cycles?\n\n3. **Hallucination Mitigation**: The paper acknowledges hallucination as a potential issue, especially in label rationalization. Are there additional methods or heuristics that the authors could explore to reduce hallucinations during both answer generation and rationalization phases?\n\n4. **Generalizability to Complex Tasks**: Video-STaR shows notable improvements in performance on diverse video tasks. Could the authors elaborate on how the method handles particularly complex reasoning tasks beyond action recognition and quality assessment? How might this approach generalize to more abstract tasks, such as subjective video analysis?\n\n5. **Computational Overhead**: Given the iterative nature of the self-training process, what are the computational requirements and trade-offs for implementing Video-STaR at scale? Would this method be feasible for larger video datasets, and if so, what optimizations could be applied?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"- **Originality**: Video-STaR introduces an innovative self-training approach for video instruction tuning, enabling the use of any labeled video dataset for instruction tuning. This represents a novel solution to a critical problem in video QA and reasoning by removing the dependency on manual dataset collection.\n\n- **Quality**: The paper is well-structured, with detailed experimental evaluations demonstrating significant improvements in video QA accuracy and adaptability to downstream tasks. It also addresses weaknesses in previous methods, such as reducing hallucinations in model answers.\n\n- **Clarity**: The methodology, including the self-training cycles of generation, filtering, and tuning, is clearly articulated, with effective visual aids to illustrate the process.\n\n- **Significance**: The paper contributes meaningfully to the field of video QA and LMM adaptation, showing improvements in both general video understanding and task-specific performance. It also creates a large and diverse dataset (VSTAR-1M), further enriching the resources available for multimodal research."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- **Computational Intensity**: The iterative nature of the self-training cycles, involving both answer generation and label rationalization, introduces significant computational overhead. This could limit the scalability and efficiency of the proposed method, especially in resource-constrained environments.\n\n- **Hallucination in Rationalization**: The reliance on label rationalization, especially for difficult tasks, may increase the likelihood of hallucinations. This undermines the robustness of the system in generating accurate explanations and answers, particularly in complex video tasks like FineDiving.\n\n- **Generalizability of Label Rationalization**: While the method assumes that all video labels require rationalization, certain straightforward labels may not benefit from this additional step. This leads to unnecessary computational load without proportional gains in performance.\n\n- **Limited Dataset Variety**: The study largely focuses on specific datasets like FineDiving, STAR-benchmark, and Kinetics700. Expanding the evaluation to a broader set of tasks could help demonstrate the full generalizability and adaptability of the approach to more diverse real-world video datasets."}},"nonreaders":[],"tmdate":1731427323239,"tcdate":1730943983207,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission826/Reviewer_MNYz"],"signatures":["ICLR.cc/2025/Conference/Submission826/Reviewer_MNYz"],"forum":"JYV2hrtFSv","number":4,"license":"CC BY 4.0","cdate":1730943983207,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission826/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427323239,"domain":"ICLR.cc/2025/Conference","replyto":"JYV2hrtFSv","id":"A8pJl6AdFC","forumContent":{"venue":{"value":"ICLR 2025 Poster"},"TLDR":{"value":"Introducing the first self-training method for video-LMMs"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video Understanding","Visual Instruction Tuning","Self-Training","Chain-of-thought reasoning"]},"supplementary_material":{"value":"/attachment/9c72d32e9bce4627af3727c58ce7afee71870795.pdf"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"The performance and reasoning capabilities of Large Multi-modal Models (LMMs) is dependent on the size and quality of their training datasets. However, collecting datasets that support chain-of-thought instruction tuning is highly challenging. Existing video instruction tuning datasets are often derived by prompting large language models with video captions to generate question-answer pairs, which makes them predominantly descriptive rather than reasoning-focused. \nMeanwhile, many labeled video datasets with diverse labels and supervision exist -- however, we find that their integration into LMMs is non-trivial. \nHerein, we present $\\underline{\\text{Video}}$ $\\underline{\\text{S}}\\text{elf}$-$\\underline{\\text{T}}\\text{raining}$ $\\text{with}$ $\\underline{\\text{a}}\\text{ugmented}$ $\\underline{\\text{R}}\\text{easoning}$ (Video-STaR), the first self-training approach for video instruction tuning. \nVideo-STaR allows the utilization of *any* labeled video dataset for video instruction tuning.\nIn Video-STaR, an LMM cycles between instruction generation and finetuning, which we show (I) improves general video understanding and (II) adapts LMMs to novel downstream tasks with existing supervision. \nDuring instruction generation, an LMM is prompted to propose an answer.  The answers are then filtered only to those that contain the original video labels, and the LMM is then re-trained on the generated dataset. \nBy training exclusively on generated answers containing the correct video labels, Video-STaR leverages these existing labels as weak supervision for video instruction tuning.\nOur results demonstrate that Video-STaR-augmented LMMs achieve notable improvements in (I) general Video QA, where TempCompass performance improved by 6.1%, *and* (II) downstream tasks, with a 9.9% increase in Kinetics700-QA accuracy and a 4.0% improvement in action quality assessment on FineDiving, while also exhibiting better interpretability."},"_bibtex":{"value":"@inproceedings{\nzohar2025videostar,\ntitle={Video-{ST}aR: Self-Training Enables Video Instruction Tuning with Any Supervision},\nauthor={Orr Zohar and Xiaohan Wang and Yonatan Bitton and Idan Szpektor and Serena Yeung-Levy},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=JYV2hrtFSv}\n}"},"title":{"value":"Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision"},"pdf":{"value":"/pdf/2dbd229d9e63576bc3553bb16d276747ace91914.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"zohar|videostar_selftraining_enables_video_instruction_tuning_with_any_supervision"},"authorids":{"value":["~Orr_Zohar1","~Xiaohan_Wang2","~Yonatan_Bitton1","~Idan_Szpektor1","~Serena_Yeung-Levy1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Orr Zohar","Xiaohan Wang","Yonatan Bitton","Idan Szpektor","Serena Yeung-Levy"]}},"version":2},{"content":{"TLDR":{"value":"This paper shows that a mammography-based breast cancer risk prediction model relies on shortcut cues such as markers, view annotations, and acquisition features, which means high accuracy does not ensure clinically grounded reasoning."},"venue":{"value":"FAIMI-BRIDGE-EPIMI 2026"},"workshop_topic":{"value":["Bias Detection/Mitigation","Analysis of challenges to clinical deployment"]},"keywords":{"value":["Breast Cancer Risk Prediction; Mammography; Shortcut Learning; Explainable Artificial Intelligence."]},"abstract":{"value":"Deep learning models can predict future breast cancer risk directly from screening mammograms, yet it remains unclear whether these models rely primarily on clinically meaningful imaging features or exploit shortcut signals. The presence of clinically meaningful biomarkers alone does not preclude shortcut learning. We investigated shortcut learning in a published mammography-based breast cancer risk prediction model (AUC = 0.69, 95% CI: 0.62–0.77) through occlusion analysis, background perturbation experiments, and representation analysis. Our results revealed multiple shortcut pathways. Occlusion analysis showed that predictive information was diffusely distributed in CC views but concentrated in the upper breast region of MLO views. Although progressively including larger portions of the background had little effect on performance, localized non-breast cues, including radiographic markers, view-position annotations, and BMI-associated anatomy, influenced predictions. Representation analysis further revealed that acquisition-related information was encoded more strongly than clinically relevant factors. Together, these findings demonstrate that mammography-based risk prediction models simultaneously exploit clinically meaningful imaging features and shortcut signals, indicating that strong predictive performance alone should not be interpreted as evidence that predictions are driven primarily by clinically meaningful imaging features."},"_bibtex":{"value":"@inproceedings{\nashmawy2026multiple,\ntitle={Multiple Shortcut Pathways in Mammography-Based Breast Cancer Risk Prediction},\nauthor={Judy Ashmawy and Habiba Mohamed and Maya Hussien and Tamer Basha and Alaa Melek},\nbooktitle={Joint FAIMI, BRIDGE, and EPIMI Workshop at MICCAI 2026},\nyear={2026},\nurl={https://openreview.net/forum?id=uG64jemYMl}\n}"},"title":{"value":"Multiple Shortcut Pathways in Mammography-Based Breast Cancer Risk Prediction"},"camera_ready_submission":{"value":"/attachment/b2b9630cfa131731c6484ae1d17b2e564c23e7a8.zip"},"pdf":{"value":"/pdf/5457d7d87ce7805f5e7ce13632c163ea35189c52.pdf"},"venueid":{"value":"MICCAI.org/2026/Workshop/FAIMI-BRIDGE-EPIMI"},"paperhash":{"value":"ashmawy|multiple_shortcut_pathways_in_mammographybased_breast_cancer_risk_prediction"},"authorids":{"value":["~Judy_Ashmawy1","~Habiba_Alaaeldin_Mohamed_Mohamed1","~Maya_Hussien1","~Tamer_Basha1","~Alaa_Melek1"]},"authors":{"value":["Judy Ashmawy","Habiba Mohamed","Maya Hussien","Tamer Basha","Alaa Melek"]}},"tmdate":1787848111226,"pdate":1786962251036,"tcdate":1783362695648,"writers":["MICCAI.org/2026/Workshop/FAIMI-BRIDGE-EPIMI","MICCAI.org/2026/Workshop/FAIMI-BRIDGE-EPIMI/Submission27/Authors"],"signatures":["MICCAI.org/2026/Workshop/FAIMI-BRIDGE-EPIMI/Submission27/Authors"],"forum":"uG64jemYMl","license":"CC BY 4.0","number":27,"cdate":1783362695648,"readers":["everyone"],"invitations":["MICCAI.org/2026/Workshop/FAIMI-BRIDGE-EPIMI/-/Submission","MICCAI.org/2026/Workshop/FAIMI-BRIDGE-EPIMI/-/Post_Submission","MICCAI.org/2026/Workshop/FAIMI-BRIDGE-EPIMI/-/Edit","MICCAI.org/2026/Workshop/FAIMI-BRIDGE-EPIMI/Submission27/-/Camera_Ready_Submission"],"mdate":1787848111226,"odate":1786962251036,"domain":"MICCAI.org/2026/Workshop/FAIMI-BRIDGE-EPIMI","id":"uG64jemYMl","version":2},{"content":{"venue":{"value":"MIDL 2024 Poster"},"keywords":{"value":["Image synthesis","misaligned pairs","GAN","generative"]},"abstract":{"value":"Medical image synthesis generates additional imaging modalities that are costly, invasive or harmful to acquire, which helps to facilitate the clinical workflow. When training pairs are substantially misaligned (e.g., lung MRI-CT pairs with respiratory motion), accurate image synthesis remains a critical challenge. Recent works explored the directional registration module to adjust misalignment in generative adversarial networks (GANs); however, substantial misalignment will lead to 1) suboptimal data mapping caused by correspondence ambiguity, and 2) degraded image fidelity caused by morphology influence on discriminators. To address the challenges, we propose a novel Deformation-aware GAN (DA-GAN) to dynamically correct the misalignment during the image synthesis based on multi-objective inverse consistency. Specifically, in the generative process, three levels of inverse consistency cohesively optimise symmetric registration and image generation for improved correspondence. In the adversarial process, to further improve image fidelity under misalignment, we design deformation-aware discriminators to disentangle the mismatched spatial morphology from the judgement of image fidelity. Experimental results show that DA-GAN achieved superior performance on a public dataset with simulated misalignments and a real-world lung MRI-CT dataset with respiratory motion misalignment. The results indicate the potential for a wide range of medical image synthesis tasks such as radiotherapy planning."},"_bibtex":{"value":"@inproceedings{\nxin2024deformationaware,\ntitle={Deformation-aware {GAN} for Medical Image Synthesis with Substantially Misaligned Pairs},\nauthor={Bowen Xin and Tony Young and Claire Wainwright and Tamara Blake and Leo Lebrat and Thomas Gaass and Thomas Benkert and Alto Stemmer and David Coman and Jason Dowling},\nbooktitle={Medical Imaging with Deep Learning},\nyear={2024},\nurl={https://openreview.net/forum?id=2iAsI3SZ8C}\n}"},"title":{"value":"Deformation-aware GAN for Medical Image Synthesis with Substantially Misaligned Pairs"},"latex_code":{"value":"/attachment/2e3eb2efe858f9e00f684cce62711ee9081a02bb.zip"},"pdf":{"value":"/pdf/f2724dc683cea59b4bf083c4b6a56dfaa8bba95d.pdf"},"copyright_form":{"value":"/attachment/fbe4d4e589056dd295762e19278ec32aaf156681.pdf"},"venueid":{"value":"MIDL.io/2024/Conference"},"paperhash":{"value":"xin|deformationaware_gan_for_medical_image_synthesis_with_substantially_misaligned_pairs"},"authorids":{"value":["~Bowen_Xin1","~Tony_Young1","~Claire_Wainwright1","~Tamara_Blake1","~Leo_Lebrat1","~Thomas_Gaass1","~Thomas_Benkert1","~Alto_Stemmer1","~David_Coman1","~Jason_Dowling1"]},"authors":{"value":["Bowen Xin","Tony Young","Claire Wainwright","Tamara Blake","Leo Lebrat","Thomas Gaass","Thomas Benkert","Alto Stemmer","David Coman","Jason Dowling"]}},"tmdate":1717666798869,"pdate":1717666798858,"tcdate":1706159319772,"writers":["MIDL.io/2024/Conference","MIDL.io/2024/Conference/Submission42/Authors"],"signatures":["MIDL.io/2024/Conference/Submission42/Authors"],"forum":"2iAsI3SZ8C","license":"CC BY 4.0","number":42,"cdate":1706159319772,"readers":["everyone"],"invitations":["MIDL.io/2024/Conference/-/Submission","MIDL.io/2024/Conference/-/Post_Submission","MIDL.io/2024/Conference/Submission42/-/Revision","MIDL.io/2024/Conference/Submission42/-/Camera_Ready","MIDL.io/2024/Conference/-/Edit"],"mdate":1717666798869,"odate":1707823859654,"domain":"MIDL.io/2024/Conference","id":"2iAsI3SZ8C","version":2},{"content":{"venue":{"value":"CoRR 2024"},"pdf":{"value":"http://arxiv.org/pdf/2412.02071v1"},"venueid":{"value":"dblp.org/journals/CORR/2024"},"paperhash":{"value":"xue|progressaware_video_frame_captioning"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Zihui_Xue:","~Joungbin_An1","https://dblp.org/search/pid/api?q=author:Xitong_Yang:","https://dblp.org/search/pid/api?q=author:Kristen_Grauman:"]},"html":{"value":"https://doi.org/10.48550/arXiv.2412.02071"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2412-02071,\n  publtype={informal},\n  author={Zihui Xue and Joungbin An and Xitong Yang and Kristen Grauman},\n  title={Progress-Aware Video Frame Captioning},\n  year={2024},\n  cdate={1704067200000},\n  journal={CoRR},\n  volume={abs/2412.02071},\n  url={https://doi.org/10.48550/arXiv.2412.02071}\n}\n"},"abstract":{"value":"While image captioning provides isolated descriptions for individual images, and video captioning offers one single narrative for an entire video clip, our work explores an important middle ground: progress-aware video captioning at the frame level. This novel task aims to generate temporally fine-grained captions that not only accurately describe each frame but also capture the subtle progression of actions throughout a video sequence. Despite the strong capabilities of existing leading vision language models, they often struggle to discern the nuances of frame-wise differences. To address this, we propose ProgressCaptioner, a captioning model designed to capture the fine-grained temporal dynamics within an action sequence. Alongside, we develop the FrameCap dataset to support training and the FrameCapEval benchmark to assess caption quality. The results demonstrate that ProgressCaptioner significantly surpasses leading captioning models, producing precise captions that accurately capture action progression and set a new standard for temporal precision in video captioning. Finally, we showcase practical applications of our approach, specifically in aiding keyframe selection and advancing video understanding, highlighting its broad utility."},"title":{"value":"Progress-Aware Video Frame Captioning"},"authors":{"value":["Zihui Xue","Joungbin An","Xitong Yang","Kristen Grauman"]}},"tmdate":1739996755121,"pdate":1704067200000,"tcdate":1739996752726,"writers":["~"],"signatures":["~Joungbin_An1"],"forum":"dETQrOyhBL","license":"CC BY-SA 4.0","number":335278,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1739996755121,"domain":"DBLP.org","id":"dETQrOyhBL","version":2},{"content":{"summary":{"value":"This paper proposes a new low-light video enhancement (LLVE) model, IS-SFD. As a zero-reference LLVE method, it does not requires low-normal light pairs for supervised learning, but it is subjected to the flickering and noise problems. To confront these challenges, the authors propose two modules, a GIE-Net that fuses multi-scale features by a gating mechanism to enhance inter-frames consistency and a SGRD-Net which combines frequency and semantic features to guide denoising."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"In the last paragraph of section 1, it states that \"the semantic features extracted from a ResNet-50 image encoder provide high-level ....\". But actually it uses CLIP model. Is this a typo?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"+ This work is well-motivatd. GIE-Net and SGRD-Net are tailored for addressing the inherent challenge of zero-reference LLVE models.\n+ The F-S pair design is insightful, which may also applys to other video enhancement tasks."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Inconsistent presentation of the best/second results in Table 1 and Table 2. It's better to use a unified rule.\n- Evaluation is not comprehensive enough. It's better to evaluate with advanced no-reference image quality assessment and video quality assessment models.\n- There are altnerative methods to extract frequency features (e.g., DCT, Laplacian pyramid, etc.). The authors should compare DWT with them.\n- Similarly, there are many other vision encoders can be used to extract semnatic features. The authors should compare CLIP with them.\n- While this work focus on video enhancement rather than image enhancement. It's better to provide video examples in supplementary material. Failing to do so decredits the promise of this work."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762931125129,"tcdate":1761545927230,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission19095/Reviewer_4Pey"],"signatures":["ICLR.cc/2026/Conference/Submission19095/Reviewer_4Pey"],"forum":"vc6pSwumm6","number":4,"license":"CC BY 4.0","cdate":1761545927230,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission19095/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762931125129,"domain":"ICLR.cc/2026/Conference","replyto":"vc6pSwumm6","id":"yufmr0o7lM","forumContent":{"TLDR":{"value":"a zero-reference LLVE method; illumination smoothness; semantic-frequency denoising; multiscale similarity"},"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["low-light video enhancement","zero-reference method","semantic feature","frequency feature"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Low-light video enhancement (LLVE) is important for real-world applications where visibility degradation impairs human perception or downstream vision tasks. While zero-reference methods do not need paired image data, they often have flickering problems and struggle to suppress noise while preserving image details. We propose Illumination Smoothness and Semantic-frequency Denoising (IS-SFD), a zero-reference framework for enhancing low-light videos through temporal illumination modeling and denoising guided by semantic and frequency features. To ensure temporal consistency, we introduce a Gated Illumination Estimation Network (GIE-Net) that adaptively fuses multi-frame features by a gating mechanism guided by multiscale similarity of adjacent video frames. For denoising, we design a Semantic-frequency Guided Reflection Denoising Network (SGRD-Net), which combines frequency features from a DWT encoder and semantic features from a frozen CLIP encoder. These features are fused to suppress noise while maintaining structural details in critical areas such as object boundaries. Experiments demonstrate that IS-SFD outperforms existing methods in visual quality and temporal consistency, establishing a new baseline for zero reference LLVE. The code will be made available upon acceptance of the paper."},"_bibtex":{"value":"@misc{\nyu2025issfd,\ntitle={{IS}-{SFD}: Illumination Smoothness and Semantic-frequency Denoising for low-light video enhancement},\nauthor={Ye Yu and yang zheng yang and Jun Yi and Deming Chen and Lu Qiang and Wei Jia and Jun Yu},\nyear={2025},\nurl={https://openreview.net/forum?id=vc6pSwumm6}\n}"},"title":{"value":"IS-SFD: Illumination Smoothness and Semantic-frequency Denoising for low-light video enhancement"},"pdf":{"value":"/pdf/00bcaa5ca64ff1517b191fcfb690c2921f006fca.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"yu|issfd_illumination_smoothness_and_semanticfrequency_denoising_for_lowlight_video_enhancement"},"authorids":{"value":["~Ye_Yu3","~yang_zheng_yang1","~Jun_Yi5","~Deming_Chen4","~Lu_Qiang2","~Wei_Jia3","~Jun_Yu3"]},"authors":{"value":["Ye Yu","yang zheng yang","Jun Yi","Deming Chen","Lu Qiang","Wei Jia","Jun Yu"]}},"version":2},{"content":{"summary":{"value":"This paper introduces Video Diffusion Transformer (VDT), a video generation model with pure transformer architecture.\nThe core idea is to replace the unet structure of the existing video generation model with a pure transformer structure.\nThe authors claimed that it's the first successful model in transformer-based video diffusion.\nThey propose a unified spatial-temporal mask modeling mechanism for VDT, enabling it to unify a diverse array of general-purpose tasks.\nThey show the effectiveness of VDT on several video-related tasks."},"presentation":{"value":"3 good"},"contribution":{"value":"2 fair"},"soundness":{"value":"3 good"},"strengths":{"value":"1. Compared with the current U-net-based network with demonic attention structures, VDT is a pure transformer network. If it is indeed the first Transformer-only video diffusion network, it would be a good baseline for this area.\n2. This paper further proposes the spatiotemporal mask mechanism, which can make VDT adapt to various video tasks, including video generation, video prediction, image-to-video generation, video completion, etc.\n3. This paper verifies the effectiveness of VDT on multiple video generation tasks.\n4. The code in this article is open source and will help others replicate this work."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The core idea of this paper is simple and reasonable. Replacing the original CNN with transformer is an effective method that has been widely verified in the whole field. In addition, this paper adopts the basic attention mechanism, and the mask strategy proposed is also commonly used in video tasks. So, from this perspective, the contribution is slightly less obvious.\n2. I feel that the work in this paper can be a good benchmark in this field. However, this paper only verifies that it can be effective on multiple video tasks. As a good benchmark, this paper should give more detailed experimental analysis of ablation. For example, why does space-time pay attention to time before space? In this paper, the influence of each module on the video generation effect, the comparison experiment of hyperparameter and so on.\n3. Tab.5, 142.3 (blacken) is not the best one."},"confidence":{"value":"4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work."},"questions":{"value":"see weakness"},"rating":{"value":"6: marginally above the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699635998096,"tcdate":1698588711926,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission710/Reviewer_cvvZ"],"signatures":["ICLR.cc/2024/Conference/Submission710/Reviewer_cvvZ"],"forum":"Un0rgm9f04","number":2,"license":"CC BY 4.0","cdate":1698588711926,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission710/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699635998096,"domain":"ICLR.cc/2024/Conference","replyto":"Un0rgm9f04","id":"Avh9COtk2R","forumContent":{"venue":{"value":"ICLR 2024 poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["diffusion model","video generation"]},"primary_area":{"value":"generative models"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"This work introduces Video Diffusion Transformer (VDT), which pioneers the use of transformers in diffusion-based video generation.\nIt features transformer blocks with modularized temporal and spatial attention modules to leverage the rich spatial-temporal representation inherited in transformers. Additionally, we propose a unified spatial-temporal mask modeling mechanism, seamlessly integrated with the model, to cater to diverse video generation scenarios.\n\nVDT offers several appealing benefits. (1) It excels at capturing temporal dependencies to produce temporally consistent video frames and even simulate the physics and dynamics of 3D objects over time. (2) It facilitates flexible conditioning information, e.g., simple concatenation in the token space, effectively unifying different token lengths and modalities. (3) Pairing with our proposed spatial-temporal mask modeling mechanism, it becomes a general-purpose video diffuser for harnessing a range of tasks, including unconditional generation, video prediction, interpolation, animation, and completion, etc. Extensive experiments on these tasks spanning various scenarios, including autonomous driving, natural weather, human action, and physics-based simulation, demonstrate the effectiveness of VDT. Moreover, we provide a comprehensive study on the capabilities of VDT in capturing accurate temporal dependencies, handling conditioning information, and the spatial-temporal mask modeling mechanism. Additionally, we present comprehensive studies on how VDT handles conditioning information with the mask modeling mechanism, which we believe will benefit future research and advance the field. Codes and models are available at the https://VDT-2023.github.io."},"_bibtex":{"value":"@inproceedings{\nlu2024vdt,\ntitle={{VDT}: General-purpose Video Diffusion Transformers via Mask Modeling},\nauthor={Haoyu Lu and Guoxing Yang and Nanyi Fei and Yuqi Huo and Zhiwu Lu and Ping Luo and Mingyu Ding},\nbooktitle={The Twelfth International Conference on Learning Representations},\nyear={2024},\nurl={https://openreview.net/forum?id=Un0rgm9f04}\n}"},"title":{"value":"VDT: General-purpose Video Diffusion Transformers via Mask Modeling"},"pdf":{"value":"/pdf/8781429d598437687744d54f5e6102be5c4ed7cd.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference"},"paperhash":{"value":"lu|vdt_generalpurpose_video_diffusion_transformers_via_mask_modeling"},"authorids":{"value":["~Haoyu_Lu1","~Guoxing_Yang3","~Nanyi_Fei1","~Yuqi_Huo1","~Zhiwu_Lu1","~Ping_Luo2","~Mingyu_Ding1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Haoyu Lu","Guoxing Yang","Nanyi Fei","Yuqi Huo","Zhiwu Lu","Ping Luo","Mingyu Ding"]}},"version":2},{"content":{"venue":{"value":"IEEE Trans. Circuits Syst. Video Technol. 2025"},"pdf":{"value":"https://ieeexplore.ieee.org/iel8/76/10989278/10811981.pdf"},"venueid":{"value":"dblp.org/journals/TCSV/2025"},"paperhash":{"value":"shao|textvideo_knowledge_guided_prompting_for_weakly_supervised_temporal_action_localization"},"authorids":{"value":["","~Feifei_Zhang2",""]},"html":{"value":"https://doi.org/10.1109/TCSVT.2024.3521125"},"_bibtex":{"value":"@article{DBLP:journals/tcsv/ShaoZX25,\n  author={Yuxiang Shao and Feifei Zhang and Changsheng Xu},\n  title={Text-Video Knowledge Guided Prompting for Weakly Supervised Temporal Action Localization},\n  year={2025},\n  month={May},\n  cdate={1746057600000},\n  journal={IEEE Trans. Circuits Syst. Video Technol.},\n  volume={35},\n  number={5},\n  pages={4184-4197},\n  url={https://doi.org/10.1109/TCSVT.2024.3521125}\n}\n"},"abstract":{"value":"Weakly supervised temporal action localization (WTAL) aims to localize action instances with only video-level labels for supervision. Recent methods convert category labels to natural language through prompting and utilize pre-trained vision-language models to generate text representation from natural language for supervision. This is because natural language can provide more prosperous and generalized semantic supervision to compensate for the lack of supervision in weakly supervised scenarios. However, it should be noted that current prompting methods face limitations in generating dynamic prompts that adapt to each video, which leads to difficulties in accurately aligning text and video representations. In this work, we propose a novel Text-Video Knowledge Guided Prompting (TVKP) framework for WTAL, which generates video-aware prompts based on text-video knowledge to enhance semantic alignment between text and video representations and introduce more video-related external category labels to enrich semantic supervision. We introduce the video-aware prompting (VAP) module to learn text-video knowledge from the joint distribution of text and video representations to generate video-aware text representation. Meanwhile, to make VAP more effectively learn text-video knowledge, a text-video contrastive loss is proposed to ensure semantic consistency between text and video representations. In addition, we propose the external knowledge prompting (EKP) module to introduce more video-related text labels from an external knowledge base to enrich prompts for accurate semantic alignment. Extensive experiments are conducted on three public datasets, THUMOS14, ActivityNet1.2, and ActivityNet1.3, demonstrating that our approach outperforms state-of-the-art methods."},"title":{"value":"Text-Video Knowledge Guided Prompting for Weakly Supervised Temporal Action Localization"},"authors":{"value":["Yuxiang Shao","Feifei Zhang","Changsheng Xu"]}},"tmdate":1772634263810,"pdate":1767139200000,"externalIds":["dblp:journals/tcsv/ShaoZX25"],"tcdate":1772634258419,"writers":["~"],"signatures":["~Feifei_Zhang2"],"forum":"y2JNgZ84mB","license":"CC BY-SA 4.0","number":847549,"cdate":1746057600000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1772634263810,"domain":"DBLP.org","id":"y2JNgZ84mB","version":2},{"content":{"summary":{"value":"This paper introduces a framework named X-PlugVid, that leverages pre-trained image-based plugins for video generation without additional retraining. The proposed method includes a spatial-temporal adapter that transforms spatial control signals from image plugins (like ControlNet) into temporally coherent guidance for video diffusion models. The author proposes a timestep remapping strategy, which maps richer image features from later timesteps to guide early stages in video generation. The proposed X-PlugVid enhances controllable video generation, allowing a single adapter to make diverse plugins compatible with different video backbones."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"Please refer to the weakness part.\n\nMy concerns were not addressed during the rebuttal, so I changed my score to reject."},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The proposed training-free design enables seamless integration of image plugins into video models, saving computational resources.\n\n2. The proposed adapter is compatible with various plugins and backbones, showing versatility across different video generation models.\n\n3. By incorporating the timestep remapping strategy, the proposed method enables better control over the video generation process, enhancing guidance and improving video quality."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The model relies on high-frequency injection, which may introduce artifacts if not well-tuned, especially when aggressive controls are applied. The author should discuss this, as well as the failure cases.\n\n2. The proposed high-pass filtering may limit performance with plugins that require low-frequency guidance. It would be great if the authors could discuss this in their paper/rebuttal.\n\n3. The discussion of the proposed method in the introduction, and method sections are all around SVD. However, the main experiments are conducted on Hotshot-XL and I2VGen-XL. There are only a few discussions about applying the proposed method to SVD in the appendix which is weird. SVD is an image-to-video method and Hotshot-XL/I2VGen-XL are text/image-to-video models, there should be some difference between them when incorporating the proposed methods.\n\n4. It is highly encouraged to upload the video by supplementary or to an anonymous account on YouTube/GitHub because not all the reviewers have a PDF reader that supports playing GIFs. Also, more video cases are encouraged. There are only 10 videos in the appendix. Please provide at least 20-30 video examples to better demonstrate the method's capabilities.\n\n5. Is there any result that supports a wide-screen aspect ratio? For example, 16:9? Both Hotshot-XL and I2VGen-XL support this.\n\n6. Why do the results in Table 2 in the Appendix have a text? The SVD method only supports image-to-video generation.\n\n7. Table 2 in the Appendix should be Figure 2.\n\n8. The name of UNet should be U-Net."}},"nonreaders":[],"tmdate":1733123751713,"tcdate":1730719543513,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission6384/Reviewer_NDeZ"],"signatures":["ICLR.cc/2025/Conference/Submission6384/Reviewer_NDeZ"],"forum":"TTWxMAwS6n","number":3,"license":"CC BY 4.0","cdate":1730719543513,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission6384/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1733123751713,"domain":"ICLR.cc/2025/Conference","replyto":"TTWxMAwS6n","id":"8ErFQUymnx","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["video generation","diffusion model","efficiency"]},"primary_area":{"value":"generative models"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"We introduce X-PlugVid, a unified framework designed to seamlessly adapt pretrained image-based plug-and-play modules for video diffusion models, facilitating controllable video generation without the need for retraining. This framework leverages a spatial-temporal adapter to effectively bridge the gap between image and video diffusion models. Specifically, we adopt a frozen copy of a large-scale\npretrained image diffusion model (e.g. Stable Diffusion v1.5) as spatial prior. Then we train a spatial-temporal adapter to convert the prior into temporally consistent guidance for video diffusion models (e.g. SVD). To further enhance the effectiveness of image plugins in guiding video models, we introduce a timestep remapping strategy. Recognizing that denoising is an entropic reduction process, this strategy selects priors from later timesteps of the image model, which contain richer information, to be injected into the video models, optimizing the quality and consistency of the generated videos. Comprehensive experimental evaluations of X-PlugVid demonstrate its broad compatibility with diverse operational conditions and different plugins, confirming that leveraging priors from a pretrained diffusion model can minimize redundant training and enable versatile controllable video generation."},"_bibtex":{"value":"@misc{\nran2024xplugvid,\ntitle={X-PlugVid: Versatile Adaptation of Image Plugins for Controllable Video Generation},\nauthor={Lingmin Ran and Chenyang Si and Xudong Lin and Jia-Wei Liu and Rui Zhao and Ziwei Liu and Jussi Keppo and Mike Zheng Shou},\nyear={2024},\nurl={https://openreview.net/forum?id=TTWxMAwS6n}\n}"},"title":{"value":"X-PlugVid: Versatile Adaptation of Image Plugins for Controllable Video Generation"},"pdf":{"value":"/pdf/7fa89672cb73393600912455bb3a8c8c0fbaa0db.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"ran|xplugvid_versatile_adaptation_of_image_plugins_for_controllable_video_generation"},"authorids":{"value":["~Lingmin_Ran1","~Chenyang_Si2","~Xudong_Lin1","~Jia-Wei_Liu1","~Rui_Zhao12","~Ziwei_Liu1","~Jussi_Keppo1","~Mike_Zheng_Shou1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Lingmin Ran","Chenyang Si","Xudong Lin","Jia-Wei Liu","Rui Zhao","Ziwei Liu","Jussi Keppo","Mike Zheng Shou"]}},"version":2},{"content":{"summary":{"value":"The manuscript proposes an ensemble framework for video understanding based on pre-trained visual-language models, which involves utilizing LLM to enhance the descriptiveness of text labels and leveraging a video-to-text model to enhance video representations. The authors conducted comprehensive experiments on action recognition, video-text retrieval, and time-sensitive video tasks, demonstrating the effectiveness of the approach in zero-shot scenarios."},"presentation":{"value":"3 good"},"contribution":{"value":"2 fair"},"soundness":{"value":"3 good"},"strengths":{"value":"The experiments and ablation studies are conducted comprehensively, validated on different pre-existing architectures, and taken into account various types of video data.\n\nEmploying the video-to-text model (not limited to VGPT and caption models) is novel and worthy to explore in the video field. Fusing the text and video representations is depicted to be beneficial in bridging the gap between video and textual labels in the embedding space.\n\nThe manuscript provides a detailed explanation and examples of prompting the GPT to refine the simple textual label, which in turn enhances reproducibility."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"In the 'video-to-text guided visual feature enhancement' (section 2.2), the adopted VGPT relies on CLIP-ViT-L and vicuna, where the computational cost of performing multiple inferences (including text embedding and filtering) far exceeds that of the basic video understanding model. This limits the practical value of the proposed approach. \n\nExcept for CLIP, an image-language pre-trained model targeted specifically for the image field, the proposed approach shows relatively limited performance gain in other video-based models (ViFi-CLIP, AIM, ActionCLIP), considering the additional computational requirements.\n\nThe configurations of adopted pre-trained models (AIM, ActionCLIP, …) remain unclear, which datasets are these models pre-trained on (e.g. K400, K700, …)? For AIM, do the authors directly remove the classification layers?"},"confidence":{"value":"4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work."},"questions":{"value":"Are the high-level action contexts mentioned in the manuscript manually designed?"},"rating":{"value":"5: marginally below the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636584524,"tcdate":1698315403732,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission5632/Reviewer_FnrC"],"signatures":["ICLR.cc/2024/Conference/Submission5632/Reviewer_FnrC"],"forum":"9F0xInGNBF","number":1,"license":"CC BY 4.0","cdate":1698315403732,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission5632/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636584524,"domain":"ICLR.cc/2024/Conference","replyto":"9F0xInGNBF","id":"EYpZBP60Ni","forumContent":{"venue":{"value":"Submitted to ICLR 2024"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video-Language models","LLM","Video Understanding","Zero-shot"]},"primary_area":{"value":"representation learning for computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Vision-language models (VLMs) classify the query video by calculating a similarity score between the visual features and text-based class label representations.\nRecently, large language models (LLMs) have been used to enrich the text-based\nclass labels by enhancing the descriptiveness of the class names. However, these\nimprovements are restricted to the text-based classifier only, and the query visual\nfeatures are not considered. In this paper, we propose a framework which combines pre-trained discriminative VLMs with pre-trained generative video-to-text\nand text-to-text models. We introduce two key modifications to the standard zero-shot setting. First, we propose language-guided visual feature enhancement and\nemploy a video-to-text model to convert the query video to its descriptive form.\nThe resulting descriptions contain vital visual cues of the query video, such as\nwhat objects are present and their spatio-temporal interactions. These descriptive cues provide additional semantic knowledge to VLMs to enhance their zero-shot performance. Second, we propose video-specific prompts to LLMs to generate more meaningful descriptions to enrich class label representations. Specifically, we introduce prompt techniques to create a Tree Hierarchy of Categories for\nclass names, offering a higher-level action context for additional visual cues, We\ndemonstrate the effectiveness of our approach in video understanding across three\ndifferent zero-shot settings: 1) video action recognition, 2) video-to-text and text-to-video retrieval, and 3) time-sensitive video tasks. Consistent improvements\nacross multiple benchmarks and with various VLMs demonstrate the effectiveness of our proposed framework. Our code will be made publicly available."},"_bibtex":{"value":"@misc{\nyousaf2024videoprompter,\ntitle={{VIDEOPROMPTER}: {AN} {ENSEMBLE} {OF} {FOUNDATIONAL} {MODELS} {FOR} {ZERO}-{SHOT} {VIDEO} {UNDERSTANDING}},\nauthor={Adeel Yousaf and Muzammal Naseer and Salman Khan and Fahad Khan and Mubarak Shah},\nyear={2024},\nurl={https://openreview.net/forum?id=9F0xInGNBF}\n}"},"title":{"value":"VIDEOPROMPTER: AN ENSEMBLE OF FOUNDATIONAL MODELS FOR ZERO-SHOT VIDEO UNDERSTANDING"},"pdf":{"value":"/pdf/29043d747f21b29024fcff3dab8bee0aa87078b8.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Rejected_Submission"},"paperhash":{"value":"yousaf|videoprompter_an_ensemble_of_foundational_models_for_zeroshot_video_understanding"},"authorids":{"value":["~Adeel_Yousaf1","~Muzammal_Naseer1","~Salman_Khan4","~Fahad_Khan1","~Mubarak_Shah3"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Adeel Yousaf","Muzammal Naseer","Salman Khan","Fahad Khan","Mubarak Shah"]}},"version":2},{"content":{"summary":{"value":"The paper proposes MLLMEraser, a test-time unlearning framework for MLLMs that leverages activation steering to erase targeted knowledge without parameter updates. Specifically, it introduces a multimodal erasure direction constructed from contrastive image-text pairs and an input-aware steering mechanism to selectively apply interventions. Experiments on LLaVA-1.5 and Qwen-2.5-VL show superior forgetting performance and lower compute cost relative to prior training-based unlearning methods."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"Please see Weaknesses"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"(1)\tThe proposed MLLMEraser introduces a novel activation steering vector that contrasts adversarial image-text pairs with their refusal-style counterparts, thereby overcoming the limitations of common text-only steering in MLLMs.\n\n(2)\tBy acting at inference rather than requiring re-training or parameter updates, it offers an efficient unlearning framework compared to traditional training-based unlearning approaches."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"(1) The meanings of $D^+$ and $D^-$ are not inconsistent throughout the main text, which significantly hampers clarity. Specifically, in Equation (3), the text states that '$D^+$  denotes the set of knowledge-recall samples and $D^-$  the corresponding knowledge-erasure samples'. In Equation (7), however, knowledge-recall pairs are assigned to the negative set $D^-$, whereas knowledge-erasure pairs are assigned to the positive set $D^+$. \n\n(2) The steering strength λ and regularization parameter γ are reported as empirical values tailored to LLaVA-1.5-7B and Qwen-2.5-VL-7B models, yet the tuning process is not adequately detailed in the main text. Thus, it remains unclear how to set these hyper-parameters on new datasets or architectures.\n\n(3) While Figures 4 provide qualitative insights into how activation distributions change before and after steering, the paper lacks a deeper quantitative analysis of these changes. Without such details, the robustness of the justification for the null-space projection constraint is not fully convincing.\n\n(4) I am curious about whether steering at different LLM layers or within the vision encoder could affect unlearning efficacy.\n\n(5) In Table 1, boldface data do not always represent the best results. In particular, for the Ret and Cele metrics, the results labeled 'Ours' are generally not better than the Vanilla method."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762915628908,"tcdate":1761813643633,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission859/Reviewer_vJ3Q"],"signatures":["ICLR.cc/2026/Conference/Submission859/Reviewer_vJ3Q"],"forum":"VmrJ8C3gxu","number":4,"license":"CC BY 4.0","cdate":1761813643633,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission859/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762915628908,"domain":"ICLR.cc/2026/Conference","replyto":"VmrJ8C3gxu","id":"E2z1HfpbC3","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["MLLM Unlearning","Test-time Unlearning","Activation Steering"]},"primary_area":{"value":"interpretability and explainable AI"},"abstract":{"value":"Multimodal large language models (MLLMs) have demonstrated remarkable capabilities across vision–language tasks, yet their large-scale deployment raises pressing concerns about memorized private data, outdated knowledge, and harmful content. Existing unlearning approaches for MLLMs typically adapt training-based strategies such as gradient ascent or preference optimization, but these methods are computationally expensive, irreversible, and often distort retained knowledge. In this work, we propose MLLMEraser, an input-aware, training-free framework for test-time unlearning. Our approach leverages activation steering to enable dynamic knowledge erasure without parameter updates. Specifically, we construct a multimodal erasure direction by contrasting adversarially perturbed, knowledge-recall image–text pairs with knowledge-erasure counterparts, capturing both textual and visual discrepancies. To prevent unnecessary interference, we further design an input-aware steering mechanism that adaptively determines when and how the erasure direction should be applied, preserving utility on retained knowledge while enforcing forgetting on designated content. Experiments on LLaVA-1.5 and Qwen-2.5-VL demonstrate that MLLMEraser consistently outperforms state-of-the-art MLLM unlearning baselines, achieving stronger forgetting performance with lower computational cost and minimal utility degradation."},"_bibtex":{"value":"@misc{\nding2026mllmeraser,\ntitle={{MLLME}raser: Achieving Test-Time Unlearning in Multimodal Large Language Models through Activation Steering},\nauthor={Chenlu Ding and Jiancan Wu and Leheng Sheng and Fan Zhang and Yancheng Yuan and Xiang Wang and Xiangnan He},\nyear={2026},\nurl={https://openreview.net/forum?id=VmrJ8C3gxu}\n}"},"title":{"value":"MLLMEraser: Achieving Test-Time Unlearning in Multimodal Large Language Models through Activation Steering"},"pdf":{"value":"/pdf/c9fcf16b5ae3f5acd2a319a32bd342741cf6564c.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"ding|mllmeraser_achieving_testtime_unlearning_in_multimodal_large_language_models_through_activation_steering"},"authorids":{"value":["~Chenlu_Ding1","~Jiancan_Wu1","~Leheng_Sheng2","~Fan_Zhang74","~Yancheng_Yuan1","~Xiang_Wang2","~Xiangnan_He1"]},"authors":{"value":["Chenlu Ding","Jiancan Wu","Leheng Sheng","Fan Zhang","Yancheng Yuan","Xiang Wang","Xiangnan He"]}},"version":2},{"content":{"summary":{"value":"This paper improves upon the original shortcut model paper by introducing a cumulative self-consistency loss. In the paper, the original shortcut model loss is first framed as an optimal control problem, where the objective is to minimize *total* accumulated error. However, the standard loss only penalizes error at the *current* timestep, and thus does not account for future error. This paper derives the optimal objective that maximizes for future error minimization, analogous to maximizing future return rather than reward in an RL framework. Practically, it is found that using a discrete number of future steps (R=2 or R=4) already leads to sizable improvement. Experiments show that the proposed cumulative SL reliably improves upon the base SL, and ablations show that the improvement can be steadily improved with additional compute allocated to R."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"See section above for further questions associated with each topic.\n\nAs a side note, the discussion in the Appendix on the connection to TD-based value learning methods is quite interesting, and could be a worthy direction for future work.\n\nIn general, the ideas in this paper can be strengthened by a more thorough clarity in writing. The introduction and background sections can be condensed, and the relation between optimal control and the shortcut loss can benefit from additional textual commentary explaining the intuition behind the equations. With clarifications to the questions in the weaknesses section, I would consider a raise to the review score.\n\nCan the cumulative self-consistency objective be applied to other few-step techniques such as consistency models or meanflow?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"This paper highlights a well-posed analysis of the shortcut modelling loss, showing that is an instantaneous approximation to an objective that should be expanded to account for future errors. Using this viewpoint, the paper derives a clear true objective, and shows that even a discrete approximation to this true objective results in performance improvement. In this way, the paper clearly ties together empirical improvement towards a theoretically-motivated insight.\n\nThe clarity in the paper is passable. Figure 2 provides a solid intuition as to why the cumulative objective is a more precise target. (See below for weaknesses).\n\nThe significance of this paper is solid, as it provides a new viewpoint which can be applied to distillation methods in general. The experiments are conducted on standard benchmarks at a reasonable network size, and wall-clock experiments are presented to take computational requirements into account. Error bars are included."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The empirical performance of this method could be better framed by comparing to recent prior works, specifically Meanflow. \n\nIn terms of clarity, the notation introduced in Section 3/4/5 can be refined. A full explanation of optimal control as given in section 3 may not be necessary, especially terms such as b, g, and h which are not used in the final relation to shortcut model loss. In equation 4, it will help to make it clear that u(X) is the *error* term which has an optimality at zero, but the network itself does not predict u(X) but rather a velocity (that must be compared to the teacher velocity). \n\nFor equation 14, is the gradient with relation to future errors taken through multiple evaluations of the network? If so, is the error term for k>1 backpropgated to all previous steps including the current step? An algorithmic explanation would strengthen this section.\n\nWhile the text argues that equations 11 and 12 should be different, the resulting integrals appear identical. What is the precise difference between these two objectives? Commentary would strengthen this section."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762931628183,"tcdate":1761683964535,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission19781/Reviewer_CsHB"],"signatures":["ICLR.cc/2026/Conference/Submission19781/Reviewer_CsHB"],"forum":"cZqAk87Lu4","number":2,"license":"CC BY 4.0","cdate":1761683964535,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission19781/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762931628183,"domain":"ICLR.cc/2026/Conference","replyto":"cZqAk87Lu4","id":"HnWhESUPbG","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["One-Step Diffusion","Optimal Control","Shortcut Diffusion Models"]},"primary_area":{"value":"generative models"},"abstract":{"value":"Although iterative denoising (i.e., diffusion/flow) methods offer strong generative performance, they suffer from low generation efficiency, requiring hundreds of steps of network forward passes to simulate a single sample. Mitigating this requires taking larger step-sizes during simulation, thereby allowing one- or few-step generation. Recently proposed shortcut model learns larger step-sizes by enforcing alignment between its direction and the path defined by a base many-step flow-matching model through a self-consistency loss. However, its generation quality is significantly lower than the base model. In this paper, we formulate few-step generation as a controlled base generative process, and show that self-consistency loss can be understood through the lens of optimal control. This perspective naturally motivates its generalization to the proposed cumulative self-consistency loss that cumulatively penalizes misalignment along the entire trajectory. This encourages larger step-sizes that not only align with the base model at the current time step but also support alignment in the subsequent steps, facilitating high-quality generation. Furthermore, we draw a connection between our approach and reinforcement learning, potentially opening the door to a new set of approaches for few-step generation. Experiments show that we significantly improve one- and few-step generation quality under the same training budget. Implementation is available at: [https://github.com/paribeshregmi/Shortcut-CSL](https://github.com/paribeshregmi/Shortcut-CSL)"},"_bibtex":{"value":"@inproceedings{\nregmi2026shortcut,\ntitle={Shortcut Diffusion Training with Cumulative Consistency Loss: An Optimal Control View},\nauthor={Paribesh Regmi and Sandesh Ghimire and Rui Li},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=cZqAk87Lu4}\n}"},"title":{"value":"Shortcut Diffusion Training with Cumulative Consistency Loss: An Optimal Control View"},"pdf":{"value":"/pdf/1b88b1373e822c88adbcbd6921499915ee5a276b.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"regmi|shortcut_diffusion_training_with_cumulative_consistency_loss_an_optimal_control_view"},"authorids":{"value":["~Paribesh_Regmi1","~Sandesh_Ghimire2","~Rui_Li3"]},"authors":{"value":["Paribesh Regmi","Sandesh Ghimire","Rui Li"]}},"version":2},{"content":{"summary":{"value":"This work deal with sample complexity of Robust MDPs. The major improvement is that, contrary to previous research that relies on generative models or pre-collected datasets, this paper focuses on RMDPs learning through interactive data collection, addressing two key challenges: distributional robustness and balancing exploration and exploitation.\n\nThree main contributions are :\n\n1)  Explaining that sample-efficient learning is unachievable without additional assumptions due to the curse of support shift, where training and testing environments may have non-overlapping distributions.\n2) Introducing of the vanishing minimal value assumption for Robust Markov Decision Processes (RMDPs) with a total-variation distance robust set, assuming the minimal value of the optimal robust value function is zero, leading to a tractable case.\n3)  Proposing an algorithm with a provable sample complexity guarantee under this framework."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"9) Do you think that vanishing minimal assumption is restrictive to derive Deep Robust RL algorithm ? \n10) Do you think it would be possible to adapt Proposition 4.2 to other norms such as $L_p$ ? \n11) To clarify thing, the main difference between generative model setting and online setting in RMDPs is to deal with support shift of kernel ? \n12) I understand that KL and $\\chi^2$ is a slightly different problem, does vanishing minimal assumption for KL or $\\chi^2$ divergences RMDPs make sense as the definition of these divergences already impose bounded support shift of transition kernel ?\n13) From a practical point of view, would the idea that sample complexity of RMDPS is smaller than MDPs (both in online and generative model setting) could lead to more sample efficient algorithms ?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"The strengths of this paper are :\n\n1) The paper is clearly written and ideas are well exposed.\n2) Authors derive a nice lower bound with counterexample on the sample complexity of RMDPs in an online setting, and tight upper bound using extra assumptions such that vanishing minimal value assumption where TV uncertainty set can be rewritten using Radon-Nikodym derivatives.\n3) The proof seems correct for me.\n4) Algorithm is quite classic but make sense to derive robust policy while balancing exploration and exploitation."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"5) It would also be interesting to extend results for $s$- rectangular case.\n6) There is no lower bound with vanishing minimal assumption to ensure that the upper bound under these assumptions is tight.\n7)  It would be nice to add the range of $\\epsilon$ for the upper bound where sample complexity is valid or give  a condition on K rather than saying \"lower order term in K\"\n8) I think it would be interesting to gives more intuition on vanishing minimal assumption."},"limitations":{"value":"No limitations."}},"nonreaders":[],"tmdate":1730879186002,"tcdate":1720372151767,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission7715/Reviewer_bEix"],"signatures":["NeurIPS.cc/2024/Conference/Submission7715/Reviewer_bEix"],"forum":"aYWtfsf3uP","number":1,"license":"CC BY 4.0","cdate":1720372151767,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission7715/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730879186002,"domain":"NeurIPS.cc/2024/Conference","replyto":"aYWtfsf3uP","id":"JSCEOSKy8l","forumContent":{"venue":{"value":"NeurIPS 2024 poster"},"keywords":{"value":["Theory of distributionally robust reinforcement learning","interactive data collection","robust Markov decision process","sample complexity","online regret"]},"primary_area":{"value":"reinforcement_learning"},"abstract":{"value":"The sim-to-real gap, which represents the disparity between training and testing environments, poses a significant challenge in reinforcement learning (RL). A promising approach to addressing this challenge is distributionally robust RL, often framed as a robust Markov decision process (RMDP). In this framework, the objective is to find a robust policy that achieves good performance under the worst-case scenario among all environments within a pre-specified uncertainty set centered around the training environment. Unlike previous work, which relies on a generative model or a pre-collected offline dataset enjoying good coverage of the deployment environment, we tackle robust RL via interactive data collection, where the learner interacts with the training environment only and refines the policy through trial and error. In this robust RL paradigm, two main challenges emerge: managing distributional robustness while striking a balance between exploration and exploitation during data collection. Initially, we establish that sample-efficient learning without additional assumptions is unattainable owing to the curse of support shift; i.e., the potential disjointedness of the distributional supports between the training and testing environments. To circumvent such a hardness result, we introduce the vanishing minimal value assumption to RMDPs with a total-variation (TV) distance robust set, postulating that the minimal value of the optimal robust value function is zero. We prove that such an assumption effectively eliminates the support shift issue for RMDPs with a TV distance robust set, and present an algorithm with a provable sample complexity guarantee. Our work makes the initial step to uncovering the inherent difficulty of robust RL via interactive data collection and sufficient conditions for designing a sample-efficient algorithm accompanied by sharp sample complexity analysis."},"_bibtex":{"value":"@inproceedings{\nlu2024distributionally,\ntitle={Distributionally Robust Reinforcement Learning with Interactive Data Collection: Fundamental Hardness and Near-Optimal Algorithms},\nauthor={Miao Lu and Han Zhong and Tong Zhang and Jose Blanchet},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=aYWtfsf3uP}\n}"},"title":{"value":"Distributionally Robust Reinforcement Learning with Interactive Data Collection: Fundamental Hardness and Near-Optimal Algorithms"},"pdf":{"value":"/pdf/d0f097dc1f14176950a572b7949309caf6e0cd71.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"lu|distributionally_robust_reinforcement_learning_with_interactive_data_collection_fundamental_hardness_and_nearoptimal_algorithms"},"authorids":{"value":["~Miao_Lu3","~Han_Zhong1","~Tong_Zhang2","~Jose_Blanchet1"]},"authors":{"value":["Miao Lu","Han Zhong","Tong Zhang","Jose Blanchet"]}},"version":2},{"content":{"summary":{"value":"This paper aims to overcome the limitations of motion generation based on pixel-level supervision in previous studies by proposing joint modeling of human appearance and motion. \n\nThe authors propose a DiT-based architecture that processes tokens from two modalities. The SMPL parameters are used to represent human poses, and to emphasize the dual-modality nature, Query, Key, and Value extracted from both the video and motion are concatenated and processed through self-attention. This structure enables attention to consider multiple modalities, which is advantageous for joint modeling. A motion-video synchronized RoPE (MVS-RoPE) which is an encoding method applicable to both modalities, is also proposed. Specifically, a diagonal extension is proposed to prevent interference between motion latents and video latents. \n\nIn addition, the authors propose the HuMoVe dataset, containing over 80,000 video-motion pairs. This dataset includes descriptive textual captions, 3D SMPL motion parameters, and video pairs, making it valuable for multi-modal generative modeling that considers vision, text, and motion jointly. \n\nThe experimental results present various metrics and human evaluations, showing performance improvements over baselines. Furthermore, ablation studies for each module are provided to analyze the effectiveness of the proposed methods."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"- p.2 L66: The authors mention that previous works are limited because, even with a 3D prior, supervision is applied after projecting it into 2D, which constrains accurate 3D (motion) generation. However, since the proposed method is also trained through a video diffusion process, hasn’t it still failed to overcome the problem of losing 3D information?\n\n- p.4 L189: Motion tokens are generated as 51 dimension. What is the specific reason for this number?\n\n- Motion tokens are added diagonally to visual tokens. Since maintaining temporal alignment is sufficient, there seems to be no strict reason for using the diagonal arrangement. Is there experimental evidence supporting the \"positional collisions\" mentioned in the text?\n\n- Eq. 6: If the reviewer's understanding is correct, the last term must be u_{\\theta}(\\phi, m_{t}, \\phi)\n\n- Quantitative comparison is provided only as self-evaluation. Although direct quantitative comparison with previous studies may be difficult, joint modeling is expected to enhance the performance of both the video and motion decoders. Therefore, a quantitative comparison between the videos and motions generated by the proposed framework and those produced by conditional generation methods (e.g., VideoJAM), given GT as condition, could better highlight the advantage of joint modeling (even if the performance does not surpass that of conditional generation).\n\n\n- Minor Comments:\n\n-- p.5 L231: i is the motion token -> i is the motion token index ?\n\n-- Fig.3: The distinction between \"noisy\" and \"clean\" is described only in text. It would be clearer and easier to understand if visual symbols were added to indicate the presence or absence of noise."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- The paper proposes the large-scale HuMoVe dataset. Since the dataset includes test captions, videos, and motion parameter pairs, it is highly useful for multi-modal modeling tasks.\n\n- MVS-RoPE that can be jointly applied to visual and motion embeddings is proposed. This encoding technique utilizes diagonal positioning to prevent interference between vision and motion latents, which is a reasonable approach (although more experimental evidence is needed to support this).\n\n- The paper is easy to follow."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The deep network structure is only a simple extension of existing networks. Except for MVS-RoPE, the network mainly uses self-attention on concatenated features for joint modeling, which is quite simple and straightforward. Discussion on whether other components could be improved to better support joint modeling would strengthen the paper.\n\n- The quantitative evaluation relies only on self-evaluation. Even if direct comparison with prior studies is difficult, the paper should include analyses comparing the video and motion decoder performance improved from joint modeling with existing conditional generation methods (e.g., VideoJAM) to show the degree of improvement or equivalence.\n\n- The explanation of how text descriptions were generated for the HuMoVe dataset needs to be clarified. In particular, since the initial data were created using an LLM, the paper should provide more detailed information about the prompts used."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762917253322,"tcdate":1761914983137,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission4250/Reviewer_T7MT"],"signatures":["ICLR.cc/2026/Conference/Submission4250/Reviewer_T7MT"],"forum":"eflUxFmIhZ","number":2,"license":"CC BY 4.0","cdate":1761914983137,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission4250/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762917253322,"domain":"ICLR.cc/2026/Conference","replyto":"eflUxFmIhZ","id":"KTYFJVN3N4","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Video Generation","Human Motion Generation;"]},"supplementary_material":{"value":"/attachment/9e34ac13e1808d69fc38b77f2d39339474d455ae.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Video generation models have advanced significantly, yet they still struggle to synthesize complex human movements due to the high degrees of freedom in human articulation. This limitation stems from the intrinsic constraints of pixel-only training objectives, which inherently bias models toward appearance fidelity at the expense of learning underlying kinematic principles. To address this, we introduce EchoMotion, a framework designed to model the joint distribution of appearance and human motion, thereby improving the quality of complex human action video generation. EchoMotion extends the DiT (Diffusion Transformer) framework with a dual-branch architecture that jointly processes tokens concatenated from different modalities.\nFurthermore, we propose MVS-RoPE (Motion-Video Syncronized RoPE), which offers unified 3D positional encoding for both video and motion tokens. By providing a synchronized coordinate system for the dual-modal latent sequence, MVS-RoPE establishes an inductive bias that fosters temporal alignment between the two modalities. We also propose a Motion-Video Two-Stage Training Strategy. This strategy enables the model to perform both the joint generation of complex human action videos and their corresponding motion sequences, as well as versatile cross-modal conditional generation tasks. \nTo facilitate the training of a model with these capabilities, we construct \\textit{HuMoVe}, a large-scale dataset of approximately 80,000 high-quality, human-centric video-motion pairs.\nOur findings reveal that explicitly representing human motion is complementary to appearance, significantly boosting the coherence and plausibility of human-centric video generation. Project page at: https://yuxiaoyang23.github.io/EchoMotion-webpage/."},"_bibtex":{"value":"@inproceedings{\nyang2026echomotion,\ntitle={EchoMotion: Unified Human Video and Motion Generation via Dual-Modality Diffusion Transformer},\nauthor={Yuxiao Yang and Hualian Sheng and Sijia Cai and Jing Lin and Jiahao Wang and Bing Deng and Junzhe Lu and Haoqian Wang and Jieping Ye},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=eflUxFmIhZ}\n}"},"title":{"value":"EchoMotion: Unified Human Video and Motion Generation via Dual-Modality Diffusion Transformer"},"pdf":{"value":"/pdf/a8c72ba1620961e63e7e3c328f0a1d9444dd1ea8.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"yang|echomotion_unified_human_video_and_motion_generation_via_dualmodality_diffusion_transformer"},"authorids":{"value":["~Yuxiao_Yang2","~Hualian_Sheng1","~Sijia_Cai2","~Jing_Lin3","~Jiahao_Wang14","~Bing_Deng1","~Junzhe_Lu1","~Haoqian_Wang1","~Jieping_Ye4"]},"authors":{"value":["Yuxiao Yang","Hualian Sheng","Sijia Cai","Jing Lin","Jiahao Wang","Bing Deng","Junzhe Lu","Haoqian Wang","Jieping Ye"]}},"version":2},{"content":{"venue":{"value":"ICLR 2027 Conference Submission"},"keywords":{"value":["4D Video Editing","Region-Aware Conditioning","Video Diffusion Models"]},"supplementary_material":{"value":"/attachment/8e198162b9cb92fec563bda78127c184ba6fe1cc.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"While 4D-driven video diffusion models can generate visually plausible videos, faithful 4D editing poses a stricter requirement: source-observed content should be preserved, while disoccluded and out-of-view regions must be synthesized naturally. We identify a key failure mode in existing methods, which we term Evidence-Role Mismatch: source-backed observations, uncertain rendered cues, and unsupported regions are entangled within a single conditioning signal, which can lead to preservation drift, ghosting, and unstable extrapolation. To mitigate this issue, we propose PREX (Preserve, Reveal, Expand), a region-aware framework that decomposes the target spatiotemporal volume into three distinct roles based on observation support and scene extent. PREX traces edited 4D points back to source frames and retrieves observed colors and pairs these appearance cues with geometry-aware confidence and region roles, and injects them into a frozen video diffusion backbone through an adapter trained with self-supervised proxy editing tasks. We further introduce PREBench, a diagnostic benchmark specifically designed for 4D video editing, to expose and quantify such region-specific failures. PREBench separately evaluates Preserve fidelity, Reveal quality, and Expand plausibility of videos generated by 4D driven video diffusion models. Experiments show that PREX substantially reduces region-specific failure modes while maintaining strong visual quality and 4D edit control compared with existing 4D-based video diffusion methods. Code will be publicly released."},"_bibtex":{"value":"@inproceedings{\nanonymous2026preserve,\ntitle={Preserve, Reveal, Expand: Towards Faithful 4D Video Editing with Region-Aware Conditioning and Benchmarking},\nauthor={Anonymous},\nbooktitle={Submitted to The Fifteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=dHsk6g3pRH},\nnote={under review}\n}"},"title":{"value":"Preserve, Reveal, Expand: Towards Faithful 4D Video Editing with Region-Aware Conditioning and Benchmarking"},"pdf":{"value":"/pdf/67a9356d071b28bda5c28190489e2d7c13a84fea.pdf"},"venueid":{"value":"ICLR.cc/2027/Conference/Submission"}},"tmdate":1791222543308,"tcdate":1788430291542,"writers":["ICLR.cc/2027/Conference","ICLR.cc/2027/Conference/Submission8493/Authors"],"signatures":["ICLR.cc/2027/Conference/Submission8493/Authors"],"forum":"dHsk6g3pRH","license":"CC BY 4.0","number":8493,"cdate":1788430291542,"readers":["everyone"],"invitations":["ICLR.cc/2027/Conference/-/Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Bidding","ICLR.cc/2027/Conference/Submission8493/-/Full_Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Reviewing","ICLR.cc/2027/Conference/-/Edit"],"mdate":1791222543308,"odate":1791031328286,"domain":"ICLR.cc/2027/Conference","id":"dHsk6g3pRH","version":2},{"content":{"venue":{"value":"CoRR 2025"},"pdf":{"value":"http://arxiv.org/pdf/2505.01406v1"},"venueid":{"value":"dblp.org/journals/CORR/2025"},"paperhash":{"value":"teymoorianfard|vidstamp_a_temporallyaware_watermark_for_ownership_and_integrity_in_video_diffusion_models"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Mohammadreza_Teymoorianfard:","~Shiqing_Ma1","https://dblp.org/search/pid/api?q=author:Amir_Houmansadr:"]},"html":{"value":"https://doi.org/10.48550/arXiv.2505.01406"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2505-01406,\n  publtype={informal},\n  author={Mohammadreza Teymoorianfard and Shiqing Ma and Amir Houmansadr},\n  title={VIDSTAMP: A Temporally-Aware Watermark for Ownership and Integrity in Video Diffusion Models},\n  year={2025},\n  month={May},\n  cdate={1746057600000},\n  journal={CoRR},\n  volume={abs/2505.01406},\n  url={https://doi.org/10.48550/arXiv.2505.01406}\n}\n"},"abstract":{"value":"The rapid rise of video diffusion models has enabled the generation of highly realistic and temporally coherent videos, raising critical concerns about content authenticity, provenance, and misuse. Existing watermarking approaches, whether passive, post-hoc, or adapted from image-based techniques, often struggle to withstand video-specific manipulations such as frame insertion, dropping, or reordering, and typically degrade visual quality. In this work, we introduce VIDSTAMP, a watermarking framework that embeds per-frame or per-segment messages directly into the latent space of temporally-aware video diffusion models. By fine-tuning the model's decoder through a two-stage pipeline, first on static image datasets to promote spatial message separation, and then on synthesized video sequences to restore temporal consistency, VIDSTAMP learns to embed high-capacity, flexible watermarks with minimal perceptual impact. Leveraging architectural components such as 3D convolutions and temporal attention, our method imposes no additional inference cost and offers better perceptual quality than prior methods, while maintaining comparable robustness against common distortions and tampering. VIDSTAMP embeds 768 bits per video (48 bits per frame) with a bit accuracy of 95.0%, achieves a log P-value of -166.65 (lower is better), and maintains a video quality score of 0.836, comparable to unwatermarked outputs (0.838) and surpassing prior methods in capacity-quality tradeoffs. Code: Code: \\url{https://github.com/SPIN-UMass/VidStamp}"},"title":{"value":"VIDSTAMP: A Temporally-Aware Watermark for Ownership and Integrity in Video Diffusion Models"},"authors":{"value":["Mohammadreza Teymoorianfard","Shiqing Ma","Amir Houmansadr"]}},"tmdate":1757256366561,"pdate":1735689600000,"externalIds":["dblp:journals/corr/abs-2505-01406"],"tcdate":1757256359515,"writers":["~"],"signatures":["~Shiqing_Ma2"],"forum":"AIB2sMaqN7","license":"CC BY-SA 4.0","number":621040,"cdate":1746057600000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1757256366561,"domain":"DBLP.org","id":"AIB2sMaqN7","version":2},{"content":{"summary":{"value":"In pair-based self-supervised representation learning methods (e.g. SimCLR), positive pairs are constructed using data augmentation. \nThis paper proposes to produce many random positive pairs (e.g. 4) for any query image and compute loss over the \"hardest\" one.\nIn a way, it is a hard positive (or view) selection method.\nIt is shown that this technique can be utilized for various methods such as DINO, SimSiam and SimCLR, yielding consistent improvements on ImageNet classification and several transfer learning scenarios."},"presentation":{"value":"2 fair"},"contribution":{"value":"1 poor"},"soundness":{"value":"2 fair"},"strengths":{"value":"The main idea behind the paper is to make the training of pair-based self-supervised methods harder. \nThis is done by producing several random positive pairs and backpropagating gradients through the hardest one.\nIt is shown that this strategy picks pairs that overlap less, hence models learn better representations on ImageNet-1K after being trained the same amount of \"epochs\".\nBetter representations mean consistent improvements on ImageNet-1K classification and various transfer learning experiments including classifcation, detection and segmentation.\n\nOverall, the paper is easy to read."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"I have two main concerns, listed below.\n\n1) As shown in Figure-3, positive pairs used for computing loss overlap less thanks to the proposed method, and this facilitates better representations. Tian 2020b already shows this phenomenon, i.e. there is a sweat spot in the IoU ratio which leads to the \"optimal\" performance. It is not clear what more this paper offers. One distinct angle in this paper is the fact that multiple positive pairs are utilized during training (although loss is computed over only 1 pair in the end). I wonder what would happen if loss was computed over all crops, by simple averaging or weighted averaging (depending on the difficulty of a pair).\n\n2) The computational overhead introduced by the proposed method (\"about a factor of 1.55× for SimSiam\") is a bit unfair for the baselines. I wonder if this extra compute time can be used in favor of other models too, e.g. by training models longer. Cause longer training schedules often bring substantial gains for self-supervised methods. Also, it should be noted that the proposed method has seen more samples (due to encoding multiple pairs) which already impacted for instance batch-norm statistics (although loss is not backpropagated over unused pairs). Then it would be nice to see baselines processed 4x more samples."},"confidence":{"value":"4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work."},"questions":{"value":"I would like the authors to address the concerns I raised in the weaknesses part."},"rating":{"value":"3: reject, not good enough"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699635955458,"tcdate":1698872496817,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission294/Reviewer_seyJ"],"signatures":["ICLR.cc/2024/Conference/Submission294/Reviewer_seyJ"],"forum":"ioBIT7gLBm","number":3,"license":"CC BY 4.0","cdate":1698872496817,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission294/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699635955458,"domain":"ICLR.cc/2024/Conference","replyto":"ioBIT7gLBm","id":"INQ8xtWEfP","forumContent":{"venue":{"value":"Submitted to ICLR 2024"},"TLDR":{"value":"We propose Hard View Selection, a simple new learning method for contrastive learning that exposes the model to harder samples during pretraining and achieves accuracy boosts on ImageNet between 0.55% and 1.9% on DINO, SimSiam, and SimCLR."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Contrastive Learning","Self-Supervised Learning","Pretraining","Data Augmentation"]},"supplementary_material":{"value":"/attachment/b42fb5b660d691af0d0fb36542af1b0dd71f248f.pdf"},"primary_area":{"value":"unsupervised, self-supervised, semi-supervised, and supervised representation learning"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Many Contrastive Learning (CL) methods train their models to be invariant to different \"views\" of an image input for which a good data augmentation pipeline is crucial. While considerable efforts were directed towards improving pre-text tasks, architectures, or robustness (e.g., Siamese networks or teacher-softmax centering), the majority of these methods remain strongly reliant on the random sampling of operations within the image augmentation pipeline, such as the random resized crop or color distortion operation. \nIn this paper, we argue that the role of the view generation and its effect on performance has so far received insufficient attention. To address this, we propose an easy, learning-free, yet powerful Hard View Selection (HVS) strategy designed to extend the random view generation to expose the pretrained model to harder samples during CL training. It encompasses the following iterative steps: 1) randomly sample multiple views and create pairs of two views, 2) run forward passes for each view pair on the currently trained model, 3) adversarially select the pair yielding the worst loss, and 4) run the backward pass with the selected pair.\nIn our empirical analysis we show that under the hood, HVS increases task difficulty by controlling the Intersection over Union of views during pretraining. With only 300-epoch pretraining, HVS is able to closely rival the 800-epoch DINO baseline which remains very favorable even when factoring in the slowdown induced by the additional forwards of HVS. Additionally, HVS consistently achieves accuracy improvements on ImageNet between 0.55% and 1.9% on linear evaluation and similar improvements on transfer tasks across multiple CL methods, such as DINO, SimSiam, and SimCLR."},"_bibtex":{"value":"@misc{\nferreira2024hard,\ntitle={Hard View Selection for Contrastive Learning},\nauthor={Fabio Ferreira and Ivo Rapant and Frank Hutter},\nyear={2024},\nurl={https://openreview.net/forum?id=ioBIT7gLBm}\n}"},"title":{"value":"Hard View Selection for Contrastive Learning"},"pdf":{"value":"/pdf/81f8b8a1842c046b9a4496c2ff4c6df47f419934.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Rejected_Submission"},"paperhash":{"value":"ferreira|hard_view_selection_for_contrastive_learning"},"authorids":{"value":["~Fabio_Ferreira1","~Ivo_Rapant2","~Frank_Hutter1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Fabio Ferreira","Ivo Rapant","Frank Hutter"]}},"version":2},{"content":{"summary":{"value":"This paper introduces Temporal Sentence Grounding with Relevance Feedback (TSG-RF) in videos, a new task that addresses the limitations of traditional Temporal Sentence Grounding (TSG), which assumes relevant segments always exist within a video. TSG-RF accounts for the possibility that a video may not include a segment related to the query, aiming to localize segments that align with the query when present and provide feedback when they are absent. The proposed Relation-aware Temporal Sentence Grounding (RaTSG) network reformulates TSG-RF as a foreground-background detection problem, assessing query-related semantics at both frame and video levels. It utilizes a multi-granularity relevance discriminator for precise relevance feedback and a relation-aware segment grounding module to adaptively ground segments. To validate RaTSG, two popular TSG datasets are reconstructed, establishing a benchmark for TSG-RF. Experimental results demonstrate the effectiveness of RaTSG for this task."},"soundness":{"value":2},"confidence":{"value":5},"questions":{"value":"Please refer to the Weaknesses section."},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"- This paper introduces the Relation-aware Temporal Sentence Grounding (RaTSG) network, which effectively uses a multi-granularity relevance discriminator for predicting relevance feedback and a relation-aware segment grounding module for selectively determining segment boundaries.\n- The paper argument the original datasets for evaluation."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The key motivation of this paper is related to another video temporal task -- highlight detection. However, this paper do not have any discussion.\n- The method is cumbersome, featuring several incremental designs.\n- The experiments utilize outdated datasets. Newer datasets, such as QV-Highlights, should be considered.\n- The baselines are no up to date which do not include sota methods. such as UMT, UniVTG, QD-DETR."},"limitations":{"value":"The Relevance feedback in Video Moment Retrieval is a trivial research question. Either we could generalize this problem to general video grounding (spatial / temporal) or general video understanding (not just VMR, but also hallucinations in VLMs)."}},"nonreaders":[],"tmdate":1730879290502,"tcdate":1720707739113,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission9129/Reviewer_dpih"],"signatures":["NeurIPS.cc/2024/Conference/Submission9129/Reviewer_dpih"],"forum":"eOonmxzzno","number":3,"license":"CC BY 4.0","cdate":1720707739113,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission9129/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730879290502,"domain":"NeurIPS.cc/2024/Conference","replyto":"eOonmxzzno","id":"prPBTmOa0e","forumContent":{"venue":{"value":"NeurIPS 2024 poster"},"keywords":{"value":["Temporal Sentence Grounding","Cross-Modal Retrieval","Vision and Language"]},"supplementary_material":{"value":"/attachment/f56fd469927308412a79f23c823be6bd2210173d.zip"},"primary_area":{"value":"machine_vision"},"abstract":{"value":"As a widely explored multi-modal task, Temporal Sentence Grounding in videos (TSG) endeavors to retrieve a specific video segment matched with a given query text from a video. The traditional paradigm for TSG generally assumes that relevant segments always exist within a given video. However, this assumption is restrictive and unrealistic in real-world applications where the existence of a query-related segment is uncertain, easily resulting in erroneous grounding. Motivated by the research gap and practical application, this paper introduces a new task, named Temporal Sentence Grounding with Relevance Feedback (TSG-RF) in videos, which accommodates the possibility that a video may or may not include a segment related to the query. This task entails localizing precise video segments that semantically align with the query text when such content is present, while delivering definitive feedback on the non-existence of related segments when absent. Moreover, we propose a novel Relation-aware Temporal Sentence Grounding (RaTSG) network for addressing this challenging task. This network first reformulates the TSG-RF task as a foreground-background detection problem by investigating whether the query-related semantics exist in both frame and video levels. Then, a multi-granularity relevance discriminator is exploited to produce precise video-query relevance feedback and a relation-aware segment grounding module is employed to selectively conduct the grounding process, dynamically adapting to the presence or absence of query-related segments in videos. To validate our RaTSG network, we reconstruct two popular TSG datasets, establishing a rigorous benchmark for TSG-RF. Experimental results demonstrate the effectiveness of our proposed RaTSG for the TSG-RF task. Our source code is available at https://github.com/HuiGuanLab/RaTSG."},"_bibtex":{"value":"@inproceedings{\ndong2024temporal,\ntitle={Temporal Sentence Grounding with Relevance Feedback in Videos},\nauthor={Jianfeng Dong and Xiaoman Peng and Daizong Liu and Xiaoye Qu and Xun Yang and Cuizhu Bao and Meng Wang},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=eOonmxzzno}\n}"},"title":{"value":"Temporal Sentence Grounding with Relevance Feedback in Videos"},"pdf":{"value":"/pdf/34afa4783abfa1d87954f9ab2f34323b062c73d9.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"dong|temporal_sentence_grounding_with_relevance_feedback_in_videos"},"authorids":{"value":["~Jianfeng_Dong2","~Xiaoman_Peng1","~Daizong_Liu1","~Xiaoye_Qu1","~Xun_Yang1","~Cuizhu_Bao1","~Meng_Wang3"]},"authors":{"value":["Jianfeng Dong","Xiaoman Peng","Daizong Liu","Xiaoye Qu","Xun Yang","Cuizhu Bao","Meng Wang"]}},"version":2},{"content":{"summary":{"value":"This paper presents Text-Grounded Trajectories (TGT), a framework for controllable text-to-video generation that links localized text descriptions with motion trajectories. The key innovation lies in the Location-Aware Cross-Attention (LACA) module, which effectively fuses spatial trajectory cues with textual semantics, and a dual-CFG strategy that decouples local (object-level) and global (scene-level) guidance."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"See Weaknesses"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- This paper has strong motivation, proposing the first paradigm that guide point-based trajectory controllable video generation with per-trajectory text prompts."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Abuse of notation: In Section 3.1, $L$ is used to denote the length of the encoded text prompt and $N$ is used to denote the token number of the video in latent space. However, in Section 3.2, $L$ is used to denote the token number of the pachified video in latent space and the length of the encoded local text prompt. This is confusing and it needs explanation that why the pachified video and local text prompt have the same token number.\n- Uninformative and inconsistent figures: Figure 2 is uninformative as it does not provide a clear overview of the proposed pipeline. What does the right half of Figure 2 mean? Does it only means that the video latent is flattened in spatial and temporal dimensions? And it is inconsistent with the proposed method since the LACA module incoporates both the local text prompt and global text prompt, but Figure 2 only shows that the LACA module takes the local text prompt as input.\n- Unknown metrics: In the experiments, the authors use EPE while there is no explanation nor reference for it. The authors need to clarify what EPE is and how it is calculated.\n- Missing experiments result: The table 2 only shows the results of baselines and the results of the proposed method are missing."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762917241727,"tcdate":1761832554254,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission4232/Reviewer_bcpG"],"signatures":["ICLR.cc/2026/Conference/Submission4232/Reviewer_bcpG"],"forum":"qUwOlwao20","number":1,"license":"CC BY 4.0","cdate":1761832554254,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission4232/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762917241727,"domain":"ICLR.cc/2026/Conference","replyto":"qUwOlwao20","id":"dqBudL1uhY","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Text-to-Video Generation; Motion Control"]},"supplementary_material":{"value":"/attachment/6e1e7fccbea1bb7e72815737a616744b7ae96035.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Text-to-video generation has advanced rapidly in visual fidelity, whereas standard methods still have limited ability to control the subject composition of generated scenes. Prior work shows that adding localized text control signals, such as bounding boxes or segmentation masks, can help. However, these methods struggle in complex scenarios and degrade in multi-object settings, offering limited precision and lacking a clear correspondence between individual trajectories and visual entities as the number of controllable objects increases. We introduce Text-Grounded Trajectories (TGT), a framework that conditions video generation on trajectories paired with localized text descriptions. We propose $\\textit{Location-Aware Cross-Attention}$ (LACA) to integrate these signals and adopt a dual-CFG scheme to separately modulate local and global text guidance. In addition, we develop a data processing pipeline that produces trajectories with localized descriptions of tracked entities, and we annotate two million high quality video clips to train TGT. Together, these components enable TGT to use point trajectories as intuitive motion handles, pairing each trajectory with text to control both appearance and motion. Extensive experiments show that TGT achieves higher visual quality, more accurate text alignment, and improved motion controllability compared with prior approaches. Website: https://textgroundedtraj.github.io."},"_bibtex":{"value":"@misc{\nzhang2025tgt,\ntitle={{TGT}: Text-Grounded Trajectories for Locally Controlled Video Generation},\nauthor={Guofeng Zhang and Angtian Wang and Jacob Zhiyuan Fang and Liming Jiang and Haotian Yang and Bo Liu and Yiding Yang and Guang Chen and Longyin Wen and Alan Yuille and Chongyang Ma},\nyear={2025},\nurl={https://openreview.net/forum?id=qUwOlwao20}\n}"},"title":{"value":"TGT: Text-Grounded Trajectories for Locally Controlled Video Generation"},"pdf":{"value":"/pdf/bb1243cca895d8c8ed42d7a0e645cf41492a66a2.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"zhang|tgt_textgrounded_trajectories_for_locally_controlled_video_generation"},"authorids":{"value":["~Guofeng_Zhang4","~Angtian_Wang2","~Jacob_Zhiyuan_Fang1","~Liming_Jiang1","~Haotian_Yang1","~Bo_Liu16","~Yiding_Yang1","~Guang_Chen5","~Longyin_Wen1","~Alan_Yuille1","~Chongyang_Ma1"]},"authors":{"value":["Guofeng Zhang","Angtian Wang","Jacob Zhiyuan Fang","Liming Jiang","Haotian Yang","Bo Liu","Yiding Yang","Guang Chen","Longyin Wen","Alan Yuille","Chongyang Ma"]}},"version":2},{"content":{"summary":{"value":"The paper introduces the novel task of Video Connecting, which aims to generate transitional video content that seamlessly links a given start clip and end clip. The authors propose VC-Bench, a new benchmark to evaluate this task, consisting of a 1579 video dataset and a three part evaluation framework (Video Quality Score, Start-End Consistency Score, and Transition Smoothness Score). The paper provides a comprehensive evaluation of several video generation models on this benchmark, identifying current limitations in start-end consistency and transition smoothness."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"- Regarding the human alignment evaluation (Section 5.3): Did the authors just check the correlation with human scores, or did they actively try to make the metrics match human preferences? For instance, how were the weights for the sub-metrics (in VQS, SECS, TSS) decided? Were they tuned to match human scores, or just set by a simple rule, like averaging?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- The paper formalizes the Video Connecting task, which, while related to existing video generation problems, presents a non-trivial challenge. The paper provides a valuable comparison by adapting and evaluating several recent state-of-the-art video generation models for this new task.\n\n- The paper offers a detailed pipeline for the VC-Bench dataset construction and the calculation of the proposed evaluation metrics."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The long-term impact of the benchmark may be limited, as Video Connecting could be viewed as a niche or minor task rather than a foundational problem. There is significant overlap with existing video generation, extension, or interpolation tasks, and few works are specifically dedicated to this problem, which may limit the benchmark's adoption.\n\n- The dataset construction pipeline (e.g., scene detection, clip filtering, captioning) and core evaluation metrics (particularly the Video Quality Score components) are largely adopted from existing benchmarks for text-to-video and image-to-video generation. The newly proposed task-specific metrics, such as the Start-End Consistency Score and Transition Smoothness Score, appear to be straightforward implementations and lack significant novelty.\n\n- The proposed SECS and TSS metrics rely on a direct comparison against the ground-truth video. For a generative task, there are potentially many plausible ways to connect two clips. Relying on ground-truth similarity may unfairly penalize novel or creative, thus making the metrics less reliable for evaluating the true generative capabilities of a model.\n\n- Related to the point above, while the paper distinguishes the VC task from First-Last Frame to Video generation, the evaluation metrics do not seem to fully capture the complexity of ensuring content consistency with the entirety of the start and end clips, instead focusing on pixel-level and optical flow comparisons which are still largely frame-based."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762926653605,"tcdate":1761890379474,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission16574/Reviewer_u86t"],"signatures":["ICLR.cc/2026/Conference/Submission16574/Reviewer_u86t"],"forum":"Ws8HwWHf8N","number":1,"license":"CC BY 4.0","cdate":1761890379474,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission16574/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762926653605,"domain":"ICLR.cc/2026/Conference","replyto":"Ws8HwWHf8N","id":"IwO3f2AeYV","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Video Generation","Video Connecting","VC-Bench"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Current video generation techniques mainly focus on creating under text or image conditioning. However, real-world applications often require seamlessly connecting two independent video clips. To address this, we introduce **Video Connecting**, an innovative task that aims to generate smooth intermediate video content between given start and end clips.\nHowever, the absence of standardized evaluation benchmarks has hindered the development of this task. To bridge this gap, we proposed **VC-Bench**, a novel benchmark specifically designed for video connecting. It includes 1,579 high-quality videos collected from public platforms, covering 15 main categories and 72 subcategories to ensure diversity and structure. VC-Bench focuses on three core aspects: **Video Quality Score** *VQS*, **Start-End Consistency Score** *SECS*, and **Transition Smoothness Score** *TSS*. Together, they form a comprehensive framework that moves beyond conventional quality-only metrics.\nWe evaluated multiple state-of-the-art video generation models on VC-Bench. Experimental results reveal significant limitations in maintaining start-end consistency and transition smoothness, with notable performance gaps among models. We expect that VC-Bench will serve as a pioneering benchmark to inspire and guide future research in video connecting. The evaluation metrics and dataset are publicly available at: https://anonymous.4open.science/r/VC-Bench-1B67/."},"_bibtex":{"value":"@misc{\nyin2026vcbench,\ntitle={{VC}-Bench: Pioneering the Video Connecting Benchmark with a Dataset and Evaluation Metrics},\nauthor={Zhiyu Yin and Zhipeng Liu and Kehai Chen and Lemao Liu and Jin Liu and Hongdong Li and Yang Xiang and Min Zhang},\nyear={2026},\nurl={https://openreview.net/forum?id=Ws8HwWHf8N}\n}"},"title":{"value":"VC-Bench: Pioneering the Video Connecting Benchmark with a Dataset and Evaluation Metrics"},"pdf":{"value":"/pdf/6f9633c337593f1dabcfb4b267d9cf94c5e66b7c.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"yin|vcbench_pioneering_the_video_connecting_benchmark_with_a_dataset_and_evaluation_metrics"},"authorids":{"value":["~Zhiyu_Yin2","~Zhipeng_Liu6","~Kehai_Chen2","~Lemao_Liu3","~Jin_Liu12","~Hongdong_Li4","~Yang_Xiang4","~Min_Zhang9"]},"authors":{"value":["Zhiyu Yin","Zhipeng Liu","Kehai Chen","Lemao Liu","Jin Liu","Hongdong Li","Yang Xiang","Min Zhang"]}},"version":2},{"content":{"summary":{"value":"This work makes attempts to build a multi-view video diffusion model conditioned on reference single-view video and reference multi-view images.  To inherit knowledge learned by previous 3D diffusion models and video diffusion models, the authors carefully design the multi-view video generation network to combine both of them. Specifically, they insert the frame-attention module of the video generation model to 3D generation model to achieve spatial and temporal consistency. To handle GPU memory problem when inferring long-videos, they propose a mixed-sampling scheme. Extensive experiments with regard to novel view video synthesis and 4D generation demonstrate the advantages of the proposed method."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"1) training dataset: I can't find the size of the training dataset in the paper. (Maybe I miss it?) However, as far as I know, many animated objects in objaverse datasets are of low-quality, so I'm curious how many models is used in this work after data filtering. \n2) I'm curious why the authors choose dynamic nerf as the 4D representations instead of dynamic gaussians."},"rating":{"value":6},"details_of_ethics_concerns":{"value":"No"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"1) The work makes attempts to train a mult-view video diffusion model for 4D generation, which is encouraging and could benifit the community.\n 2) The authors test their model on both real-world and synthetic objects to prove the generalization ability of their model.\n 3) The authors conduct extensive compairsons with previous mehtods."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1) network design: According to Figure 2, frame-attention is inserted after view-attention, and the output of frame-attention directly serves as the final result. However, since frame-attention only operates on time axis and is not aware of multi-view information, latents processed by it might suffer from multi-view inconsistency, and cannot directly serve as the final result. It is observed that the 4D reconstruction results suffer from blurry effect, and I suspect it is due to multi-view inconsistency of the generated videos.\n2) problem setting: The proposed model takes generated reference video and generated multi-view images of the first frame as the input. However, it is possible that the two inputs have conflict with each other since they are generated independently. For example, if the video depicts a object is turning around, it may not be consisent with the generated multi-view images espicially in side views and back views. I don't see the authors show objects with such movements in the experiments.\n3) experiments: I'm curious about the real-world video experiements. The authors only train their model on synthetic data, so the generailization to real-world cases might be difficult. I don't find many visualizations of real-world generated objects. Could the authors point out where the visualizations are?\n4) Some citations regarding to concurrent works are not accurate, for example (Line245 Wang et al, 2024a) and (Line 149 4Diffusion Yang et al, 2024). The authors are suggested to further check them.\n5) The authors claim they can handle arbitrary length videos; however, when the input video is very long, the anchor frames will have little temporal coherence and be out-of-distribution for the proposed model."}},"nonreaders":[],"tmdate":1731427767507,"tcdate":1730517190784,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission3209/Reviewer_p3AT"],"signatures":["ICLR.cc/2025/Conference/Submission3209/Reviewer_p3AT"],"forum":"tJoS2d0Onf","number":3,"license":"CC BY 4.0","cdate":1730517190784,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission3209/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427767507,"domain":"ICLR.cc/2025/Conference","replyto":"tJoS2d0Onf","id":"Vrkzl8XRoZ","forumContent":{"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Generative models","4D generation"]},"supplementary_material":{"value":"/attachment/b4e84bd9f074f62bbc85fe923aa7466f705aa3f7.zip"},"primary_area":{"value":"generative models"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"We present Stable Video 4D (SV4D) — a latent video diffusion model for multi-frame and multi-view consistent dynamic 3D content generation. Unlike previous methods that rely on separately trained generative models for video generation and novel view synthesis, we design a unified diffusion model to generate novel view videos of dynamic 3D objects.  Specifically, given a monocular reference video, SV4D generates novel views for each video frame that are temporally consistent. We then use the generated novel view videos to optimize an implicit 4D representation (dynamic NeRF) efficiently, without the need for cumbersome SDS-based optimization used in most prior works. To train our unified novel view video generation model, we curate a dynamic 3D object dataset from the existing Objaverse dataset. Extensive experimental results on multiple datasets and user studies demonstrate SV4D's state-of-the-art performance on novel-view video synthesis as well as 4D generation compared to prior works. Project page: https://sv4d.github.io."},"_bibtex":{"value":"@inproceedings{\nxie2025svd,\ntitle={{SV}4D: Dynamic 3D Content Generation with Multi-Frame and Multi-View Consistency},\nauthor={Yiming Xie and Chun-Han Yao and Vikram Voleti and Huaizu Jiang and Varun Jampani},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=tJoS2d0Onf}\n}"},"title":{"value":"SV4D: Dynamic 3D Content Generation with Multi-Frame and Multi-View Consistency"},"pdf":{"value":"/pdf/b83d91732b4667fd132c57892a41027047a460f7.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"xie|sv4d_dynamic_3d_content_generation_with_multiframe_and_multiview_consistency"},"authorids":{"value":["~Yiming_Xie2","~Chun-Han_Yao1","~Vikram_Voleti1","~Huaizu_Jiang1","~Varun_Jampani2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Yiming Xie","Chun-Han Yao","Vikram Voleti","Huaizu Jiang","Varun Jampani"]}},"version":2},{"content":{"venue":{"value":"ISCAS 2025"},"pdf":{"value":"https://ieeexplore.ieee.org/iel8/11043142/11042930/11044286.pdf"},"venueid":{"value":"dblp.org/conf/ISCAS/2025"},"paperhash":{"value":"liao|ivca_interrelationaware_video_complexity_analyzer"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Junqi_Liao:","https://dblp.org/search/pid/api?q=author:Yao_Li_0016:","~Zhuoyuan_Li2","https://dblp.org/search/pid/api?q=author:Li_Li_0040:","https://dblp.org/search/pid/api?q=author:Dong_Liu_0002:"]},"html":{"value":"https://doi.org/10.1109/ISCAS56072.2025.11044286"},"_bibtex":{"value":"@inproceedings{DBLP:conf/iscas/LiaoLLLL25,\n  author={Junqi Liao and Yao Li and Zhuoyuan Li and Li Li and Dong Liu},\n  title={IVCA: Inter-relation-aware Video Complexity Analyzer},\n  year={2025},\n  cdate={1735689600000},\n  pages={1-5},\n  url={https://doi.org/10.1109/ISCAS56072.2025.11044286},\n  booktitle={ISCAS},\n  crossref={conf/iscas/2025}\n}\n"},"abstract":{"value":"To address the real-time analysis requirements of video streaming applications, we propose an innovative inter-relation-aware video complexity analyzer (IVCA) to enhance the existing video complexity analyzer (VCA). The IVCA overcomes the limitations of the VCA by incorporating inter-frame relations, focusing on inter motion and reference structure. To begin with, we improve the accuracy of temporal features by integrating feature-domain motion estimation into the IVCA framework, which allows for a more nuanced understanding of motion across frames. Furthermore, inspired by the hierarchical reference structures utilized in modern codecs, we introduce layer-aware weights that effectively adjust the contributions of frame complexity across different layers, ensuring a more balanced representation of video characteristics. In addition, we broaden the analysis of temporal features by considering reference frames rather than relying solely on the preceding frame, thereby enriching the contextual understanding of video content. Experimental results demonstrate a significant enhancement in complexity estimation accuracy achieved by the IVCA, coupled with a negligible increase in time complexity, indicating its potential for real-time applications in video streaming scenarios. This advancement not only improves video processing efficiency but also paves the way for more sophisticated analytical tools in video technology."},"title":{"value":"IVCA: Inter-relation-aware Video Complexity Analyzer"},"authors":{"value":["Junqi Liao","Yao Li","Zhuoyuan Li","Li Li","Dong Liu"]}},"tmdate":1757877958067,"pdate":1735689600000,"externalIds":["dblp:conf/iscas/LiaoLLLL25"],"tcdate":1757877956382,"writers":["~"],"signatures":["~Zhuoyuan_Li2"],"forum":"dSPLA4UsiA","license":"CC BY-SA 4.0","number":622503,"cdate":1735689600000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1757877958067,"domain":"DBLP.org","id":"dSPLA4UsiA","version":2},{"content":{"summary":{"value":"This paper presents StereoCrafter-Zero, a zero-shot stereo video generation framework that synthesizes consistent stereo video pairs (left-right views) from a single image and text prompt, without requiring stereo training data. Previous work only involved generating stereo images, whereas this paper extends to video for the first time. The method builds upon video diffusion priors (e.g., DynamiCrafter) and introduces three core innovations:\n1.\tNoisy Restart — a latent initialization and controlled noise injection strategy to improve temporal and inter-view coherence;\n2.\tIterative Refinement — repeated re-denoising of occluded regions to harmonize latent representations;\n3.\tDissolved Depth Maps — low-frequency depth representations designed to enhance latent-space warping stability."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"The authors compare against stereo conversion methods but not (or insufficiently) against established stereo generation or stereo-view synthesis benchmarks. Could the authors report results on one or more standard stereo datasets (or convert to them) to allow more objective quantitative comparison?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1.\tThe stereo video generation task remains underexplored yet highly relevant, addressing the growing demand for VR and immersive content creation. The paper provides a clear motivation for the necessity of joint spatial-temporal-stereo consistency, extending beyond previous single-view or 2D-to-3D conversion studies. The authors creatively identify a novel research problem and propose a training-free solution through StereoCrafter-Zero.\n2.\tTechnical innovations are clear. The Noisy Restart strategy is conceptually elegant, reusing the diffusion process' stochasticity for structural stabilization. The Dissolved Depth Map idea is intuitive yet effective-reducing high-frequency depth noise aligns well with the latent-space nature of diffusion models. Iterative Refinement is a simple but practical scheme to correct occluded regions without excessive computational overhead."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1.\tWhile the stereo video generation task is inherently complex, most evaluation metrics are adapted from standard video generation benchmarks, which lack stereo-specific evaluation criteria. Currently, only subjective user studies demonstrate the performance of StereoCrafter-Zero, which limits the objectivity of the validation. It is recommended to conduct quantitative comparisons with stereo diffusion models on established stereo conversion benchmarks to substantiate the approach further.\n2.\tThe generated results exhibit certain artifacts. For example, in the teaser video, the left hand of Wukong shows spatial inconsistency, and the butterflies display temporal incoherence between left-right views. These issues highlight the challenges in maintaining stereo consistency for training-free frameworks."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762916955115,"tcdate":1761890620365,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission3739/Reviewer_4dv8"],"signatures":["ICLR.cc/2026/Conference/Submission3739/Reviewer_4dv8"],"forum":"gE29pT2T3e","number":1,"license":"CC BY 4.0","cdate":1761890620365,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission3739/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762916955115,"domain":"ICLR.cc/2026/Conference","replyto":"gE29pT2T3e","id":"xgYTQCm8xF","forumContent":{"venue":{"value":"ICLR 2026 Conference Desk Rejected Submission"},"keywords":{"value":["Stereo Synthesis","Video Generation"]},"supplementary_material":{"value":"/attachment/dec0359d1ce2be48f844b0f1aa2fe0fb85e7cff0.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Generating high-quality stereo videos requires consistent depth perception and temporal coherence across frames. Despite advances in image and video synthesis using diffusion models, producing high-quality stereo videos remains a challenging task due to the difficulty of maintaining consistent temporal and spatial coherence between left and right views.\nWe introduce \\textit{StereoCrafter-Zero}, a novel framework for zero-shot stereo video generation that leverages video diffusion priors without requiring paired training data. Our key innovations include a noisy restart strategy to initialize stereo-aware latent representations and an iterative refinement process that progressively harmonizes the latent space, addressing issues like temporal flickering and view inconsistencies.\nIn addition, we propose the use of dissolved depth maps to streamline latent space operations by reducing high-frequency depth information.\nOur comprehensive evaluations, including quantitative metrics and user studies, demonstrate that \\textit{StereoCrafter-Zero} produces high-quality stereo videos with enhanced depth consistency and temporal smoothness. In terms of epipolar consistency, our method achieves an $11.7\\%$ improvement in MEt3R score over the current state-of-the-art. Furthermore, user studies indicate strong perceptual gains over the previous arts, with an $8.0\\%$ higher perceived frame quality and $10.9\\%$ higher perceived temporal coherence.\nOur code will be made publicly available upon acceptance of this manuscript."},"_bibtex":{"value":"@misc{\nanonymous2026stereocrafterzero,\ntitle={StereoCrafter-Zero: Zero-Shot Stereo Video Generation with Noisy Restart},\nauthor={Anonymous},\nyear={2026},\nurl={https://openreview.net/forum?id=gE29pT2T3e}\n}"},"title":{"value":"StereoCrafter-Zero: Zero-Shot Stereo Video Generation with Noisy Restart"},"pdf":{"value":"/pdf/70b020464e4b00247fb19a943acd74feed8866e9.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Desk_Rejected_Submission"},"paperhash":{"value":"shi|stereocrafterzero_zeroshot_stereo_video_generation_with_noisy_restart"},"authorids":{"value":["~Jian_Shi4","~Qian_Wang19","~Zhenyu_Li3","~Ramzi_Idoughi1","~Wenqing_Cui1","~Peter_Wonka1"]},"authors":{"value":["Jian Shi","Qian Wang","Zhenyu Li","Ramzi Idoughi","Wenqing Cui","Peter Wonka"]}},"version":2},{"content":{"summary":{"value":"This paper tackles the central trade-off in personalized video generation between identity consistency and action realism. The authors propose ReactID, a framework that coordinates advances in data curation, training strategy, and action conditioning to balance identity fidelity and motion naturalness. They introduce ReactID-Data, a large-scale dataset built via a high-precision pipeline. To stabilize learning, they analyze difficulty along axes such as subject size, appearance similarity, and sampling, and adopt a progressive curriculum from easy to hard to mitigate identity overfitting and copy-paste artifacts. For action modeling, they propose a timeline-based conditioning scheme that augments text prompts with structured multi-action sequences annotated with timestamps, integrated through two components: a subject-aware cross-attention module and a temporally-adaptive RoPE mechanism. Experiments report state-of-the-art performance on both identity preservation and action realism, suggesting the method effectively balances these competing objectives."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"See Weaknesses"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The manuscript is clearly written and easy to follow.\n2. The motivation is well defined, pinpointing three key challenges in video customization: inaccurate identity preservation, unstable convergence, and compromised action naturalness.\n3. The work is solid, introducing a new dataset and method components, including Subject-Aware Cross-Attention and Temporally-Adaptive RoPE.\n4. Extensive experiments convincingly demonstrate the effectiveness of the approach."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Will the dataset, code, and model weights be released? Open-sourcing would significantly benefit community research and reproducibility.\n2. How well does the method generalize to multi-subject customization beyond two subjects (e.g., three or more)?\n3. For multi-subject scenarios, can the approach handle more complex, independent action timelines per subject (e.g., multiple subjects performing distinct actions concurrently)?\n4. The results shown in the paper are about human subjects. Can the method support customizing animals and general objects?\n5. Please consider citing closely related work:\n    - CustomCrafter: Customized Video Generation with Preserving Motion and Concept Composition Abilities\n    - ReVersion: Diffusion-Based Relation Inversion from Images\n    - DreamRelation: Relation-Centric Video Customization"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762928354681,"tcdate":1761984860887,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission18647/Reviewer_PVcC"],"signatures":["ICLR.cc/2026/Conference/Submission18647/Reviewer_PVcC"],"forum":"yn0Wu7NsTa","number":4,"license":"CC BY 4.0","cdate":1761984860887,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission18647/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762928354681,"domain":"ICLR.cc/2026/Conference","replyto":"yn0Wu7NsTa","id":"2cr2xXj87S","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Video Generation; Identity Preserving; Diffusion Models"]},"primary_area":{"value":"generative models"},"abstract":{"value":"Personalized video generation faces a fundamental trade-off between identity consistency and action realism: overly rigid identity preservation often leads to unnatural motion, while emphasis on action dynamics can compromise subject fidelity. This tension stems from three interrelated challenges: imprecise subject-video alignment, unstable training due to varying sample difficulties, and inadequate modeling of fine-grained actions. To address this, we propose ReactID, a comprehensive framework that harmonizes identity accuracy and motion naturalness through coordinated advances in data, training, and action modeling. First, we construct ReactID-Data, a large-scale dataset annotated with a high-precision pipeline combining vision-based entity label extraction, MLLM-based subject detection, and post-verification to ensure reliable subject-video correspondence. Second, we analyze learning difficulty along dimensions such as subject size, appearance similarity, and sampling strategy, and devise a progressive training curriculum that evolves from easy to hard samples, ensuring stable convergence while avoiding identity overfitting and copy-paste artifacts. Third, ReactID introduces a novel timeline-based conditioning mechanism that supplements monolithic text prompts with structured multi-action sequences. Each sub-action is annotated with precise timestamps and descriptions, and integrated into the diffusion model via two novel components: subject-aware cross-attention module to bind sub-action to the specific subject of interest and temporally-adaptive RoPE to embed the rescaled temporal coordinates invariant to action duration. Experiments show that ReactID achieves state-of-the-art performance in both identity preservation and action realism, effectively balancing the two objectives."},"_bibtex":{"value":"@inproceedings{\nli2026reactid,\ntitle={React{ID}: Synchronizing Realistic Actions and Identity in Personalized Video Generation},\nauthor={Wei Li and Yiheng Zhang and Fuchen Long and Zhaofan Qiu and Ting Yao and Xiaoyan Sun and Tao Mei},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=yn0Wu7NsTa}\n}"},"title":{"value":"ReactID: Synchronizing Realistic Actions and Identity in Personalized Video Generation"},"pdf":{"value":"/pdf/7df250483d4eeea7bd9934f739b47f0bc2db55ab.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"li|reactid_synchronizing_realistic_actions_and_identity_in_personalized_video_generation"},"authorids":{"value":["~Wei_Li111","~Yiheng_Zhang1","~Fuchen_Long1","~Zhaofan_Qiu2","~Ting_Yao1","~Xiaoyan_Sun1","~Tao_Mei3"]},"authors":{"value":["Wei Li","Yiheng Zhang","Fuchen Long","Zhaofan Qiu","Ting Yao","Xiaoyan Sun","Tao Mei"]}},"version":2},{"content":{"venue":{"value":"ICLR 2027 Conference Submission"},"keywords":{"value":["Autoregressive Video Generation","KV Cache"]},"primary_area":{"value":"generative models"},"abstract":{"value":"Autoregressive (AR) video generation has emerged as a promising paradigm for long-horizon video synthesis, where each frame is generated conditioned on previously generated tokens. To accelerate inference, the KV cache is used to avoid redundant recomputation across generation steps. Nevertheless, its growth with generation length introduces increasing memory and error accumulation, limiting the scalability of AR models to even longer sequences. Existing KV cache compression methods mitigate this issue by selectively retaining only video tokens deemed important. However, most existing methods assess token importance using short-horizon signals derived from the current or historical generation context, making these methods prone to overlooking tokens that appear unimportant at early steps but later become critical for future frames. In this work, we identify an important property of trained AR video models: although RoPE-modulated queries evolve across autoregressive steps, the underlying canonical pre-RoPE query distribution remains remarkably stable throughout the video generation process. This approximate stationarity implies that future query distributions are estimable from historical statistics, enabling principled future-aware cache decisions without any additional training. Building on this insight, we propose Future Forcing, a training-free future-aware KV cache policy for AR video generation. Specifically, Future Forcing first constructs a future query proxy from historical statistics, then scores KV cache tokens by their importance under this proxy, and finally merges redundant token pairs within the affine subspace induced by the future query. Extensive experiments show that Future Forcing improves long-horizon consistency under limited KV caches, achieving up to 1.49 improvement in subject consistency on VBench-Long for 60s generation over existing AR video KV cache policies."},"_bibtex":{"value":"@inproceedings{\nanonymous2026future,\ntitle={Future Forcing: Future-aware Training-free {KV} Cache Policy for Autoregressive Video Generation},\nauthor={Anonymous},\nbooktitle={Submitted to The Fifteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=z6I3thBB6N},\nnote={under review}\n}"},"title":{"value":"Future Forcing: Future-aware Training-free KV Cache Policy for Autoregressive Video Generation"},"pdf":{"value":"/pdf/4f0b8e97a528c9164fa7622f0fe46698cb1e8fd1.pdf"},"venueid":{"value":"ICLR.cc/2027/Conference/Submission"}},"tmdate":1791227537565,"tcdate":1789311623713,"writers":["ICLR.cc/2027/Conference","ICLR.cc/2027/Conference/Submission16163/Authors"],"signatures":["ICLR.cc/2027/Conference/Submission16163/Authors"],"forum":"z6I3thBB6N","license":"CC BY 4.0","number":16163,"cdate":1789311623713,"readers":["everyone"],"invitations":["ICLR.cc/2027/Conference/-/Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Bidding","ICLR.cc/2027/Conference/Submission16163/-/Full_Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Reviewing","ICLR.cc/2027/Conference/-/Edit"],"mdate":1791227537565,"odate":1791031328286,"domain":"ICLR.cc/2027/Conference","id":"z6I3thBB6N","version":2},{"content":{"venue":{"value":"CoRR 2026"},"pdf":{"value":"https://arxiv.org/pdf/2605.30083v1"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"luo|future_forcing_futureaware_trainingfree_kv_cache_policy_for_autoregressive_video_generation"},"html":{"value":"https://doi.org/10.48550/arXiv.2605.30083"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2605-30083,\n  publtype={informal},\n  author={Jiayi Luo and Qiyan Liu and Tengyang Wang and JunHao Liu and Jiayu Chen and Cong Wang and Hanxin Zhu and Chen Gao and Xiaobin Hu and Qingyun Sun and Zhibo Chen},\n  title={Future Forcing: Future-aware Training-free KV Cache Policy for Autoregressive Video Generation},\n  year={2026},\n  month={May},\n  cdate={1777593600000},\n  journal={CoRR},\n  volume={abs/2605.30083},\n  url={https://doi.org/10.48550/arXiv.2605.30083}\n}\n"},"abstract":{"value":"Autoregressive (AR) video generation has emerged as a promising paradigm for long-horizon video synthesis, where each frame is generated conditioned on previously generated tokens. To accelerate inference, the KV cache is used to avoid redundant recomputation across generation steps. Nevertheless, its growth with generation length introduces increasing memory and error accumulation, limiting the scalability of AR models to even longer sequences. Existing KV cache compression methods mitigate this issue by selectively retaining only video tokens deemed important. However, most existing methods assess token importance using short-horizon signals derived from the current or historical generation context, making these methods prone to overlooking tokens that appear unimportant at early steps but later become critical for future frames. In this work, we identify an important property of trained AR video models: although RoPE-modulated queries evolve across autoregressive steps, the underlying canonical pre-RoPE query distribution remains remarkably stable throughout the video generation process. This approximate stationarity implies that future query distributions are estimable from historical statistics, enabling principled future-aware cache decisions without any additional training. Building on this insight, we propose Future Forcing, a training-free future-aware KV cache policy for AR video generation. Specifically, Future Forcing first constructs a future query proxy from historical statistics, then scores KV cache tokens by their importance under this proxy, and finally merges redundant token pairs within the affine subspace induced by the future query. Extensive experiments show that Future Forcing improves long-horizon consistency under limited KV caches, achieving up to 1.49 improvement in subject consistency on VBench-Long for 60s generation over existing AR video KV cache policies."},"title":{"value":"Future Forcing: Future-aware Training-free KV Cache Policy for Autoregressive Video Generation"},"authors":{"value":[{"fullname":"Jiayi Luo","username":"~Jiayi_Luo1"},{"fullname":"Qiyan Liu","username":""},{"fullname":"Tengyang Wang","username":""},{"fullname":"JunHao Liu","username":""},{"fullname":"Jiayu Chen","username":""},{"fullname":"Cong Wang","username":""},{"fullname":"Hanxin Zhu","username":""},{"fullname":"Chen Gao","username":""},{"fullname":"Xiaobin Hu","username":""},{"fullname":"Qingyun Sun","username":"~Qingyun_Sun2"},{"fullname":"Zhibo Chen","username":""}]}},"tmdate":1784621874884,"pdate":1798675200000,"externalIds":["dblp:journals/corr/abs-2605-30083"],"tcdate":1784621614373,"writers":["~","OpenReview.net/Public_Article/DBLP.org","OpenReview.net/Support"],"signatures":["~Jiayi_Luo1"],"forum":"B5m49VhXxu","license":"CC BY-SA 4.0","number":80739,"cdate":1777593600000,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/DBLP.org/-/Record","OpenReview.net/Public_Article/-/Edit","OpenReview.net/Public_Article/-/Authorship_Claim"],"mdate":1784621874884,"domain":"OpenReview.net/Public_Article","id":"B5m49VhXxu","version":2},{"content":{"summary":{"value":"This paper introduces an online dense video captioning approach that dynamically retrieves and integrates relevant action phrases processing video segments incrementally, combining both text prefix and embedding fusion strategies. In addition, the paper proposes an image-based video pretraining method that reduces reliance on video datasets while achieving SOTA performance across multiple benchmarks."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"- Why use CLIP embeddings vs other embedding approaches? Have you tried VLMs?\n- What's the optimal trade-off between action phrase simplicity and informativeness? Can we include multiple actions and/or multiple objects?\n- How sensitive is the model to segment length?\n- What's the process to find the optimal ratio between augmented/non-augmented training? and How does this affect model convergence?\n- What's the impact of action phrase quality? Have you done any study?\n- What causes temporal misalignment? Can you provide some examples?"},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- The paper introduces a novel online retrieval-augmented approach for dense video captioning.\n- The use of dynamic action phrase retrieval and integration is well-motivated.\n- The image-based simulated video pretraining approach reduces reliance on video datasets."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"* The paper lacks a strong theoretical justification for why action phrase retrieval works better in comparison to just full captions.\n\n* The comparison with Streaming (Zhou et al., 2024) may not be completely fair as it focuses on memory management and not semantics.\n\n* The approach relies on precomputed embeddings for the action phrases.\n\n* Lines 110-111: \"Unlike existing methods that use either text prefixing or embedding fusion...\" lack concrete explanation.\n\n* Lines 112-113: \"We simplify action phrases into compact forms, such as action-object pairs...\" need further elaboration on failure cases e.g. :\n\n 1. Actions that can't be reduced to simple action-object pairs:\n    \"gradually stirring while slowly pouring in liquid\",\n    \"adjusting equipment settings while monitoring readings\"\n\n 2. Actions requiring temporal context:\n    \"continuing to mix until consistency changes\"\n\n\n 3. Actions with multiple objects or relationships:\n    \"transferring contents from bowl to pan\"\n* I am not sure about the training stability here as switching between modes i.e. Action retrieval-augmented training and \nstandard training happens."}},"nonreaders":[],"tmdate":1731427794708,"tcdate":1730720636633,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission3294/Reviewer_KcsL"],"signatures":["ICLR.cc/2025/Conference/Submission3294/Reviewer_KcsL"],"forum":"oO3oXJ19Pb","number":4,"license":"CC BY 4.0","cdate":1730720636633,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission3294/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427794708,"domain":"ICLR.cc/2025/Conference","replyto":"oO3oXJ19Pb","id":"RHTNOsxP0N","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Dense video captioning","Online dense video captioning"]},"supplementary_material":{"value":"/attachment/22be9efeac9199c2db56442e763f88fbc635769f.pdf"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Dense video captioning requires solving the challenging tasks of temporally localizing events and generating descriptive captions within long video sequences. Existing methods often struggle to capture the evolving context within video streams and to produce accurate temporal alignment. To address this, we propose an online retrieval-augmented approach that processes video segments incrementally while dynamically retrieving relevant action phrases from a pre-constructed action-text corpus. This enriches the contextual information for both the video representation and the subsequent text decoder, improving the caption generation. Additionally, we present image-based simulated video pretraining, which mitigates the reliance on extensive video datasets by using image-level text-paired data aligned with the online video captioning format. Our experiments on the ViTT, YouCook2, and ActivityNet benchmarks demonstrate that our model significantly outperforms both existing global and online methods, validating its effectiveness."},"_bibtex":{"value":"@misc{\nkim2024actions,\ntitle={Actions Inspire Every Moment: Online Action-Augmented Dense Video Captioning},\nauthor={Dahun Kim and AJ Piergiovanni and Anelia Angelova},\nyear={2024},\nurl={https://openreview.net/forum?id=oO3oXJ19Pb}\n}"},"title":{"value":"Actions Inspire Every Moment: Online Action-Augmented Dense Video Captioning"},"pdf":{"value":"/pdf/ced2509c52424b1ede3d852681c4a02933dc58a7.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"kim|actions_inspire_every_moment_online_actionaugmented_dense_video_captioning"},"authorids":{"value":["~Dahun_Kim1","~AJ_Piergiovanni1","~Anelia_Angelova1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Dahun Kim","AJ Piergiovanni","Anelia Angelova"]}},"version":2},{"content":{"venue":{"value":"ICLR 2027 Conference Submission"},"keywords":{"value":["Image Segmentation","Machine Unlearning","Shortcut Learning"]},"primary_area":{"value":"alignment, fairness, safety, privacy, and societal considerations"},"abstract":{"value":"Many medical AI imaging applications rely on segmentation-classification pipelines.\nIn practice, coarse or imprecise segmentation masks often admit *spurious* background cues\nand the model learns the associations between these cues and the pathological regions of interest (organs, lesions, etc.).\nSuch *shortcuts* tend to severely affect the accuracy of the downstream task.\nPrior art relies on fine-grained pixel-level annotations to separate background cues from pathology regions.\nUnfortunately, high-quality annotations are expensive.\nWe aim to ensure high performance on downstream tasks without relying on expensive annotations.\nWe propose using machine unlearning to mitigate shortcut learning.\nIn contrast to relying on heuristic methods, we propose a principled approach that, for the first time, provides guarantees for shortcut unlearning.\nWe develop a definition of certified unlearning for segmentation tasks and formally connect it to its canonical definition.\nFurther, we introduce an information-theoretic objective function, *global information reduction*, and prove that any $(\\varepsilon,\\delta)$-certified machine unlearning operator bounds it relative to the fine-mask reference. Importantly, the bound further limits the accuracy any decision rule can extract from the artefact. We discover existing certified operators collapse at the scale of real medical decoders, and introduce a new machine unlearning algorithm that challenges the SOTA in segmentation at tight $(\\varepsilon,\\delta)$ budgets. \nOn ISIC melanoma trap sets and SIIM-ACR pneumothorax with a naturally occurring chest-tube shortcut, we compare unlearning with state-of-the-art refinement methods on segmentation, shortcut robustness, fine-mask budget and compute. Unlearning spans the Pareto front on both datasets, no refinement method reaches it, and the certified operator is ahead of exact retraining, the only other method with a guarantee, at every fine-mask budget."},"_bibtex":{"value":"@inproceedings{\nanonymous2026towards,\ntitle={Towards Certified Shortcut Unlearning in Medical Imaging},\nauthor={Anonymous},\nbooktitle={Submitted to The Fifteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=zNncWjK2hJ},\nnote={under review}\n}"},"title":{"value":"Towards Certified Shortcut Unlearning in Medical Imaging"},"pdf":{"value":"/pdf/5964273d4ca787ff31932ac73a148439f11fedc8.pdf"},"venueid":{"value":"ICLR.cc/2027/Conference/Submission"}},"tmdate":1791231604767,"tcdate":1789660208023,"writers":["ICLR.cc/2027/Conference","ICLR.cc/2027/Conference/Submission32841/Authors"],"signatures":["ICLR.cc/2027/Conference/Submission32841/Authors"],"forum":"zNncWjK2hJ","license":"CC BY 4.0","number":32841,"cdate":1789660208023,"readers":["everyone"],"invitations":["ICLR.cc/2027/Conference/-/Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Bidding","ICLR.cc/2027/Conference/Submission32841/-/Full_Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Reviewing","ICLR.cc/2027/Conference/-/Edit"],"mdate":1791231604767,"odate":1791031328286,"domain":"ICLR.cc/2027/Conference","id":"zNncWjK2hJ","version":2},{"content":{"summary":{"value":"This paper proposes RIVAL, a video understanding framework built on small open-source LLMs (≤72B), aiming to rival GPT-4-level proprietary methods. RIVAL consists of (1) MSRP (Multi-stage ReAct Planner) for structured reasoning with explicit sub-states (OBSERVE → THINK → ACT) and tool-calling, and (2) MADR (Multi-Agent Debate Refinement) for adversarial multi-role answer refinement. The system retrieves key frames via CLIP and performs iterative information augmentation plus debate-based correction. Experiments on EgoSchema and Next-QA show strong results, surpassing GPT-4 baselines on subsets, and showing robustness on extremely long (28h) concatenated video."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1. Can the authors include ablations isolating (a) no MSRP, (b) no MADR, (c) no CLIP key-frame retrieval, to quantify contribution of each component?\n2. Could the authors report results on free-form open-ended summarization tasks to illustrate generality beyond MCQ style QA?\n3. Given that the video is often reduced to textual descriptions, does RIVAL degrade on videos with non-linguistically describable cues (e.g., spatial geometry, implicit physics)?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. Strong empirical results. RIVAL achieves substantial gains over prior GPT-4-based VideoAgent/LLoVi on EgoSchema subset (+6.6%) and competitive Next-QA performance.\n\n2. Multi-agent debate refinement is effective and well-motivated. MADR empirically corrects initial errors and is demonstrated clearly with case study.\n\n3. Very long video case study is interesting. Handling 28h concatenated input with minimal degradation is a good stress test."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Clarity & ablations missing. The current writing does not sufficiently quantify how much performance comes from CLIP retrieval, MSRP decomposition, and MADR debate individually. Ablation will greatly strengthen causal attribution.\n\n2. Significant engineering heuristics. Many parts of MSRP are manually structured and rely on prompt templates / tool definitions — unclear robustness to domain shift or tasks not fitting stepwise logic.\n\n3. Scalability beyond QA not validated. RIVAL is only evaluated on video QA benchmarks; unclear if this paradigm generalizes to open-ended summarization / event boundary detection / reasoning beyond MCQ.\n\n4. Some baselines may not be strictly comparable. For Next-QA, several older entries are pre-CLIP/2024-era; more recent strong open models could be added for fairness."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762925491780,"tcdate":1761991534922,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission15188/Reviewer_DJH9"],"signatures":["ICLR.cc/2026/Conference/Submission15188/Reviewer_DJH9"],"forum":"mlgRKaosrj","number":4,"license":"CC BY 4.0","cdate":1761991534922,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission15188/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762925491780,"domain":"ICLR.cc/2026/Conference","replyto":"mlgRKaosrj","id":"xnbQH25yra","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Video Understanding","Large Language Model","Agent"]},"primary_area":{"value":"foundation or frontier models, including LLMs"},"abstract":{"value":"The rapid development of large language models (LLMs) has brought new perspectives to the field of video understanding. However, existing methods often rely on large-scale proprietary models, such as GPT-4, to achieve competitive performance. This paper challenges the notion that scale is the primary driver of capability by introducing RIVAL, a framework demonstrating how multi-agent collaboration enables smaller open-source models (72B or fewer) to rival their large-scale counterparts. RIVAL consists of two key components: a Multi-stage React Planner (MSRP) for structured stepwise reasoning and Multi-agent Debate Refinement (MADR) for collaborative answer generation. MSRP enhances instruction-following through precise control, while MADR improves answer quality via multi-perspective debate. Using a 72B model, our framework sets a new state-of-the-art on the EgoSchema subset with 66.8\\% accuracy, surpassing prior GPT-4 based methods by 6.6\\%. Furthermore, we demonstrate that even smaller open-source models (0.6B to 32B) across the Qwen 2.5 and 3 series achieve competitive performance with RIVAL. We also demonstrate competitive performance on the Next QA benchmark. Highlighting its efficiency, RIVAL can process over 28 hours of continuous video input using limited computational resources."},"_bibtex":{"value":"@misc{\nxi2026rethinking,\ntitle={Rethinking Scale: How Multi-Agent Collaboration Enables Smaller Models to Rival {GPT}-4 in Video Understanding},\nauthor={Xing Xi and P.C.Wen and Yushe Cao and Yu Cheng and Weiqiang Wang and Xing Fu and Yicheng Lu and Ronghua Luo},\nyear={2026},\nurl={https://openreview.net/forum?id=mlgRKaosrj}\n}"},"title":{"value":"Rethinking Scale: How Multi-Agent Collaboration Enables Smaller Models to Rival GPT-4 in Video Understanding"},"pdf":{"value":"/pdf/fab2f34137ce23f8b0b69b7dd48b7f4466109bf3.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"xi|rethinking_scale_how_multiagent_collaboration_enables_smaller_models_to_rival_gpt4_in_video_understanding"},"authorids":{"value":["~Xing_Xi1","~P.C.Wen1","~Yushe_Cao1","~Yu_Cheng15","~Weiqiang_Wang4","~Xing_Fu1","~Yicheng_Lu5","~Ronghua_Luo1"]},"authors":{"value":["Xing Xi","P.C.Wen","Yushe Cao","Yu Cheng","Weiqiang Wang","Xing Fu","Yicheng Lu","Ronghua Luo"]}},"version":2},{"content":{"summary":{"value":"The Learning Latent Causal Processes (LLCP) framework introduces a novel approach to Video Question Answering (VideoQA) by focusing on causal reasoning rather than traditional cross-modality matching. LLCP utilizes a multivariate generative model to analyze spatial-temporal dynamics and trains through self-supervised local auto-regression, thus eliminating the need for annotated question-answer pairs. It adeptly handles accident attribution and counterfactual prediction tasks, identifying root causes and forecasting potential outcomes through modifications in variable embeddings."},"presentation":{"value":"3 good"},"contribution":{"value":"2 fair"},"soundness":{"value":"2 fair"},"strengths":{"value":"1. The idea of learning causality from video is interesting.\n2. The paper is overall presented clearly.\n3, The authors made efforts on providing fair comparisons with existing methods."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The main concern is that the current setting is too far from realistic settings. The current evaluation setting is really more like hacking parts of existing datasets. The reviewer encourage the authors to make this work more complete.\n\na. It is too constraint to only evaluate the proposed method when it has no access to QA labels. If the obtain model really captures the casual relationship in videos, plugging it into existing methods in the supervised training setting can show a much more broad application of the proposed method.\n\nb. It is also not realistic to exclude the text query from the training process since the fusion between visual and textual input is also the crucial design to really solve the VQA problem. For example, if there are two accidents going on in the video, the current framework will have systematical flaw as the model is not conditioned on the question text which specifies which accident it is about.\n\nc. Despite the effort of re-training many existing methods, it is not well-justified why it is necessary to discard important features like motion or object as used in Causal-Vid-QA, which brings all the models to a low-performance scheme.\n\nd. There is no proper comparison with methods that do not require QA data like but not including to [a,b,c]. The authors should also acknowledge and at least provide comparison with some of these relevant methods to really provide the audience a correct and comprehensive understanding of the relevant solution to this setting.\n\ne. Once a and b are done, it is also necessary to provide additional comparison on broader VideoQA datasets to understand the importance of causal learning process in videos for broader videoQA tasks, which is really beneficial for the community.\n\n\n\n[a] 🦩Flamingo: a Visual Language Model for Few-Shot Learning\n[b] Language Models with Image Descriptors are Strong Few-Shot Video-Language Learners\n[c] Language Models are Causal Knowledge Extractors for Zero-shot Video Question Answering\n\n\nMinor:\n1. Title, related work: Video Question Answer -> Video Question Answering."},"confidence":{"value":"4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work."},"questions":{"value":"Please check weakness for details."},"rating":{"value":"5: marginally below the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636017853,"tcdate":1698806436537,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission910/Reviewer_dMTZ"],"signatures":["ICLR.cc/2024/Conference/Submission910/Reviewer_dMTZ"],"forum":"Cu5wJa5LGO","number":4,"license":"CC BY 4.0","cdate":1698806436537,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission910/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636017853,"domain":"ICLR.cc/2024/Conference","replyto":"Cu5wJa5LGO","id":"8l59fdwInP","forumContent":{"venue":{"value":"ICLR 2024 poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video Question Answer","Visual Reasoning","Causal Represetation Learning"]},"supplementary_material":{"value":"/attachment/6bc923ab1e364483305123c8939eca6a0a8b7801.zip"},"primary_area":{"value":"representation learning for computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Current approaches to Video Question Answering (VideoQA) primarily focus on cross-modality matching, which is limited by the requirement for extensive data annotations and the insufficient capacity for causal reasoning (e.g. attributing accidents). To address these challenges, we introduce a causal framework for video reasoning, termed Learning Latent Causal Processes (LLCP). At the heart of LLCP lies a multivariate generative model designed to analyze the spatial-temporal dynamics of objects within events. Leveraging the inherent modularity of causal mechanisms, we train the model through self-supervised local auto-regression eliminating the need for annotated question-answer pairs. During inference, the model is applied to answer two types of reasoning questions: accident attribution, which infers the cause from observed effects, and counterfactual prediction, which predicts the effects of counterfactual conditions given the factual evidence. In the first scenario, we identify variables that deviate from the established distribution by the learned model, signifying the root cause of accidents. In the second scenario, we replace embeddings of previous variables with counterfactual ones, enabling us to forecast potential developments. Once we have identified these cause/effect variables, natural language answers are derived through a combination of grammatical parsing and a pre-trained vision-language model. We assess the efficacy of LLCP on both synthetic and real-world data, demonstrating comparable performance to supervised methods despite our framework using no paired textual annotations."},"_bibtex":{"value":"@inproceedings{\nchen2024llcp,\ntitle={{LLCP}: Learning Latent Causal Processes for Reasoning-based Video Question Answer},\nauthor={Guangyi Chen and Yuke Li and Xiao Liu and Zijian Li and Eman Al Suradi and Donglai Wei and Kun Zhang},\nbooktitle={The Twelfth International Conference on Learning Representations},\nyear={2024},\nurl={https://openreview.net/forum?id=Cu5wJa5LGO}\n}"},"title":{"value":"LLCP: Learning Latent Causal Processes for Reasoning-based Video Question Answer"},"pdf":{"value":"/pdf/e4886678d1fb61b3be5fd3c9001c2465abbc8e48.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference"},"paperhash":{"value":"chen|llcp_learning_latent_causal_processes_for_reasoningbased_video_question_answer"},"authorids":{"value":["~Guangyi_Chen1","~Yuke_Li1","~Xiao_Liu23","~Zijian_Li1","~Eman_Al_Suradi1","~Donglai_Wei1","~Kun_Zhang1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Guangyi Chen","Yuke Li","Xiao Liu","Zijian Li","Eman Al Suradi","Donglai Wei","Kun Zhang"]}},"version":2},{"content":{"venue":{"value":"IEEE Trans. Pattern Anal. Mach. Intell. 2007"},"pdf":{"value":"https://ieeexplore.ieee.org/iel5/34/4204153/04204162.pdf"},"venueid":{"value":"dblp.org/journals/PAMI/2007"},"paperhash":{"value":"goldlücke|weighted_minimal_hypersurface_reconstruction"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Bastian_Goldlücke:","~Ivo_Ihrke1","https://dblp.org/search/pid/api?q=author:Christian_Linz:","https://dblp.org/search/pid/api?q=author:Marcus_A._Magnor:"]},"html":{"value":"https://doi.org/10.1109/TPAMI.2007.1146"},"_bibtex":{"value":"@article{DBLP:journals/pami/GoldluckeILM07,\n  author={Bastian Goldlücke and Ivo Ihrke and Christian Linz and Marcus A. Magnor},\n  title={Weighted Minimal Hypersurface Reconstruction},\n  year={2007},\n  cdate={1167609600000},\n  journal={IEEE Trans. Pattern Anal. Mach. Intell.},\n  volume={29},\n  number={7},\n  pages={1194-1208},\n  url={https://doi.org/10.1109/TPAMI.2007.1146}\n}\n"},"abstract":{"value":"Many problems in computer vision can be formulated as a minimization problem for an energy functional. If this functional is given as an integral of a scalar-valued weight function over an unknown hypersurface, then the sought-after minimal surface can be determined as a solution of the functional's Euler-Lagrange equation. This paper deals with a general class of weight functions that may depend on surface point coordinates as well as surface orientation. We derive the Euler-Lagrange equation in arbitrary dimensional space without the need for any surface parameterization, generalizing existing proofs. Our work opens up the possibility of solving problems involving minimal hypersurfaces in a dimension higher than three, which were previously impossible to solve in practice. We also introduce two applications of our new framework: We show how to reconstruct temporally coherent geometry from multiple video streams, and we use the same framework for the volumetric reconstruction of refractive and transparent natural phenomena, bodies of flowing water."},"title":{"value":"Weighted Minimal Hypersurface Reconstruction"},"authors":{"value":["Bastian Goldlücke","Ivo Ihrke","Christian Linz","Marcus A. Magnor"]}},"tmdate":1741193403028,"pdate":1167609600000,"tcdate":1741193379729,"writers":["~"],"signatures":["~Ivo_Ihrke1"],"forum":"xaPRY2xF6t","license":"CC BY-SA 4.0","number":354816,"cdate":1167609600000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1741193403028,"domain":"DBLP.org","id":"xaPRY2xF6t","version":2},{"content":{"summary":{"value":"This paper studies how noisy preference data in RLHF harms reward model generalization and proposes Collaborative Reward Modeling (CRM), where two reward models perform peer review with curriculum learning to filter noisy pairs. Experiments report higher preference accuracy and improved win rates over standard training and robust preference learning baselines, under intentionally added label-flip noise."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"Please see the weakness."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":1},"strengths":{"value":"1. The selected problem setting, i.e., noisy preference data in RLHF is interesting, and is known to impact downstream alignment performance. This paper gives a demonstration through the definition of robust preference pairs, incorrect preference pairs and ambiguous preference pairs, though the the provided results in Figure 2 and 3 may be inaccurate because of \"self-loss\".\n2. The proposed CRM framework is very simple and easy to implement: two reward models select low-loss pairs for each other and follow an easy-to-hard curriculum learning method, which though inevitably brings additional training complexity and costs into the pipeline."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The proposed CRM is limited in theory, and its designs including using self-loss, data filtering strategy and etc. are mostly heuristic, which indicates it may not be effective or extendable in other settings. For example, why choose two models evolve instead of 3 or more?\n2. The empirical settings may have fetal flaw that the noise in the data is manually added by symmetric label flipping, which cannot reflect the complex situation in alignment of real world. The used \"self-loss\" for distinguishing the noised preference pairs may not effective when applied for realistic settings. The effectiveness and performance gains may come from the chosen label flipping. In another way, the self-loss for identifying noisy data is not fair or practical. I would like to know the ratio of noise data in the original preference dataset/benchmark often labeled by human or superior LLMs, and whether how many noisy pairs the self-loss could identity in this situation. In addition, the authors mention that real datasets already include noise, the evaluation does not convincingly separate gains from intentional corruption versus truly realistic noise.\n3. The proposed method depends on a prior noise-rate estimate \\(\\eta\\) to set the selection ratio, but this paper gives limited practical guidance on how to estimate this quantity in real deployment.\n4. The proposed method treats ambiguous and non-robust pairs as noise to be suppressed, which may introduce additional bias, since ambiguous pairs may not be ambiguous if handled by more powerful LLM or human. Also, the provided cases are very few and limited. Again, it would be better to see the statistics in original dataset.\n5. The evaluation is constrained to Llama-3-3B backbones and a single alignment pipeline per setting, so it is unclear whether CRM still brings clear gains at larger model sizes. To my understanding, the reward model should not become the computation burden especially in offline RL settings (e.g., DPO). Also, it will be interesting to see the two models are from different sizes and different model families. I also want to know whether it can be applied to LLM-as-a-Judge setting. It is also suggested to use reasoning models for alignment evaluation instead of relatively outdated GPT-4.\n6. To understand the method in depth, It would be useful to see diagnostics or qualitative examples of pairs that are frequently discarded, to understand any systematic bias introduced by the filtering."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762917852385,"tcdate":1762200746759,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission5067/Reviewer_wBDe"],"signatures":["ICLR.cc/2026/Conference/Submission5067/Reviewer_wBDe"],"forum":"BB1aypUDAF","number":4,"license":"CC BY 4.0","cdate":1762200746759,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission5067/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762917852385,"domain":"ICLR.cc/2026/Conference","replyto":"BB1aypUDAF","id":"AZVMIeyPZl","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Preference Learning","Reward Modeling","RLHF"]},"primary_area":{"value":"alignment, fairness, safety, privacy, and societal considerations"},"abstract":{"value":"Reward models (RMs) play a pivotal role in aligning large language models (LLMs) with human values. \nHowever, noisy preferences in human feedback can lead to reward misgeneralization – a phenomenon where reward models learn spurious correlations or overfit to noisy preferences, which poses important challenges to the generalization of RMs.\nThis paper systematically analyzes the characteristics of preference pairs and aims to identify how noisy preferences differ from human-aligned preferences in reward modeling. \nOur analysis reveals that noisy preferences are difficult for RMs to fit, as they cause sharp training fluctuations and irregular gradient updates.\nThese distinctive dynamics suggest the feasibility of identifying and excluding such noisy preferences.\nEmpirical studies clarify that policy LLM optimized with a reward model trained on the full preference dataset, which includes substantial noise, performs worse than the one trained on a subset of exclusively high-quality data.\nTo address this challenge, we propose an online Collaborative Reward Modeling (CRM) framework to achieve robust preference learning through peer review and curriculum learning. \nIn particular, CRM maintains two RMs that collaboratively filter potential noisy preferences by peer-reviewing each other’s data selections.\nCurriculum learning establishes a well-defined learning trajectory to synchronize the capabilities of two RMs, further promoting the utility of peer review.\nExtensive experiments demonstrate that CRM significantly enhances RM generalization, with up to $9.94$-points improvement on RewardBench under an extreme 40\\% noise. \nMoreover, CRM can seamlessly extend to implicit-reward alignment methods, offering a robust and versatile alignment strategy."},"_bibtex":{"value":"@misc{\nzhang2026two,\ntitle={Two Minds Better Than One: Collaborative Reward Modeling for {LLM} Alignment},\nauthor={Jiazheng Zhang and Wenqing Jing and Zizhuo Zhang and Zhiheng Xi and Shihan Dou and Rongxiang Weng and Jiahuan Li and Jingang Wang and Mingxu Chai and Shibo Hong and Yuming Yang and Xuanjing Huang and Tao Gui and Qi Zhang},\nyear={2026},\nurl={https://openreview.net/forum?id=BB1aypUDAF}\n}"},"title":{"value":"Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment"},"pdf":{"value":"/pdf/8861d3995eb27a7e994a0f01f704dbaa6cfed478.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"zhang|two_minds_better_than_one_collaborative_reward_modeling_for_llm_alignment"},"authorids":{"value":["~Jiazheng_Zhang1","~Wenqing_Jing1","~Zizhuo_Zhang1","~Zhiheng_Xi1","~Shihan_Dou1","~Rongxiang_Weng2","~Jiahuan_Li1","~Jingang_Wang1","~Mingxu_Chai1","~Shibo_Hong1","~Yuming_Yang1","~Xuanjing_Huang1","~Tao_Gui1","~Qi_Zhang8"]},"authors":{"value":["Jiazheng Zhang","Wenqing Jing","Zizhuo Zhang","Zhiheng Xi","Shihan Dou","Rongxiang Weng","Jiahuan Li","Jingang Wang","Mingxu Chai","Shibo Hong","Yuming Yang","Xuanjing Huang","Tao Gui","Qi Zhang"]}},"version":2},{"content":{"venue":{"value":"Neurocomputing 2025"},"venueid":{"value":"dblp.org/journals/IJON/2025"},"paperhash":{"value":"wang|on_the_shortcut_learning_in_multilingual_neural_machine_translation"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Wenxuan_Wang_0001:","~Wenxiang_Jiao1","","https://dblp.org/search/pid/api?q=author:Zhaopeng_Tu:","https://dblp.org/search/pid/api?q=author:Michael_R._Lyu:"]},"html":{"value":"https://doi.org/10.1016/j.neucom.2024.128833"},"_bibtex":{"value":"@article{DBLP:journals/ijon/WangJHTL25,\n  author={Wenxuan Wang and Wenxiang Jiao and Jen-tse Huang and Zhaopeng Tu and Michael R. Lyu},\n  title={On the shortcut learning in multilingual neural machine translation},\n  year={2025},\n  cdate={1735689600000},\n  journal={Neurocomputing},\n  volume={615},\n  pages={128833},\n  url={https://doi.org/10.1016/j.neucom.2024.128833}\n}\n"},"abstract":{"value":"In this study, we revisit the commonly-cited off-target issue in multilingual neural machine translation (MNMT). By carefully designing experiments on different MNMT scenarios and models, we attribute the off-target issue to the overfitting of the shortcuts of (non-centric, centric) language mappings. Specifically, the learned shortcuts biases MNMT to mistakenly translate non-centric languages into the centric language instead of the expected non-centric language for zero-shot translation. Analyses on learning dynamics show that the shortcut learning generally occurs in the later stage of model training, and multilingual pretraining accelerates and aggravates the shortcut learning. Based on these observations, we propose a simple and effective training strategy to eliminate the shortcuts in MNMT models by leveraging the forgetting nature of model training. The only difference from the standard training is that we remove the training instances that may induce the shortcut learning in the later stage of model training. Without introducing any additional data and computational costs, our approach can consistently and significantly improve the zero-shot translation performance by alleviating the shortcut learning for different MNMT models and benchmarks."},"title":{"value":"On the shortcut learning in multilingual neural machine translation"},"authors":{"value":["Wenxuan Wang","Wenxiang Jiao","Jen-tse Huang","Zhaopeng Tu","Michael R. Lyu"]}},"tmdate":1790880608869,"pdate":1735689600000,"tcdate":1735057796720,"writers":["~"],"signatures":["~Jen-tse_Huang1"],"forum":"HtnFleZIe5","license":"CC BY-SA 4.0","number":259388,"cdate":1735689600000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1790880608869,"domain":"DBLP.org","id":"HtnFleZIe5","version":2},{"content":{"venue":{"value":"CoRR 2024"},"pdf":{"value":"http://arxiv.org/pdf/2411.10581v1"},"venueid":{"value":"dblp.org/journals/CORR/2024"},"paperhash":{"value":"wang|on_the_shortcut_learning_in_multilingual_neural_machine_translation"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Wenxuan_Wang_0001:","~Wenxiang_Jiao1","","https://dblp.org/search/pid/api?q=author:Zhaopeng_Tu:","https://dblp.org/search/pid/api?q=author:Michael_R._Lyu:"]},"html":{"value":"https://doi.org/10.48550/arXiv.2411.10581"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2411-10581,\n  publtype={informal},\n  author={Wenxuan Wang and Wenxiang Jiao and Jen-tse Huang and Zhaopeng Tu and Michael R. Lyu},\n  title={On the Shortcut Learning in Multilingual Neural Machine Translation},\n  year={2024},\n  cdate={1704067200000},\n  journal={CoRR},\n  volume={abs/2411.10581},\n  url={https://doi.org/10.48550/arXiv.2411.10581}\n}\n"},"abstract":{"value":"In this study, we revisit the commonly-cited off-target issue in multilingual neural machine translation (MNMT). By carefully designing experiments on different MNMT scenarios and models, we attribute the off-target issue to the overfitting of the shortcuts of (non-centric, centric) language mappings. Specifically, the learned shortcuts biases MNMT to mistakenly translate non-centric languages into the centric language instead of the expected non-centric language for zero-shot translation. Analyses on learning dynamics show that the shortcut learning generally occurs in the later stage of model training, and multilingual pretraining accelerates and aggravates the shortcut learning. Based on these observations, we propose a simple and effective training strategy to eliminate the shortcuts in MNMT models by leveraging the forgetting nature of model training. The only difference from the standard training is that we remove the training instances that may induce the shortcut learning in the later stage of model training. Without introducing any additional data and computational costs, our approach can consistently and significantly improve the zero-shot translation performance by alleviating the shortcut learning for different MNMT models and benchmarks."},"title":{"value":"On the Shortcut Learning in Multilingual Neural Machine Translation"},"authors":{"value":["Wenxuan Wang","Wenxiang Jiao","Jen-tse Huang","Zhaopeng Tu","Michael R. Lyu"]}},"tmdate":1790880608481,"pdate":1704067200000,"externalIds":["dblp:journals/corr/abs-2411-10581"],"tcdate":1756903014682,"writers":["~"],"signatures":["~Jen-tse_Huang1"],"forum":"cvKmJi6ujB","license":"CC BY-SA 4.0","number":620045,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1790880608481,"domain":"DBLP.org","id":"cvKmJi6ujB","version":2},{"content":{"summary":{"value":"The paper investigates video-text representation alignment, focusing on modern video and language encoders. It demonstrates how cross-modal alignment depends on the richness of visual and textual data provided at test time. The study introduces test-time scaling laws, showing how adjustments to the number of frames and captions can improve alignment scores. The paper further explores the relationship between alignment quality and downstream task performance, including temporal reasoning and general video understanding. The findings suggest that alignment could serve as a valuable zero-shot metric for evaluating video models."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"N/A"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":1},"strengths":{"value":"**Comprehensive Approach**: The paper provides the first comprehensive study of video-text representation alignment, extending the Platonic Representation Hypothesis to the temporal domain, making it a significant contribution.\n\n**Correlation with Downstream Tasks**: The correlation between alignment scores and performance on semantic and non-semantic tasks demonstrates the practical value of alignment as a metric."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The idea of probing visual representation with video-text alignment is not such convincing. This evaluation is fair for models proposed on cross-modal tasks, but visual ability is not only cross-modal alignment.   \nFor example, in tasks such as video object detection and video object tracking, the vision model only need to detect pixel-level difference in the picture, without the need to be aware of textual semantics.   \nThe DINO-series [1], SAM-seris [2], I-JEPA [3] and V-JEPA [4] are some evidence that model can excel in visual tasks without the need of textual semantic. \n\n[1] Caron, M., Touvron, H., Misra, I., Jégou, H., Mairal, J., Bojanowski, P. and Joulin, A., 2021. Emerging properties in self-supervised vision transformers. In Proceedings of the IEEE/CVF international conference on computer vision (pp. 9650-9660).    \n[2] Kirillov, Alexander, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao et al. \"Segment anything.\" In Proceedings of the IEEE/CVF international conference on computer vision, pp. 4015-4026. 2023.  \n[3] Assran, M., Duval, Q., Misra, I., Bojanowski, P., Vincent, P., Rabbat, M., LeCun, Y. and Ballas, N., 2023. Self-supervised learning from images with a joint-embedding predictive architecture. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 15619-15629).   \n[4] Assran, Mido, Adrien Bardes, David Fan, Quentin Garrido, Russell Howes, Matthew Muckley, Ammar Rizvi et al. \"V-jepa 2: Self-supervised video models enable understanding, prediction and planning.\" arXiv preprint arXiv:2506.09985 (2025)."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1764358840989,"tcdate":1761555412430,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission15905/Reviewer_AP4E"],"signatures":["ICLR.cc/2026/Conference/Submission15905/Reviewer_AP4E"],"forum":"gE17TwVMNh","number":3,"license":"CC BY 4.0","cdate":1761555412430,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission15905/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1764358840989,"domain":"ICLR.cc/2026/Conference","replyto":"gE17TwVMNh","id":"1HxCv1IiWR","forumContent":{"TLDR":{"value":"Our study of video-text representation alignment demonstrates that alignment is dramatically improved by using richer test-time data, such as multiple video frames and diverse captions."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Platonic Representation hypothesis","video understanding","video-text alignment"]},"primary_area":{"value":"unsupervised, self-supervised, semi-supervised, and supervised representation learning"},"abstract":{"value":"The alignment of representations from different modalities has recently been shown to provide insights on the structural similarities and downstream capabilities of different encoders across diverse data types.\nWhile significant progress has been made in aligning images with text, the temporal nature of _video_ data remains largely unexplored in this context. \nIn this work, we conduct the first comprehensive study of video-text representation alignment, probing the capabilities of modern video and language encoders. \nOur findings reveal several key insights. \nFirst, we demonstrate that cross-modal alignment highly depends on the richness of both visual (static images vs. multi-frame videos) and text (single caption vs. a collection) data _provided at test time_, especially when using state-of-the-art video encoders. \nWe propose parametric test-time scaling laws that capture this behavior and show remarkable predictive power against empirical observations.\nSecondly, we investigate the correlation between semantic alignment and performance on both semantic and non-semantic downstream tasks, providing initial evidence that strong alignment against text encoders may be linked to _general-purpose_ video representation and understanding.\nFinally, we correlate temporal reasoning with cross-modal alignment providing a challenging test-bed for vision and language models. \nOverall, our work introduces video-text alignment as an informative zero-shot way to probe the representation power of different encoders for spatio-temporal data."},"_bibtex":{"value":"@inproceedings{\novsjanikov2026dynamic,\ntitle={Dynamic Reflections: Probing Video Representations with Text Alignment},\nauthor={Maks Ovsjanikov and Viorica Patraucean and Leonidas Guibas and Tyler Zhu and Tengda Han},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=gE17TwVMNh}\n}"},"title":{"value":"Dynamic Reflections: Probing Video Representations with Text Alignment"},"pdf":{"value":"/pdf/1d77288a69f5acaef130462f119cba6cebe6bdbe.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"ovsjanikov|dynamic_reflections_probing_video_representations_with_text_alignment"},"authorids":{"value":["~Maks_Ovsjanikov1","~Viorica_Patraucean1","~Leonidas_Guibas1","~Tyler_Zhu2","~Tengda_Han1"]},"authors":{"value":["Maks Ovsjanikov","Viorica Patraucean","Leonidas Guibas","Tyler Zhu","Tengda Han"]}},"version":2},{"content":{"review":{"value":"The most convincing contribution is the analysis of resampling. Showing that a label round trip at 1.5 mm spacing has a maximum achievable DSC of 0.864 gives a concrete explanation for an otherwise puzzling performance ceiling. The improvement after correcting spacing supports this diagnosis. The brightness-shortcut observation for video and the reporting of several negative findings are also practically useful.\n\nThe empirical analysis remains limited in scope. Task 2 is absent, and multiple changes are grouped in the internal experiments, so the contributions of normalization, augmentation, loss, patch size, and training length cannot be separated. The final video DSC is reasonable, but very large HD and ASD suggest severe case-level failures that deserve visual and quantitative analysis. Because the development process relies substantially on repeated hidden-test feedback, the manuscript should disclose the number of submissions and the rule used to choose the final model; otherwise, hidden-test iteration may act as adaptive tuning. The six-video training set also makes patient-level splitting and leakage prevention especially important. Exact seeds, fallback behavior, and code availability should be documented.\n\nDespite the limited scope, the paper is honest and technically useful. I support acceptance as a challenge report, with stronger disclosure of the experimental selection process."},"confidence":{"value":4},"rating":{"value":7},"title":{"value":"This concise challenge report studies CT and surgical-video mitral-valve segmentation. It identifies target spacing as a limiting factor for thin CT anatomy and image brightness as a shortcut in video, then reports improved hidden-test performance after targeted corrections."}},"parentInvitations":"MICCAI.org/2026/Workshop/MWM/-/Official_Review","nonreaders":[],"tmdate":1785779920790,"tcdate":1785774438422,"writers":["MICCAI.org/2026/Workshop/MWM","MICCAI.org/2026/Workshop/MWM/Submission113/Reviewer_ZdGM"],"signatures":["MICCAI.org/2026/Workshop/MWM/Submission113/Reviewer_ZdGM"],"forum":"SVX3aNHR3R","number":2,"license":"CC BY 4.0","cdate":1785774438422,"readers":["everyone"],"invitations":["MICCAI.org/2026/Workshop/MWM/Submission113/-/Official_Review","MICCAI.org/2026/Workshop/MWM/-/Edit"],"mdate":1785779920790,"domain":"MICCAI.org/2026/Workshop/MWM","replyto":"SVX3aNHR3R","id":"hFNtJ7DUf9","forumContent":{"venue":{"value":"MWM2026"},"pdf":{"value":"/pdf/82e9cbe324606672de89c783d0affff62e6b49a3.pdf"},"keywords":{"value":["mitral valve","segmentation","nnU-Net","challenge","semi-supervised learning","cardiac CT","surgical video"]},"venueid":{"value":"MICCAI.org/2026/Workshop/MWM"},"paperhash":{"value":"nahata|mitral_valve_segmentation_in_cardiac_ct_and_surgical_video_mvaa_2026_challenge_report"},"authorids":{"value":["~Valmik_Nahata1"]},"abstract":{"value":"We present our solution to the MVAA 2026 challenge for mitral valve segmentation. Our final submission achieves a Task 1 (cardiac CT) hidden test score of DSC 0.811, HD 6.86 mm, and ASD 0.810 mm. For Task 3 (surgical video), we achieve DSC 0.775, HD 361 mm, and ASD 233 mm. We adopted a diagnostic methodology focused on identifying pipeline bottlenecks. We found that resampling target spacing imposes a hard ceiling on achievable accuracy for thin structures. A perfect prediction resampled to 1.5 mm and back scores only 0.864 DSC, whereas the mitral valve is a thin sheet (median maximum inscribed radius 1.82 mm). Correcting the spacing to the dataset median improved our real hidden-test DSC from 0.651 to 0.789. We also show that our Task 3 model exploited a brightness shortcut, which we mitigated using targeted augmentations."},"_bibtex":{"value":"@inproceedings{\nnahata2026mitral,\ntitle={Mitral Valve Segmentation in Cardiac {CT} and Surgical Video: {MVAA} 2026 Challenge Report},\nauthor={Valmik Nahata},\nbooktitle={The 1st MICCAI Workshop on Medical World Models},\nyear={2026},\nurl={https://openreview.net/forum?id=SVX3aNHR3R}\n}"},"title":{"value":"Mitral Valve Segmentation in Cardiac CT and Surgical Video: MVAA 2026 Challenge Report"},"authors":{"value":["Valmik Nahata"]}},"version":2},{"content":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["generative models"]},"supplementary_material":{"value":"/attachment/d3f073fd3c20a63d9ef676d79183c3c325fb563a.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Sampling from unnormalized densities presents a fundamental challenge with wide-ranging applications, from posterior inference to molecular dynamics simulations. Continuous flow-based neural samplers offer a promising approach, learning a velocity field that satisfies key principles of marginal density evolution (e.g., the continuity equation) to generate samples. However, this learning procedure requires accurate estimation of intractable terms linked to the computationally challenging partition function, for which existing estimators often suffer from high variance or low accuracy. To overcome this, we introduce an improved estimator for these challenging quantities, employing a velocity-driven Sequential Monte Carlo method enhanced with control variates. Furthermore, we introduce a shortcut consistency model to boost the runtime efficiency of the flow-based neural sampler by minimizing its required sampling steps. Our proposed Neural Flow Shortcut Sampler empirically outperforms existing flow-based neural samplers on both synthetic datasets and complex n-body system targets"},"_bibtex":{"value":"@misc{\nchen2025neural,\ntitle={Neural Flow Samplers with Shortcut Models},\nauthor={Wuhao Chen and Zijing Ou and Yingzhen Li},\nyear={2025},\nurl={https://openreview.net/forum?id=hHfUwjl3hF}\n}"},"title":{"value":"Neural Flow Samplers with Shortcut Models"},"pdf":{"value":"/pdf/6bb36d45cc2323b345570fb49768c9ab441cf3df.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"chen|neural_flow_samplers_with_shortcut_models"},"authorids":{"value":["~Wuhao_Chen1","~Zijing_Ou1","~Yingzhen_Li1"]},"authors":{"value":["Wuhao Chen","Zijing Ou","Yingzhen Li"]}},"tmdate":1764326049484,"tcdate":1758206875349,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission12288/Authors"],"signatures":["ICLR.cc/2026/Conference/Submission12288/Authors"],"forum":"hHfUwjl3hF","license":"CC BY 4.0","number":12288,"cdate":1758206875349,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/-/Submission","ICLR.cc/2026/Conference/-/Post_Submission","ICLR.cc/2026/Conference/Submission12288/-/Full_Submission","ICLR.cc/2026/Conference/-/Withdrawn_Submission"],"mdate":1764326049484,"odate":1759896705795,"domain":"ICLR.cc/2026/Conference","id":"hHfUwjl3hF","version":2},{"content":{"venue":{"value":"CoRR 2025"},"pdf":{"value":"http://arxiv.org/pdf/2502.07337v2"},"venueid":{"value":"dblp.org/journals/CORR/2025"},"paperhash":{"value":"chen|neural_flow_samplers_with_shortcut_models"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Wuhao_Chen:","https://dblp.org/search/pid/api?q=author:Zijing_Ou:","~Yingzhen_Li1"]},"html":{"value":"https://doi.org/10.48550/arXiv.2502.07337"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2502-07337,\n  publtype={informal},\n  author={Wuhao Chen and Zijing Ou and Yingzhen Li},\n  title={Neural Flow Samplers with Shortcut Models},\n  year={2025},\n  month={February},\n  cdate={1738368000000},\n  journal={CoRR},\n  volume={abs/2502.07337},\n  url={https://doi.org/10.48550/arXiv.2502.07337}\n}\n"},"abstract":{"value":"Sampling from unnormalized densities presents a fundamental challenge with wide-ranging applications, from posterior inference to molecular dynamics simulations. Continuous flow-based neural samplers offer a promising approach, learning a velocity field that satisfies key principles of marginal density evolution (e.g., the continuity equation) to generate samples. However, this learning procedure requires accurate estimation of intractable terms linked to the computationally challenging partition function, for which existing estimators often suffer from high variance or low accuracy. To overcome this, we introduce an improved estimator for these challenging quantities, employing a velocity-driven Sequential Monte Carlo method enhanced with control variates. Furthermore, we introduce a shortcut consistency model to boost the runtime efficiency of the flow-based neural sampler by minimizing its required sampling steps. Our proposed Neural Flow Shortcut Sampler empirically outperforms existing flow-based neural samplers on both synthetic datasets and complex n-body system targets."},"title":{"value":"Neural Flow Samplers with Shortcut Models"},"authors":{"value":["Wuhao Chen","Zijing Ou","Yingzhen Li"]}},"tmdate":1760082803287,"pdate":1735689600000,"externalIds":["dblp:journals/corr/abs-2502-07337"],"tcdate":1760082799263,"writers":["~"],"signatures":["~Yingzhen_Li1"],"forum":"O2PTua1u60","license":"CC BY-SA 4.0","number":641916,"cdate":1738368000000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1760082803287,"domain":"DBLP.org","id":"O2PTua1u60","version":2},{"content":{"summary":{"value":"This paper presents an image-to-video generation method with mask guidance. Thanks to this design, the method can generate video with very limited data. This method needs frame-specific masks for training and testing. The method also has a first frame-sharing noise method to enable better temporal consistency and a mask-aware attention model. This method compares several zero-shot/one-shot methods."},"soundness":{"value":2},"confidence":{"value":5},"questions":{"value":"1. Why do we need mask-aware video generation? How to get diverse results when the mask is provided.\n2. Where is the video demo?\n3. There is only a visual ablation of this method in the paper, what about the numerical results in the larger scale?"},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"- Mask-guided is novel for me as a video generation condition.\n- This method requires only a small dataset for training."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- I am curious about the motivation behind this paper. The mask is hard to generate when inference, which makes this method impractical.\n- There are no video results for this paper.\n- The comparison methods are too old to evaluate the performance of this method.\n- Generating videos using small datasets from stable diffusion (text-to-image model) is out of fashion. Current state-of-the-art methods directly generate videos from large-scale training."}},"nonreaders":[],"tmdate":1731427876113,"tcdate":1729779103602,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission3540/Reviewer_nq54"],"signatures":["ICLR.cc/2025/Conference/Submission3540/Reviewer_nq54"],"forum":"9GNTtaIZh6","number":1,"license":"CC BY 4.0","cdate":1729779103602,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission3540/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427876113,"domain":"ICLR.cc/2025/Conference","replyto":"9GNTtaIZh6","id":"wnSYTfruY1","forumContent":{"TLDR":{"value":"Our model efficiently utilizes resources and achieves controllability and consistency in video generation through motion sequences obtained from drawings or extracted masks."},"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Diffusion models","video generation"]},"supplementary_material":{"value":"/attachment/4f7b0c4acf3ccb77f02f0bb06589b2225ce3d1ed.zip"},"primary_area":{"value":"generative models"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Recent advancements in diffusion models have brought new vitality into visual content creation. However, current text-to-video generation models still face challenges such as high training costs, substantial data requirements, and difficulties in maintaining consistency between given text and motion of the foreground object. To address these challenges, we propose mask-guided video generation, which requires only a small amount of data and is trained on a single GPU. Furthermore, to mitigate the impact of background interference on controllable text-to-video generation, we utilize mask  sequences obtained through drawing or extraction, along with the first-frame content, to guide video generation. Specifically, our model introduces foreground masks into existing architectures to learn region-specific attention, precisely matching text features and the motion of the foreground object.  Subsequently, video generation is guided by the mask sequences to prevent the sudden disappearance of foreground objects. Our model also incorporates a first-frame sharing strategy during inference, leading to better stability in the video generation. Additionally, our approach allows for incrementally  generation of longer video sequences. By employing this method, our model achieves efficient resource utilization and ensures controllability and consistency in video generation using mask sequences. Extensive qualitative and quantitative experiments demonstrate that this approach excels in various video generation tasks, such as video editing and generating artistic videos, outperforming previous methods in terms of consistency and quality."},"_bibtex":{"value":"@misc{\nfeng2024maskguided,\ntitle={Mask-Guided Video Generation: Enhancing Motion Control and Quality with Limited Data},\nauthor={SiCong Feng and Li Peng and Jielong Yang},\nyear={2024},\nurl={https://openreview.net/forum?id=9GNTtaIZh6}\n}"},"title":{"value":"Mask-Guided Video Generation: Enhancing Motion Control and Quality with Limited Data"},"pdf":{"value":"/pdf/7c432419930262ab73467f66c7cf0ae95aa0d7c6.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"feng|maskguided_video_generation_enhancing_motion_control_and_quality_with_limited_data"},"authorids":{"value":["~SiCong_Feng1","~Li_Peng2","~Jielong_Yang1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["SiCong Feng","Li Peng","Jielong Yang"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a new video generation framework based on extracting the global features of the video and conditional diffusion model to predict frame features, leading to the frame. The paper argues the proposed method outperforms prior video generation methods on various benchmarks, including UCF-101, Taichi-HD, and SkyTimelapse. "},"presentation":{"value":"1 poor"},"contribution":{"value":"2 fair"},"soundness":{"value":"2 fair"},"strengths":{"value":"- Compared with the prior video generation methods, the proposed method considers non-autoregressive approach for generating the video, which can improve the efficiency in inference time.\n- The proposed method shows better performance compared with prior works. "},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The overall framework is quite complex, including so many notations, and a bit difficult to follow. For instance, Why $I_j$ is put into the video decoder model as well as $I_j$ in Figure 2? Moreover, are KL-VAE and video encode/decoder, discriminator jointly trained or not? What is the intuition of letting the video decoder network as a conditional diffusion model instead of letting simple 2D CNNs? Why do we need to consider \"keyframes\" for extracting global features from a given video? Does DiT for modeling global features is trained in post-hoc manner after the training of the entire framework? \n- The paper misses an efficiency comparison with recent latent video diffusion models to improve the efficiency in training and efficiency: e.g., LVDM [He et al., 2023] and PVDM. Compared with these frameworks, what is the advantage and disadvantages of the method?\n- Typo: Specificcally -> Specifically in L197. \n\n---\n[He et al., 2023] Latent Video Diffusion Models for High-Fidelity Long Video Generation   \n[Yu et al., 2023] Video Probabilistic Diffusion Models in Projected Latent Space, CVPR 2023"},"confidence":{"value":"5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully."},"questions":{"value":"- I guess the proposed method may show a worse performance if the targeting video length becomes large because the quality of global features might have a limitation and the decoder that synthesizes a frame in frame-index conditioned manner has a limited capacity. What is the (empirical) maximum length for high-quality modeling with this framework? "},"rating":{"value":"4: Borderline reject: Technically solid paper where reasons to reject, e.g., limited evaluation, outweigh reasons to accept, e.g., good evaluation. Please use sparingly."},"code_of_conduct":{"value":"Yes"},"limitations":{"value":"The paper adequately addresses the limitations in Conclusion section. "}},"nonreaders":[],"tmdate":1702411056622,"tcdate":1688719799170,"writers":["NeurIPS.cc/2023/Conference","NeurIPS.cc/2023/Conference/Submission6343/Reviewer_P49b"],"signatures":["NeurIPS.cc/2023/Conference/Submission6343/Reviewer_P49b"],"forum":"TRbklCR2ZW","number":1,"license":"CC BY 4.0","cdate":1688719799170,"mdate":1702411056622,"readers":["everyone"],"invitations":["NeurIPS.cc/2023/Conference/Submission6343/-/Official_Review","NeurIPS.cc/2023/Conference/-/Edit"],"domain":"NeurIPS.cc/2023/Conference","replyto":"TRbklCR2ZW","id":"Dk2FKQ0Ja4","forumContent":{"venue":{"value":"NeurIPS 2023 poster"},"keywords":{"value":["Video Generation","Video Autoencoder","Diffusion Probabilistic Model"]},"supplementary_material":{"value":"/attachment/a3c747157f50b4d8abdce74cdef258f0d8ef154b.pdf"},"_bibtex":{"value":"@inproceedings{\nsun2023glober,\ntitle={{GLOBER}: Coherent Non-autoregressive Video Generation via {GLOB}al Guided Video Decod{ER}},\nauthor={Mingzhen Sun and Weining Wang and Zihan Qin and Jiahui Sun and Sihan Chen and Jing Liu},\nbooktitle={Thirty-seventh Conference on Neural Information Processing Systems},\nyear={2023},\nurl={https://openreview.net/forum?id=TRbklCR2ZW}\n}"},"title":{"value":"GLOBER: Coherent Non-autoregressive Video Generation via GLOBal Guided Video DecodER"},"paperhash":{"value":"sun|glober_coherent_nonautoregressive_video_generation_via_global_guided_video_decoder"},"TLDR":{"value":"This work presents a novel non-autoregressive method GLOBER for video generation tasks, which generates coherent videos with global guidance and maximum flexibility."},"abstract":{"value":"Video generation necessitates both global coherence and local realism. This work presents a novel non-autoregressive method GLOBER, which first generates global features to obtain comprehensive global guidance and then synthesizes video frames based on the global features to generate coherent videos. Specifically, we propose a video auto-encoder, where a video encoder encodes videos into global features, and a video decoder, built on a diffusion model, decodes the global features and synthesizes video frames in a non-autoregressive manner. To achieve maximum flexibility, our video decoder perceives temporal information through normalized frame indexes, which enables it to synthesize arbitrary sub video clips with predetermined starting and ending frame indexes. Moreover, a novel adversarial loss is introduced to improve the global coherence and local realism between the synthesized video frames. Finally, we employ a diffusion-based video generator to fit the global features outputted by the video encoder for video generation. Extensive experimental results demonstrate the effectiveness and efficiency of our proposed method, and new state-of-the-art results have been achieved on multiple benchmarks."},"pdf":{"value":"/pdf/4a8260ebb7e55420c951627d7fa38e9764afd26e.pdf"},"venueid":{"value":"NeurIPS.cc/2023/Conference"},"authorids":{"value":["~Mingzhen_Sun1","~Weining_Wang3","~Zihan_Qin1","~Jiahui_Sun2","~Sihan_Chen3","~Jing_Liu1"]},"authors":{"value":["Mingzhen Sun","Weining Wang","Zihan Qin","Jiahui Sun","Sihan Chen","Jing Liu"]}},"version":2},{"content":{"summary":{"value":"Authors propose a long-video QnA evaluation benchmark consisting of human annotated video question answer pairs. They evaluate multiple state-of-the-art approaches on this dataset, highlighting the difficulty of the benchmark for existing approaches. The authors clearly highlight the distinctions of this benchmark compared to prior work, focussing on video length, diversity, and question complexity. Several interesting analysis is provided on the dataset."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. The additional explanations regarding various sections under weaknesses\n2. How would alternate evaluation schemes like likelihood based selection affect performance? This ties into Q5 above: is performance on the dataset actually correlated to long video understanding ability or instruction following ability of evaluated models?\n3. Possibly related to LLM filtering idea; how would the naive baselines (suggested under weaknesses) perform on the benchmark? \n\nThis proposed dataset / benchmark would definitely be valuable for the community to better evaluate long video understanding approaches. It would be really great if the authors could clarify / address the above points."},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. Timely and useful benchmark for long-video QnA containing diverse and long videos \n2. Paper is quite well written clearly outlining the reasoning for building this benchmark and how it differentiates from existing works. Table 1 in particular is quite useful for the latter. \n3. Several interesting analysis on dataset statistics."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. L143: Consider adding hours, i.e.  … 4101 seconds, … -> … 4101 seconds ( > one hour), … \n\n2. **On L316 (answer matching to choices):** The authors use the following setup for selecting the correct choice with a given VLM: “After obtaining the model responses, we first attempted to extract the answers using regular expression matching. For questions where the matching process was unsuccessful, we employed a GLM-4 model to extract the answers from the responses.” \n    1. Could you try likelihood selection / cloze-prompting similar to most NLP benchmarks (see likelihood selection in [1] or cloze-prompting in [2])? This could be more efficient to evaluate (as opposed to generating tokens iteratively) and also avoid any negatives due to difficulty in answer matching. \n    2. Could you explain how the GLM-4 model is used exactly to extract answers? Are you directly prompting GLM-4 to select one answer given the outputs of the model? \n    3. If yes (to 2 above), could you also try using some text-encoder feature-space similarity between generated answer and choices (i.e. standard zero-shot setup in CLIP)? \n\n3. **Naive Baselines:** Can we include two naive baselines with only text and a single center frame? These may give some interesting insights about the benchmark.\n    1. For only text, evaluate with a SOTA LLM (feeding no images, using only the question and choices) similar to what is done in [1, 3]. See Just LLM in [1] and blind variant in Table 4 of [3]. \n    2. For single center frame, using some SOTA VLM and feed only the single center frame (this is motivated by Single Frame VLM in [1]). \n\n4. Table 2: Could you repeat the abbreviation names (ER EU KIR TG Rea Sum) in the caption? This would make it easier for a reader. \n\n5. Is performance on the dataset actually correlated to long video understanding ability or instruction following ability of evaluated models? \n    1. This is motivated by L379-382 mentioning “MovieChat and LWM exhibited a strong bias towards selecting option A, regardless of the question. In contrast, InternVL240B demonstrated the strongest instruction-following capability, never generating responses outside the constrained options and producing a nearly uniform distribution over the answer choices.”\n    2. Maybe a secondary set of results using likelihood selection (see point 2.1 above) may resolve this problem. Integrating such an evaluation scheme into the benchmark could help it better stand out from existing long-video benchmarks making evaluation both faster and more robust.  \n\n6. Can you explain LLM filtering (Section 4.4) in more detail? If this is related to point 3, adding a full row of that to Table 2 could provide a lot more insights, strengthening the benchmark. \n\n7. “Clue Duration” : can you explain how this is calculated? Is this part of the dataset ground-truth provided by humans? \n\n8. Similar to Figure 2, maybe consider adding visual examples for each question type in the appendix. \n \n9. Also consider adding visual examples for each of the 6 major categories in the appendix. And possibly link to examples of all 21 sub-categories also in some external website. \n\n10. L072 mentions “Through meticulous human annotation and multi-stage quality control processes …” for dataset curation. Please explain this process in more detail. \n\n11. Compute, Inference Time, and other details\n    1. For a reader, it will be useful to know the compute required to evaluate a model on this benchmark. Consider elaborating on compute used for evaluation and time taken. \n    2. Details on dataset size (i.e. storage required) and also any copyrights details for any videos used will be useful for those using benchmark in future.   \n\n&nbsp;\n\n[1] “Understanding Long Videos in One Multimodal Language Model Pass.” ArXiv abs/2403.16998\n\n[2] “Leveraging Large Language Models for Multiple Choice Question Answering.” ICLR 2023\n\n[3] “Memory Consolidation Enables Long-Context Video Understanding.” ArXiv abs/2402.05861"}},"nonreaders":[],"tmdate":1731427872157,"tcdate":1730582682466,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission8887/Reviewer_qFeo"],"signatures":["ICLR.cc/2025/Conference/Submission8887/Reviewer_qFeo"],"forum":"uHgVrGF2Wn","number":1,"license":"CC BY 4.0","cdate":1730582682466,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission8887/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427872157,"domain":"ICLR.cc/2025/Conference","replyto":"uHgVrGF2Wn","id":"k2xHxCE1I6","forumContent":{"TLDR":{"value":"We introduce LVBench, a benchmark specifically designed for long video understanding."},"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video understanding","Multimodal learning","Visual question answering","Long-form video","Datasets and benchmarking"]},"primary_area":{"value":"datasets and benchmarks"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Recent progress in multimodal large language models has markedly enhanced the understanding of short videos (typically under one minute), and several evaluation datasets have emerged accordingly. However, these advancements fall short of meeting the demands of real-world applications such as embodied intelligence for long-term decision-making, in-depth movie reviews and discussions, and live sports commentary, all of which require comprehension of long videos spanning several hours. To address this gap, we introduce LVBench, a benchmark specifically designed for long video understanding. Our dataset comprises publicly sourced videos and encompasses a diverse set of tasks aimed at long video comprehension and information extraction. LVBench is designed to challenge multimodal models to demonstrate long-term memory and extended comprehension capabilities. Our extensive evaluations reveal that current multimodal models still underperform on these demanding long video understanding tasks. Through LVBench, we aim to spur the development of more advanced models capable of tackling the complexities of long video comprehension."},"_bibtex":{"value":"@misc{\nwang2024lvbench,\ntitle={{LVB}ench: An Extreme Long Video Understanding Benchmark},\nauthor={Weihan Wang and Zehai He and Wenyi Hong and Yean Cheng and Xiaohan Zhang and Ji Qi and Ming Ding and Xiaotao Gu and Shiyu Huang and Bin Xu and Yuxiao Dong and Jie Tang},\nyear={2024},\nurl={https://openreview.net/forum?id=uHgVrGF2Wn}\n}"},"title":{"value":"LVBench: An Extreme Long Video Understanding Benchmark"},"pdf":{"value":"/pdf/841779fddd82bcb7aa920cc3028ce329537a9552.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"wang|lvbench_an_extreme_long_video_understanding_benchmark"},"authorids":{"value":["~Weihan_Wang2","~Zehai_He1","~Wenyi_Hong1","~Yean_Cheng1","~Xiaohan_Zhang6","~Ji_Qi3","~Ming_Ding1","~Xiaotao_Gu1","~Shiyu_Huang2","~Bin_Xu1","~Yuxiao_Dong1","~Jie_Tang1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Weihan Wang","Zehai He","Wenyi Hong","Yean Cheng","Xiaohan Zhang","Ji Qi","Ming Ding","Xiaotao Gu","Shiyu Huang","Bin Xu","Yuxiao Dong","Jie Tang"]}},"version":2},{"content":{"summary":{"value":"The paper proposes a variation of Chamfer distance for point cloud completion. It focuses on understanding the point cloud in a distribution way. Specifically, the similarity of the distribution of the paired point should be considered in the point cloud metrics. Unlike previous loss functions, such as CD and other CD variants, the proposed LandauCD considers both the near-distance and far-distance point pairs. The proposed loss takes more consideration of larger distance values to prioritize the long-distance point pairs by using Landau distribution to approximate a re-weighting function. In a way, the proposed loss function can (1) dynamically re-distribute paired distance in the paired point sets; and (2) make use of both near and far points. Experiments on synthetic datasets ShapeNet show promising results when plugging in different point cloud completion models."},"presentation":{"value":"2 fair"},"contribution":{"value":"2 fair"},"soundness":{"value":"2 fair"},"strengths":{"value":"- Chamfer distance is a widely used loss function in the 3D point cloud and geometry learning. An improvement to such loss function may have a significant impact on the 3D vision community. The proposed method considers the usually neglected part in the CD loss---the long-distance point pairs, which is interesting and worth exploring. \n\n- The proposed method considers the 3D point in a distribution view. Instead of directly solving for the weighting term, Landau distribution is proposed to approximate the re-weighting function for the point distance. In a way, the re-weighting term is no longer a static scalar, but a dynamically-changed function."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- A common strategy to deal with the outliers is to remove point distance values larger than some thresholds. I wondered if some comparisons to this truncated CD loss could be presented to further validate the effect of prioritizing long-distance point pairs.\n\n- Is the searched Gaussian distribution general to different datasets? The authors could add experiments.\n\n- Figure 2 does not show a clear pattern of whether the large point distance is reduced when prioritizing the long-distance point pairs. \n\n- I wondered if some theoretical explanation on choosing Landau distribution to approximate the re-weighting function in eq.4 could be presented.\n\n- All these loss functions, including NeurIPS 2021 paper Density-aware CD, are tested on ModelNet, ShapeNet, etc. However, they lack practical use in the real world. I wondered if the authors could provide experiments on some real-world datasets.\n\n- Experimental validation of using two composed distributions over using one distribution (on near or far-distance point pairs) should be provided to further validate the effect of the proposed method.\n\n- As the paper proposed a general loss function, I wondered if some broader use of the loss function could be discussed in the paper."},"confidence":{"value":"4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work."},"questions":{"value":"I have detailed comments above in the weakness section. I hope the authors could consider improving the paper by adding more theoretical and empirical evidence to validate the effect of the proposed method."},"rating":{"value":"3: reject, not good enough"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636148657,"tcdate":1699622739546,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission2158/Reviewer_1Bi4"],"signatures":["ICLR.cc/2024/Conference/Submission2158/Reviewer_1Bi4"],"forum":"mkfvssecfl","number":4,"license":"CC BY 4.0","cdate":1699622739546,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission2158/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636148657,"domain":"ICLR.cc/2024/Conference","replyto":"mkfvssecfl","id":"xmJsyfn613","forumContent":{"venue":{"value":"ICLR 2024 Conference Withdrawn Submission"},"TLDR":{"value":"We propose a new distribution-based loss function for point cloud completion, namely LandauCD"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Point Cloud Completion","Landau Distribution","Loss Function Design"]},"primary_area":{"value":"unsupervised, self-supervised, semi-supervised, and supervised representation learning"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Point clouds are fundamental discrete representations used in computer vision, robotics, etc. Chamfer Distance (CD) is widely adopted as a metric and training loss to evaluate the similarity between two point clouds. However, the vanilla CD is sensitive to outliers, which means a few widely distributed points can disproportionately affect the final similarity score. Besides, CD calculates the simple average of distances of matched point pairs between two sets, which does not take into account the underlying point-wise distance distribution across two point clouds (same weights assigned for  short- and  long-distance pairs by using uniform distribution). To mitigate these issues, we analyze the effect of prioritizing short- and long-distance pairs with Gaussian distributions obtained with grid search, and based on the findings, we take an indirect approach to find Landau distribution, out of many distributions, fits in the form of bimodal Gaussian mixture model which balances two types of pairs. Based on this observation, we propose LandauCD, an innovative loss function grounded in the Landau distribution. We conduct comprehensive experiments using LandauCD and observe significant improvements consistently over all the popular baseline networks trained with CD-based losses, leading to new state-of-the-art results on several benchmarks (PCN, Shapet-55/34, ShapeNet-Part). We also delve into the theoretical explanation behind the consistent improvements of LandauCD.  Code and weights will be released upon acceptance."},"_bibtex":{"value":"@misc{\nlin2024point,\ntitle={Point Cloud Completion with Landau Distribution: A Probabilistic View},\nauthor={Fangzhou Lin and Songlin Hou and Haotian Liu and Haoying Zhou and Xuechu Yu and Kazunori Yamada and Ziming Zhang},\nyear={2024},\nurl={https://openreview.net/forum?id=mkfvssecfl}\n}"},"title":{"value":"Point Cloud Completion with Landau Distribution: A Probabilistic View"},"pdf":{"value":"/pdf/2152a19a90908b82075b0fd9b5ec0c00ac307f1e.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Withdrawn_Submission"},"paperhash":{"value":"lin|point_cloud_completion_with_landau_distribution_a_probabilistic_view"},"authorids":{"value":["~Fangzhou_Lin1","~Songlin_Hou1","~Haotian_Liu6","~Haoying_Zhou1","~Xuechu_Yu2","~Kazunori_Yamada1","~Ziming_Zhang4"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Fangzhou Lin","Songlin Hou","Haotian Liu","Haoying Zhou","Xuechu Yu","Kazunori Yamada","Ziming Zhang"]}},"version":2},{"content":{"summary":{"value":"STARFlow-V is an end-to-end, strictly causal video generator based on normalizing flows, extending STARFlow from images to videos. It operates in a compressed spatiotemporal latent space via a 3D causal VAE and is trained with exact maximum likelihood. The core Global–Local design uses a deep–shallow autoregressive-flow hierarchy: a shallow, frame-local invertible stack refines within-frame details and contributes the Jacobian term, while a deep autoregressive flow performs Gaussian next-token prediction over frame latents causally across time. This mitigates pixel-space error propagation and enables streamable generation.  \n\nTraining stability and output cleanliness are addressed via noise-augmented learning and Flow-Score Matching: a lightweight, near-causal denoiser learns the model score with a one-frame look-ahead and applies a single Tweedie update at inference, avoiding non-causal gradients and artifacts. Efficiency is improved by casting inversion as a nonlinear fixed-point problem, enabling blockwise Jacobi iteration with video-aware initialization and pipelined decoding. A single backbone supports T2V, I2V, and controllable generation, achieving competitive VBench performance and strong long-horizon consistency."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. It would be much better if the paper provided video demos for side-by-side comparisons, including long-horizon and fast-motion cases, to make temporal coherence and artifact differences clearer and more convincing."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. A strictly causal, flow-based video generator with a Global–Local design and deep–shallow AF hierarchy. The near-causal Flow-Score Matching denoiser (one-frame look-ahead) enables single-step Tweedie cleanup without non-causal gradients. Efficient inference via fixed-point (blockwise Jacobi), video-aware init, and pipelined decoding reduces latency while preserving causality.  \n\n2. Broad coverage: 480p/81f T2V/I2V, streaming up to 30s, and control. Better than several AR diffusion baselines with lower sampling cost. ~10× latency reduction and CFG-compatible streaming show practical viability."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Claims of reduced exposure bias and improved temporal stability rely heavily on qualitative demos and VBench. There is no dedicated metric suite or failure-mode analysis for >30s sequences.  \n\n2. The paper reports ~10× speedup but does not deeply analyze its sources or limits. There is no component-wise attribution, nor end-to-end profiling across resolution, clip length and block size. Crucially, standardized compute and resource metrics are missing, e.g., FLOPs per frame/video, parameter count, throughput. Without these, claims of “scalable” and \"efficient\" are hard to reproduce, compare, or operationalize. The paper would benefit from detailed complexity tables, comprehensive hyperparameter sweeps, and component-level ablations that quantify individual and combined contributions to speed and quality."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762925401281,"tcdate":1761833824822,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission15077/Reviewer_ASDa"],"signatures":["ICLR.cc/2026/Conference/Submission15077/Reviewer_ASDa"],"forum":"qsffecsbJg","number":2,"license":"CC BY 4.0","cdate":1761833824822,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission15077/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762925401281,"domain":"ICLR.cc/2026/Conference","replyto":"qsffecsbJg","id":"0MJMOfDaQB","forumContent":{"TLDR":{"value":"We show for the first time that normalizing flows can be scaled for high-quality video synthesis"},"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["normalizing flow","video generation","generative model","autoregressive model"]},"primary_area":{"value":"generative models"},"abstract":{"value":"High-quality video generation at scale requires models that are strictly causal, robust over long horizons, and fast at inference. We present STARFlow-V, a flow-based autoregressive video generator that operates in compressed spatiotemporal latents and is trained with exact likelihood end-to-end. Two design choices ensure causality for autoregressive prediction while mitigating error propagation and enabling end-to-end training: (i) Global–Local architecture, which constrains each token to depend only on the past along time while preserving rich within-frame interactions; and (ii) noise-augmented training jointly with \\emph{flow-score matching}, a lightweight causal denoiser that recovers clean samples from noisy generation. To improve efficiency, STARFlow-V employs a video-aware fixed-point iteration scheme that reformulates inner updates as parallelizable iterations without violating causal structure, yielding substantially faster inference. A deep–shallow autoregressive-flow hierarchy further balances capacity and stability over long videos. The same model natively supports both text-to-video (T2V) and text-/image-to-video (TI2V) generation via unified conditioning, avoiding separate pipelines. Empirically, STARFlow-V achieves strong visual fidelity and temporal consistency with markedly lower sampling cost compared to diffusion-only or discrete AR baselines. By marrying causality, likelihood, and efficiency in a single architecture, STARFlow-V helps pave the way toward a flow-based, scalable paradigm for world modeling."},"_bibtex":{"value":"@misc{\ngu2025endtoend,\ntitle={End-to-End Video Generative Modeling with Scalable Normalizing Flows},\nauthor={Jiatao Gu and Ying Shen and Tianrong Chen and Laurent Dinh and Yuyang Wang and Miguel {\\'A}ngel Bautista and David Berthelot and Joshua M. Susskind and Shuangfei Zhai},\nyear={2025},\nurl={https://openreview.net/forum?id=qsffecsbJg}\n}"},"title":{"value":"End-to-End Video Generative Modeling with Scalable Normalizing Flows"},"pdf":{"value":"/pdf/41b6c84d08c5ccd4938bb0798e893d6167abd660.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"gu|endtoend_video_generative_modeling_with_scalable_normalizing_flows"},"authorids":{"value":["~Jiatao_Gu1","~Ying_Shen4","~Tianrong_Chen1","~Laurent_Dinh1","~Yuyang_Wang3","~Miguel_Ángel_Bautista1","~David_Berthelot1","~Joshua_M._Susskind1","~Shuangfei_Zhai3"]},"authors":{"value":["Jiatao Gu","Ying Shen","Tianrong Chen","Laurent Dinh","Yuyang Wang","Miguel Ángel Bautista","David Berthelot","Joshua M. Susskind","Shuangfei Zhai"]}},"version":2},{"content":{"venue":{"value":"ICPR 2010"},"pdf":{"value":"https://ieeexplore.ieee.org/iel5/5595335/5595735/05597410.pdf"},"venueid":{"value":"dblp.org/conf/ICPR/2010"},"paperhash":{"value":"shin|corecognition_of_actions_in_video_pairs"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Young_Min_Shin:","~Minsu_Cho1","~Kyoung_Mu_Lee1"]},"html":{"value":"https://doi.org/10.1109/ICPR.2010.120"},"_bibtex":{"value":"@inproceedings{DBLP:conf/icpr/ShinCL10,\n  author={Young Min Shin and Minsu Cho and Kyoung Mu Lee},\n  title={Co-recognition of Actions in Video Pairs},\n  year={2010},\n  cdate={1262304000000},\n  pages={456-459},\n  url={https://doi.org/10.1109/ICPR.2010.120},\n  booktitle={ICPR},\n  crossref={conf/icpr/2010}\n}\n"},"abstract":{"value":"In this paper, we present a method that recognizes single or multiple common actions between a pair of video sequences. We establish an energy function that evaluates geometric and photometric consistency, and solve the action recognition problem by optimizing the energy function. The proposed stochastic inference algorithm based on the Monte Carlo method explores the video pair from the local spatio-temporal interest point matches to find the common actions. Our algorithm works in unsupervised way without prior knowledge about the type and the number of common actions. Experiments show that our algorithm produces promising results on single and multiple action recognition."},"title":{"value":"Co-recognition of Actions in Video Pairs"},"authors":{"value":["Young Min Shin","Minsu Cho","Kyoung Mu Lee"]}},"tmdate":1747294269738,"pdate":1262304000000,"tcdate":1729113612294,"writers":["~"],"signatures":["~Minsu_Cho1"],"forum":"dLuKHgMCO1","license":"CC BY-SA 4.0","number":153526,"cdate":1262304000000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1747294269738,"domain":"DBLP.org","id":"dLuKHgMCO1","version":2},{"content":{"summary":{"value":"- The paper introduces Youku-mPLUG, the largest public Chinese video-language dataset and benchmarks, collected from Youku, a Chinese video-sharing website, with strict criteria of safety, diversity, and quality.\n- The paper also proposes mPLUG-video, a decoder-only video-language model that leverages a frozen large language model and a visual abstractor module to reduce the computation burden and improve the performance.\n- The paper evaluates mPLUG-video and other models on three downstream tasks: video category classification, video captioning, and video-text retrieval. The results show that mPLUG-video achieves good results in video category classification and video captioning, and demonstrates impressive zero-shot video instruction understanding ability."},"presentation":{"value":"3 good"},"contribution":{"value":"2 fair"},"soundness":{"value":"2 fair"},"strengths":{"value":"- The paper introduces a novel and large-scale Chinese video-language dataset and benchmarks, which can facilitate the research and development of video-language models for the Chinese language and culture. The paper also proposes a decoder-only model that leverages a frozen large language model and a visual abstractor module, which is a creative combination of existing ideas that reduces the computation burden and improves the performance.\n- The paper is well-written and organized, with clear figures and tables. The paper provides details and analysis on the proposed method and dataset. \n- The paper explains the problem statement, the motivation, the challenges, and the gap in the existing literature clearly in the abstract and introduction. The paper also describes the dataset collection, annotation, and preprocessing process, and provides some statistics and examples of the data. The paper also explains the model architecture, training, and fine-tuning process, and provides some examples.\n- The paper makes a significant contribution to the field of video-language modeling, especially for the Chinese language and culture. The paper presents a large-scale and diverse dataset that can enable various downstream tasks, such as video category classification, video captioning, video-text retrieval, and video instruction understanding. The paper also presents a state-of-the-art model that can achieve impressive results on these tasks."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- After downloading the dataset, it was found that there were many duplicate clips from the same source and static clips. Does the situation exist where these 400 million video clips come from the same original video? If so, during the filtering process, how is the quality of the selected videos ensured given the lack of quantifiable performance measures, such as CLIP similarity?\n- There is a lack of exploration into the status of text annotation in the dataset. Chinese and Latin languages such as English have significant differences in vocabulary, grammar, and sentence structure. The diversity of the text part of this dataset is not sufficiently demonstrated, and the text quality is slightly lower compared to the WebVid10M dataset. The paper should also compare the dataset with other existing video-language datasets, such as translated HowTo100M, WebVid10M or CNVid-3.5M[1], and discuss the advantages and limitations of the dataset.\n- This paper only explores the zero-shot capability in instruction understanding. Why not further investigate the zero-shot performance in video classification, retrieval, and description?\n- In instruction understanding, does VideoLLaMA also receive Chinese prompts? Has it been trained on Chinese instruction data? Comparing a MLLM trained on English datasets with one training in Chinese is unfair.\n- During data collection, the online model achieved a performance of about 94% in video category classification. However, in Table 4, the model trained by Youku-mPLUG actually performs worse than the unfiltered online model.\n\n----\n\nReference:\n[1] https://github.com/CNVid/CNVid-3.5M"},"confidence":{"value":"4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work."},"questions":{"value":"see weaknesses"},"rating":{"value":"3: reject, not good enough"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1701141593996,"tcdate":1698802395655,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission4734/Reviewer_kj3u"],"signatures":["ICLR.cc/2024/Conference/Submission4734/Reviewer_kj3u"],"forum":"mzxKLZNbrQ","number":1,"license":"CC BY 4.0","cdate":1698802395655,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission4734/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1701141593996,"domain":"ICLR.cc/2024/Conference","replyto":"mzxKLZNbrQ","id":"Qc212gAv2q","forumContent":{"venue":{"value":"Submitted to ICLR 2024"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Chinese","Video-language pre-training","video-language benchmarks","mPLUG","Youku","video captioning","video classification"]},"primary_area":{"value":"datasets and benchmarks"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"We firstly release the largest public Chinese high-quality video-language dataset named Youku-mPLUG, which is collected from Youku, a well-known Chinese video-sharing website, with strict criteria of safety, diversity, quality, and copyright. Youku-mPLUG contains 10 million Chinese video-text pairs filtered from 400 million raw videos across a wide range of 45 diverse categories for large-scale pre-training. In addition, to facilitate a comprehensive evaluation of video-language models, we carefully build the largest human-annotated Chinese benchmarks covering three popular video-language tasks across cross-modal retrieval, video captioning, and video category classification. \nWe also provide comprehensive benchmark evaluations of models across different architectures including encoder-only (i.e., ALPRO), encoder-decoder (i.e., mPLUG-2), and decoder-only (i.e., mPLUG-Video) for comparison. Especially, we train the first Chinese Multimodal LLM with only 1.7% trainable parameters for video understanding. Experiments show that models pre-trained on Youku-mPLUG gain up to 23.1% improvement in video category classification. Besides, mPLUG-video achieves a new state-of-the-art result on these benchmarks with 80.5% top-1 accuracy in video category classification and 68.9 CIDEr score in video captioning, respectively. Finally, the 2.7B version of mPLUG-video demonstrates impressive instruction and video understanding ability. The zero-shot instruction understanding experiment indicates that pretraining with Youku-mPLUG can enhance the ability to comprehend overall and detailed visual semantics, recognize scene text, and leverage open-domain knowledge."},"_bibtex":{"value":"@misc{\nxu2024youkumplug,\ntitle={Youku-m{PLUG}: A 10 Million Large-scale Chinese Video-Language Dataset for Pre-training and Benchmarks},\nauthor={Haiyang Xu and Qinghao Ye and Xuan Wu and Ming Yan and Yuan Miao and Jiabo Ye and Anwen Hu and Guohai Xu and Yaya Shi and Guangwei Xu and Chenliang Li and Qi Qian and Maofei Que and Ji Zhang and Xiao Zeng and Fei Huang},\nyear={2024},\nurl={https://openreview.net/forum?id=mzxKLZNbrQ}\n}"},"title":{"value":"Youku-mPLUG: A 10 Million Large-scale Chinese Video-Language Dataset for Pre-training and Benchmarks"},"pdf":{"value":"/pdf/614c835115f1443393e0caec779fcaaf14e25811.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Rejected_Submission"},"paperhash":{"value":"xu|youkumplug_a_10_million_largescale_chinese_videolanguage_dataset_for_pretraining_and_benchmarks"},"authorids":{"value":["~Haiyang_Xu1","~Qinghao_Ye1","~Xuan_Wu5","~Ming_Yan2","~Yuan_Miao1","~Jiabo_Ye1","~Anwen_Hu1","~Guohai_Xu1","~Yaya_Shi1","~Guangwei_Xu2","~Chenliang_Li2","~Qi_Qian1","~Maofei_Que1","~Ji_Zhang3","~Xiao_Zeng4","~Fei_Huang1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Haiyang Xu","Qinghao Ye","Xuan Wu","Ming Yan","Yuan Miao","Jiabo Ye","Anwen Hu","Guohai Xu","Yaya Shi","Guangwei Xu","Chenliang Li","Qi Qian","Maofei Que","Ji Zhang","Xiao Zeng","Fei Huang"]}},"version":2},{"content":{"summary":{"value":"The paper presents a new benchmark for understanding multi-shot videos, named Shot2Story. The benchmark provides a dataset of 42,958 short videos, each consisting of an average of 4.4 shots, with detailed textual descriptions, video summaries, and question-answering pairs for multi-modal understanding. It aims to enhance the understanding of videos by providing shot-level captions, human narrations, and generated summaries. Several distinct tasks are defined using this dataset: single-shot video captioning, multi-shot video summarization, and multi-shot video question answering. The paper also presents baseline models and demonstrates the challenges of long, comprehensive video summaries."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"1. Is it an optimal approach to use large language models (LLMs) to annotate datasets and then use those annotations for reasoning tasks by the LLMs themselves? This self-referential process might introduce biases or limitations in understanding. Is there a better approach to improving the depth of reasoning and multi-modal alignment, perhaps involving a more collaborative method of training distinct specialized models (e.g., one focused on annotation and another on reasoning)?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The paper is well-written, providing clear explanations of the benchmark, tasks, and dataset construction, making complex ideas accessible.\n2. Shot2Story sets a new standard for multi-shot video understanding, with detailed shot-level captions, audio narrations, and summaries that advance research in video analysis and multi-modal understanding.\n3. The benchmark's unique focus on multi-shot videos with rich shot-level annotations differentiates it from previous benchmarks, contributing to a better understanding of event transitions in multi-shot videos.\n4. The authors conduct thorough experiments using baseline models like MiniGPT-4 and VideoChat2, providing valuable insights into current models' strengths and limitations, and establishing a solid foundation for benchmarking future models."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. There is a format issue in Table 1 that affects its readability. Consistent formatting is essential for clarity and for the accurate presentation of comparative data.\n2. The motivation for annotating audio text descriptions is not clearly articulated. Understanding audio alongside video is crucial for comprehensive multimedia understanding, especially in tasks involving multi-modal large language models (MLLMs). Merely using textual descriptions without integrating auditory context reduces the potential for deep reasoning capabilities across modalities, which limits the advancement of MLLM applications.\n3. The experimental section lacks baselines from well-regarded models such as the Qwen-VL series. Including such comparisons would provide a stronger benchmark and enable a more effective evaluation of the proposed model’s performance.\n4. The paper does not propose any self-developed model tailored specifically for the Shot2Story benchmark, which would demonstrate the specific advantages and limitations of the benchmark through a purpose-built approach."}},"nonreaders":[],"tmdate":1731428922042,"tcdate":1730522570655,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission7580/Reviewer_YdUF"],"signatures":["ICLR.cc/2025/Conference/Submission7580/Reviewer_YdUF"],"forum":"FZv3kPHTtB","number":1,"license":"CC BY 4.0","cdate":1730522570655,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission7580/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428922042,"domain":"ICLR.cc/2025/Conference","replyto":"FZv3kPHTtB","id":"1zsd3ed3Ij","forumContent":{"venue":{"value":"ICLR 2025 Poster"},"TLDR":{"value":"Shot2Story presents a large-scale dataset with 43,000 multi-shot videos and 188,000 manually annotated shots, offering detailed visual/audio captions, summaries, and QA pairs to advance multi-shot video understanding."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["vision language model","video question answering","video captioning","multi-shot videos"]},"supplementary_material":{"value":"/attachment/19e43f5040b7d5e255d15b50136dc016857cf8ad.pdf"},"primary_area":{"value":"datasets and benchmarks"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"A short clip of video may contain progression of multiple events and an interesting story line. A human need to capture both the event in every shot and associate them together to understand the story behind it. In this work, we present a new multi-shot video understanding benchmark \\dataset with detailed shot-level captions, comprehensive video summaries and question-answering pairs. To facilitate better semantic understanding of videos, we provide captions for both visual signals and human narrations. We design several distinct tasks including single-shot video captioning, multi-shot video summarization, and multi-shot video question answering. Preliminary experiments show some challenges to generate a long and comprehensive video summary for multi-shot videos. Nevertheless, the generated imperfect summaries can already achieve competitive performance on existing video understanding tasks such as video question-answering, promoting an under-explored setting of video understanding with detailed summaries."},"_bibtex":{"value":"@inproceedings{\nhan2025shotstory,\ntitle={Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos},\nauthor={Mingfei Han and Linjie Yang and Xiaojun Chang and Lina Yao and Heng Wang},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=FZv3kPHTtB}\n}"},"title":{"value":"Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos"},"pdf":{"value":"/pdf/316c2e51532fadb7d781d28a19d13e78f4aa4ca0.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"han|shot2story_a_new_benchmark_for_comprehensive_understanding_of_multishot_videos"},"authorids":{"value":["~Mingfei_Han1","~Linjie_Yang4","~Xiaojun_Chang4","~Lina_Yao2","~Heng_Wang2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Mingfei Han","Linjie Yang","Xiaojun Chang","Lina Yao","Heng Wang"]}},"version":2},{"content":{"summary":{"value":"This paper introduces LatentWarp, a framework for zero-shot video-to-video translation using image diffusion models. It addresses the challenge of maintaining temporal consistency in generated video frames. LatentWarp focuses on constraining query tokens for temporal consistency. It achieves this by warping the latent features from the previous frame to align with the current frame using optical flow information.  Extensive experiments confirm the superiority of LatentWarp in achieving high-quality video-to-video translation with temporal coherence."},"presentation":{"value":"2 fair"},"contribution":{"value":"2 fair"},"soundness":{"value":"2 fair"},"strengths":{"value":"The introduction of the LatentWarp framework offers a novel approach to zero-shot video-to-video translation. LatentWarp  emphasizes on preserving temporal consistency during video frame generation, achieved through optical flow and warping operations, significantly enhances temporal coherence,  which is a crucial aspect of video generation. The writing is good, and the structure of the paper is clear."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"I find this method to be quite intuitive. My concerns mainly pertain to the experimental aspects:\n\n1. Data-related issues: The authors do not compare their method to datasets used in their previous work like \"tune-a-video.\" This omission may undermine the fairness of the experimental results.\n\n2. Base model choice: The authors employ ControlNet as the base model instead of using LDM directly. ControlNet offers strong structural control, which might make the improvement from LatentWarp seem relatively small. It would be beneficial to provide experimental configurations with LatentWarp combined with LDM or Tune-a-Video to showcase this point.\n\n3. Quantitative data details: The authors seem to have omitted reporting the sizes of the datasets used in their quantitative experiments.\n\n4. Compared methods: It is advisable for the authors to compare their method with a broader range of existing approaches, such as Video-P2P and more recent methods.\n\n5. Supplementary material: The authors have not provided corresponding video supplementary materials to visually assess temporal consistency.\n\n6. User surveys: The authors did not provide user surveys as prior works have done. This is important to evaluate the visual effect.\n\n7. Running costs: Editing time and GPU resource consumption, should be reported and compared to help readers understand the resource requirements and efficiency.\n\nThese concerns should be addressed to enhance the completeness and rigor of the experimental evaluation in the study."},"confidence":{"value":"5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully."},"questions":{"value":"See weaknesses. I will rejudge the rating according to the rebuttal."},"rating":{"value":"6: marginally above the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1701427579203,"tcdate":1698581098189,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission936/Reviewer_9TQZ"],"signatures":["ICLR.cc/2024/Conference/Submission936/Reviewer_9TQZ"],"forum":"ZJHdiYDD5k","number":1,"license":"CC BY 4.0","cdate":1698581098189,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission936/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1701427579203,"domain":"ICLR.cc/2024/Conference","replyto":"ZJHdiYDD5k","id":"6MlUzKmMe1","forumContent":{"venue":{"value":"Submitted to ICLR 2024"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["diffusion model","video generation"]},"primary_area":{"value":"generative models"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Leveraging the generative ability of image diffusion models offers great potential for zero-shot video-to-video translation. The key lies in how to maintain temporal consistency across generated video frames by image diffusion models. Previous methods typically adopt cross-frame attention, i.e., sharing the key and value tokens across attentions of different frames, to encourage the temporal consistency. However, in those works, temporal inconsistency issue may not be thoroughly solved, rendering the fidelity of generated videos limited. In this paper, we find the bottleneck lies in the unconstrained query tokens and propose a new zero\u0002shot video-to-video translation framework, named LatentWarp. Our approach is simple: to constrain the query tokens to be temporally consistent, we further incorporate a warping operation on the latent space to constrain the query tokens. Specifically, based on the optical flow obtained from the original video, we warp the generated latent features of last frame to align with the current frame during the denoising process. As a result, the corresponding regions across the adjacent frames can share closely-related query tokens and attention outputs, which can further improve latent-level consistency to enhance visual temporal coherence of generated videos. Extensive experiment results demonstrate the superiority of LatentWarp in achieving video-to-video translation with temporal coherence."},"_bibtex":{"value":"@misc{\nbao2024latentwarp,\ntitle={LatentWarp: Consistent Diffusion Latents for Zero-Shot Video-to-Video Translation},\nauthor={Yuxiang Bao and Di Qiu and Guoliang Kang and Baochang Zhang and Bo Jin and Kaiye Wang and Pengfei Yan},\nyear={2024},\nurl={https://openreview.net/forum?id=ZJHdiYDD5k}\n}"},"title":{"value":"LatentWarp: Consistent Diffusion Latents for Zero-Shot Video-to-Video Translation"},"pdf":{"value":"/pdf/6e1de8a0488cc4e317f5632c26f19294de23df54.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Rejected_Submission"},"paperhash":{"value":"bao|latentwarp_consistent_diffusion_latents_for_zeroshot_videotovideo_translation"},"authorids":{"value":["~Yuxiang_Bao1","~Di_Qiu2","~Guoliang_Kang1","~Baochang_Zhang1","~Bo_Jin3","~Kaiye_Wang2","~Pengfei_Yan1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Yuxiang Bao","Di Qiu","Guoliang Kang","Baochang Zhang","Bo Jin","Kaiye Wang","Pengfei Yan"]}},"version":2},{"content":{"venue":{"value":"CVPR 2023"},"pdf":{"value":"https://ieeexplore.ieee.org/iel7/10203037/10203050/10204453.pdf"},"venueid":{"value":"dblp.org/conf/CVPR/2023"},"paperhash":{"value":"lee|shapeaware_textdriven_layered_video_editing"},"authorids":{"value":["~Yao-Chih_Lee1","https://dblp.org/search/pid/api?q=author:Ji-Ze_Genevieve_Jang:","~Yi-Ting_Chen3","https://dblp.org/search/pid/api?q=author:Elizabeth_Qiu:","https://dblp.org/search/pid/api?q=author:Jia-Bin_Huang_0001:"]},"html":{"value":"https://doi.org/10.1109/CVPR52729.2023.01376"},"_bibtex":{"value":"@inproceedings{DBLP:conf/cvpr/LeeJCQ023,\n  author={Yao-Chih Lee and Ji-Ze Genevieve Jang and Yi-Ting Chen and Elizabeth Qiu and Jia-Bin Huang},\n  title={Shape-Aware Text-Driven Layered Video Editing},\n  year={2023},\n  cdate={1672531200000},\n  pages={14317-14326},\n  url={https://doi.org/10.1109/CVPR52729.2023.01376},\n  booktitle={CVPR},\n  crossref={conf/cvpr/2023}\n}\n"},"abstract":{"value":"Temporal consistency is essential for video editing applications. Existing work on layered representation of videos allows propagating edits consistently to each frame. These methods, however, can only edit object appearance rather than object shape changes due to the limitation of using a fixed UV mapping field for texture atlas. We present a shape-aware, text-driven video editing method to tackle this challenge. To handle shape changes in video editing, we first propagate the deformation field between the input and edited keyframe to all frames. We then leverage a pre-trained text-conditioned diffusion model as guidance for refining shape distortion and completing unseen regions. The experimental results demonstrate that our method can achieve shape-aware consistent video editing and compare favorably with the state-of-the-art."},"title":{"value":"Shape-Aware Text-Driven Layered Video Editing"},"authors":{"value":["Yao-Chih Lee","Ji-Ze Genevieve Jang","Yi-Ting Chen","Elizabeth Qiu","Jia-Bin Huang"]}},"tmdate":1728852392360,"pdate":1672531200000,"tcdate":1728170939680,"writers":["~"],"signatures":["~Yi-Ting_Chen3"],"forum":"O3q6SEgwtI","license":"CC BY-SA 4.0","number":140383,"cdate":1672531200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1728852392360,"domain":"DBLP.org","id":"O3q6SEgwtI","version":2},{"content":{"summary":{"value":"This paper presents a novel method for neural video compression (NVC) that adapts to different video contents effectively. The proposed approach, named Group-aware Parameter-efficient Updating (GPU), addresses two main challenges in neural video compression: error accumulation and the complexity of updating large encoder parameters. The GPU method uses a group-aware strategy for parameter updates and introduces lightweight adapter modules for efficient optimization. The method segments videos into patch-based groups of pictures (GoPs) and updates them sequentially, reducing memory demands and improving compression efficiency. The approach is validated on HEVC test sequences, showing superior performance compared to existing methods."},"suitability":{"value":3},"strengths":{"value":"1. The use of group-aware parameter-efficient updating and lightweight adapters represents a significant advancement in content-adaptive NVC.\n2. Group-aware parameter-efficient updating analyzes the video sequence and adjust GOP size dynamically.\n3. The adapter is very light weight, only 3 conv layers and 2 LeakyRelu.\n4. A patch-based GoP updating strategy significantly reduces computational complexity and memory requirements."},"confidence":{"value":4},"rating":{"value":4},"limitations":{"value":"1. The method involves complex segmentation and updating strategies, including online training (adapter) which might be challenging to implement and integrate into existing systems when 2-pass coding is disabled.\n2. The process of fine-tuning and adapting the model to new content can be computationally intensive, limiting its use to Random Access mode, it cannot be use in LDB mode as future frames are not available in LDB mode."}},"nonreaders":[],"tmdate":1721657127368,"tcdate":1716327701724,"writers":["acmmm.org/ACMMM/2024/Conference","acmmm.org/ACMMM/2024/Conference/Submission202/Reviewer_KMFd"],"signatures":["acmmm.org/ACMMM/2024/Conference/Submission202/Reviewer_KMFd"],"forum":"KOIoLNLTAW","number":1,"license":"CC BY 4.0","cdate":1716327701724,"readers":["everyone"],"invitations":["acmmm.org/ACMMM/2024/Conference/Submission202/-/Official_Review","acmmm.org/ACMMM/2024/Conference/-/Edit","acmmm.org/ACMMM/2024/Conference/Submission202/Official_Review1/-/Final_Rating"],"mdate":1721657127368,"domain":"acmmm.org/ACMMM/2024/Conference","replyto":"KOIoLNLTAW","id":"ltR0pf10MX","forumContent":{"venue":{"value":"MM2024 Poster"},"supplementary_material":{"value":"/attachment/6222af4b69936d9e3d8e2055a4d6c82ccdeb4af3.zip"},"abstract":{"value":"Content-adaptive compression is crucial for enhancing the adaptability of the pre-trained neural codec for various contents. Although these methods have been very practical in neural image compression (NIC), their application in neural video compression (NVC) is still limited due to two main aspects: 1), video compression relies heavily on temporal redundancy, therefore updating just one or a few frames can lead to significant errors accumulating over time; 2), NVC frameworks are generally more complex, with many large components that are not easy to update quickly during encoding. To address the previously mentioned challenges, we have developed a content-adaptive NVC technique called Group-aware Parameter-Efficient Updating (GPU). Initially, to minimize error accumulation, we adopt a group-aware approach for updating encoder parameters. This involves adopting a patch-based Group of Pictures (GoP) training strategy to segment a video into patch-based GoPs, which will be updated to facilitate a globally optimized domain-transferable solution. Subsequently, we introduce a parameter-efficient delta-tuning strategy, which is achieved by integrating several light-weight adapters into each coding component of the encoding process by both serial and parallel configuration. Such architecture-agnostic modules stimulate the components with large parameters, thereby reducing both the update cost and the encoding time. We incorporate our GPU into the latest NVC framework and conduct comprehensive experiments, whose results showcase outstanding video compression efficiency across four video benchmarks and adaptability of one medical image benchmark."},"relevance_to_conference":{"value":"Video coding is a crucial technique for the transport and delivery of multimedia content. Developing such a system presents challenges but has attracted significant attention from multimedia researchers. In this submission, we investigate innovative methods for content-adaptive NVC. We propose a novel strategy, termed GPU, notable for its patch-based GoP updating mechanism and encoder-side adaptor modules, tailored for diverse configurations. Our extensive experiments validate our GPU's effectiveness and adaptability, demonstrating its integration with contemporary NVC framework for the compression of both standard videos and specialized medical MRI sequences. Remarkably, it consistently surpasses state-of-the-art video compression algorithms, such as H.266/VVC and DCVC_DC, thereby establishing a formidable benchmark for both video and MRI compression. These achievements underscore our GPU approach as a robust and efficient baseline for content-adaptive NVC methods. As the landscape of NVC evolves to incorporate more intricate techniques, our architecture-agnostic updating strategy is anticipated to increase, offering a solution to reduce computational requirements while maximizing efficiency. Moreover, such efficient content-adaptive NVC is poised to broaden the scope of NVC technologies, enabling them to address more extensive applications of complex content types beyond standard video. This research heralds new avenues for the compression of more diverse content."},"_bibtex":{"value":"@inproceedings{\nchen2024groupaware,\ntitle={Group-aware Parameter-efficient Updating for Content-Adaptive Neural Video Compression},\nauthor={Zhenghao Chen and Luping Zhou and Zhihao Hu and Dong Xu},\nbooktitle={ACM Multimedia 2024},\nyear={2024},\nurl={https://openreview.net/forum?id=KOIoLNLTAW}\n}"},"title":{"value":"Group-aware Parameter-efficient Updating for Content-Adaptive Neural Video Compression"},"secondary_subject_area":{"value":["[Experience] Multimedia Applications"]},"pdf":{"value":"/pdf/28da6583071959912251156602d8e5020ebc7ced.pdf"},"venueid":{"value":"acmmm.org/ACMMM/2024/Conference"},"paperhash":{"value":"chen|groupaware_parameterefficient_updating_for_contentadaptive_neural_video_compression"},"primary_subject_area":{"value":"[Systems] Transport and Delivery"},"authorids":{"value":["~Zhenghao_Chen2","~Luping_Zhou3","~Zhihao_Hu1","~Dong_Xu2"]},"authors":{"value":["Zhenghao Chen","Luping Zhou","Zhihao Hu","Dong Xu"]}},"version":2},{"content":{"summary":{"value":"This work focuses on efficient video processing framework for video VLMs. It proposes a pipeline to first generates a high-level overview of the entire video and then adaptively zooms in on specific parts based on the content being generated. Experiments show effectiveness on the video detailed description dataset."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"My main concern focuses on the experiment part. Current experiments are not enought to support the general efficient framework for VLM."},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The efficient video compression for VLM is promising and useful for practical usage.\n2. The course-to-fine design for video understanding is interesting and seems to be useful in caption.\n3. Overall, the writing is clear and easy to follow."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. From Figure 2, the proposed ZoomVLM utilizes LLM twice for image-level and token level selection. But the efficency in Table 1 is seems better for vanilla manner in Token/sec. Is the efficiency mainly from reduction in video tokens? Considering provide the whole inference time and time spent on each component (video overview generation, token adjustment, etc.) for clear comparison. It could be better to compare other metrics like memory usage or FLOPs.\n2. The evalution in Video description is far from enough. The authors are recommended to conduct experiments on VideoMME [A] and EgoSchema [B], which could be more close to real-life scenes. Of course, it's better to provide some analysis or limitation on different benchmarks.\n3. One of the main drawback in current pipeline could be the multi-round QA, which is more useful in practical applications. Because the Vanilla or Slowfast do no need to generate the token again for different round. Are the authors have any solutions or ideas to this tasks or potential optimizations for repeated queries on the same video? \n4. Because the video overview augmenter and adaptive token adjustment are all target to reduce reduent tokens, why not only keep the adaptive token adjustment with larger reduction rate? The authors are recommended to discuss any potential synergies or trade-offs between the two approaches.\n\n[A] \"Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis\", arXiv:2405.21075, 2024\n\n[B] \"Egoschema: A diagnostic benchmark for very long-form video language understanding\", NeurIPS, 2023"}},"nonreaders":[],"tmdate":1731429609754,"tcdate":1730142879591,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission12635/Reviewer_hSzd"],"signatures":["ICLR.cc/2025/Conference/Submission12635/Reviewer_hSzd"],"forum":"689MfSyeNz","number":2,"license":"CC BY 4.0","cdate":1730142879591,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission12635/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731429609754,"domain":"ICLR.cc/2025/Conference","replyto":"689MfSyeNz","id":"rDd3kk6lXS","forumContent":{"TLDR":{"value":"We propose a tuning-free framework that boosts video vision-language model efficiency without sacrificing accuracy by adaptively zooming parts based on attention."},"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Vision Language Model","Multi-modal"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Recent advances in vision-language models (VLMs) have led to impressive progress in video understanding. However, despite their promising performance, existing state-of-the-art (SOTA) solutions require an excessive number of tokens (e.g., up to 6,272 tokens in the Llava-OneVision model) to represent input videos, leading to a non-negligible bottleneck in inference efficiency. Motivated by findings in human perception, where individuals first focus on high-level overviews and then zoom into specific areas for detailed information, we hypothesize that a similar approach can enhance the inference efficiency of VLMs by reducing the number of tokens needed to represent videos. Based on this hypothesis, we propose ZoomVLM, a tuning-free, plug-and-play efficient video processing framework for video VLMs. ZoomVLM first generates an overview of the entire video and then adaptively zooms in and out on different parts based on the content being generated. Our key insight is that the attention distributions in the Large Language Model (LLM) within the VLM can provide sensible guidance on where to focus (by allocating more tokens) and where to discard (by dropping tokens) during inference. Specifically, ZoomVLM integrates two key components: (1) a Video Overview Augmenter, which enables cost-effective high-level understanding by augmenting downsampled video overview with a few high-resolution keyframes; and (2) an Adaptive Token Adjustment, which predicts the importance of different video parts in the upcoming generation process and adjusts the number of tokens allocated to each part according to their importance. Extensive experiments and ablation studies across two challenging open-ended video understanding benchmarks and four models validate that ZoomVLM effectively improves inference efficiency by reducing the number of tokens and boosting throughput in terms of the number of generated tokens per second without degradation in achievable accuracy. Specifically, when applying ZoomVLM to Llava-Next-Video-7B-DPO, ZoomVLM achieves a 30\\% higher token generation rate with a 0.259 improvement in the Video Detail Description score."},"_bibtex":{"value":"@misc{\nyu2025zoomvlm,\ntitle={Zoom{VLM}: A Tuning-Free Framework for Efficient Video Understanding via Adaptive Zooming in Vision-Language Models},\nauthor={Zhongzhi Yu and Zheng Wang and Zhenyang Chen and Chaojian Li and Hyewon Suh and Yonggan Fu and Dachuan Shi and Hongxu Yin and Jan Kautz and Pavlo Molchanov and Yingyan Celine Lin},\nyear={2025},\nurl={https://openreview.net/forum?id=689MfSyeNz}\n}"},"title":{"value":"ZoomVLM: A Tuning-Free Framework for Efficient Video Understanding via Adaptive Zooming in Vision-Language Models"},"pdf":{"value":"/pdf/9f3ba3b65cad1d5703e74cf486afe230131d4e60.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"yu|zoomvlm_a_tuningfree_framework_for_efficient_video_understanding_via_adaptive_zooming_in_visionlanguage_models"},"authorids":{"value":["~Zhongzhi_Yu1","~Zheng_Wang38","~Zhenyang_Chen1","~Chaojian_Li1","~Hyewon_Suh1","~Yonggan_Fu1","~Dachuan_Shi2","~Hongxu_Yin2","~Jan_Kautz1","~Pavlo_Molchanov1","~Yingyan_Celine_Lin1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Zhongzhi Yu","Zheng Wang","Zhenyang Chen","Chaojian Li","Hyewon Suh","Yonggan Fu","Dachuan Shi","Hongxu Yin","Jan Kautz","Pavlo Molchanov","Yingyan Celine Lin"]}},"version":2},{"content":{"venue":{"value":"EMNLP 2025"},"pdf":{"value":"https://aclanthology.org/2025.emnlp-main.1308.pdf"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"jeon|sali4vid_saliencyaware_video_reweighting_and_adaptive_caption_retrieval_for_dense_video_captioning"},"html":{"value":"https://doi.org/10.18653/v1/2025.emnlp-main.1308"},"_bibtex":{"value":"@inproceedings{DBLP:conf/emnlp/JeonKKKK25,\n  author={MinJu Jeon and Si-Woo Kim and Ye-Chan Kim and HyunGee Kim and Dong-Jin Kim},\n  title={Sali4Vid: Saliency-Aware Video Reweighting and Adaptive Caption Retrieval for Dense Video Captioning},\n  year={2025},\n  cdate={1735689600000},\n  pages={25777-25790},\n  url={https://doi.org/10.18653/v1/2025.emnlp-main.1308},\n  booktitle={EMNLP},\n  crossref={conf/emnlp/2025}\n}\n"},"abstract":{"value":"Dense video captioning aims to temporally localize events in video and generate captions for each event. While recent works propose end-to-end models, they suffer from two limitations: (1) applying timestamp supervision only to text while treating all video frames equally, and (2) retrieving captions from fixed-size video chunks, overlooking scene transitions. To address these, we propose **Sali4Vid**, a simple yet effective saliency-aware framework. We introduce Saliency-aware Video Reweighting, which converts timestamp annotations into sigmoid-based frame importance weights, and Semantic-based Adaptive Caption Retrieval, which segments videos by frame similarity to capture scene transitions and improve caption retrieval. Sali4Vid achieves state-of-the-art results on YouCook2 and ViTT, demonstrating the benefit of jointly improving video weighting and retrieval for dense video captioning."},"title":{"value":"Sali4Vid: Saliency-Aware Video Reweighting and Adaptive Caption Retrieval for Dense Video Captioning"},"authors":{"value":[{"fullname":"MinJu Jeon","username":"~MinJu_Jeon1"},{"fullname":"Si-Woo Kim","username":""},{"fullname":"Ye-Chan Kim","username":""},{"fullname":"HyunGee Kim","username":""},{"fullname":"Dong-Jin Kim","username":""}]}},"tmdate":1779763960646,"pdate":1767139200000,"externalIds":["dblp:conf/emnlp/JeonKKKK25"],"tcdate":1779763957633,"writers":["~","OpenReview.net/Public_Article/DBLP.org","OpenReview.net/Support"],"signatures":["~MinJu_Jeon1"],"forum":"B7BQQXWyYS","license":"CC BY-SA 4.0","number":25879,"cdate":1735689600000,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/DBLP.org/-/Record","OpenReview.net/Public_Article/-/Edit"],"mdate":1779763960646,"domain":"OpenReview.net/Public_Article","id":"B7BQQXWyYS","version":2},{"content":{"venue":{"value":"CoRR 2025"},"pdf":{"value":"http://arxiv.org/pdf/2509.04602v1"},"venueid":{"value":"dblp.org/journals/CORR/2025"},"paperhash":{"value":"jeon|sali4vid_saliencyaware_video_reweighting_and_adaptive_caption_retrieval_for_dense_video_captioning"},"authorids":{"value":["~MinJu_Jeon1","","https://dblp.org/search/pid/api?q=author:Ye-Chan_Kim:","https://dblp.org/search/pid/api?q=author:HyunGee_Kim:","https://dblp.org/search/pid/api?q=author:Dong-Jin_Kim:"]},"html":{"value":"https://doi.org/10.48550/arXiv.2509.04602"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2509-04602,\n  publtype={informal},\n  author={MinJu Jeon and Si-Woo Kim and Ye-Chan Kim and HyunGee Kim and Dong-Jin Kim},\n  title={Sali4Vid: Saliency-Aware Video Reweighting and Adaptive Caption Retrieval for Dense Video Captioning},\n  year={2025},\n  month={September},\n  cdate={1756684800000},\n  journal={CoRR},\n  volume={abs/2509.04602},\n  url={https://doi.org/10.48550/arXiv.2509.04602}\n}\n"},"abstract":{"value":"Dense video captioning aims to temporally localize events in video and generate captions for each event. While recent works propose end-to-end models, they suffer from two limitations: (1) applying timestamp supervision only to text while treating all video frames equally, and (2) retrieving captions from fixed-size video chunks, overlooking scene transitions. To address these, we propose Sali4Vid, a simple yet effective saliency-aware framework. We introduce Saliency-aware Video Reweighting, which converts timestamp annotations into sigmoid-based frame importance weights, and Semantic-based Adaptive Caption Retrieval, which segments videos by frame similarity to capture scene transitions and improve caption retrieval. Sali4Vid achieves state-of-the-art results on YouCook2 and ViTT, demonstrating the benefit of jointly improving video weighting and retrieval for dense video captioning"},"title":{"value":"Sali4Vid: Saliency-Aware Video Reweighting and Adaptive Caption Retrieval for Dense Video Captioning"},"authors":{"value":["MinJu Jeon","Si-Woo Kim","Ye-Chan Kim","HyunGee Kim","Dong-Jin Kim"]}},"tmdate":1779763957814,"pdate":1735689600000,"externalIds":["dblp:journals/corr/abs-2509-04602"],"tcdate":1762312176996,"writers":["~"],"signatures":["~Si-Woo_Kim1"],"forum":"mO5zOGIB3q","license":"CC BY-SA 4.0","number":655537,"cdate":1756684800000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1779763957814,"domain":"DBLP.org","id":"mO5zOGIB3q","version":2},{"content":{"venue":{"value":"Empirical Methods in Natural Language Processing"},"pdf":{"value":"/pdf/2ea23a1340724b0e05fe0abf1a4ca504c70737ad.pdf"},"venueid":{"value":"OpenReview.net/Archive"},"paperhash":{"value":"jeon|sali4vid_saliencyaware_video_reweighting_and_adaptive_caption_retrieval_for_dense_video_captioning"},"authorids":{"value":["~MinJu_Jeon1","~Si-Woo_Kim1","~Ye-Chan_Kim1","~HyunGee_Kim1","~Dong-Jin_Kim1"]},"html":{"value":"https://arxiv.org/pdf/2509.04602"},"abstract":{"value":"Dense video captioning aims to temporally localize events in video and generate captions for each event. While recent works propose end-to-end models, they suffer from two limitations: (1) applying timestamp supervision only to text while treating all video frames equally, and (2) retrieving captions from fixed-size video chunks, overlooking scene transitions. To address these, we propose Sali4Vid, a simple yet effective saliency-aware framework. We introduce Saliency-aware Video Reweighting, which converts timestamp annotations into sigmoid-based frame importance weights, and Semantic-based Adaptive Caption Retrieval, which segments videos by frame similarity to capture scene transitions and improve caption retrieval. Sali4Vid achieves state-of-the-art results on YouCook2 and ViTT, demonstrating the benefit of jointly improving video weighting and retrieval for dense video captioning."},"title":{"value":"Sali4Vid: Saliency-Aware Video Reweighting and Adaptive Caption Retrieval for Dense Video Captioning"},"authors":{"value":["MinJu Jeon","Si-Woo Kim","Ye-Chan Kim","HyunGee Kim","Dong-Jin Kim"]}},"tmdate":1762944949156,"pdate":1762182000000,"tcdate":1762944949156,"writers":["~MinJu_Jeon1","~Si-Woo_Kim1","~Ye-Chan_Kim1","~HyunGee_Kim1","~Dong-Jin_Kim1"],"signatures":["~MinJu_Jeon1"],"forum":"EuaZYct9X7","license":"CC BY 4.0","number":40836,"cdate":1762944949156,"readers":["everyone"],"invitations":["OpenReview.net/Archive/-/Direct_Upload"],"mdate":1762944949156,"domain":"OpenReview.net/Archive","id":"EuaZYct9X7","version":2},{"content":{"summary":{"value":"The paper introduces Trans4D for generating large deformation within the text-to-4D generative models. The authors divide the process into three steps. Firstly, they propose a physics-aware 4D transition planning based on multi-modal large language models, providing the base for 4D scene initialization. Secondly, Trans4D includes a geometry-aware 4D scene transition module. Specifically, this module will determine whether a Gaussian point will appear at each timestep. Finally, a refining process has also be used to further improve the result's quality. Extensive experiments have been conducted and provided to demonstrate the performance and effectiveness of the proposed network."},"soundness":{"value":2},"confidence":{"value":5},"questions":{"value":"Please see the above strengths and weaknesses"},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"+ The paper tackles the interesting and important task of bringing more deformation to current text-to-4D generative pipelines.\n\n+ The paper is well-written and easy to follow.\n\n+ The proposed geometry-aware 4D transition planning and geometry-aware 4D scene transition module is reasonable and showcased to be useful."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The quality is still far away from satisfaction. See results in Fig.1 and Fig.3. The results are not good when two different objects interact with each other. The same issues can also be found in the video provided in the supplementary. I recognize it as the most severe issue of this paper.\n\n- The geometry-aware 4D transition network will only predict the appearance of each point, which may be the cause for the previous weakness. Meanwhile, it may also make this method not suitable for modeling articulate deformation, limiting the applications of the proposed Trans4D.\n\n- More comparisons with image-to-4D or video-to-4D pipelines would be better."}},"nonreaders":[],"tmdate":1731427196394,"tcdate":1730743924895,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission106/Reviewer_cfQT"],"signatures":["ICLR.cc/2025/Conference/Submission106/Reviewer_cfQT"],"forum":"gkOtsxD6fr","number":4,"license":"CC BY 4.0","cdate":1730743924895,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission106/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427196394,"domain":"ICLR.cc/2025/Conference","replyto":"gkOtsxD6fr","id":"aYIrN82Sx0","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Text-to-4D generation","Scene transition"]},"supplementary_material":{"value":"/attachment/025c8c52c4ee617f8741560fbbd1e697257575b1.zip"},"primary_area":{"value":"generative models"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Recent advances in diffusion models have demonstrated exceptional capabilities in image and video generation, further improving the effectiveness of 4D synthesis. Existing 4D generation methods can generate high-quality 4D objects or scenes based on user-friendly conditions, benefiting the gaming and video industries. However, these methods struggle to synthesize significant object deformation of complex 4D transitions and interactions within scenes. To address this challenge, we propose **Trans4D**, a novel text-to-4D synthesis framework that enables realistic complex scene transitions. Specifically, we first use multi-modal large language models (MLLMs) to produce a physic-aware scene description for 4D scene initialization and effective transition timing planning. Then we propose a geometry-aware 4D transition network to realize a complex scene-level 4D transition based on the plan, which involves expressive geometrical object deformation. Extensive experiments demonstrate that **Trans4D** consistently outperforms existing state-of-the-art methods in generating 4D scenes with accurate and high-quality transitions, validating its effectiveness."},"_bibtex":{"value":"@misc{\nzeng2024transd,\ntitle={Trans4D: Realistic Geometry-Aware Transition for Compositional Text-to-4D Synthesis},\nauthor={Bohan Zeng and Ling Yang and Siyu Li and Jiaming Liu and Zixiang Zhang and Juanxi Tian and Kaixin Zhu and yongzhen.gyz and Fu-Yun Wang and Minkai Xu and Stefano Ermon and Wentao Zhang},\nyear={2024},\nurl={https://openreview.net/forum?id=gkOtsxD6fr}\n}"},"title":{"value":"Trans4D: Realistic Geometry-Aware Transition for Compositional Text-to-4D Synthesis"},"pdf":{"value":"/pdf/05fa6e6ac8bc3964f0b47a19762037431ab92170.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"zeng|trans4d_realistic_geometryaware_transition_for_compositional_textto4d_synthesis"},"authorids":{"value":["~Bohan_Zeng1","~Ling_Yang1","~Siyu_Li2","~Jiaming_Liu7","~Zixiang_Zhang1","~Juanxi_Tian1","~Kaixin_Zhu1","~yongzhen.gyz1","~Fu-Yun_Wang1","~Minkai_Xu1","~Stefano_Ermon1","~Wentao_Zhang1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Bohan Zeng","Ling Yang","Siyu Li","Jiaming Liu","Zixiang Zhang","Juanxi Tian","Kaixin Zhu","yongzhen.gyz","Fu-Yun Wang","Minkai Xu","Stefano Ermon","Wentao Zhang"]}},"version":2},{"content":{"summary":{"value":"This paper presents a dense temporal reasoning benchmark, named Vinoground, for checking CLIP-based and video LLM's capabilities on fundamental temporal counterfactual reasoning tasks. Vinoground contains 1000 video-caption pairs. The authors discuss how they collect counterfactual captions and curate videos for this benchmark. Also, the evaluation can be splitted into three main catergories: object, action, and viewpoint, or four minor categories: interaction, cyclical, spatial, contextual. The authors then discuss. how to evaluate models, either CLIP-based or LLM-based, on this benchmark, and find that today's sota models are struggling to do such temporal counterfactual reasoning tasks, while average crowdsourcing people perform well. The authors present detailed discussions about this results as well."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"See weakness."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"This is overall a good paper investigating the dark side of today's video language models, especially video large language models, on dense temporal counterfactual tasks. The findings are quit interesting: 1. CLIP-based models perform worse than random guessing; 2. Chain-of-thoughts prompting can improve GPT4o but do not help most open-source video LLMs; 3. More frames (larger than 32 for near all models) do not lead to a better result on this dense temporal reasoning task. These may draw to a conclusion: today's video LLMs are performing simple pattern matching  (rather than reasoning). In my view, these findings are valuable and will have impact on further researches on this domain."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"I have no major concerns but it will be more complete that the authors present detailed studies of sota video LLM results on Vinoground, including: 1. on which samples models are prone to fail (hard samples, and their things in common), and easy samples vice versa; 2. why CLIP-based models are strugging with random guessing? Is it due to incapable language encoders? Or incapable visual encoders? 3. Results of temporal random-permuted frames as these results can be more representative than `random chance` in the table since Vinoground is about temporal reasoning."}},"nonreaders":[],"tmdate":1731427456054,"tcdate":1730640457729,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission1636/Reviewer_ySRv"],"signatures":["ICLR.cc/2025/Conference/Submission1636/Reviewer_ySRv"],"forum":"a1P5kh2oo8","number":4,"license":"CC BY 4.0","cdate":1730640457729,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission1636/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427456054,"domain":"ICLR.cc/2025/Conference","replyto":"a1P5kh2oo8","id":"xu4T7t1Gk4","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"TLDR":{"value":"Modern SoTA LMMs still demonstrates subpar performance at temporal reasoning with our temporal counterfactual benchmark composed of natural videos."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["temporal reasoning; counterfactual reasoning; short video comprehension"]},"supplementary_material":{"value":"/attachment/fd093e5ac3c36911b68e297a2932dcf6ce1ed019.zip"},"primary_area":{"value":"datasets and benchmarks"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"There has been growing sentiment recently that modern large multimodal models (LMMs) have addressed most of the key challenges related to short video comprehension. As a result, both academia and industry are gradually shifting their attention towards the more complex challenges posed by understanding long-form videos. \nHowever, is this really the case?  Our studies indicate that LMMs still lack many fundamental reasoning capabilities even when dealing with short videos.  We introduce Vinoground, a temporal counterfactual LMM evaluation benchmark encompassing 1000 short and natural video-caption pairs. We demonstrate that existing LMMs severely struggle to distinguish temporal differences between different actions and object transformations.  For example, the best model GPT-4o only obtains $\\sim$50\\% on our text and video scores, showing a large gap compared to the human baseline of $\\sim$90\\%. All open-source multimodal models and CLIP-based models perform much worse, producing mostly random chance performance. Through this work, we shed light onto the fact that temporal reasoning in short videos is a problem yet to be fully solved. We will make our benchmark publicly available."},"_bibtex":{"value":"@misc{\nzhang2025vinoground,\ntitle={Vinoground: Scrutinizing {LMM}s over Dense Temporal Reasoning with Short Videos},\nauthor={Jianrui Zhang and Mu Cai and Yong Jae Lee},\nyear={2025},\nurl={https://openreview.net/forum?id=a1P5kh2oo8}\n}"},"title":{"value":"Vinoground: Scrutinizing LMMs over Dense Temporal Reasoning with Short Videos"},"pdf":{"value":"/pdf/2e0dd590d0c11d11d7071fa71d7eaea83ded91ac.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"zhang|vinoground_scrutinizing_lmms_over_dense_temporal_reasoning_with_short_videos"},"authorids":{"value":["~Jianrui_Zhang1","~Mu_Cai1","~Yong_Jae_Lee2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Jianrui Zhang","Mu Cai","Yong Jae Lee"]}},"version":2},{"content":{"summary":{"value":"This paper introduces Rolling Forcing, a novel framework for real-time autoregressive long-video diffusion. The key idea is to mitigate error accumulation in streaming generation by performing joint denoising over a rolling window of consecutive frames rather than processing single frames sequentially. Three main contributions are proposed: (1) a rolling-window joint denoising strategy enabling mutual refinement across neighboring frames while keeping real-time throughput; (2) integration of an attention-sink mechanism that anchors initial frames as a global context to maintain long-term temporal consistency; and (3) an efficient few-step distillation algorithm over non-overlapping windows that conditions on self-generated histories to reduce exposure bias.\nExperimental results show that Rolling Forcing can stream multi-minute videos at 16 FPS on a single GPU with minimal drift, achieving state-of-the-art performance on the VBench benchmark and outperforming prior autoregressive baselines such as Self Forcing and CausVid in both visual quality and temporal stability."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1. What motivates the focus on long video generation, and how is this challenge tackled?  \n2. How does this paper's contribution distinguish itself from the limitations of related works?  \n3. How do the results compare to FramePack?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. Rolling forcing is a robust and effective framework. Components like progressively increasing noise levels, attention sink, and few-step distillation enhance long video generation.  \n2. Compared to baselines like causvid and self-forcing, rolling forcing significantly reduces error accumulation and extends video length while maintaining content consistency."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The weaknesses primarily stem from the novelty and experiments, as outlined below:  \n1. The proposed components are not entirely novel, as they have been used in video generation. For example, progressively increasing noise is used in Rolling Diffusion[1], MAGI-1[2], and FIFO-Diffusion[3]. Attention sink has been applied in LLMs[4], and step distillation is utilized in self-forcing[5]. This paper seems to integrate these components into an autoregressive video generation paradigm.  \n2. The paper does not compare its approach to related works in autoregressive video generation, such as FramePack[6], which it cites.  \n3. The training length in the paper is set to 27, with a KV cache length of 24 frames. It raises questions about how the attention sink performs with such a short training length.  \n4. In Table 2, the second row (without RF training) achieves strong performance, though still lower than rolling forcing. This raises the question of the key benefits of RF training.\n\n[1] Rolling Diffusion Models.\n[2] MAGI-1: Autoregressive Video Generation at Scale.\n[3] FIFO-Diffusion: Generating Infinite Videos from Text without Training.\n[4] When Attention Sink Emerges in Language Models\n[5] Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion\n[6] Frame Context Packing and Drift Prevention in Next-Frame-Prediction Video Diffusion Models."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762917369092,"tcdate":1761750788532,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission4445/Reviewer_M4Wp"],"signatures":["ICLR.cc/2026/Conference/Submission4445/Reviewer_M4Wp"],"forum":"IAyzXjbfwo","number":2,"license":"CC BY 4.0","cdate":1761750788532,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission4445/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762917369092,"domain":"ICLR.cc/2026/Conference","replyto":"IAyzXjbfwo","id":"g0FyFb94U6","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"TLDR":{"value":"REAL-TIME streaming generation of MULTI-MINUTE videos!"},"keywords":{"value":["autoregressive video generation","long video generation","real-time video generation"]},"primary_area":{"value":"generative models"},"abstract":{"value":"Streaming video generation as one fundamental component in interactive world models and neural game engines aims to generate high-quality, low-latency, and temporally coherent long stream videos. However, most existing work suffers from severe error accumulation that often significantly degrades the generated stream videos over long horizons. We design Rolling Forcing, a novel video generation technique that enables streaming long videos with minimal error accumulation. Rolling Forcing comes with three novel designs. First, instead of iteratively sampling individual frames which accelerates error propagation, we design a joint denoising scheme that simultaneously denoises multiple frames with progressively increasing noise levels. This design relaxes the strict causality across adjacent frames, effectively suppressing error growth. Second, we introduce the attention sink mechanism into the long-horizon stream video generation task, which allows the model to keep key–value states of initial frames as a global context anchor and thereby enhances long-term global consistency. Third, we design an efficient training algorithm that enables few-step distillation over largely extended denoising windows. This algorithm operates on non-overlapping windows and mitigates exposure bias conditioned on self-generated histories. Extensive experiments show that Rolling Forcing enables real-time streaming generation of multi-minute videos on a single GPU, with substantially reduced error accumulation."},"_bibtex":{"value":"@inproceedings{\nliu2026rolling,\ntitle={Rolling Forcing: Autoregressive Long Video Diffusion in Real Time},\nauthor={Kunhao Liu and Wenbo Hu and Jiale Xu and Ying Shan and Shijian Lu},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=IAyzXjbfwo}\n}"},"title":{"value":"Rolling Forcing: Autoregressive Long Video Diffusion in Real Time"},"pdf":{"value":"/pdf/75046738baed876ab95731f249eaa20cb17c28fd.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"liu|rolling_forcing_autoregressive_long_video_diffusion_in_real_time"},"authorids":{"value":["~Kunhao_Liu1","~Wenbo_Hu2","~Jiale_Xu1","~Ying_Shan2","~Shijian_Lu1"]},"authors":{"value":["Kunhao Liu","Wenbo Hu","Jiale Xu","Ying Shan","Shijian Lu"]}},"version":2},{"content":{"summary":{"value":"This paper identifies a shortcut in current inductive KGC datasets, where simple methods like Personalized PageRank can achieve strong performance without using relational information. The authors propose a new dataset construction strategy to eliminate this shortcut and benchmark popular KGC methods on these datasets."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"Please refer to Weaknesses."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. This work provides a clearer understanding of the capabilities and challenges in inductive KGC and the benchmark construction.\n\n2. The manuscript is well-organized and easy to follow.\n\n3. Experiments have verified the effectiveness of the proposed benchmarks."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The technical contribution is kind of limited for a long paper. The applied graph partitioning strategy is straightforward and the research challenge in this paper is not significant.\n\n2. There is a confusing point in the benchmark construction. Why the differences in SPD are beneficial from graph partitioning sampling? Since “the different in mean SPD” (also a typo in line 237) causes this shortcut, why not improve the negative sampling strategy directly? For example, select entities having a much shorter distance from the query entity as negative samples. It could be independent of the graph structure.\n\n3. In my opinion, the shortcut proposed in this paper is mainly caused by the negative-sampled evaluation (distinguishing the positive entity from a fixed number of negative ones). Recent studies, including RED-GNN and NBFNet, already employ the full evaluation (finding the positive entity from the entire entity set). This might be the reason for the superior performance of the two models as well as ULTRA. From this point, have the authors compared the performance differences between the two evaluation settings on new benchmarks?"}},"nonreaders":[],"tmdate":1732516374453,"tcdate":1730452156853,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission7457/Reviewer_mSV8"],"signatures":["ICLR.cc/2025/Conference/Submission7457/Reviewer_mSV8"],"forum":"npBAHV5BJI","number":2,"license":"CC BY 4.0","cdate":1730452156853,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission7457/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732516374453,"domain":"ICLR.cc/2025/Conference","replyto":"npBAHV5BJI","id":"6u08WZyApt","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Knowledge graphs","graphs","link prediction"]},"primary_area":{"value":"datasets and benchmarks"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Knowledge Graph Completion (KGC) attempts to predict missing facts in a Knowledge Graph (KG). Recently, there's been an increased focus on designing KGC methods that can excel in the *inductive setting*, where a portion or all of the entities and relations seen in inference are unobserved during training. Numerous benchmark datasets have been proposed for inductive KGC, all of which are subsets of existing KGs used for transductive KGC. However, we find that the current procedure for constructing inductive KGC datasets inadvertently creates a shortcut that can be exploited even while disregarding the relational information. Specifically, we observe that the Personalized PageRank (PPR) score can achieve strong or near SOTA performance on most inductive datasets. In this paper, we study the root cause of this problem. Using these insights, we propose an alternative strategy for constructing inductive KGC datasets that helps mitigate the PPR shortcut. We then benchmark multiple popular methods using the newly constructed datasets and analyze their performance. The new benchmark datasets help promote a better understanding of the capabilities and challenges of inductive KGC by removing any shortcuts that obfuscate performance."},"_bibtex":{"value":"@misc{\nshomer2025towards,\ntitle={Towards Better Benchmark Datasets for Inductive Knowledge Graph Completion},\nauthor={Harry Shomer and Jay Revolinsky and Jiliang Tang},\nyear={2025},\nurl={https://openreview.net/forum?id=npBAHV5BJI}\n}"},"title":{"value":"Towards Better Benchmark Datasets for Inductive Knowledge Graph Completion"},"pdf":{"value":"/pdf/9f23a17b441a9d4a9a2583f3c1edfaf35a535244.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"shomer|towards_better_benchmark_datasets_for_inductive_knowledge_graph_completion"},"authorids":{"value":["~Harry_Shomer1","~Jay_Revolinsky1","~Jiliang_Tang1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Harry Shomer","Jay Revolinsky","Jiliang Tang"]}},"version":2},{"content":{"summary":{"value":"The authors propose a new Self-Supervised Learning method for surgical video. Assuming that pixels/things that move similarly are likely to be correlated or even part of the same object, they propose a novel dense pixel representation driven contrastive loss that leverages optical flow information of different views. Specifically, given an anchor pixel in a view, it is compared to all other pixels in another view by computing the differences between the optical flow of the respective pixels. From these optical flow differences, they can extract a set of positive and negative pixels for the given anchor and finally compute the proposed dense contrastive loss."},"final_rating":{"value":4,"readers":["everyone"]},"justification_of_final_rating":{"value":"The authors have adressed most of my concerns about the related works issues and performed the suggested experiment. The revised version is also clearer and better organized. I am increasing my grade to weak accept.","readers":["everyone"]},"justification_of_the_preliminary_rating":{"value":"Even if I am not very familiar with the literature on self-supervised video representation learning, the claim that the authors are the first to study \"the use of optical flow during self-supervision to contrast parts of images\" seems a bit too strong. I expect them to give a more exhaustive literature review about self-supervised representation learning based on optical flow and at least discuss the differences with the references I have provided. Additionally, in my opinion, the experiments section is a bit messy and would require additional work to be made clearer and more concise (clear distinction between final performances and ablation studies, some parts could be moved to the supplementary like detailed network architecture, ...). Since the idea is clever but I am not sure about the novelty of this work while I think the paper could be improved, I give a Borderline. However, I am willing to change the grade if the authors address all my concerns."},"strengths":{"value":"* Even If I am not familiar with self-supervised learning for video representation learning and surgical video in general, I have found the paper well-written, pleasing to read, and quite clear.\n* The main assumption of the dense contrastive loss formulation is intuitive and seems reasonable while leading to robust performances on different downstream tasks of surgical video datasets."},"weaknesses":{"value":"* The authors did not respect the mandatory 8 pages limit.\n* Even if I am not very familiar with the literature on self-supervised video representation learning, the claim that the authors are the first to study \"the use of optical flow during self-supervision to contrast parts of images\" seems a bit too strong. For instance, papers [1, 2, 3] seem related to the self-supervised video representation learning based on optical flow. Therefore, I think a more complete related works would be necessary.\n* It is more a discussion than a weakness, it has been shown that including more negative pairs when contrasting results in better representations. In the authors' loss formulation, it seems that they use triplets of pixel representations (1 anchor, 1 positive, 1 negative). Therefore, my question is why have authors chosen this kind of contrastive loss instead of contrastive losses that include more negative pairs such as NT-Xent loss [4]?\n* It would have been interesting to see a visualization to confirm the main hypothesis that the loss formulation relies on (i.e.: \"things that move similarly are likely to be correlated, perhaps even part of the same object\"). For instance, for annotated frames with semantic masks, it would have been possible to measure for each anchor pixel to which extent the set of positive pixels contains pixels of the same class (purity metric). I think it would strengthen the paper and would demonstrate the soundness of this hypothesis. \n* The experiments section is a bit hard to follow. I think it would make this section clearer if the authors would make a clear separation between comparisons of the final model + concurrent works and the ablation studies.\n\nReferences:\n* [1] Takahashi, T., Yashima, S., Ishikawa, K., Sato, I., & Yokota, R. (2023). Pixel-Level Contrastive Learning of Driving Videos With Optical Flow. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 3179-3186).\n* [2] Sharma, Y., Zhu, Y., Russell, C., & Brox, T. (2022). Pixel-level Correspondence for Self-Supervised Learning from Video. arXiv preprint arXiv:2207.03866.\n* [3] Xiong, Y., Ren, M., Zeng, W., & Urtasun, R. (2021). Self-supervised representation learning from flow equivariance. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 10191-10200).\n* [4] Chen, T., Kornblith, S., Norouzi, M., & Hinton, G. (2020, November). A simple framework for contrastive learning of visual representations. In International conference on machine learning (pp. 1597-1607). PMLR."},"confidence":{"value":2},"detailed_comments":{"value":"* The network architecture section is too detailed and could be shifted to the appendices. \n* Conversely, I think Figure 2 which is the main novelty of the method would have been better in the main paper.\n* Table 1 and Table 2  could be resized to avoid margin overflows."},"special_issue":{"value":"No"},"questions_to_address_in_the_rebuttal":{"value":"Please see the weaknesses section"},"preliminary_rating":{"value":3}},"nonreaders":[],"tmdate":1712410958994,"tcdate":1709638815786,"writers":["MIDL.io/2024/Conference","MIDL.io/2024/Conference/Submission98/Reviewer_fcmR"],"signatures":["MIDL.io/2024/Conference/Submission98/Reviewer_fcmR"],"forum":"YN4aoXelxg","number":2,"license":"CC BY 4.0","cdate":1709638815786,"readers":["everyone"],"invitations":["MIDL.io/2024/Conference/Submission98/-/Official_Review","MIDL.io/2024/Conference/-/Edit","MIDL.io/2024/Conference/Submission98/Official_Review2/-/Review_Revision"],"mdate":1712410958994,"domain":"MIDL.io/2024/Conference","replyto":"YN4aoXelxg","id":"Oo2IxpM86c","forumContent":{"venue":{"value":"MIDL 2024 Poster"},"keywords":{"value":["Data Efficient Learning","Laparoscopic Semantic Segmentation","Robotic Instrument Pose Estimation","Self-Supervised Representation Learning"]},"abstract":{"value":"During minimally invasive surgery, surgeons monitor their actions and the relevant tissue through a camera. This provides an ideal environment for artificial intelligence (AI) assisted surgery. For the development of such AI components, the need for expert annotations remains a key bottleneck. In this paper, we study the application of self-supervised learning (SSL) on surgical data. In a self-supervised setting, a representation backbone is trained on information that is inherently present in the data. There is no need for annotations, leaving the backbone free to train on all recordings, not just labeled ones. We leveraged optical flow for weighting pairs in a view-contrastive self-supervised learning loss. Constructed as an Info Noise-Contrastive Estimation (InfoNCE) loss, it contrasted the pixel representations of two differently, photometrically and geometrically transformed views. The importance of each contrasted pixel pair is determined by computing the difference between the optical flows of the respective pixels. In this way, the optical flow guided the representations of pixels that move together to similar vectors. We tested the usefulness of the representation vectors by training simple networks for semantic segmentation or robotic instrument key point detection. These networks showed competitive performance, even when using over 92% fewer annotated samples than other works. For semantic segmentation, we used as little as 99.73% fewer samples for training, originating from the m2caiSeg dataset, and remained competitive even when testing on the unseen cholecSeg8k dataset."},"_bibtex":{"value":"@inproceedings{\nmoens2024laparoflowssl,\ntitle={Laparoflow-{SSL}: Image Analysis From a Tiny Dataset Through Self-Supervised Transformers Leveraging Unlabeled Surgical Video},\nauthor={Karel Moens and Jonas De Vylder and Matthew B. Blaschko and Tinne Tuytelaars},\nbooktitle={Medical Imaging with Deep Learning},\nyear={2024},\nurl={https://openreview.net/forum?id=YN4aoXelxg}\n}"},"title":{"value":"Laparoflow-SSL: Image Analysis From a Tiny Dataset Through Self-Supervised Transformers Leveraging Unlabeled Surgical Video"},"latex_code":{"value":"/attachment/c7ebeae5b0b19ff2a5b83ecdd2c6eb0dacdca340.zip"},"pdf":{"value":"/pdf/84790db75530ccd4a46931951da0bbc60af6bedc.pdf"},"copyright_form":{"value":"/attachment/07d190c4bf833f9025f7791facc7c03f7e6453ec.pdf"},"venueid":{"value":"MIDL.io/2024/Conference"},"paperhash":{"value":"moens|laparoflowssl_image_analysis_from_a_tiny_dataset_through_selfsupervised_transformers_leveraging_unlabeled_surgical_video"},"authorids":{"value":["~Karel_Moens1","jonas.devylder@barco.com","~Matthew_B._Blaschko1","~Tinne_Tuytelaars1"]},"authors":{"value":["Karel Moens","Jonas De Vylder","Matthew B. Blaschko","Tinne Tuytelaars"]}},"version":2},{"content":{"summary":{"value":"This work introduces Animal-Bench, a novel benchmark for evaluating multimodal video models in animal-centric video understanding. The benchmark covers 13 tasks spanning 7 major animal categories and 822 species. It proposes an automated pipeline for data filtering and question-answer pair generation, reducing human effort and potential biases. To simulate real-world shooting conditions, it employs video editing methods based on diffusion models to evaluate model robustness under various scenarios. This work evaluates 8 popular multimodal video models on Animal-Bench, identifying considerable room for improvement on animal-centric tasks."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"See [Weaknesses]"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1) This work introduces a comprehensive animal-centric benchmark covering a diverse range of tasks, including several that have been previously under-explored in the field.\n2) The authors claim to open source the code and data, which could be beneficial for the research community.\n3) By evaluating multiple recent multimodal video models on Animal-Bench, the work provides insights into current model capabilities and limitations, and highlights potential directions for future research and development."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The answer accuracy of the QA pairs. For example, in the \"Reasoning\" task illustrated in Figure 2, the correct answer appears to be \"to fight with dog\" rather than \"cat\".\n\n2. The question quality needs further improvement. 1) Ambiguity: in the \"Time\" task shown in Figure 2, the presence of multiple objects in the video frames renders the subject of the action ambiguous. 2) Inconsistency between video frames and question description: in the \"Object Count\" task in Figure 13, the setting appears to be a \"grassland\" rather than a \"forest\".\n\n3. The simulated changes intended to mimic real-world shooting scenarios exhibit noticeable artifacts and unrealistic situations. For example, Figure 11 shows visible boundaries from outpainting and implausible weather conditions (e.g., snow added to scenes with green grassland). To address this, the authors could consider implementing an aesthetic score-based filter or a specially trained discriminator to get rid of data with severe artifacts.\n\n4. Section 4.1 mentions resizing input videos to 224. For non-square videos (particularly those with highly disproportionate size), it's unclear whether additional operations (such as padding or cropping) were employed to accommodate the inputs. If such operations were used, an analysis of their potential impact on model performance across various tasks would be beneficial."},"limitations":{"value":"The uniform set of parameters used for all models in the evaluation, as mentioned in Table 4, may not align with each model's recommended settings, such as temperature. It could potentially prevent from fully leveraging the capabilities of individual models."}},"nonreaders":[],"tmdate":1730879964509,"tcdate":1722045053064,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission18302/Reviewer_xc4z"],"signatures":["NeurIPS.cc/2024/Conference/Submission18302/Reviewer_xc4z"],"forum":"DexM7d1H6e","number":4,"license":"CC BY 4.0","cdate":1722045053064,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission18302/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730879964509,"domain":"NeurIPS.cc/2024/Conference","replyto":"DexM7d1H6e","id":"LGJVTUFTuL","forumContent":{"venue":{"value":"NeurIPS 2024 poster"},"keywords":{"value":["Multimodal video model","Evaluation benchmark","Robustness testing"]},"primary_area":{"value":"evaluation"},"abstract":{"value":"With the emergence of large pre-trained multimodal video models, multiple benchmarks have been proposed to evaluate model capabilities. However, most of the benchmarks are human-centric, with evaluation data and tasks centered around human applications. Animals are an integral part of the natural world, and animal-centric video understanding is crucial for animal welfare and conservation efforts. Yet, existing benchmarks overlook evaluations focused on animals, limiting the application of the models. To address this limitation, our work established an animal-centric benchmark, namely Animal-Bench, to allow for a comprehensive evaluation of model capabilities in real-world contexts, overcoming agent-bias in previous benchmarks. Animal-Bench includes 13 tasks encompassing both common tasks shared with humans and special tasks relevant to animal conservation, spanning 7 major animal categories and 819 species, comprising a total of 41,839 data entries. To generate this benchmark, we defined a task system centered on animals and proposed an automated pipeline for animal-centric data processing. To further validate the robustness of models against real-world challenges, we utilized a video editing approach to simulate realistic scenarios like weather changes and shooting parameters due to animal movements. We evaluated 8 current multimodal video models on our benchmark and found considerable room for improvement. We hope our work provides insights for the community and opens up new avenues for research in multimodal video models. Our data and code will be released at https://github.com/PRIS-CV/Animal-Bench."},"_bibtex":{"value":"@inproceedings{\njing2024animalbench,\ntitle={Animal-Bench: Benchmarking Multimodal Video Models for Animal-centric Video Understanding},\nauthor={Yinuo Jing and Ruxu Zhang and Kongming Liang and Yongxiang Li and Zhongjiang He and Zhanyu Ma and Jun Guo},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=DexM7d1H6e}\n}"},"title":{"value":"Animal-Bench: Benchmarking Multimodal Video Models for Animal-centric Video Understanding"},"pdf":{"value":"/pdf/1385b842d017357564468802fe0fd83c961c61a1.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"jing|animalbench_benchmarking_multimodal_video_models_for_animalcentric_video_understanding"},"authorids":{"value":["~Yinuo_Jing1","~Ruxu_Zhang2","~Kongming_Liang2","~Yongxiang_Li2","~Zhongjiang_He1","~Zhanyu_Ma1","~Jun_Guo1"]},"authors":{"value":["Yinuo Jing","Ruxu Zhang","Kongming Liang","Yongxiang Li","Zhongjiang He","Zhanyu Ma","Jun Guo"]}},"version":2},{"content":{"summary":{"value":"This paper introduces OpenBiomedVid, a biomedical video-text dataset (1,031 hours) curated from YouTube educational videos, along with two expert-curated benchmarks—SurgeryVideoQA and MIMICEchoQA—for evaluating biomedical video understanding. The authors further fine-tune Qwen2-VL and InternVL3 models on this dataset, demonstrating improvements on both video and image benchmarks."},"soundness":{"value":4},"confidence":{"value":5},"questions":{"value":"See weaknesses. I may consider raising my score depending on the authors’ response, as I am genuinely interested in this work."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":4},"strengths":{"value":"1. Clear presentation: The paper clearly describes the construction of instruction data, evaluation datasets, and model training procedures, making the methodology easy to follow.\n\n2. Thoughtful data curation: The dataset creation process includes several reasonable and interesting design choices, such as leveraging SigLIP-Medical and Whisper models, to ensure the reliability and quality of the collected data.\n\n3. Practical and novel contribution: While prior works have explored YouTube data for research, the focus on video instruction tuning and the introduction of corresponding evaluation benchmarks fills an important gap in the current biomedical multimodal landscape. The dataset and benchmarks are likely to be of practical use to the community."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Although the paper targets video-level instruction tuning, prior works (e.g., Quilt-1M) have already shown success with image-level instruction data in the medical domain. It remains unclear whether video-level supervision provides a substantial advantage over image-text data for medical VQA tasks. A valuable follow-up experiment could involve creating a subset of the dataset where temporal information is minimized (e.g., few-frame clips or frame-level QA) to empirically assess whether videos or images contribute to the performance gains.\n2. In lines 343–346, the authors state that “performance on video benchmarks remains significantly lower than on text and image benchmarks.” However, the reported results show only modest improvements on text benchmarks (with MedQA even decreasing) and limited gains for PathVQA with the 7B model. The discussion should be more nuanced.\n3. The dataset mainly focuses on videos, but I noticed certain performance improvements on the image-text datasets VQA-RAD and SLAKE. If possible, I encourage the authors to derive and release a medical image instruction tuning dataset from the existing collection. For example, by selecting clips with fewer frames or converting segments that do not require strict temporal encoding into multi-image samples, since many QA pairs may not rely on temporal information. This would further advance the development of medical vision-language models."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762942037386,"tcdate":1761654023104,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission22059/Reviewer_QWpu"],"signatures":["ICLR.cc/2026/Conference/Submission22059/Reviewer_QWpu"],"forum":"u4PmZOmtko","number":1,"license":"CC BY 4.0","cdate":1761654023104,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission22059/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762942037386,"domain":"ICLR.cc/2026/Conference","replyto":"u4PmZOmtko","id":"Jc4vZR6JW6","forumContent":{"TLDR":{"value":"Instruction-tuning Qwen-2-VL on 1,031 hours of pedagogical biomedical videos dramatically boosts video and image understanding and includes new expert-curated benchmarks, with all data and code released."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["vision-language models","biomedicine","datasets","evaluations"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Publicly available biomedical videos, such as those on YouTube, serve as valuable educational resources for medical students. Unlike standard machine learning datasets, these videos are designed for human learners, often mixing medical imagery with narration, explanatory diagrams, and contextual framing. In this work, we investigate whether such pedagogically rich, yet non-standardized and heterogeneous videos can effectively teach general-domain vision-language models biomedical knowledge. To this end, we introduce OpenBiomedVid, a biomedical video instruction tuning dataset comprising 1031 hours of video-caption and Q/A pairs, curated through a multi-step human-in-the-loop pipeline. Diverse biomedical video datasets are rare, and OpenBiomedVid fills an important gap by providing instruction-style supervision grounded in real-world educational content. Surprisingly, despite the informal and heterogeneous nature of these videos, the fine-tuned Qwen-2-VL models exhibit substantial performance improvements across most benchmarks. The 2B model achieves gains of 98.7% on video tasks, 71.2% on image tasks, and 0.2% on text tasks. The 7B model shows improvements of 37.09% on video and 11.2% on image tasks, with a slight degradation of 2.7% on text tasks compared to their respective base models. To address the lack of standardized biomedical video evaluation datasets, we also introduce two new expert curated benchmarks, MIMICEchoQA and SurgeryVideoQA. On these benchmarks, the 2B model achieves gains of 99.1% and 98.1%, while the 7B model shows gains of 22.5% and 52.1%, respectively, demonstrating the models' ability to generalize and perform biomedical video understanding on cleaner and more standardized datasets than those seen during training. These results suggest that educational videos created for human learning offer a surprisingly effective training signal for biomedical VLMs. We release OpenBiomedVid, MIMICEchoQA, SurgeryVideoQA, the fine-tuned models, and the complete codebase to support future research."},"_bibtex":{"value":"@misc{\nthapa2026how,\ntitle={How Well Can General Vision-Language Models Learn Medicine By Watching Public Educational Videos?},\nauthor={Rahul Thapa and Andrew Li and Qingyang Wu and Bryan He and Yuki Sahashi and Christina Binder and Angela Zhang and Ben Athiwaratkun and Shuaiwen Leon Song and David Ouyang and James Zou},\nyear={2026},\nurl={https://openreview.net/forum?id=u4PmZOmtko}\n}"},"title":{"value":"How Well Can General Vision-Language Models Learn Medicine By Watching Public Educational Videos?"},"pdf":{"value":"/pdf/fcf643dd9df9746824583d9bb992dfdc21d7b0d9.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"thapa|how_well_can_general_visionlanguage_models_learn_medicine_by_watching_public_educational_videos"},"authorids":{"value":["~Rahul_Thapa1","~Andrew_Li4","~Qingyang_Wu1","~Bryan_He1","~Yuki_Sahashi1","~Christina_Binder1","~Angela_Zhang1","~Ben_Athiwaratkun1","~Shuaiwen_Leon_Song1","~David_Ouyang1","~James_Zou1"]},"authors":{"value":["Rahul Thapa","Andrew Li","Qingyang Wu","Bryan He","Yuki Sahashi","Christina Binder","Angela Zhang","Ben Athiwaratkun","Shuaiwen Leon Song","David Ouyang","James Zou"]}},"version":2},{"content":{"summary":{"value":"PickStyle introduces a video-to-video style transfer framework that preserves motion and context while rendering stylized frames from one of nine trained styles. The method’s primary innovation seems to be a context-style classifier-free guidance mechanism, allowing explicit control over content and style conditioning during diffusion. Additionally, a tunable noise initialization strategy enables improved temporal coherence and perceptual fidelity. The paper includes reasonable experiments demonstrate improvements over prior methods in both qualitative and quantitative metrics, including a standard battery of metrics across content, video quality, etc."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Q1 What would happen if multiple style prompts were given as input?  What would happen if an out of set style prompt were given?\n\nQ2 What are the limitations of applying this to video?  Is there any reason to expect degradation for longer videos, for example?\n\nQ3 Is the dataset created here publicly available?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"S1 The paper builds on the Wan2.1 generation backbone to include style adaptation and seems to be able capture nine different styles in a way that generalizes from image-pairs to video.  The paper makes reasonable technical innovations to accomplish this that seem original and relevant, such as CS-CFG which creates a tunable trade-off between fidelty and stylization (although this trade-off seems not analyzed in the paper).\n\nS2 The paper includes a motion augmentation strategy that enables the use of image pairing as training data.\n\nS3 Results seem compelling and meaningful analyses are included.  It seems clear that PickStyle is best at being able to match the style prompt, at least according R Precision score (although it is curious that this score only uses one frame from the video).  It also seems that the video quality aspects are strong."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"W1 The key aspects of the proposed method seem to the LoRA adapters that modulate the attention to capture style, the CS-CFG method and the noise initialization.  Yet, none of these are really thoroughly analyzed in any way.  For example although Fig. 8 captures one example of where CS-CFG helps, the interplay between $t_\\text{guide}$ and $c_\\text{guide}$ is not studied.  Similarly, we have no evidence about to what degree the context-based initialized is necessary.  Hence it is impossible to actually assess whether the technical innovations align with the observed results improvements, or if it is from other sources (e.g., the different data, the augmentation approach, etc.)\n\nW2 It is not clear whether the comparisons are fair.  Considering this paper creates a composite dataset with nine styles, have the other methods to which the paper compares been retrained on this dataset?  The paper does not sufficiently describe this critical point.\n\nW3 The approach to augment the paired image samples with some motion to generating pair training videos is not well described in the text and therefor hard to analyze.  It would seem, for example, that the types of augmentations used are not able to capture realistic motions in video resulting from 3D content and perspective effects.  This implies that perhaps the datasets used and results shown, however compelling they may be, may not be indicative of utility on more general video. \n\nW4 It seems that PickStyle is the most computationally expensive of the methods evaluated.\n\n\nMinor things\n- The manner in which the references are typically cited, e.g., \"VACE Jiang et al. (2025)\" is not proper, at least not for this style of including the author name.  These should be in parenthesis or better incorporated directly into the text.  VACE by Jiang et al. (2025) or VACE (Jiang et al. 2025)."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762942340700,"tcdate":1762009879946,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission22696/Reviewer_eovg"],"signatures":["ICLR.cc/2026/Conference/Submission22696/Reviewer_eovg"],"forum":"NRWI7NRaFD","number":3,"license":"CC BY 4.0","cdate":1762009879946,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission22696/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762942340700,"domain":"ICLR.cc/2026/Conference","replyto":"NRWI7NRaFD","id":"CcQNFcKIdi","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Video style transfer","video diffusion model","video-to-video translation"]},"supplementary_material":{"value":"/attachment/162b27ba572c9a4a02c33c945a7f82d34af98557.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"We address the task of video style transfer with diffusion models, where the goal is to preserve the context of an input video while rendering it in a target style specified by a text prompt. A major challenge is the lack of paired video data for supervision. We propose PickStyle, a video-to-video style transfer framework that augments pretrained video diffusion backbones with style adapters and benefits from paired still image data with source–style correspondences for training. PickStyle inserts low-rank adapters into the self-attention layers of conditioning modules, enabling efficient specialization for motion–style transfer while maintaining strong alignment between video content and style. To bridge the gap between static image supervision and dynamic video, we construct synthetic training clips from paired images by applying shared augmentations that simulate camera motion, ensuring temporal priors are preserved. In addition, we introduce Context–Style Classifier-Free Guidance (CS–CFG), a novel factorization of classifier-free guidance into independent text (style) and video (context) directions. CS–CFG ensures that context is preserved in generated video while the style is effectively transferred. Experiments across benchmarks show that our approach achieves temporally coherent, style-faithful, and content-preserving video translations, outperforming existing baselines both qualitatively and quantitatively."},"_bibtex":{"value":"@misc{\nmehraban2026pickstyle,\ntitle={PickStyle: Video-to-Video Style Transfer with Context-Style Adapters},\nauthor={Soroush Mehraban and Vida Adeli and Jacob Rommann and Babak Taati and Kyryl Truskovskyi},\nyear={2026},\nurl={https://openreview.net/forum?id=NRWI7NRaFD}\n}"},"title":{"value":"PickStyle: Video-to-Video Style Transfer with Context-Style Adapters"},"pdf":{"value":"/pdf/90c492daa5a0ac4597ec109d40f071a44866bec4.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"mehraban|pickstyle_videotovideo_style_transfer_with_contextstyle_adapters"},"authorids":{"value":["~Soroush_Mehraban1","~Vida_Adeli1","~Jacob_Rommann1","~Babak_Taati1","~Kyryl_Truskovskyi1"]},"authors":{"value":["Soroush Mehraban","Vida Adeli","Jacob Rommann","Babak Taati","Kyryl Truskovskyi"]}},"version":2},{"content":{"venue":{"value":"ECCV (32) 2022"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-031-19824-3_21.pdf"},"venueid":{"value":"dblp.org/conf/ECCV/2022"},"paperhash":{"value":"cavalli|nefsac_neurally_filtered_minimal_samples"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Luca_Cavalli:","~Marc_Pollefeys1","https://dblp.org/search/pid/api?q=author:Daniel_Barath:"]},"html":{"value":"https://doi.org/10.1007/978-3-031-19824-3_21"},"_bibtex":{"value":"@inproceedings{DBLP:conf/eccv/CavalliPB22,\n  author={Luca Cavalli and Marc Pollefeys and Daniel Barath},\n  title={NeFSAC: Neurally Filtered Minimal Samples},\n  year={2022},\n  cdate={1640995200000},\n  pages={351-366},\n  url={https://doi.org/10.1007/978-3-031-19824-3_21},\n  booktitle={ECCV (32)},\n  crossref={conf/eccv/2022-32}\n}\n"},"abstract":{"value":"Since RANSAC, a great deal of research has been devoted to improving both its accuracy and run-time. Still, only a few methods aim at recognizing invalid minimal samples early, before the often expensive model estimation and quality calculation are done. To this end, we propose NeFSAC, an efficient algorithm for neural filtering of motion-inconsistent and poorly-conditioned minimal samples. We train NeFSAC to predict the probability of a minimal sample leading to an accurate relative pose, only based on the pixel coordinates of the image correspondences. Our neural filtering model learns typical motion patterns of samples which lead to unstable poses, and regularities in the possible motions to favour well-conditioned and likely-correct samples. The novel lightweight architecture implements the main invariants of minimal samples for pose estimation, and a novel training scheme addresses the problem of extreme class imbalance. NeFSAC can be plugged into any existing RANSAC-based pipeline. We integrate it into USAC and show that it consistently provides strong speed-ups even under extreme train-test domain gaps – for example, the model trained for the autonomous driving scenario works on PhotoTourism too. We tested NeFSAC on more than 100 k image pairs from three publicly available real-world datasets and found that it leads to one order of magnitude speed-up, while often finding more accurate results than USAC alone. The source code is available at https://github.com/cavalli1234/NeFSAC."},"title":{"value":"NeFSAC: Neurally Filtered Minimal Samples"},"authors":{"value":["Luca Cavalli","Marc Pollefeys","Daniel Barath"]}},"tmdate":1762679692809,"pdate":1640995200000,"externalIds":["dblp:conf/eccv/CavalliPB22"],"tcdate":1762679664630,"writers":["~"],"signatures":["~Marc_Pollefeys2"],"forum":"PzKNCc3FaK","license":"CC BY-SA 4.0","number":680993,"cdate":1640995200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1762679692809,"domain":"DBLP.org","id":"PzKNCc3FaK","version":2},{"content":{"summary":{"value":"This paper proposes CAT-LVDM, a corruption-aware training framework for Latent Video Diffusion Models designed to improve robustness against imperfect or noisy text prompts. The core of the framework consists of two structured noise injection strategies:\n\n1. BCNI: Perturbs text embeddings along intra-batch semantic directions to increase entropy while maintaining semantic alignment.\n\n2. SACN: Injects noise along dominant spectral modes (via SVD) to enhance low-frequency smoothness and temporal coherence.\n\nThe authors provide theoretical justification for both methods, analyzing conditional entropy and Wasserstein distance bounds. Experiments are conducted on several datasets (WebVid-2M, MSR-VTT, MSVD, UCF-101), where the proposed methods demonstrate quantitative improvements (e.g., reduced FVD) compared to uncorrupted baselines and simpler noise injection techniques."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"See above"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The paper provides a theoretical analysis for its proposed methods (BCNI and SACN)\n\n2. The paper addresses the practical problem of training video models on large-scale, noisy web data. The idea of corruption-aware training for video diffusion is considered interesting and worthwhile.\n\n3. The method demonstrates clear quantitative performance gains (e.g., FVD, SSIM, PSNR) over uncorrupted baselines and naive (Gaussian/Uniform) noise baselines on the tested datasets.\n\n4. The paper is generally well-written and easy to follow."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. A major concern is that the method was only validated on an older LVDM architecture (DEMO). Its effectiveness and scalability on modern, state-of-the-art models (e.g., DiT-based) are unproven\n\n2. Despite quantitative gains, the practical impact is undermined by the generated videos. the visual results not unsatisfactory. \n\n3. he paper lacks a clear justification for applying BCNI primarily to caption-rich datasets (WebVid-2M, MSR-VTT) and SACN only to a class-labeled dataset (UCF-101). \n\n4. The evaluation is missing key comparisons. It fails to compare against stronger, modern LVDM baselines and does not empirically validate the necessity of \"video-specific\" corruption. \n\n5. The core idea of perturbing conditions to improve robustness has been explored in the image domain, limiting the paper's technical novelty."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762931543011,"tcdate":1761840203927,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission19699/Reviewer_wNJ3"],"signatures":["ICLR.cc/2026/Conference/Submission19699/Reviewer_wNJ3"],"forum":"unZhwukf0T","number":1,"license":"CC BY 4.0","cdate":1761840203927,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission19699/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762931543011,"domain":"ICLR.cc/2026/Conference","replyto":"unZhwukf0T","id":"6DfOOCwEmC","forumContent":{"TLDR":{"value":"We introduce CAT-Video, a corruption-aware training framework that improves robustness and temporal coherence in video diffusion models through structured, data-aligned noise injection."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["video diffusion","corruption-aware training","robust video generation","structured noise injection","multimodal robustness","temporal coherence"]},"supplementary_material":{"value":"/attachment/c99cdbeea034fc1e74ad38310569a2906228cc13.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Latent Video Diffusion Models (LVDMs) have achieved state-of-the-art generative quality for image and video generation; however, they remain brittle under noisy conditioning, where small perturbations in text or multimodal embeddings can cascade over timesteps and cause semantic drift. Existing corruption strategies from image diffusion (Gaussian, Uniform) fail in video settings because static noise disrupts temporal fidelity. In this paper, we propose **CAT-Video**, a corruption-aware training framework with structured, data-aligned noise injection tailored for video diffusion. Our two operators—*Batch-Centered Noise Injection (BCNI)* and *Spectrum-Aware Contextual Noise (SACN)*—align perturbations with batch semantics or spectral dynamics to preserve coherence. CAT-Video yields substantial gains: BCNI reduces FVD by **31.9%** on WebVid-2M, MSR-VTT, and MSVD, while SACN improves UCF-101 by **12.3%**, outperforming Gaussian, Uniform, and even large diffusion baselines like DEMO (2.3B) and Lavie (3B) despite training on $\\mathbf{5}\\times$ less data. Ablations confirm the unique value of low-rank, data-aligned noise, and theory establishes why these operators tighten robustness and generalization bounds. CAT-Video thus sets a new framework for robust video diffusion, and our experiments show that it can also be extended to autoregressive generation and multimodal video understanding LLMs."},"_bibtex":{"value":"@misc{\nmaduabuchi2026catvideo,\ntitle={{CAT}-{VIDEO}: {CORRUPTION}-{AWARE} {TRAINING} {FOR} {ROBUST} {VIDEO} {DIFFUSION} {MODELS}},\nauthor={Chika Maduabuchi and Hao Chen and Yujin Han and Jindong Wang},\nyear={2026},\nurl={https://openreview.net/forum?id=unZhwukf0T}\n}"},"title":{"value":"CAT-VIDEO: CORRUPTION-AWARE TRAINING FOR ROBUST VIDEO DIFFUSION MODELS"},"pdf":{"value":"/pdf/36b779e3109b1503084b84e3c1c30d2bc4985918.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"maduabuchi|catvideo_corruptionaware_training_for_robust_video_diffusion_models"},"authorids":{"value":["~Chika_Maduabuchi1","~Hao_Chen15","~Yujin_Han1","~Jindong_Wang4"]},"authors":{"value":["Chika Maduabuchi","Hao Chen","Yujin Han","Jindong Wang"]}},"version":2},{"content":{"decision":{"value":"Accept (poster)"},"comment":{"value":"This paper introduces a shadow-aware video outpainting framework that explicitly models object–shadow pairs to reveal the spatial extent and motion of temporarily invisible instances, coupling instance-aware optical-flow completion with a diffusion outpainting backbone and a video-aware shadow-alignment discriminator. Reviewers find the core idea novel and well-motivated, with coherent modular design and strong empirical gains (PSNR/SSIM/FVD) plus a user study; after rebuttal, all reviewers are positive (one accept, three borderline accept). The rebuttal further strengthens the case with robustness analyses (insensitivity to shadow-detector thresholds), camera-motion splits showing stability under fast motion, and an efficiency comparison indicating competitive inference without per-video test-time retraining (unlike MOTIA). Thus, the AC decides to accept the submission."},"title":{"value":"Paper Decision"}},"parentInvitations":"NeurIPS.cc/2025/Conference/-/Decision","nonreaders":[],"tmdate":1761718359123,"tcdate":1758113421932,"writers":["NeurIPS.cc/2025/Conference","NeurIPS.cc/2025/Conference/Program_Chairs"],"signatures":["NeurIPS.cc/2025/Conference/Program_Chairs"],"forum":"irni27kAeP","number":1,"license":"CC BY 4.0","cdate":1758113421932,"readers":["everyone"],"invitations":["NeurIPS.cc/2025/Conference/Submission19821/-/Decision","NeurIPS.cc/2025/Conference/-/Edit"],"mdate":1761718359123,"domain":"NeurIPS.cc/2025/Conference","replyto":"irni27kAeP","id":"cgBscvyMpV","forumContent":{"venue":{"value":"NeurIPS 2025 poster"},"keywords":{"value":["Video Shadow Processing","Video Generation"]},"supplementary_material":{"value":"/attachment/529fc44bb31fa5b21e33d1c711424c0c49b92383.zip"},"primary_area":{"value":"deep_learning"},"abstract":{"value":"Conventional video outpainting methods primarily focus on maintaining coherent textures and visual consistency across frames.\nHowever, they often fail at handling dynamic scenes due to the complex motion of objects or camera movement, leading to temporal incoherence and visible flickering artifacts across frames. This is primarily because they lack instance-aware modeling to accurately separate and track individual object motions throughout the video. In this paper, we propose a novel video outpainting framework that explicitly takes shadow-object pairs into consideration to enhance the temporal and spatial consistency of instances, even when they are temporarily invisible. Specifically, we first track the shadow-object pairs across frames and predict the instances in the scene to unveil the spatial regions of invisible instances. Then, these prediction results are fed to guide the instance-aware optical flow completion to unveil the temporal motion of invisible instances. Next, these spatiotemporal guidances of instances are used to guide the video outpainting process. Finally, a video-aware discriminator is implemented to enhance alignment among dynamic shadows and the extended semantics in the scene. Comprehensive experiments underscore the superiority of our approach, outperforming existing state-of-the-art methods in widely recognized benchmarks."},"_bibtex":{"value":"@inproceedings{\nli2025dynamic,\ntitle={Dynamic Shadow Unveils Invisible Semantics for Video Outpainting},\nauthor={Ruilin Li and Hang Yu and Jiayan Qiu},\nbooktitle={The Thirty-ninth Annual Conference on Neural Information Processing Systems},\nyear={2025},\nurl={https://openreview.net/forum?id=irni27kAeP}\n}"},"title":{"value":"Dynamic Shadow Unveils Invisible Semantics for Video Outpainting"},"pdf":{"value":"/pdf/23f320173fb0cd8864e1f7ef7f33c8b9c9ddbc54.pdf"},"venueid":{"value":"NeurIPS.cc/2025/Conference"},"paperhash":{"value":"li|dynamic_shadow_unveils_invisible_semantics_for_video_outpainting"},"authorids":{"value":["~Ruilin_Li2","~Hang_Yu9","~Jiayan_Qiu1"]},"authors":{"value":["Ruilin Li","Hang Yu","Jiayan Qiu"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a video editing method that enhances first-frame-guided approaches with more granular temporal control. The core contribution is a mask-aware LoRA fine-tuning strategy for pre-trained Image-to-Video models. This technique uses a spatiotemporal mask to teach the model to selectively preserve background content while generating new content in specified regions. This dual approach allows the model to learn consistent motion from the source video and new appearances from user-provided reference frames, enabling complex edits like a flower blooming into a different color."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"See weaknesses, but my main questions are based on its practical applicability and efficiency."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"- Video results are pretty good (videos demonstrated in the supplementary and the demo in the supplementary showcase).\n\n- Application-wise it's interesting and definitely be useful. However, still, efficiency is a key problem.\n\n- Clearly outperforms previous approaches in terms of both quantitative and qualitative metrics."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The primary limitation is the need of fine-tuning per-video= The method requires 100 training steps to learn motion and potentially another 100 steps to learn new appearance (Sec 4.1). While this is more efficient than full model training, it is still a significant computational step for every single video edit compared to zero-shot or inference-time methods. How do you compare it in terms of computataional efficiency? I think putting a comparison in term of run-time memory give some insights.\n\n- The method relies on a precise \"spatiotemporal mask\" to separate the edited region from the background (Sec 3.3.1). The paper does not specify how this mask is acquired. Does this require the user to perform manual video segmentation for all frames? If so, this is still a requirement that would make the method impractical for most users. If the mask is generated automatically, its quality would be critical to the final edit. What could be the possible ways to automatically get these masks?\n\n- The full-frame method requires 20GB of GPU VRAM (Sec 4.1), which is inaccessible to many users. The \"low-cost training strategy\" (Appendix C) is a good alternative. Can you elaborate on this, for example what ways can be done to even make it more efficient?"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762918437233,"tcdate":1761853280978,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission6060/Reviewer_636m"],"signatures":["ICLR.cc/2026/Conference/Submission6060/Reviewer_636m"],"forum":"xkRMJ1Y7Um","number":1,"license":"CC BY 4.0","cdate":1761853280978,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission6060/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762918437233,"domain":"ICLR.cc/2026/Conference","replyto":"xkRMJ1Y7Um","id":"uiikqbFGW2","forumContent":{"TLDR":{"value":"The paper introduces a mask-based LoRA tuning method for highly flexible video editing using the pre-trained Image-to-Video model."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Video Editing"]},"supplementary_material":{"value":"/attachment/034635334583436f67f6ff70e86f40fbc1976fcf.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Video editing using diffusion models has achieved remarkable results in generating high-quality edits for videos. However, current methods often rely on large-scale pretraining, limiting flexibility for specific edits. First-frame-guided editing provides control over the first frame, but lacks fine-grained control over the edit's subsequent temporal evolution. To address this, we propose a mask-based LoRA (Low-Rank Adaptation) tuning method that adapts pretrained Image-to-Video models for flexible video editing.\nOur key innovation is using a spatiotemporal mask to strategically guide the LoRA fine-tuning process. This teaches the model two distinct skills: first, to interpret the mask as a command to either preserve content from the source video or generate new content in designated regions. Second, for these generated regions, LoRA learns to synthesize either temporally consistent motion inherited from the video or novel appearances guided by user-provided reference frames.\nThis dual-capability LoRA grants users control over the edit's entire temporal evolution, allowing complex transformations like an object rotating or a flower blooming. Experimental results show our method achieves superior video editing performance compared to baseline methods. The code and video results are available at our project website: https://cjeen.github.io/LoRAEdit."},"_bibtex":{"value":"@inproceedings{\ngao2026controllable,\ntitle={Controllable First-Frame-Guided Video Editing via Mask-Aware Lo{RA} Fine-Tuning},\nauthor={Chenjian Gao and Lihe Ding and Xin Cai and Zhanpeng Huang and Zibin Wang and Tianfan Xue},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=xkRMJ1Y7Um}\n}"},"title":{"value":"Controllable First-Frame-Guided Video Editing via Mask-Aware LoRA Fine-Tuning"},"pdf":{"value":"/pdf/a7fc2c198d3de6a267b714b5b442ae140e592d4c.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"gao|controllable_firstframeguided_video_editing_via_maskaware_lora_finetuning"},"authorids":{"value":["~Chenjian_Gao1","~Lihe_Ding1","~Xin_Cai2","~Zhanpeng_Huang1","~Zibin_Wang1","~Tianfan_Xue2"]},"authors":{"value":["Chenjian Gao","Lihe Ding","Xin Cai","Zhanpeng Huang","Zibin Wang","Tianfan Xue"]}},"version":2},{"content":{"summary":{"value":"This paper introduces ROSI (Rank-One Safety Injection), a white-box method for enhancing safety alignment in LLMs through permanent rank-one weight modifications. The approach extracts a \"safety direction\" from harmful/harmless instruction pairs using difference-in-means, then injects this direction into residual stream write matrices via the update rule W'_out ← W_out + α·ŝ·w̄^T. Experiments across aligned models (LLAMA, QWEN, GEMMA, YI) and uncensored models (DOLPHIN series) demonstrate improved harm refusal rates and jailbreak robustness with minimal utility degradation on standard benchmarks."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. Which layers benefit most from ROSI? Did you try layer-specific \\alpha values or applying ROSI to only a subset of layers?\n2. Does a safety direction extracted from one model transfer to architecturally similar models? This could have interesting implications for safety."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. ROSI provides a lightweight alternative to expensive fine-tuning, requiring only 50 instruction pairs and simple weight modifications\n\n2. The paper tests across 13 models, multiple safety benchmarks (CATQA, HARMBENCH, WILDJAILBREAK), utility benchmarks (MMLU, HELLASWAG, ARC, etc.), and attack scenarios\n\n3. Demonstrating effectiveness on both aligned and uncensored models broadens the method's utility\n\n4. Tables 3 and 6 show remarkably stable performance across capability benchmarks (typically <0.5% average change)\n\n5. The method maintains transparency about what is being modified and why, unlike black-box fine-tuning approaches"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Why is w̄·\\hat{s}^T the right rank-one update? The paper doesn't justify this choice over alternatives like random projections or learned directions. An ablation comparing different rank-one formulations would strengthen the claims.\n\n2. The paper states l* is \"selected based on a validation set\" but provides no details about this validation procedure, what metrics were optimized, or how many layers were tested.\n\n3. Only 50 harmful/harmless pairs seems quite small. What's the variance across different samples?\n\n4. The safety system prompt approach (Figure 2, Appendix A) seems somewhat circular—you're using a prompt to elicit safety behavior, then trying to make that permanent. How robust is this to variations in the prompt? The ❢ ablations suggest this is fragile for smaller models."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762918288045,"tcdate":1761988655924,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission5826/Reviewer_jJC3"],"signatures":["ICLR.cc/2026/Conference/Submission5826/Reviewer_jJC3"],"forum":"8c2SbG5PLj","number":4,"license":"CC BY 4.0","cdate":1761988655924,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission5826/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762918288045,"domain":"ICLR.cc/2026/Conference","replyto":"8c2SbG5PLj","id":"rfurpkqjcs","forumContent":{"TLDR":{"value":"This paper introduces ROSI, a lightweight training-free method that amplifies safety in LLMs without training. The method can also be used realign uncensored LLMs."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Large Language Models","Alignment","Safety","Refusal"]},"primary_area":{"value":"alignment, fairness, safety, privacy, and societal considerations"},"abstract":{"value":"Safety alignment in Large Language Models (LLMs) often involves mediating internal representations to refuse harmful requests. Recent research has demonstrated that these safety mechanisms can be bypassed by ablating or removing specific representational directions within the model. In this paper, we propose the opposite approach: ***Rank-One Safety Injection (ROSI)***, a white-box method that amplifies a model's safety alignment by permanently steering its activations toward the refusal-mediating subspace. **ROSI** operates as a simple, fine-tuning-free rank-one weight modification applied to all residual stream write matrices. The required safety direction can be computed from a small set of harmful and harmless instruction pairs. We show that **ROSI** consistently increases safety refusal rates - as evaluated by Llama Guard 3 - while preserving the utility of the model on standard benchmarks such as MMLU, HellaSwag, and Arc. Furthermore, we show that **ROSI** can also re-align 'uncensored' models by amplifying their own latent safety directions, demonstrating its utility as an effective last-mile safety procedure. Our results suggest that targeted, interpretable weight steering is a cheap and potent mechanism to improve LLM safety, complementing more resource-intensive fine-tuning paradigms."},"_bibtex":{"value":"@misc{\nshairah2026turning,\ntitle={Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection},\nauthor={Harethah Abu Shairah and Hasan Abed Al Kader Hammoud and George Turkiyyah and Bernard Ghanem},\nyear={2026},\nurl={https://openreview.net/forum?id=8c2SbG5PLj}\n}"},"title":{"value":"Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection"},"pdf":{"value":"/pdf/389be1197099e0b718638a61bee614b5a6f5310a.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"shairah|turning_the_spell_around_lightweight_alignment_amplification_via_rankone_safety_injection"},"authorids":{"value":["~Harethah_Abu_Shairah1","~Hasan_Abed_Al_Kader_Hammoud1","~George_Turkiyyah2","~Bernard_Ghanem1"]},"authors":{"value":["Harethah Abu Shairah","Hasan Abed Al Kader Hammoud","George Turkiyyah","Bernard Ghanem"]}},"version":2},{"content":{"venue":{"value":"ICLR 2027 Conference Submission"},"TLDR":{"value":"We introduce a field-aware audit of nine video anomaly understanding benchmarks that reveals video–text alignment failures and guides annotation revision to improve downstream performance."},"keywords":{"value":["Video Anomaly Understanding","Benchmark Auditing","Annotation Quality","Video–Text Alignment","Vision–Language Models"]},"supplementary_material":{"value":"/attachment/c2e3b08f19fa17eeac41eb0f492957dc638f59e0.zip"},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Language annotations in video anomaly understanding benchmarks guide model training and evaluation. Their quality must be assessed in relation to field responsibilities, temporal scopes, and granularity. We introduce a contract-aware audit framework that translates native field requirements into diagnostic questions and specifies the permitted evidence. Four tasks assess video quality, standalone text quality, native cross-text consistency, and video--text alignment. To determine whether a video--text pair meets multiple quality requirements simultaneously, we link video, standalone text, and alignment judgments to compute conditional joint usability. Our study includes 7,928 video entries and 56,103 video--text pairs across nine benchmarks. In the model-assisted audit, conditional joint usability ranges from 22.07% to 74.20% across benchmarks under field-macro aggregation. The predominant failure pattern is that video and text pass standalone checks, but their pairing fails alignment. To test whether audit-guided annotation revision improves performance, we revise CUVA descriptions and conduct controlled experiments with LVLMs, evaluating both description generation and multiple-choice question answering on VALU. The results show downstream benefits from revision, with effects varying by model and evaluation metric. Field-level audits further reveal quality differences within individual benchmarks. These diagnostics identify annotations requiring review and, together with native task coverage, guide the selection of benchmarks and fields suited to the evaluation goal and temporal granularity."},"_bibtex":{"value":"@inproceedings{\nanonymous2026auditing,\ntitle={Auditing the Ground Truth: A Field-Aware Audit of Video Anomaly Understanding Benchmarks},\nauthor={Anonymous},\nbooktitle={Submitted to The Fifteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=NfCZJAaQdb},\nnote={under review}\n}"},"title":{"value":"Auditing the Ground Truth: A Field-Aware Audit of Video Anomaly Understanding Benchmarks"},"pdf":{"value":"/pdf/bc070099ab03829a10c93306f45bb51c8e5147ec.pdf"},"venueid":{"value":"ICLR.cc/2027/Conference/Submission"}},"tmdate":1791223329485,"tcdate":1788525820982,"writers":["ICLR.cc/2027/Conference","ICLR.cc/2027/Conference/Submission9191/Authors"],"signatures":["ICLR.cc/2027/Conference/Submission9191/Authors"],"forum":"NfCZJAaQdb","license":"CC BY 4.0","number":9191,"cdate":1788525820982,"readers":["everyone"],"invitations":["ICLR.cc/2027/Conference/-/Submission","ICLR.cc/2027/Conference/-/Edit","ICLR.cc/2027/Conference/-/Submission_Change_Before_Bidding","ICLR.cc/2027/Conference/Submission9191/-/Full_Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Reviewing"],"mdate":1791223329485,"odate":1791031328286,"domain":"ICLR.cc/2027/Conference","id":"NfCZJAaQdb","version":2},{"content":{"venue":{"value":"IEEE J. Sel. Areas Commun. 2009"},"pdf":{"value":"https://ieeexplore.ieee.org/iel5/49/5072299/05072357.pdf"},"venueid":{"value":"dblp.org/journals/JSAC/2009"},"paperhash":{"value":"seferoglu|videoaware_opportunistic_network_coding_over_wireless_networks"},"authorids":{"value":["~Hulya_Seferoglu1","https://dblp.org/search/pid/api?q=author:Athina_Markopoulou:"]},"html":{"value":"https://doi.org/10.1109/JSAC.2009.090612"},"_bibtex":{"value":"@article{DBLP:journals/jsac/SeferogluM09,\n  author={Hulya Seferoglu and Athina Markopoulou},\n  title={Video-aware opportunistic network coding over wireless networks},\n  year={2009},\n  cdate={1230768000000},\n  journal={IEEE J. Sel. Areas Commun.},\n  volume={27},\n  number={5},\n  pages={713-728},\n  url={https://doi.org/10.1109/JSAC.2009.090612}\n}\n"},"abstract":{"value":"In this paper, we study video streaming over wireless networks with network coding capabilities. We build upon recent work, which demonstrated that network coding can increase throughput over a broadcast medium, by mixing packets from different flows into a single packet, thus increasing the information content per transmission. Our key insight is that, when the transmitted flows are video streams, network codes should be selected so as to maximize not only the network throughput but also the video quality. We propose video-aware opportunistic network coding schemes that take into account both the decodability of network codes by several receivers and the importance and deadlines of video packets. Simulation results show that our schemes significantly improve both video quality and throughput. This work is a first step towards content-aware network coding."},"title":{"value":"Video-aware opportunistic network coding over wireless networks"},"authors":{"value":["Hulya Seferoglu","Athina Markopoulou"]}},"tmdate":1747126439027,"pdate":1230768000000,"tcdate":1747126299863,"writers":["~"],"signatures":["~Hulya_Seferoglu1"],"forum":"boiiJ3FSxy","license":"CC BY-SA 4.0","number":443110,"cdate":1230768000000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1747126439027,"domain":"DBLP.org","id":"boiiJ3FSxy","version":2},{"content":{"TLDR":{"value":"We propose Pair-Aware Contextual Tokens (PACT) framework, a relation representation that preserves global scene context while being aware of the target pair in Video Visual Relation Detection"},"venue":{"value":"ICLR 2027 Conference Submission"},"keywords":{"value":["Video Understanding","Video Visual Relation Detection"]},"primary_area":{"value":"learning on time series and dynamical systems"},"abstract":{"value":"Video Visual Relation Detection (VidVRD) focuses on detecting ⟨subject, predicate, object⟩ relations together with their temporal extents in videos. Most existing methods construct relation representations from features extracted from detected object regions. Consequently, contextual information from the surrounding scene is largely overlooked, and the representation remains dependent on the quality of the object detector. However, many predicates are determined by how two objects interact within the broader scene. We propose Pair-Aware Contextual Tokens (PACT), a relation representation that preserves global scene context while being aware of the target pair. PACT represents each frame using all patch tokens from a frozen image encoder. Since the patch tokens are shared across different pairs, PACT injects information about the target pair into them. Specifically, each token encodes its spatial overlap with the subject and object regions, together with the semantic embeddings of the two object categories. Experiments on ImageNet-VidVRD and VidOR show that PACT achieves state-of-the-art mAP performance on both benchmarks. These results show that preserving the global scene context while making the representation aware of the target pair is an effective way to learn relation representations."},"_bibtex":{"value":"@inproceedings{\nanonymous2026pact,\ntitle={{PACT}: Pair-Aware Contextual Tokens for Video Visual Relation Detection},\nauthor={Anonymous},\nbooktitle={Submitted to The Fifteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=020jLJb71S},\nnote={under review}\n}"},"title":{"value":"PACT: Pair-Aware Contextual Tokens for Video Visual Relation Detection"},"pdf":{"value":"/pdf/4cf7a14a13c65da8828b01141d45b9f895e46823.pdf"},"venueid":{"value":"ICLR.cc/2027/Conference/Submission"}},"tmdate":1791230204967,"tcdate":1789543409005,"writers":["ICLR.cc/2027/Conference","ICLR.cc/2027/Conference/Submission23866/Authors"],"signatures":["ICLR.cc/2027/Conference/Submission23866/Authors"],"forum":"020jLJb71S","license":"CC BY 4.0","number":23866,"cdate":1789543409005,"readers":["everyone"],"invitations":["ICLR.cc/2027/Conference/-/Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Bidding","ICLR.cc/2027/Conference/Submission23866/-/Full_Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Reviewing","ICLR.cc/2027/Conference/-/Edit"],"mdate":1791230204967,"odate":1791031328286,"domain":"ICLR.cc/2027/Conference","id":"020jLJb71S","version":2},{"content":{"summary":{"value":"This paper proposes a framework for multi-action long video generation, Instruct-Video-Continuation (InstructVC) that contains two control stages: Temporal Action Binding and Causal Video Continuation. The authors further introduce an inference-time instance SteinsGate."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Please see above."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. The paper is well organized.  \n\n2. The topic of long video generation is worth exploring in the research community. This paper targets this important problem.\n\n3. The paper provides some video demos to help reviewers better evaluate the performance of the proposed method."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"There are some concerns and questions about this paper:\n\n1.\tThe third paragraph of the introduction is a bit confusing. The first part mentions two methods for generating long videos: temporal expanding and temporal decomposition. So, to which category does the latter part of the paragraph, I2V-AR, and another work, belong?\n\n2.\tIt is recommended that the author provide a corresponding video for Fig. 1. A few frames are insufficient to convey the author's intended meaning.\n\n3.\tThe author repeatedly emphasizes the term \"temporal causality.\" What exactly does this term mean? Why does its absence lead to phenomena such as temporal inconsistency and action direction conflict in the generated video?\n\n4.\tThe two paragraphs following the introduction cover a wide range of topics, such as Temporal Action Binding, Causal Video Continuation, Guidance Interval, History-aligned Redistribution, and Path Convergence Guidance. This can be quite confusing, as it raises the question: which part is the core of the proposed method? Which part is the key to solving long-action video generation?\n\n5.\tThe video demos in the supplementary materials all seem to involve very simple actions and are basically long videos within a single content. I think focusing on story-based long video generation would be better. Additionally, the video \"woman_gestures\" is clearly discontinuous, exhibiting obvious splicing artifacts from multiple video clips."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762928263512,"tcdate":1761974870793,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission18554/Reviewer_8cn5"],"signatures":["ICLR.cc/2026/Conference/Submission18554/Reviewer_8cn5"],"forum":"8WS5nDWIWE","number":2,"license":"CC BY 4.0","cdate":1761974870793,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission18554/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762928263512,"domain":"ICLR.cc/2026/Conference","replyto":"8WS5nDWIWE","id":"AFDPXk4Wej","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Generative Models","Video Generation","Diffusion Guidance"]},"supplementary_material":{"value":"/attachment/325404a03779eedb98dfecccd83ad4d515610887.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Video generation has advanced rapidly, but current models remain limited to short clips, far from the length and complexity of real-world narratives. Long video generation is thus both important and challenging. Existing approaches either attempt to extend the modeling length of video diffusion models directly or merge short clips via shared frames. However, due to the lack of temporal causality modeling for video data, they achieve only limited extensions, suffer from discontinuous or even contradictory actions, and fail to support flexible and fine-grained temporal control. Thus, we propose Instruct-Video-Continuation (InstructVC), combining Temporal Action Binding for fine-grained temporal control and Causal Video Continuation for natural long-term simulation. Temporal Action Binding decomposes complex long videos by temporal causality into scene descriptions and action sequences with predicted durations, while Causal Video Continuation autoregressively generates coherent video narratives from the text story. We further introduce SteinsGate, an inference-time instance of InstructVC that uses an MLLM for Temporal Action Binding and Video Path Integral to enforce causality between actions, converting a pre-trained TI2V diffusion model into an autoregressive video continuation model. Benchmark results demonstrate the advantages of SteinsGate and InstructVC in achieving accurate temporal control and generating natural, smooth multi-action long videos."},"_bibtex":{"value":"@inproceedings{\nhuang2026steinsgate,\ntitle={SteinsGate: Adding Causality to Diffusions for Long Video Generation via Path Integral},\nauthor={Yufei Huang and Liangyu Yuan and Changxi Chi and Yunfan Liu and Cheng Tan and Siyuan Li and Jingbo Zhou and Haitao Lin and Chang Yu and Stan Z. Li},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=8WS5nDWIWE}\n}"},"title":{"value":"SteinsGate: Adding Causality to Diffusions for Long Video Generation via Path Integral"},"pdf":{"value":"/pdf/72d2476c464517c6b0b3d5b1e46565ebd799bdf9.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"huang|steinsgate_adding_causality_to_diffusions_for_long_video_generation_via_path_integral"},"authorids":{"value":["~Yufei_Huang4","~Liangyu_Yuan1","~Changxi_Chi1","~Yunfan_Liu2","~Cheng_Tan1","~Siyuan_Li6","~Jingbo_Zhou2","~Haitao_Lin2","~Chang_Yu1","~Stan_Z._Li2"]},"authors":{"value":["Yufei Huang","Liangyu Yuan","Changxi Chi","Yunfan Liu","Cheng Tan","Siyuan Li","Jingbo Zhou","Haitao Lin","Chang Yu","Stan Z. Li"]}},"version":2},{"content":{"summary":{"value":"This paper introduces StreamDiT, a 4-billion parameter streaming video generation model that addresses the limitations of existing text-to-video systems which only produce short clips offline. StreamDiT uses flow matching with a moving buffer, mixed training with different frame partitioning schemes, and adaptive layer normalization DiT architecture with varying time embeddings and window attention to achieve real-time video generation. Through a novel multistep distillation method that reduces function evaluations to match the number of buffer chunks, the model achieves 16 FPS performance on a single GPU at 512p resolution, enabling interactive applications like streaming generation, real-time interaction, and video-to-video transformation while maintaining content consistency and visual quality."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"Please refer to weaknesses."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. This paper enables real-time streaming video generation at 16 FPS on a single GPU, overcoming the offline-only limitation of existing text-to-video models.\n\n2. It employs mixed training with different frame partitioning schemes to ensure both content consistency and high visual quality in generated video streams.\n\n3. It introduces a tailored multistep distillation method that significantly reduces computational cost, making interactive applications practically feasible."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The paper only compares with U-Net-based methods such as ReuseDiffuse and FIFO. Please provide qualitative and quantitative comparison results with more advanced DiT-based methods.\n2. When generating longer videos, such as the car video in \"Real-Time Streaming Video Generation\" from the supplementary materials, the video quality noticeably deteriorates as time progresses.\n3. The examples of \"Interactive Video Generation\" provided in the paper all involve scene or appearance changes. How does StreamDiT perform when it comes to motion changes? Please provide relevant video results."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762942979396,"tcdate":1761816082332,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission24183/Reviewer_sacD"],"signatures":["ICLR.cc/2026/Conference/Submission24183/Reviewer_sacD"],"forum":"ayAx2YnfmD","number":2,"license":"CC BY 4.0","cdate":1761816082332,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission24183/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762942979396,"domain":"ICLR.cc/2026/Conference","replyto":"ayAx2YnfmD","id":"CqRfkhOdt6","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Diffusion Models","Video Generation","Real-time"]},"supplementary_material":{"value":"/attachment/2f40e17f1b00153ba62c3174e3c26b2c3af1c626.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Recently, great progress has been achieved in text-to-video (T2V) generation by scaling transformer-based diffusion models to billions of parameters, which can generate high-quality videos. However, existing models typically produce only short clips offline, restricting their use cases in interactive and real-time applications. This paper addresses these challenges by proposing StreamDiT, a streaming video generation model. StreamDiT training is based on flow matching by adding a moving buffer. We design mixed training with different partitioning schemes of buffered frames to boost both content consistency and visual quality. StreamDiT modeling is based on adaLN DiT with varying time embedding and window attention. To practice the proposed method, we train a StreamDiT model with 4B parameters. In addition, we propose a multistep distillation method tailored for StreamDiT. Sampling distillation is performed in each segment of a chosen partitioning scheme. After distillation, the total number of function evaluations (NFEs) is reduced to the number of chunks in a buffer. Finally, our distilled model reaches real-time performance at 16 FPS on one GPU, which can generate video streams at 512p resolution. We evaluate our method through both quantitative metrics and human evaluation. Our model enables real-time applications, e.g. streaming generation, interactive generation, and video-to-video."},"_bibtex":{"value":"@misc{\nkodaira2025streamdit,\ntitle={StreamDiT: Real-Time Streaming Text-to-Video Generation},\nauthor={Akio Kodaira and Tingbo Hou and Ji Hou and Markos Georgopoulos and Felix Juefei-Xu and Masayoshi Tomizuka and Yue Zhao},\nyear={2025},\nurl={https://openreview.net/forum?id=ayAx2YnfmD}\n}"},"title":{"value":"StreamDiT: Real-Time Streaming Text-to-Video Generation"},"pdf":{"value":"/pdf/03e9f73fab31abce981b0a4655273f351b6ee218.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"kodaira|streamdit_realtime_streaming_texttovideo_generation"},"authorids":{"value":["~Akio_Kodaira1","~Tingbo_Hou2","~Ji_Hou1","~Markos_Georgopoulos1","~Felix_Juefei-Xu1","~Masayoshi_Tomizuka2","~Yue_Zhao23"]},"authors":{"value":["Akio Kodaira","Tingbo Hou","Ji Hou","Markos Georgopoulos","Felix Juefei-Xu","Masayoshi Tomizuka","Yue Zhao"]}},"version":2},{"content":{"summary":{"value":"This paper proposes MagicTryOn, a DiT-based framework for video virtual try-on. It decomposes garments into semantic/structure/appearance cues, injects them via two cross-attention modules, and extends RoPE to garment-aware spatiotemporal positional encoding. The proposed method is evaluated on the ViViD dataset."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"See above"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. Experiments are conducted on both image-based and video-based datasets.\n2. Ablation studies are conducted to evaluate the effectiveness of each component."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The claim of correctly maintaining “compositional relationships” in multi-garment try-on is only supported by qualitative Fig.4, with no quantitative metrics (e.g., VFID-I3D, SSIM for multi-garment sequences) or statistical analysis. This leaves the performance of multi-garment handling unsubstantiated.\n2. The main comparison (Table 1) excludes recent state-of-the-art methods like DreamVVT (Zuo et al., 2025) or SwiftTry (Nguyen et al., 2025), which also focus on temporal consistency. This incomplete benchmarking makes it hard to contextualize MagicTryOn’s true standing in the current VVT landscape.\n3. Some qualitative examples are not promising. For instance, in Figure 4, the shorts in the generated video underwent deformation to align with the mask shape of the skirt from the original video.\n4. The paper lacks analysis of scenarios where the method may underperform."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762927234074,"tcdate":1761827792982,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission17298/Reviewer_XFjK"],"signatures":["ICLR.cc/2026/Conference/Submission17298/Reviewer_XFjK"],"forum":"JGiSS7hLRT","number":4,"license":"CC BY 4.0","cdate":1761827792982,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission17298/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762927234074,"domain":"ICLR.cc/2026/Conference","replyto":"JGiSS7hLRT","id":"cwykVv5Wkl","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Video Virtual Try-On","Diffusion Model","Video Generation"]},"supplementary_material":{"value":"/attachment/11697514ce3d41656d51492563e88f52edb2457e.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Video Virtual Try-On (VVT) aims to synthesize garments that appear natural across consecutive video frames, capturing both their dynamics and interactions with human motion. Despite recent progress, existing VVT methods still suffer from inadequate garment fidelity and limited spatiotemporal consistency. The reasons are (i) under-exploitation of garment information, with limited garment cues being injected, resulting in weaker fine-detail fidelity, and (ii) the lack of spatiotemporal modeling, which hampers cross-frame identity consistency and causes temporal jitter and appearance drift. In this paper, we present MagicTryOn, a diffusion transformer–based framework for garment-preserving video virtual try-on. To preserve fine-grained garment details, we propose a fine-grained garment-preservation strategy that disentangles garment cues and injects these decomposed priors into the denoising process. To improve temporal garment consistency and suppress jitter, we introduce a garment-aware spatiotemporal rotary positional embedding (RoPE) that extends RoPE within full self-attention, using spatiotemporal relative positions to modulate garment tokens. We further impose a mask-aware loss during training to enhance fidelity within garment regions. Moreover, we adopt distribution-matching distillation to compress the sampling trajectory to four steps, enabling real-time inference without degrading garment fidelity. Extensive quantitative and qualitative experiments demonstrate that MagicTryOn outperforms existing methods, delivering superior garment-detail fidelity and temporal stability in unconstrained settings. Code will be made publicly available."},"_bibtex":{"value":"@misc{\nli2026magictryon,\ntitle={MagicTryOn: Harnessing Diffusion Transformer for Garment-Preserving Video Virtual Try-on},\nauthor={Guangyuan Li and Siming Zheng and Hao Zhang and Jinwei Chen and Junsheng Luan and Binkai Ou and Lei Zhao and Bo Li and Peng-Tao Jiang},\nyear={2026},\nurl={https://openreview.net/forum?id=JGiSS7hLRT}\n}"},"title":{"value":"MagicTryOn: Harnessing Diffusion Transformer for Garment-Preserving Video Virtual Try-on"},"pdf":{"value":"/pdf/5dbbd79b43084f5b68d8a72d876d9d5572ff0f22.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"li|magictryon_harnessing_diffusion_transformer_for_garmentpreserving_video_virtual_tryon"},"authorids":{"value":["~Guangyuan_Li1","~Siming_Zheng3","~Hao_Zhang52","~Jinwei_Chen3","~Junsheng_Luan1","~Binkai_Ou1","~Lei_Zhao3","~Bo_Li20","~Peng-Tao_Jiang1"]},"authors":{"value":["Guangyuan Li","Siming Zheng","Hao Zhang","Jinwei Chen","Junsheng Luan","Binkai Ou","Lei Zhao","Bo Li","Peng-Tao Jiang"]}},"version":2},{"content":{"venue":{"value":"CoRR 2025"},"pdf":{"value":"https://arxiv.org/pdf/2509.00396v1"},"venueid":{"value":"dblp.org/journals/CORR/2025"},"paperhash":{"value":"seshimo|daovi_distortionaware_omnidirectional_video_inpainting"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Ryosuke_Seshimo:","~Mariko_Isogawa2"]},"html":{"value":"https://doi.org/10.48550/arXiv.2509.00396"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2509-00396,\n  publtype={informal},\n  author={Ryosuke Seshimo and Mariko Isogawa},\n  title={DAOVI: Distortion-Aware Omnidirectional Video Inpainting},\n  year={2025},\n  month={September},\n  cdate={1756684800000},\n  journal={CoRR},\n  volume={abs/2509.00396},\n  url={https://doi.org/10.48550/arXiv.2509.00396}\n}\n"},"abstract":{"value":"Omnidirectional videos that capture the entire surroundings are employed in a variety of fields such as VR applications and remote sensing. However, their wide field of view often causes unwanted objects to appear in the videos. This problem can be addressed by video inpainting, which enables the natural removal of such objects while preserving both spatial and temporal consistency. Nevertheless, most existing methods assume processing ordinary videos with a narrow field of view and do not tackle the distortion in equirectangular projection of omnidirectional videos. To address this issue, this paper proposes a novel deep learning model for omnidirectional video inpainting, called Distortion-Aware Omnidirectional Video Inpainting (DAOVI). DAOVI introduces a module that evaluates temporal motion information in the image space considering geodesic distance, as well as a depth-aware feature propagation module in the feature space that is designed to address the geometric distortion inherent to omnidirectional videos. The experimental results demonstrate that our proposed method outperforms existing methods both quantitatively and qualitatively."},"title":{"value":"DAOVI: Distortion-Aware Omnidirectional Video Inpainting"},"authors":{"value":["Ryosuke Seshimo","Mariko Isogawa"]}},"tmdate":1771222809477,"pdate":1767139200000,"externalIds":["dblp:journals/corr/abs-2509-00396"],"tcdate":1771222800764,"writers":["~"],"signatures":["~Mariko_Isogawa2"],"forum":"z82ySL0Xqi","license":"CC BY-SA 4.0","number":824649,"cdate":1756684800000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1771222809477,"domain":"DBLP.org","id":"z82ySL0Xqi","version":2},{"content":{"summary":{"value":"This paper introduces CVD-STORM, which employs STORM-VAE augmented with a Gaussian splatting reconstruction for a cross-view video diffusion model. The approach enables long-horizon, controllable multi-view generation and supports direct 4D spatiotemporal scene reconstruction from the generated latents. On the nuScenes dataset, the authors report markedly lower FID and FVD than recent multi-view generation baselines such as UniMLVG and DiVE, and they showcase depth and geometry obtained through a jointly trained Gaussian decoder."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"Please see the weaknesses."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"-  STORM-VAE, augmented by a 4D reconstruction auxiliary task, helps learn a geometry-aware latent space. \n- Beyond video generation, CVD-STORM can directly reconstruct dynamic 3D Gaussian Splatting from the latents.\n- Stable convergence with single-stage joint training of MM-DiT, temporal, and cross-view modules, offering practical value."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The method is more accurately characterized as an “enhanced representation learning” integration. In this approach, STORM’s 4D reconstruction head is migrated onto the SD 3.5 VAE and used for video diffusion. A more systematic comparison of “geometry-aware latent spaces” (for instance, using VAEs trained with depth or occupancy supervision) is recommended. Additional quantitative comparisons and discussions with generation+reconstruction paradigms such as MagicDrive3D and ReconDreamer, evaluated on metrics like novel-view quality and multiview consistency, would further strengthen the work.\n\n- CVD-STORM employs three frames as reference frames, but it is unclear from Tables 1 and 2(a) whether the initial frame uses a GT reference. If the GT frame is indeed used as the initial reference, the comparison presented in Table 1 may not be fair. Furthermore, when compared to UniMLVG, the downstream metrics do not demonstrate a clear advantage.\n\n- Multiview consistency metrics are missing. In addition, for Table 3(b), it would be beneficial to include the results from CVD-STORM + STORM (i.e., using the videos generated by CVD-STORM) to provide a fairer comparison. Evaluating only depth is limiting; reporting novel-view image quality metrics and additional comparisons would offer a more comprehensive assessment.\n\n- Finally, training details for STORM-VAE are not provided. Given the strong methodological and architectural similarities between STORM-VAE and STORM, an explanation of why STORM-VAE shows slightly better performance in Table 3(a) would be very helpful."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762918023412,"tcdate":1761789958906,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission5359/Reviewer_PbTC"],"signatures":["ICLR.cc/2026/Conference/Submission5359/Reviewer_PbTC"],"forum":"V66pMNOVC2","number":1,"license":"CC BY 4.0","cdate":1761789958906,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission5359/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762918023412,"domain":"ICLR.cc/2026/Conference","replyto":"V66pMNOVC2","id":"Br08lU8M2k","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["generative model","world model"]},"supplementary_material":{"value":"/attachment/8db641432b916000a33319eb6724b71614039954.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Generative models have been widely applied to world modeling for environment simulation and future state prediction. With advancements in autonomous driving, there is a growing demand not only for high-fidelity video generation under various controls, but also for producing diverse and meaningful information such as depth estimation. To address this, we propose CVD-STORM, a cross-view video diffusion model utilizing a spatial-temporal reconstruction Variational Autoencoder (VAE) that generates long-term, multi-view videos with 4D reconstruction capabilities under various control inputs. Our approach first fine-tunes the VAE with an auxiliary 4D reconstruction task, enhancing its ability to encode 3D structures and temporal dynamics. Subsequently, we integrate this VAE into the video diffusion process to significantly improve generation quality. Experimental results demonstrate that our model achieves substantial improvements in both FID\nand FVD metrics. Additionally, the jointly-trained Gaussian Splatting Decoder effectively reconstructs dynamic scenes, providing valuable geometric information for comprehensive scene understanding."},"_bibtex":{"value":"@misc{\nzhang2025cvdstorm,\ntitle={{CVD}-{STORM}: Cross-View Video Diffusion with Spatial-Temporal Reconstruction Model for Autonomous Driving},\nauthor={Tianrui ZHANG and Yichen Liu and Zilin Guo and Yuxin Guo and Jingcheng Ni and Chenjing Ding and Dan Xu and Lewei Lu and Zehuan Wu},\nyear={2025},\nurl={https://openreview.net/forum?id=V66pMNOVC2}\n}"},"title":{"value":"CVD-STORM: Cross-View Video Diffusion with Spatial-Temporal Reconstruction Model for Autonomous Driving"},"pdf":{"value":"/pdf/036b30bcf7a10a18f82dbdbd455b36aa5b97914c.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"zhang|cvdstorm_crossview_video_diffusion_with_spatialtemporal_reconstruction_model_for_autonomous_driving"},"authorids":{"value":["~Tianrui_ZHANG4","~Yichen_Liu3","~Zilin_Guo3","~Yuxin_Guo8","~Jingcheng_Ni1","~Chenjing_Ding1","~Dan_Xu4","~Lewei_Lu1","~Zehuan_Wu1"]},"authors":{"value":["Tianrui ZHANG","Yichen Liu","Zilin Guo","Yuxin Guo","Jingcheng Ni","Chenjing Ding","Dan Xu","Lewei Lu","Zehuan Wu"]}},"version":2},{"content":{"summary":{"value":"This paper proposes E3-PRUNER, which is a layer-pruning framework for large language models that learns which layers to keep or remove through a differentiable mask optimized with a Gumbel-TopK sampler. It further uses entropy-aware adaptive knowledge distillation to retain performance. Experiments show it achieves up to 2.18× speedup with minimal accuracy loss, offering an efficient, economical, and effective solution for LLM compression."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"This is the first academic paper I have seen that fine-tunes the 671B DeepSeek-R1 model. How many GPUs do you use? This information could be included in the Settings section of the paper."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. This paper proposes a differentiable mask learning framework (Gumbel-TopK) for layer pruning, enabling efficient gradient-based layer selection.\n2. The method achieves a good pruned model performance, outperforming prior pruning methods.\n3. The paper further introduces entropy-aware adaptive knowledge distillation, effectively preserving key reasoning tokens.\n4. The experiments demonstrate consistent and superior results across multiple LLMs with minimal accuracy loss."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The paper provides a limited theoretical explanation of why the Gumbel-TopK mask search is able to identify the optimal layers.\n2. The paper does not clarify whether the performance gain comes from the layer pruning method or the Adaptive Knowledge Distillation. It would be better to compare the zero-shot performance of the pruned model without fine-tuning or apply Adaptive KD to baseline pruning methods to evaluate their relative effects."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762918226701,"tcdate":1762094316363,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission5733/Reviewer_iaGe"],"signatures":["ICLR.cc/2026/Conference/Submission5733/Reviewer_iaGe"],"forum":"zYqpnm20jB","number":3,"license":"CC BY 4.0","cdate":1762094316363,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission5733/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762918226701,"domain":"ICLR.cc/2026/Conference","replyto":"zYqpnm20jB","id":"wKNLQbfVrW","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Large language model","model compression","layer pruning"]},"primary_area":{"value":"foundation or frontier models, including LLMs"},"abstract":{"value":"With the increasing size of large language models, layer pruning has gained increased attention as a hardware-friendly approach for model compression. However, existing layer pruning methods struggle to simultaneously address key practical deployment challenges, including performance degradation, high training costs, and limited acceleration. To overcome these limitations, we propose \\name, a task-\\underline{E}ffective, training-\\underline{E}conomical and inference-\\underline{E}fficient layer pruning framework. \\namespace introduces two key innovations: (1) a differentiable mask optimization method using a Gumbel-TopK sampler, enabling efficient and precise pruning mask search; and (2) an entropy-aware adaptive knowledge distillation strategy that enhances task performance. Extensive experiments over  diverse model architectures and benchmarks demonstrate the superiority of our method over state-of-the-art approaches. Notably, \\namespace achieves 96\\% accuracy, a mere 0.8\\% drop from the original model (96.8\\%) on MATH-500 when pruning 25\\% layers of Qwen3-32B, outperforming existing SOTA (95\\%), with a 1.33$\\times$ inference speedup by consuming merely 0.5B tokens (0.5\\% of the post-training data volume)."},"_bibtex":{"value":"@misc{\nyuan2026epruner,\ntitle={E\\${\\textasciicircum}3\\$-Pruner: Towards Efficient, Economical, and Effective Layer Pruning for Large Language Models},\nauthor={Tao Yuan and Haoli Bai and PanYinfei and Xuyang Cao and Tianyu Zhang and Lu Hou and Ting Hu and Xianzhi Yu},\nyear={2026},\nurl={https://openreview.net/forum?id=zYqpnm20jB}\n}"},"title":{"value":"E$^3$-Pruner: Towards Efficient, Economical, and Effective Layer Pruning for Large Language Models"},"pdf":{"value":"/pdf/c6e1cf3cd3d8da07a69f7c5dd50febad3e83d4f5.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"yuan|e^3pruner_towards_efficient_economical_and_effective_layer_pruning_for_large_language_models"},"authorids":{"value":["~Tao_Yuan4","~Haoli_Bai2","~PanYinfei1","~Xuyang_Cao4","~Tianyu_Zhang12","~Lu_Hou2","~Ting_Hu3","~Xianzhi_Yu1"]},"authors":{"value":["Tao Yuan","Haoli Bai","PanYinfei","Xuyang Cao","Tianyu Zhang","Lu Hou","Ting Hu","Xianzhi Yu"]}},"version":2},{"content":{"summary":{"value":"In this paper, it proposed a HOI synthesis method, termed as ArtHOI. It utilize a video diffusion model to generate a HOI video. Then, it utilize part segmentation for the object. Finally, the articulated object and human motion is reconstructed from the HOI video. The proposed method has been validated in different scenes and shows better performance than some previous methods."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"1. The quantitative analysis is insufficient. Whether it is possible to compare the proposed method with CHOIS according to the experimental settings of CHOIS."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. This paper utilize video diffusion motion for human-object interaction synthesis, making the synthesized HOI motion corresponding to video prior.\n2. The articulated object can be modeled according to the video and text description.\n3. The HOI synthesis is accompanied by human 3DGS modeling."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Articulated object modeling and motion capture has been widely researched. In this paper, it is utilized for HOI synthesis through a video diffusion model. This may diminish the significance of your contribution.\n2. Rather than HOI synthesis, it is more like 4D reconstruction after a video diffusion model.\n3. According to the video, the refrigerator is completely suspended in mid-air.\n4. There no detailed description of the datasets for evaluation."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762915657251,"tcdate":1761711172325,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission1021/Reviewer_YiKG"],"signatures":["ICLR.cc/2026/Conference/Submission1021/Reviewer_YiKG"],"forum":"NE1yczn1Qz","number":2,"license":"CC BY 4.0","cdate":1761711172325,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission1021/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762915657251,"domain":"ICLR.cc/2026/Conference","replyto":"NE1yczn1Qz","id":"xXauq489IR","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"TLDR":{"value":"ArtHOI enables zero-shot synthesis of realistic human interactions with articulated objects."},"keywords":{"value":["Articulated Human-object Interaction","Zero-shot Synthesis","Dynamics Distillation"]},"supplementary_material":{"value":"/attachment/14bf1f3f689f1371d9b3eed15165d834decb7dd7.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Synthesizing realistic articulated human-object interactions is challenging, especially when explicit 3D/4D supervision is unavailable. Recent zero-shot methods distill dynamics priors from pretrained video diffusion models, but this setting inherently provides only monocular evidence. That makes articulated part motion highly ambiguous and tightly coupled with human actions, so prior work falls back to rigid-object assumptions and fails on everyday articulated scenes (e.g., containing doors, fridges, cabinets). We introduce **ArtHOI**, the first zero-shot framework for synthesizing articulated human-object interactions via dynamics distillation from monocular video priors. We make two critical designs: **1)** *Flow-based part segmentation*: we use optical-flow cues to separate dynamic from static regions, because motion is the most reliable signal when multi-view information is absent. **2)** *Decoupled dynamics distillation*: joint optimization of human motion and object articulation is unstable under monocular ambiguity, so we first recover object articulation, then synthesize human motion conditioned on the reconstructed object states. ArtHOI distills dynamics from monocular 2D video priors without any 3D/4D ground truth. Across diverse scenes, ArtHOI yields physically plausible articulated interactions, improving contact quality and reducing penetration while enabling behaviors beyond rigid-only baselines. This extends zero-shot HOI synthesis from rigid manipulation to articulated dynamics. Code will be available."},"_bibtex":{"value":"@misc{\nhuang2025arthoi,\ntitle={Art{HOI}: Articulated Human-Object Interaction Synthesis via Dynamics Distillation},\nauthor={Zihao Huang and Tianqi Liu and Zhaoxi Chen and Shaocong Xu and Saining Zhang and Lixing Xiao and Zhiguo Cao and Wei Li and Hao Zhao and Ziwei Liu},\nyear={2025},\nurl={https://openreview.net/forum?id=NE1yczn1Qz}\n}"},"title":{"value":"ArtHOI: Articulated Human-Object Interaction Synthesis via Dynamics Distillation"},"pdf":{"value":"/pdf/3814acaa05dcaa5c1854460bfd3a5062ae416421.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"huang|arthoi_articulated_humanobject_interaction_synthesis_via_dynamics_distillation"},"authorids":{"value":["~Zihao_Huang2","~Tianqi_Liu3","~Zhaoxi_Chen1","~Shaocong_Xu1","~Saining_Zhang3","~Lixing_Xiao3","~Zhiguo_Cao1","~Wei_Li51","~Hao_Zhao1","~Ziwei_Liu1"]},"authors":{"value":["Zihao Huang","Tianqi Liu","Zhaoxi Chen","Shaocong Xu","Saining Zhang","Lixing Xiao","Zhiguo Cao","Wei Li","Hao Zhao","Ziwei Liu"]}},"version":2},{"content":{"summary":{"value":"Following the growing importance of compact video representations for video generation, this paper revisits video codec–inspired video VAE by designing a model that explicitly separates keyframe and inter-frame dynamic compression. Starting from a pre-trained image VAE, which is efficient, the authors introduce a Temporal Dynamic Difference Convolution (TDC) operator to learn sparse motion residuals from inter-frame differences. The quantitative results in Table 1 and Table 3 of text and human-face datasets are interesting and demonstrate strong reconstruction quality.\n\nHowever, the paper does not sufficiently discuss related prior works and lacks comparisons with relevant baseline models, which weakens the positioning of the claimed contributions. In addition, key experimental results such as quantitative evaluation for video generation and validation of common issues (e.g., flickering, reconstruction consistency) are missing."},"soundness":{"value":2},"confidence":{"value":5},"questions":{"value":"- How is the number of keyframes per clip chosen? Is it fixed or adaptively determined based on content?\n- How does the model handle dynamic backgrounds where motion and content separation is ambiguous?\n- Why does the ImageNet result outperform an image-only VAE baseline? What architectural or training differences explain this?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"- The video codec–inspired video encoding approach is interesting and shows strong reconstruction performance demonstrating that the potential benefits of incorporating codec-inspired structures into Video VAEs.\n- The observation that the first frame of a video sequence often has poorer reconstruction quality than later frames is insightful and could inspire future work on temporal consistency and keyframe modeling."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"While the core idea and experimental observations are interesting, the paper currently lacks clear validation and sufficient discussion to substantiate its main claims. Specifically, the issues described below (regarding the overstated key claim, unclear problem motivation, and insufficient quantitative validation) limit the overall strength of the contribution.\n\n1. **Key claim is somewhat overstated**\n- The paper claims that leveraging video codec design principles to improve Video VAEs remains a sparsely explored area. However, in the video compression domain, many recent works have already designed autoencoder-based architectures inspired by video codecs (e.g., [1], [2]). While most VAE-based video generative models focus on compression without explicitly disentangling keyframes and motion as in codecs, the authors should clarify how their approach fundamentally differs from these prior works. Please discuss these differences explicitly in the paper.\n\n- Moreover, the idea of separating static content and dynamic motion has already been explored in prior works such as [3] and [4]. Although these methods may not explicitly follow codec structures, the concept of content–motion disentanglement is not new. The paper should discuss and compare against these methods in more detail to clarify its novelty.\n\n2. **Unclear motivation for the implicit modeling problem**\n- The paper claims that implicit modeling is problematic but does not provide specific failure cases or quantitative evidence to support this argument. Some experimental validation (e.g., ablations or comparisons) would help clarify what concrete problem the proposed method is solving.\n- In addition, the method assumes that separating motion from a redundant static background improves efficiency, yet it is unclear how the approach performs under dynamic backgrounds, such as with camera motion, crowded scenes, or complex natural textures. Evaluating this would strengthen the motivation and validity of the proposed formulation.\n\n3. **Experimental limitations and lack of validation**\n- No quantitative evaluation for video generation. Although the method is motivated by generation via reconstruction, no FVD, KVD, or other generative quality metrics are reported. The qualitative results in the appendix also do not clearly demonstrate improvements.\n- The paper mentions that we identified a common issue, but provides no formal validation protocol or statistical evidence. This remains anecdotal. Figure 4 shows only one example and lacks context: how many videos were tested, whether the phenomenon holds across the entire dataset, and whether other metrics were considered.\n- Missing runtime analysis. Encoding and decoding times are not reported, making it difficult to assess the practical efficiency of the proposed approach compared to existing baselines.\n\n[1] DVC: An End-to-end Deep Video Compression Framework, Lu et al., CVPR 2019\\\n[2] Neural Inter-Frame Compression for Video Coding, Djelouah et al., ICCV 2019\\\n[3] CMD: Efficient Video Diffusion Models via Content-Frame Motion-Latent Decomposition, Yu et al., ICLR 2024\\\n[4] Video Probabilistic Diffusion Models in Projected Latent Space, Yu et al., CVPR 2023\\\n\nIf all of these issues are properly addressed through clearer positioning, additional experiments, and more rigorous analysis, I would be willing to raise my score."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762915460498,"tcdate":1761803818892,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission162/Reviewer_maVj"],"signatures":["ICLR.cc/2026/Conference/Submission162/Reviewer_maVj"],"forum":"UBsmQXhXg8","number":2,"license":"CC BY 4.0","cdate":1761803818892,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission162/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762915460498,"domain":"ICLR.cc/2026/Conference","replyto":"UBsmQXhXg8","id":"eqjs0M9Ysp","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Video VAE"]},"primary_area":{"value":"generative models"},"abstract":{"value":"Video Variational Auto-Encoders (Video VAEs) compress video data from the highly redundant pixel space into a compact latent representation, playing an important role in state-of-the-art video generation models. However, existing methods typically learn inter-frame correlations implicitly, overlooking the potential of breaking down video compression into two separate parts: keyframe encoding and inter-frame dynamic encoding, which is a fundamental design of traditional video codecs. To address this, we incorporate traditional video codec standard design into the Video VAE and introduce VC-VAE, a model that explicitly separates keyframe and inter-frame dynamic compression. We start by establishing a high-fidelity static keyframe anchor through initialization from a powerful pre-trained image VAE. Then, to explicitly model dynamic relative to this anchor, we introduce the Temporal Dynamic Difference Convolution (TDC), an operator designed to learn sparse motion residuals from inter-frame differences while maintaining a separate pathway for static content. Qualitative and quantitative experiments show that our proposed VC-VAE significantly outperforms baseline models in reconstruction quality, dynamic modelling, and training efficiency."},"_bibtex":{"value":"@misc{\nge2025vcvae,\ntitle={{VC}-{VAE}: Enhancing Video {VAE} with Video Codec Standard for Latent Video Diffusion Model},\nauthor={Xinxu Ge and Shang Chai and Litong Gong and Zitong YU and Xin Liu and Tiezheng Ge},\nyear={2025},\nurl={https://openreview.net/forum?id=UBsmQXhXg8}\n}"},"title":{"value":"VC-VAE: Enhancing Video VAE with Video Codec Standard for Latent Video Diffusion Model"},"pdf":{"value":"/pdf/6bc69a0444086367135743c6092fe44ec0a7ab3f.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"ge|vcvae_enhancing_video_vae_with_video_codec_standard_for_latent_video_diffusion_model"},"authorids":{"value":["~Xinxu_Ge1","~Shang_Chai1","~Litong_Gong1","~Zitong_YU3","~Xin_Liu25","~Tiezheng_Ge3"]},"authors":{"value":["Xinxu Ge","Shang Chai","Litong Gong","Zitong YU","Xin Liu","Tiezheng Ge"]}},"version":2},{"content":{"summary":{"value":"This paper proposes Self-Discriminative Optimization (SDO), a new post-training method for video diffusion models that improves quality without human labels or reward models.\nSDO introduces self-degradation, perturbing real video latents in the frequency domain to create controlled “degraded” samples that resemble low-quality generations. The model is then fine-tuned to distinguish real vs. degraded pairs, providing a stable, label-free supervision signal.\nApplied to CogVideoX-2B/5B, SDO improves structural fidelity, motion consistency, and semantic alignment with minimal data (≈2k videos) and 500 LoRA steps.\nCompared to LoRA tuning, DPO, and DDO, it achieves higher VBench and lower FVD/KVD scores, showing better temporal smoothness and fewer artifacts.\nThe method is efficient, avoids overfitting to score models, and stabilizes optimization, though its performance depends on the diversity of fine-tuning data."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"Please see the weakness."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":1},"contribution":{"value":2},"strengths":{"value":"- Proposes a novel label-free optimization method that learns from real videos and their degraded counterparts generated via the diffusion model itself. This design eliminates the need for human annotations, significantly reducing alignment cost while providing reliable supervision.\n- Demonstrates consistent improvements across different model scales (CogVideoX-2B and 5B), showing the method’s scalability and general applicability to various diffusion backbones."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The paper provides almost no quantitative results, and the evaluation relies mainly on classical metrics such as FVD, which are poorly correlated with human perceptual quality and therefore lack persuasiveness for assessing modern video generation models [1]. The only relatively new metric used is VBench, but only the total score is reported. Reporting only the total VBench score is insufficient because it is biased toward temporal consistency, meaning that results that sacrifice dynamics can still achieve deceptively high scores [2], [3], [4]. Therefore, please present all individual VBench metric values.\n- The absence of any human evaluation further weakens the argument. Based on the reported quantitative results, the performance gains appear marginal from a human perception standpoint. To address this concern, the authors should conduct a Gold Human pairwise comparison experiment against prior methods, as this has become the standard in recent video generation papers [5], [6], [7] [8].\n- The choice of base models is also insufficient. The paper does not discuss whether the proposed method could yield performance gains on state-of-the-art video diffusion models, such as Wan 2.1 or HunyuanVideo. Without such discussion or experiments, it remains unclear whether the method is effective and scalable to the latest, large-scale models, limiting the generality and practical significance of the results.  \n- Finally, the ablation studies are far from adequate. While it is clear that FFT is used to decompose frequency components and assign weights, the paper does not discuss the thresholding strategy. In addition, does this method actually achieve better performance efficiency in terms of training steps or computation cost, compared to prior methods? Moreover, it would be valuable to analyze not only the overall generation quality measured by benchmarks, but also how the proposed alignment method affects generation diversity.\n\n[1] Ge, et al. On the Content Bias in Fr ́echet Video Distance. CVPR2024.  \n[2] Liao, et al. Evaluation of Text-to-Video Generation Models: A Dynamics Perspective. NeurIPS2024.  \n[3] Liu, et al. VideoDPO: Omni-Preference Alignment for Video Diffusion Generation. CVPR2025.  \n[4] Oshima, et al. Inference-Time Text-to-Video Alignment with Diffusion Latent Beam Search. NeurIPS2025.  \n[5] Wu, et al. Boosting Text-to-Video Generative Model with MLLMs Feedback. NeurIPS2024.  \n[6] Hila, et al. VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion\nGeneration in Video Models. ICML2025.  \n[7] Shaulov, et al. FlowMo: Variance-Based Flow Guidance for Coherent Motion in Video Generation. NeurIPS2025.  \n[8] Wu, et al. DenseDPO: Fine-Grained Temporal Preference Optimization for Video Diffusion Models. NeurIPS2025."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762921421106,"tcdate":1761792115965,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission9992/Reviewer_mPBn"],"signatures":["ICLR.cc/2026/Conference/Submission9992/Reviewer_mPBn"],"forum":"I4jBCglUOI","number":1,"license":"CC BY 4.0","cdate":1761792115965,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission9992/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762921421106,"domain":"ICLR.cc/2026/Conference","replyto":"I4jBCglUOI","id":"R7czfKxzqx","forumContent":{"venue":{"value":"ICLR 2026 Conference Desk Rejected Submission"},"keywords":{"value":["Video Generation; Post-training; Diffusion Models"]},"supplementary_material":{"value":"/attachment/4a0027ff1916872daebb868d488c775154c3baaa.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Recent preference alignment strategies have gained traction in large language models (LLMs) and are now being extended to broader generative domains. Approaches such as Direct Preference Optimization have been adapted to diffusion models by leveraging human-labeled preferences or auxiliary score models to distinguish ``winner'' from ``loser''. \nHowever, these methods face two key challenges: (1) the optimization process often overfits to the score model, resulting in suboptimal generation quality; and (2) the results generated from the same text prompt exhibit significant divergence, resulting in limited effective gradients and reduced training efficiency. These limitations are further exacerbated in video generation, where evaluation is more complex and inference is slower. In this work, we introduce Self-Discriminative Optimization that using only a handful of real samples, unlocks markedly higher-quality generation. First, we introduce self-degradation that applies frequency-domain reweighting to the latent representations from real samples, yielding degraded samples that more closely match the model’s original output distribution. This leads to controlled distortions such as low-quality, temporal inconsistency and object deformation.\nWe then use these real/degraded pairs as positive and negative examples to fine-tune the pretrained model discriminatively with automatically assigned, reliable labels. \nBy exploiting the richer gradients from these controllable degradation pairs, our experiments demonstrate substantial gains in structural quality and semantic alignment using only a handful of high-quality samples and minimal fine-tuning."},"_bibtex":{"value":"@misc{\nanonymous2026selfdiscriminative,\ntitle={Self-Discriminative Optimization for Video Diffusion Models},\nauthor={Anonymous},\nyear={2026},\nurl={https://openreview.net/forum?id=I4jBCglUOI}\n}"},"title":{"value":"Self-Discriminative Optimization for Video Diffusion Models"},"pdf":{"value":"/pdf/2c3b65a573548517220076d4f075149428b867a1.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Desk_Rejected_Submission"},"paperhash":{"value":"guo|selfdiscriminative_optimization_for_video_diffusion_models"},"authorids":{"value":["~Lanqing_Guo1","~Yichen_Liu17","~Yufei_Wang5","~Yingqing_He1","~Yan_Zheng4","~Hezhen_Hu2","~Junbo_Li3","~Xiaoyu_Wang1","~Zhangyang_Wang1"]},"authors":{"value":["Lanqing Guo","Yichen Liu","Yufei Wang","Yingqing He","Yan Zheng","Hezhen Hu","Junbo Li","Xiaoyu Wang","Zhangyang Wang"]}},"version":2},{"content":{"summary":{"value":"The paper introduces InstanceAnimator, a diffusion transformer (DiT)-based framework for multi-instance sketch video colorization. Unlike existing methods that rely heavily on a single reference frame, InstanceAnimator incorporates a Canvas Guidance Condition to allow flexible placement of multiple reference elements, an Instance Matching Mechanism to enforce consistency between sketches and references, and an Adaptive Decoupled Control Module to preserve fine-grained semantic details for both characters and backgrounds. The paper presents thorough experimental results—including quantitative metrics, qualitative figures, and ablations—demonstrating improved controllability, identity preservation, temporal consistency, and usability compared to several strong baselines."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"In the anime production process, multi-view character design sheets are often used as character references. Why wasn't this form of reference sheet used?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. Unlike previous image-to-animation architectures, this paper provides a reference-to-animation solution, which simplifies the animation production pipeline.\n2. Allows users to freely position and edit reference elements, enabling more realistic animation workflows\n3. The framework is extensively evaluated across four strong baselines with appropriate metrics (FVD, SSIM, LPIPS, CLIP, Temporal) as summarized in Table 1, and ablation studies (Table 2) convincingly establish the impact of each module."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The key problem of reference-to-video is actually the data pre-process. For example, in MovieGen, the authors find that training solely on the above paired data makes the model easily learn a copy-paste shortcut solution, i.e., the generated video always follows the expression or the head pose from the reference face. Thus, they proposed to collect reference data from outside of the video clip. However, in this paper, the authors use SAM to crop the reference instances from the first frames. I suspect that the shortcut problem exists in the proposed method and the authors ignore it.\n2. Another problem of reference-to-image(video) is the instance matching. The position of the instances would interchange with each other. Although the authors proposed an Instance Matching Mechanism, it is a learning-based algorithm. It cannot convince me that the model can make sure that the characters would not exchange.\n3. The third problem is the background extraction. Directly matting from the first frame will create silhouette, which need image inpainting to fix. However, the imperfect inpainting will still leak sketch information during training, which would also cause a shortcut problem. This problem is not mentioned in this paper."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1764360152520,"tcdate":1761556272512,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission9426/Reviewer_rYRy"],"signatures":["ICLR.cc/2026/Conference/Submission9426/Reviewer_rYRy"],"forum":"9zVvlSKDZx","number":1,"license":"CC BY 4.0","cdate":1761556272512,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission9426/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1764360152520,"domain":"ICLR.cc/2026/Conference","replyto":"9zVvlSKDZx","id":"3Yjv6QoGny","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Sketch","Colorization","Video Generation"]},"supplementary_material":{"value":"/attachment/862957c06e2601ca23524c247c68d7420f9489d1.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"We propose InstanceAnimator, a novel Diffusion Transformer (DiT)-based framework for multi-instance sketch video colorization.\nExisting animation colorization methods rely heavily on a single initial reference frame, resulting in fragmented workflows and limited customizability. To eliminate these constraints, we introduce a Canvas Guidance Condition that allows users to freely place reference elements on a blank canvas, enabling flexible user control. To address the misalignment and quality degradation issues of DiT-based approaches, we design an Instance Matching Mechanism that integrates the instances with the sketch and noise channels, ensuring visual consistency across different sequences while maintaining controllability. Additionally, to mitigate the degradation of fine-grained details, we propose an Adaptive Decoupled Control Module that injects semantic features from characters, backgrounds, and text conditions into the diffusion model, significantly enhancing detail fidelity.  Extensive experimental results demonstrate that InstanceAnimator effectively enables better user control in multi-instance colorization, producing high-fidelity results with strong temporal consistency."},"_bibtex":{"value":"@misc{\nzhang2026instanceanimator,\ntitle={InstanceAnimator: Multi-Instance Sketch Video Colorization},\nauthor={Yinhan Zhang and Yue Ma and Bingyuan Wang and Kunyu Feng and Qifeng Chen and Zeyu Wang},\nyear={2026},\nurl={https://openreview.net/forum?id=9zVvlSKDZx}\n}"},"title":{"value":"InstanceAnimator: Multi-Instance Sketch Video Colorization"},"pdf":{"value":"/pdf/13507534cd6e459efa9714d537de15890f12ec59.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"zhang|instanceanimator_multiinstance_sketch_video_colorization"},"authorids":{"value":["~Yinhan_Zhang1","~Yue_Ma2","~Bingyuan_Wang1","~Kunyu_Feng1","~Qifeng_Chen1","~Zeyu_Wang15"]},"authors":{"value":["Yinhan Zhang","Yue Ma","Bingyuan Wang","Kunyu Feng","Qifeng Chen","Zeyu Wang"]}},"version":2},{"content":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["vision-language model","diagram understanding","multimodal reasoning","visual question answering","shortcut learning","knowledge grounding","benchmark dataset","multimodal evaluation"]},"primary_area":{"value":"interpretability and explainable AI"},"abstract":{"value":"Diagrams convey symbolic information in a visual format rather than a linear stream of words, making them especially challenging for AI models to process. While recent evaluations suggest that vision-language models (VLMs) perform well on diagram-related benchmarks, their reliance on knowledge, reasoning, or modality shortcuts raises concerns about whether they genuinely understand and reason over diagrams.\nTo address this gap, we introduce Chimera, a comprehensive test suite comprising 7,500 high-quality diagrams sourced from Wikipedia; each diagram is annotated with its symbolic content represented by semantic triples along with multi-level questions designed to assess four fundamental aspects of diagram comprehension: entity recognition, relation understanding, knowledge grounding, and visual reasoning.\nWe use Chimera to measure the presence of three types of shortcuts in visual question answering: \n(1) the visual-memorization shortcut, where VLMs rely on memorized visual patterns;\n(2) the knowledge-recall shortcut, where models leverage memorized factual knowledge instead of interpreting the diagram; and\n(3) the Clever-Hans shortcut, where models exploit superficial language patterns or priors without true comprehension. We evaluate 15 open-source VLMs from 7 model families on Chimera and find that their seemingly strong performance largely stems from shortcut behaviors: visual-memorization shortcuts have slight impact, knowledge-recall shortcuts play a moderate role, and Clever-Hans shortcuts contribute significantly.\nThese findings expose critical limitations in current VLMs and underscore the need for more robust evaluation protocols that benchmark genuine comprehension of complex visual inputs (e.g., diagrams) rather than question-answering shortcuts."},"_bibtex":{"value":"@misc{\nchi2026chimera,\ntitle={Chimera: Diagnosing Shortcut Learning in Visual-Language Understanding},\nauthor={Ziheng Chi and Yifan Hou and Chenxi Pang and Shaobo Cui and Mubashara Akhtar and Mrinmaya Sachan},\nyear={2026},\nurl={https://openreview.net/forum?id=q3eB3PhtqD}\n}"},"title":{"value":"Chimera: Diagnosing Shortcut Learning in Visual-Language Understanding"},"pdf":{"value":"/pdf/f92e46808334228be1f33c8c27812a84e924f889.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"chi|chimera_diagnosing_shortcut_learning_in_visuallanguage_understanding"},"authorids":{"value":["~Ziheng_Chi1","~Yifan_Hou1","~Chenxi_Pang1","~Shaobo_Cui1","~Mubashara_Akhtar1","~Mrinmaya_Sachan3"]},"authors":{"value":["Ziheng Chi","Yifan Hou","Chenxi Pang","Shaobo Cui","Mubashara Akhtar","Mrinmaya Sachan"]}},"tmdate":1767710394250,"tcdate":1758218335445,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission13473/Authors"],"signatures":["ICLR.cc/2026/Conference/Submission13473/Authors"],"forum":"q3eB3PhtqD","license":"CC BY 4.0","number":13473,"cdate":1758218335445,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/-/Submission","ICLR.cc/2026/Conference/-/Post_Submission","ICLR.cc/2026/Conference/Submission13473/-/Full_Submission","ICLR.cc/2026/Conference/-/Withdrawn_Submission"],"mdate":1767710394250,"odate":1759896705795,"domain":"ICLR.cc/2026/Conference","id":"q3eB3PhtqD","version":2},{"content":{"summary":{"value":"The paper aims to segment a wide range of concepts in video through a natural language. The paper proposes to rely on the powerful representation capabilities of the video diffusion model. Besides, to facilitate the research community, the paper introduces a new benchmark for Referral Video Process Segmentation (RVPS), which captures dynamic phenomena that exist at the intersection of video and language."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"The definition of RVOS in the paper is **Referral** Video Object Segmentation in the introduction section, but it should be **Referring** Video Object Segmentation."},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. The focus on segmenting video processes presents a new research direction, which is valuable for exploring and evaluating model performance in this context.\n2. The application of video diffusion to Referring Video Object Segmentation (RVOS) tasks is an interesting and promising approach."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The integration of video diffusion into referring segmentation appears incremental and lacks novelty, as it primarily adapts an existing video diffusion model to RVOS without specific tailoring for the task.\n2. A detailed computational complexity analysis is needed to better understand the efficiency of the proposed method.\n3. Both the figures and the overall writing quality require further refinement for better clarity and presentation."}},"nonreaders":[],"tmdate":1732975156683,"tcdate":1729307472919,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission581/Reviewer_BAgM"],"signatures":["ICLR.cc/2025/Conference/Submission581/Reviewer_BAgM"],"forum":"Ir6JxcuP6H","number":1,"license":"CC BY 4.0","cdate":1729307472919,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission581/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732975156683,"domain":"ICLR.cc/2025/Conference","replyto":"Ir6JxcuP6H","id":"kirZKVchun","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Referring Video Segmentation","Video Diffusion Models"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"We present REM, a framework for segmenting a wide variety of concepts in video that can be described through natural language. To achieve this level of generalization, our method capitalizes on visual-language representations learned by video diffusion models on Internet-scale datasets. A key insight of our approach is preserving as much of the generative model’s original representation as possible, while fine-tuning it on narrow-domain Referral Object Segmentation datasets. As a result, despite being exclusively trained on object masks from a limited set of categories, our framework is able to accurately segment and track both rare, unseen objects and non-object, dynamic concepts, such as waves crushing in the ocean. To better quantify the generalization capabilities of our model, we introduce a new benchmark for Referral Video Process Segmentation (RVPS), which captures dynamic phenomena that exist at the intersection of video and language. Our experiments show that REM performs comparably to state-of-the-art approaches on in-domain datasets while outperforming them by up to 28\\% out-of-domain, leveraging the power of Internet-scale pre-training."},"_bibtex":{"value":"@misc{\nbagchi2025refereverything,\ntitle={ReferEverything: Towards segmenting everything we can speak of in videos},\nauthor={Anurag Bagchi and Zhipeng Bao and Yu-Xiong Wang and Pavel Tokmakov and Martial Hebert},\nyear={2025},\nurl={https://openreview.net/forum?id=Ir6JxcuP6H}\n}"},"title":{"value":"ReferEverything: Towards segmenting everything we can speak of in videos"},"pdf":{"value":"/pdf/bd7d631ffb089b4987fd8dff3b7b73ecbf21660c.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"bagchi|refereverything_towards_segmenting_everything_we_can_speak_of_in_videos"},"authorids":{"value":["~Anurag_Bagchi1","~Zhipeng_Bao1","~Yu-Xiong_Wang1","~Pavel_Tokmakov2","~Martial_Hebert1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Anurag Bagchi","Zhipeng Bao","Yu-Xiong Wang","Pavel Tokmakov","Martial Hebert"]}},"version":2},{"content":{"venue":{"value":"MM2024 Poster"},"supplementary_material":{"value":"/attachment/cea39e01b84a4e7111e8115600239ca474eb1482.zip"},"abstract":{"value":"Current state-of-the-art video quality assessment (VQA) models typically integrate various perceptual features to comprehensively represent video quality degradation.  These models either directly concatenate features or fuse different perceptual scores while ignoring the domain gaps between cross-aware features, thus failing to adequately learn the correlations and interactions between different perceptual features. To this end, we analyze the independent effects and information gaps of quality- and semantic-aware features on video quality. Based on an analysis of the spatial and temporal differences between two aware features, we proposed a semantic-**A**ware and quality-**A**ware **I**nteraction **Net**work (**A$^2$INet**) for blind VQA (BVQA). For spatial gaps, we introduce a cross-aware guided interaction module to enhance the interaction between semantic- and quality-aware features in a local-to-global manner. Considering temporal discrepancies, we design a cross-aware temporal modeling module to further perceive temporal content variation and quality saliency information, and perceptual features are regressed into quality score by a temporal network and a temporal pooling.  Extensive experiments on six benchmark VQA datasets show that our model achieves state-of-the-art performance, and ablation studies further validate the effectiveness of each module. We also present a simple video sampling strategy to balance the effectiveness and efficiency of the model. The code for the proposed method will be released."},"relevance_to_conference":{"value":"Video quality assessment (VQA) is to enable the model to perceive the visual quality of videos and produce results consistent with human subjective opinions, making it a popular research topic in multimedia. Our work contributes significantly to the field of multimedia processing, aligning closely with the themes and objectives of the conference, \"Experience\". Specifically, our work improves the model's performance in predicting video quality by designing a quality-aware and quality-aware interaction network that enhances the spatial and temporal interactions between the two aware features. To the best of our knowledge, this is the first attempt to explore the gaps between various perceptual features and utilize cross-aware feature learning to enhance the perceptual capabilities of the blind VQA model. Extensive experiments validate that the proposed method outperforms the state-of-the-art VQA model on all six benchmark VQA datasets."},"_bibtex":{"value":"@inproceedings{\nxiang2024semanticaware,\ntitle={Semantic-Aware and Quality-Aware Interaction Network for Blind Video Quality Assessment},\nauthor={Jianjun Xiang and Yuanjie Dang and Peng Chen and Ronghua Liang and Ruohong Huan and Nan Gao},\nbooktitle={ACM Multimedia 2024},\nyear={2024},\nurl={https://openreview.net/forum?id=D4d0sy7TFJ}\n}"},"title":{"value":"Semantic-Aware and Quality-Aware Interaction Network for Blind Video Quality Assessment"},"secondary_subject_area":{"value":["[Experience] Multimedia Applications"]},"pdf":{"value":"/pdf/2367d6effd9b16f68f30ef05f3b3bdae5f9d8a23.pdf"},"venueid":{"value":"acmmm.org/ACMMM/2024/Conference"},"paperhash":{"value":"xiang|semanticaware_and_qualityaware_interaction_network_for_blind_video_quality_assessment"},"primary_subject_area":{"value":"[Experience] Interactions and Quality of Experience"},"authorids":{"value":["~Jianjun_Xiang1","~Yuanjie_Dang1","~Peng_Chen5","~Ronghua_Liang1","~Ruohong_Huan1","~Nan_Gao4"]},"authors":{"value":["Jianjun Xiang","Yuanjie Dang","Peng Chen","Ronghua Liang","Ruohong Huan","Nan Gao"]}},"tmdate":1721528142247,"pdate":1721513485312,"tcdate":1711613174084,"writers":["acmmm.org/ACMMM/2024/Conference","acmmm.org/ACMMM/2024/Conference/Submission1181/Authors"],"signatures":["acmmm.org/ACMMM/2024/Conference/Submission1181/Authors"],"forum":"D4d0sy7TFJ","license":"CC BY 4.0","number":1181,"cdate":1711613174084,"readers":["everyone"],"invitations":["acmmm.org/ACMMM/2024/Conference/-/Submission","acmmm.org/ACMMM/2024/Conference/-/Post_Submission","acmmm.org/ACMMM/2024/Conference/Submission1181/-/Revision","acmmm.org/ACMMM/2024/Conference/Submission1181/-/Supplementary_Material","acmmm.org/ACMMM/2024/Conference/-/Edit"],"mdate":1721528142247,"odate":1721513485312,"domain":"acmmm.org/ACMMM/2024/Conference","id":"D4d0sy7TFJ","version":2},{"content":{"summary":{"value":"In this paper, the authors present a framework that generate novel view scene from a single image by their proposed geometry-aware and temporal modeling. Specifically, they introduce two modules, a geometry-aware video depth refinement and a object-consistent temporal modeling mechanism."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"- Lack of visualizations in supplement video. In the main paper, the author compared Recamaster and TrajectoryCrafter, but did not compare them in the supplementary video. I am curious about their dynamic effects.\n- Lack of depth visualizations before and after depth refinement."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"- The structure of this paper is complete.\n- The authors provide source code, which should be encouraged."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Unfair settings for 3D methods. From the supplementary video, it can be seen that the 3D methods exhibit significant flickering in the newly filled areas. The correct approach should be to maintain the filling result of the i-th frame when filling in the i+1 frame, rather than completely refilling every frame from scratch.\n- Comparisons with 4D method. The authors mainly compare 3D methods, but currently there are many 4D methods that can be compared, such as Free4D[1], CAT4D [2], GenXD [3], and DimensionX [4].\n-  Wrong claims. The authors mentioned that semantic regions share similar depth distributions, but for a river, there is a significant difference in the depth of its foreground and background. Therefore, using Median Filtering may is not a good approach. In addition, Median Filtering sacrifices a lot of fine variations in real geometry (such as slopes/roads with depth gradients, thin structures, and surfaces), making it easy to flatten non planar surfaces into segmented parallel ones.\n-  The effectiveness of semantic segmentation. The authors use SAM in both Geometry-aware Video Depth Refinement and Object-consistent Temporal Modeling. But this will introduce more complexity, constraints, and limitations. The author should present and analyze more intermediate results to demonstrate the effectiveness of incorporating semantic segmentation. For example, for a scene with hundreds or thousands of people, how should semantic segmentation be performed\n-  Writing. This paper contains a lot of repetitive narratives in the same paragraph. For example, in line 79 to line 95, the authors repeatedly talk about the geometry-aware video depth refinement and the object-consistent temporal modeling. Can the authors explain its role in the model all at once.\n\n[1] Liu, Tianqi, et al. \"Free4D: Tuning-free 4D Scene Generation with Spatial-Temporal Consistency.\" arXiv preprint arXiv:2503.20785 (2025).\n\n[2] Wu, Rundi, et al. \"Cat4d: Create anything in 4d with multi-view video diffusion models.\" Proceedings of the Computer Vision and Pattern Recognition Conference. 2025.\n\n[3] Zhao, Yuyang, et al. \"Genxd: Generating any 3d and 4d scenes.\" arXiv preprint arXiv:2411.02319 (2024).\n\n[4] Sun, Wenqiang, et al. \"Dimensionx: Create any 3d and 4d scenes from a single image with controllable video diffusion.\" arXiv preprint arXiv:2411.04928 (2024)."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762918404403,"tcdate":1761443858660,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission5999/Reviewer_dk65"],"signatures":["ICLR.cc/2026/Conference/Submission5999/Reviewer_dk65"],"forum":"guUaZN0kyC","number":1,"license":"CC BY 4.0","cdate":1761443858660,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission5999/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762918404403,"domain":"ICLR.cc/2026/Conference","replyto":"guUaZN0kyC","id":"CVcJXCMlfc","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Scene Generation","Video Generation"]},"supplementary_material":{"value":"/attachment/b14a16ad6bae67e7e5b500c0d788de074931df51.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"We present WorldCrafter, a novel framework that enables interactive dynamic scene generation from a single image by leveraging geometry-aware and temporal modeling. Existing methods often suffer from texture distortion, structural inaccuracies, and temporal flickering under large viewpoint changes. These issues mainly caused by explicit pixel-wise reprojection strategies. To address these challenges, WorldCrafter introduces two complementary modules: 1) Geometry-aware Video Depth Refinement, which enhances structural fidelity by refining depth with multi-frame geometric priors and semantic cues; and  2) Object-consistent Temporal Modeling, which disentangles video frames into object-level layers to improve coherence between static backgrounds and dynamic foregrounds. These components form a unified rendering-inpainting framework for photorealistic and camera-controllable dynamic scene generation. Experiments demonstrate that WorldCrafter produces geometrically accurate and temporally coherent results across diverse scenes and camera trajectories."},"_bibtex":{"value":"@misc{\ndong2025worldcrafter,\ntitle={WorldCrafter: Dynamic Scene Generation from a Single Image with Geometric and Temporal Consistency},\nauthor={Haoye Dong and Gim Hee Lee},\nyear={2025},\nurl={https://openreview.net/forum?id=guUaZN0kyC}\n}"},"title":{"value":"WorldCrafter: Dynamic Scene Generation from a Single Image with Geometric and Temporal Consistency"},"pdf":{"value":"/pdf/bda6aad4d682f68ba69bd1ef445627742737bfa4.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"dong|worldcrafter_dynamic_scene_generation_from_a_single_image_with_geometric_and_temporal_consistency"},"authorids":{"value":["~Haoye_Dong1","~Gim_Hee_Lee1"]},"authors":{"value":["Haoye Dong","Gim Hee Lee"]}},"version":2},{"content":{"summary":{"value":"This paper introduces a new Motion Feature (MOFT) that can effectively capture motion information in video diffusion models. The authors reveal that robust motion-aware features already exist in video diffusion models, allowing to encode comprehensive motion information with clear interpretability. They present MOFT, which can be extracted without the need for training and is generalizable across diverse architectures."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"Clarification needed:\n\n- Figure 1:\n\n  a. It is unclear whether the motion feature in 1(a) is extracted semantically or spatially. Clarification is needed on how the similarity with other videos in 1(b) is calculated. Additionally, an explanation of what the higher score represents and why motion features from different videos could influence each other would be helpful.\n\n  b. In 1(c), the motion direction seems to be manually defined. If so, why does the paper state that MOFT serves as guidance for controlling motion direction? If MOFT controls the motion, what is the source video for that motion?\n\n- Figure 6: Why the comparison is presented in the form of a point for DIFT and a segment for MOFT."},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"- Training-free strategy effectively extracts motion information encoded in the features of the video diffusion model, demonstrating its ability to capture and leverage the inherent motion representations learned by the model.\n\n- The method presents a clean and straightforward solution for extracting motion encoding from video diffusion models, making it a ready and practical technique for various applications involving motion analysis or synthesis."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The paper lacks clarity on the training process. While it claims to be training-free, it defines loss functions for other tasks (Equations 3 and 4). It would be helpful to clarify which stages are trained and which are not.\n\n- The PCA analysis is based on a small number of videos (only 2 videos in Figure 2), which limits the generalizability of the results.\n\n- While the motion in the qualitative videos looks good, the differences compared to other alterations appear subtle and hard to recognize. Other methods show too poor results, were they tuned correctly?\n\n- The paper should report the runtime and resolution for better understanding of the method's computational requirements and output quality.\n\n- The idea is heavily inspired by DIFT and utilized for video applications, then novelty seems limited."},"limitations":{"value":"Limitations are addressed adequately."}},"nonreaders":[],"tmdate":1730878971800,"tcdate":1719729532374,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission5026/Reviewer_vDWU"],"signatures":["NeurIPS.cc/2024/Conference/Submission5026/Reviewer_vDWU"],"forum":"ZvQ4Bn75kN","number":1,"license":"CC BY 4.0","cdate":1719729532374,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission5026/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730878971800,"domain":"NeurIPS.cc/2024/Conference","replyto":"ZvQ4Bn75kN","id":"rx9LitT5gZ","forumContent":{"venue":{"value":"NeurIPS 2024 poster"},"keywords":{"value":["video generation","motion control"]},"supplementary_material":{"value":"/attachment/760725ae0ee571e480746606f3ffbbca8b516f4f.zip"},"primary_area":{"value":"generative_models"},"abstract":{"value":"Video generation primarily aims to model authentic and customized motion across frames, making understanding and controlling the motion a crucial topic. Most diffusion-based studies on video motion focus on motion customization with training-based paradigms, which, however, demands substantial training resources and necessitates retraining for diverse models. Crucially, these approaches do not explore how video diffusion models encode cross-frame motion information in their features, lacking interpretability and transparency in their effectiveness. To answer this question, this paper introduces a novel perspective to understand, localize, and manipulate motion-aware features in video diffusion models. Through analysis using Principal Component Analysis (PCA), our work discloses that robust motion-aware feature already exists in video diffusion models. We present a new MOtion FeaTure (MOFT) by eliminating content correlation information and filtering motion channels. MOFT provides a distinct set of benefits, including the ability to encode comprehensive motion information with clear interpretability, extraction without the need for training, and generalizability across diverse architectures. Leveraging MOFT, we propose a novel training-free video motion control framework. Our method demonstrates competitive performance in generating natural and faithful motion, providing architecture-agnostic insights and applicability in a variety of downstream tasks."},"_bibtex":{"value":"@inproceedings{\nxiao2024video,\ntitle={Video Diffusion Models are Training-free Motion Interpreter and Controller},\nauthor={Zeqi Xiao and Yifan Zhou and Shuai Yang and Xingang Pan},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=ZvQ4Bn75kN}\n}"},"title":{"value":"Video Diffusion Models are Training-free Motion Interpreter and Controller"},"pdf":{"value":"/pdf/655c3c980fb10954cccd67f1942955a6f177a0b8.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"xiao|video_diffusion_models_are_trainingfree_motion_interpreter_and_controller"},"authorids":{"value":["~Zeqi_Xiao2","~Yifan_Zhou11","~Shuai_Yang3","~Xingang_Pan1"]},"authors":{"value":["Zeqi Xiao","Yifan Zhou","Shuai Yang","Xingang Pan"]}},"version":2},{"content":{"summary":{"value":"X-PlugVid is a framework designed to adapt pretrained image-based plug-and-play modules for use in video diffusion models, allowing controllable video generation. The framework uses a spatial-temporal adapter to bridge the gap between image and video diffusion, with a frozen image diffusion model (e.g., Stable Diffusion v1.5) providing spatial priors. A timestep remapping strategy injects information from later timesteps of the image model into earlier timesteps of the video model, enhancing quality and temporal consistency. Experimental results show X-PlugVid’s compatibility with various video models and image plugins, and its adaptability for controllable video generation, supported by ablation studies and qualitative results."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"No questions"},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The use of an image model as a prior to support SVD for controllable video generation without additional training on either the video or image model is an interesting approach.\n2. The paper is well-organized and easy to follow."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. **Lack of Video Demonstrations**: There is no video demonstration provided to verify the effectiveness of the proposed method, which is a significant limitation for assessing its impact. Side-by-side comparisons with baselines, such as VideoComposer, ControlVideo, and Control-A-Video, are essential.\n\n2. **Inference Time Concerns**: There is a concern regarding inference speed, as the image model seems to be involved at each inference step. Specifically, how long does it take to generate a 48-frame (2-second) video clip?\n\n3. **Temporal Interpolation Challenge**: A key gap between controllable image generation and controllable video generation lies in producing reasonable interpolations when conditions vary significantly between frames. Appendix Table 2 lacks examples that address this challenge.\n\n4. **Influence of Prompts**: The effect of prompts on the outcomes is not clearly analyzed, and there appears to be a lack of ablation experiments assessing this aspect."}},"nonreaders":[],"tmdate":1731428641658,"tcdate":1730680255074,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission6384/Reviewer_xYJz"],"signatures":["ICLR.cc/2025/Conference/Submission6384/Reviewer_xYJz"],"forum":"TTWxMAwS6n","number":2,"license":"CC BY 4.0","cdate":1730680255074,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission6384/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428641658,"domain":"ICLR.cc/2025/Conference","replyto":"TTWxMAwS6n","id":"VDBIbrmOtD","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["video generation","diffusion model","efficiency"]},"primary_area":{"value":"generative models"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"We introduce X-PlugVid, a unified framework designed to seamlessly adapt pretrained image-based plug-and-play modules for video diffusion models, facilitating controllable video generation without the need for retraining. This framework leverages a spatial-temporal adapter to effectively bridge the gap between image and video diffusion models. Specifically, we adopt a frozen copy of a large-scale\npretrained image diffusion model (e.g. Stable Diffusion v1.5) as spatial prior. Then we train a spatial-temporal adapter to convert the prior into temporally consistent guidance for video diffusion models (e.g. SVD). To further enhance the effectiveness of image plugins in guiding video models, we introduce a timestep remapping strategy. Recognizing that denoising is an entropic reduction process, this strategy selects priors from later timesteps of the image model, which contain richer information, to be injected into the video models, optimizing the quality and consistency of the generated videos. Comprehensive experimental evaluations of X-PlugVid demonstrate its broad compatibility with diverse operational conditions and different plugins, confirming that leveraging priors from a pretrained diffusion model can minimize redundant training and enable versatile controllable video generation."},"_bibtex":{"value":"@misc{\nran2024xplugvid,\ntitle={X-PlugVid: Versatile Adaptation of Image Plugins for Controllable Video Generation},\nauthor={Lingmin Ran and Chenyang Si and Xudong Lin and Jia-Wei Liu and Rui Zhao and Ziwei Liu and Jussi Keppo and Mike Zheng Shou},\nyear={2024},\nurl={https://openreview.net/forum?id=TTWxMAwS6n}\n}"},"title":{"value":"X-PlugVid: Versatile Adaptation of Image Plugins for Controllable Video Generation"},"pdf":{"value":"/pdf/7fa89672cb73393600912455bb3a8c8c0fbaa0db.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"ran|xplugvid_versatile_adaptation_of_image_plugins_for_controllable_video_generation"},"authorids":{"value":["~Lingmin_Ran1","~Chenyang_Si2","~Xudong_Lin1","~Jia-Wei_Liu1","~Rui_Zhao12","~Ziwei_Liu1","~Jussi_Keppo1","~Mike_Zheng_Shou1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Lingmin Ran","Chenyang Si","Xudong Lin","Jia-Wei Liu","Rui Zhao","Ziwei Liu","Jussi Keppo","Mike Zheng Shou"]}},"version":2},{"content":{"comment":{"value":"Thank you for your follow-up questions and comments. We truly appreciate the time and effort you’ve taken to engage with our work in such detail. Your feedback has been very valuable in helping us improve the manuscript.\n\nBelow, we provide detailed responses to your follow-up queries:\n\n***\n\n> **Q1: Benefit from in-domain dataset**\n\n**Response:**\n\nFirst, as outlined in Lines 294-295, our selection of VideoChatGPT and NextQA is motivated by their question diversity. While VideoChatGPT [1] may share some raw video content with ActivityNetQA [2] (ANQA), the question-answer pairs in VideoChatGPT are entirely new, as stated in their abstract: \"we introduce a new dataset of 100,000 video-instruction pairs used to train Video-ChatGPT.\" This indicates a distinct data distribution between VideoChatGPT (our training dataset) and ActivityNet QA.\n\nSecond, it is important to note that Video Question Answering performance is fundamentally determined by the inference model—in our case, the frozen VILA. Frame-Voyager's architecture precludes direct intervention in the final results, thereby inherently limiting potential shortcut effects caused by the in-domain data.\n\nThird, our experimental validation encompasses multiple \"zero-overlapped\" benchmarks. The main paper presents results on two widely adopted benchmarks: Video-MME and MLVU. During the rebuttal phase, we extend our evaluation to include two additional benchmarks: MVBench and STAR. Notably, we have verified that MVBench has no overlap with ANQA and NextQA (detailed clarification provided in the response below).\n\nIn conclusion, our experimental evidence, derived from four entirely independent benchmarks, demonstrates the effectiveness of Frame-Voyager on out-of-distribution benchmarks while exhibiting minimal susceptibility to in-domain data influence.\n\n[1] Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models, ACL 2024.\n\n[2] ActivityNet-QA: A Dataset for Understanding Complex Web Videos via Question Answering, AAAI 2019.\n\n***\n\n> **Q2: Results on Next-QA**\n\n**Response:**\n\nWe have thoroughly validated our implementation of NextQA and confirmed its correctness. The code will be publicly released to ensure reproducibility.\n\nMeanwhile, we conduct analysis on the cases which are predicted correctly due to the higher number of frames. \n* The distribution of question types in the overall NextQA test set is shown in row 1 of Table 1.\n* The second row shows the percentage of different question types in the samples that are answered correctly due to the higher number of frames. For example, by increasing the number of frames from 8 to 16, there are additional 10 samples being predicted correctly, and there are only 1 sample asking why within the 10 samples, thus the percentage of “why” is 10%.\n\nFrom the results, we can see that increasing the number of frames will be more helpful for “why” questions and temporal reasoning problems asking about “before & after”. It demonstrates that increasing frame sampling particularly enhances performance on questions requiring causal and temporal reasoning.\n\n\n|                      | Why   | Before & After | How   | When  | Count | Location | Other |\n|----------------------|-------|---------|-------|-------|-------|----------|-------|\n| Overall Distribution       | 38.9% | 17.4%   | 13.7% | 13.6% | 3.8%  | 5.6%     | 7.0%  |\n| Correct Sample Ratio Benefiting from 8-frame to 16-frame | 48.0% | 19.1%   | 11.6% | 13.2% | 1.6%  | 2.4%     | 4.2%  |"},"title":{"value":"Response to the Reviewer ULTS (Second Round-1)"}},"tmdate":1732548455030,"tcdate":1732548455030,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission13344/Authors"],"signatures":["ICLR.cc/2025/Conference/Submission13344/Authors"],"forum":"LNL7zKvm7e","number":25,"license":"CC BY 4.0","cdate":1732548455030,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission13344/-/Official_Comment"],"mdate":1732548455030,"domain":"ICLR.cc/2025/Conference","replyto":"piC65CHPii","id":"uSC0UyyVN8","forumContent":{"venue":{"value":"ICLR 2025 Poster"},"TLDR":{"value":"Frame-Voyager queries optimal frame combinations for Video-LLMs based on task-specific queries, outperforming existing SOTA methods and achieving best results in Video Question Answering benchmarks as a plug-and-play solution."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video-LLM","Adaptive Frame Sampling"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Video Large Language Models (Video-LLMs) have made remarkable progress in video understanding tasks. However, they are constrained by the maximum length of input tokens, making it impractical to input entire videos. Existing frame selection approaches, such as uniform frame sampling and text-frame retrieval, fail to account for the information density variations in the videos or the complex instructions in the tasks, leading to sub-optimal performance. In this paper, we propose Frame-Voyager that learns to query informative frame combinations, based on the given textual queries in the task. To train Frame-Voyager, we introduce a new data collection and labeling pipeline, by ranking frame combinations using a pre-trained Video-LLM. Given a video of M frames, we traverse its T-frame combinations, feed them into a Video-LLM, and rank them based on Video-LLM's prediction losses. Using this ranking as supervision, we train Frame-Voyager to query the frame combinations with lower losses. In experiments, we evaluate Frame-Voyager on four Video Question Answering benchmarks by plugging it into two different Video-LLMs. The experimental results demonstrate that Frame-Voyager achieves impressive results in all settings, highlighting its potential as a plug-and-play solution for Video-LLMs."},"_bibtex":{"value":"@inproceedings{\nyu2025framevoyager,\ntitle={Frame-Voyager: Learning to Query Frames for Video Large Language Models},\nauthor={Sicheng Yu and CHENGKAI JIN and Huanyu Wang and Zhenghao Chen and Sheng Jin and ZHONGRONG ZUO and XU XIAOLEI and Zhenbang Sun and Bingni Zhang and Jiawei Wu and Hao Zhang and Qianru Sun},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=LNL7zKvm7e}\n}"},"title":{"value":"Frame-Voyager: Learning to Query Frames for Video Large Language Models"},"pdf":{"value":"/pdf/154af3dfa9ab115b4d4a01fa334c6ab45c7ad3af.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"yu|framevoyager_learning_to_query_frames_for_video_large_language_models"},"authorids":{"value":["~Sicheng_Yu3","~CHENGKAI_JIN2","~Huanyu_Wang2","~Zhenghao_Chen2","~Sheng_Jin3","~ZHONGRONG_ZUO1","~XU_XIAOLEI1","~Zhenbang_Sun1","~Bingni_Zhang1","~Jiawei_Wu9","~Hao_Zhang3","~Qianru_Sun2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Sicheng Yu","CHENGKAI JIN","Huanyu Wang","Zhenghao Chen","Sheng Jin","ZHONGRONG ZUO","XU XIAOLEI","Zhenbang Sun","Bingni Zhang","Jiawei Wu","Hao Zhang","Qianru Sun"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a novel video interpretation method based on perturbation. It considers the temporal dimension separately through dual perturbation, and finds a spatio-temporal visual explanation by conducting the TIS-aware spatial analysis for each frame based on extremal perturbation."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"Please address the questions in the Weaknesses."},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"- This method improves the interpretation effect through the dual strategy of temporal and spatial perturbations, and performs particularly well in identifying key frames in dynamic videos.\n\n- The experimental results show the effectiveness of TIEM in video interpretation."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Although TIEM has made improvements in explaining the temporal dynamics of video models, its core idea is based on existing extreme perturbation methods and does not show fundamental innovation in methodology. This hardly meets ICLR’s high standards for methodological innovation.\n\n- TIEM's experiments mainly focus on action recognition (UCF101-24 dataset) and artificially synthesized white-box models. Although the results are significant on these tasks, the generalization ability and applicability to a wide range of real-world application scenarios (such as medical video analysis, autonomous driving videos, etc.) have not been deeply explored.\n\n- The baseline methods compared in the paper are not novel enough and lack comparison with the latest methods, such as AOSA in related work and other recently published methods."}},"nonreaders":[],"tmdate":1731427863524,"tcdate":1730642206726,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission3509/Reviewer_X9Md"],"signatures":["ICLR.cc/2025/Conference/Submission3509/Reviewer_X9Md"],"forum":"TEjXRrhqtJ","number":2,"license":"CC BY 4.0","cdate":1730642206726,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission3509/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427863524,"domain":"ICLR.cc/2025/Conference","replyto":"TEjXRrhqtJ","id":"sWl46hfHAB","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"TLDR":{"value":"This paper proposes TIEM, a novel video interpretation method that enhances explainability through dual perturbation that evaluates temporal importance across frames and generates spatio-temporal masks explicitly using this importance."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["XAI","visual explanation","extremal mask","dual perturbation","video prediction"]},"supplementary_material":{"value":"/attachment/4d590237bd717ac596a69e5687de1fab7bf231be.zip"},"primary_area":{"value":"interpretability and explainable AI"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Explaining video data predictions is challenging due to the complex spatio-temporal information in videos. In particular, the existing perturbation-based methods for video interpretation often fail to consider different temporal contexts, making them ineffective for dynamic videos where the important regions change rapidly or appear ephemerally across frames. To address this, we propose a novel video interpretation method, time importance score-aware extremal perturbation masks (TIEM), that enhances explainability by focusing on temporal dynamics in videos. TIEM exploits a dual perturbation process: first, it evaluates temporal importance across frames via temporal perturbation and then generates spatio-temporal extremal perturbation masks using the temporal importance explicitly. Our experimental results demonstrate that TIEM resolves the key challenges of the existing methods, providing more precise explanations across the time domain in synthetic white-box models and black-box models for real-world videos."},"_bibtex":{"value":"@misc{\nje-gal2024tiem,\ntitle={{TIEM}: Enhancing Explanation of Video Prediction via Temporal Dynamics-Focused Dual Perturbation},\nauthor={Hong Je-Gal and Hyun-Suk Lee},\nyear={2024},\nurl={https://openreview.net/forum?id=TEjXRrhqtJ}\n}"},"title":{"value":"TIEM: Enhancing Explanation of Video Prediction via Temporal Dynamics-Focused Dual Perturbation"},"pdf":{"value":"/pdf/ba9ded8ac2b175ad9e9b09d701a83d91e7c69278.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"jegal|tiem_enhancing_explanation_of_video_prediction_via_temporal_dynamicsfocused_dual_perturbation"},"authorids":{"value":["~Hong_Je-Gal1","~Hyun-Suk_Lee1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Hong Je-Gal","Hyun-Suk Lee"]}},"version":2},{"content":{"summary":{"value":"The paper proposes a post-training compression technique that combines recent advancements in quantization and compression. It extends the OPTQ framework by incorporating a rate-distortion trade-off, achieved through the addition of a bit-cost function inspired by NNCodec’s entropy model. The resulting OPTQ-RD method effectively balances compression strength and inference speed, achieving high compression ratios with minimal performance degradation across various CNNs."},"soundness":{"value":2},"confidence":{"value":5},"questions":{"value":"How does the choice of entropy model impact compression outcomes, and could other entropy models be seamlessly integrated into the method?\n\nHow does the method handle outliers in model weights, especially in architectures with significant variations in weight distributions? While this may seem out of scope, it could provide insight into the method's effectiveness in quantizing activations.\n\nIf the method is compatible with various quantization schemes, could you evaluate the impact of different schemes on large language models (LLMs)? A model with 1 billion parameters should be sufficient for this validation."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":1},"strengths":{"value":"The method can be applied without modifying network architectures, broadening its utility for different models.\n\nThe method’s ability to operate in a post-training setting with minimal calibration data makes it practical for real-world deployment.\n\nThe evaluation demonstrates that the proposed method consistently outperforms the baselines in terms of weight compression while preserving accuracy on par with them."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"## Claims:\n\nThe paper does not reference previous methods that address the same problem [1,2].\n\nThe authors claim that OPTQ-RD achieves fast inference time; however, as I understand, only the weights are quantized, while activations remain in the floating point, typically resulting in a minimal reduction in inference time. This approach may reduce memory usage or storage, particularly in the context of compression and quantization for large language models (LLMs). However, it has a limited impact on convolutional neural networks (CNNs), especially those used in the experiments.\n\n\n## Experiments:\n\nThe experimental setting is quite limited. VGG, which is not commonly used in practice, is known to have sparse weights, making it difficult to generalize the findings to more widely adopted network architectures.\n\n\nAlthough bits per weight appears promising, Figure 1 and Table 1 provide little insight into its relevance to practical metrics, such as latency or energy consumption. Additionally, the plots suggest that, in intermediate cases, there is only a narrow range where the proposed method outperforms the vanilla baseline of OPTQ+DeepCABAC.\n\n\n## Writing\n\nThe authors begin the abstract with a statement on pre-trained large models and discuss recent work on LLM quantization, yet these topics appear unrelated to the core of this study. Conducting evaluations only on CNNs without addressing LLMs or even ViTs is entirely valid; however, if they are not central to the paper, it raises the question of why these topics are mentioned at all.\n\n### Refrences\n\n[1]CAT: Compression-Aware Training for bandwidth reduction, Baskin et al., JMLR 2021\n\n[2] Feature Map Transform Coding for Energy-Efficient CNN Inference, Chmiel et al. IJCNN 2020"}},"nonreaders":[],"tmdate":1732383144487,"tcdate":1730745100398,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission11036/Reviewer_awtQ"],"signatures":["ICLR.cc/2025/Conference/Submission11036/Reviewer_awtQ"],"forum":"LnKDcqOfgy","number":4,"license":"CC BY 4.0","cdate":1730745100398,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission11036/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732383144487,"domain":"ICLR.cc/2025/Conference","replyto":"LnKDcqOfgy","id":"VknpV2uHKh","forumContent":{"TLDR":{"value":"We propose a compression method for pre-trained neural networks that combines quantization and entropy based neural network compression."},"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["quantization","model compression","rate-distortion theory","compression"]},"supplementary_material":{"value":"/attachment/567abe088344c115009f505edc126ad762338bab.zip"},"primary_area":{"value":"other topics in machine learning (i.e., none of the above)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"The proliferation of large pre-trained neural networks has recently revived research in both quantization of network weights (for faster inference), and in their\ncompression (to reduce file sizes). However, there has so far been little idea transfer between the two lines of research. In this paper, we combine techniques from\nquantization and compression to propose an efficient and highly effective post-training compression method for large neural networks. Our method extends the\nrecently published quantization method OPTQ (Frantar et al., 2023) with a tunable\nrate/distortion trade-off by introducing a cost per bit into OPTQ's rounding\noperation. Crucially, we estimate the bit rate based on the predictive model used\nin the state-of-the-art neural network compression method NNCodec (Becking\net al., 2023). In our experiments with several standard pre-trained networks from\nthe computer vision community, our method leads to significantly (up to 2.7x)\nsmaller file sizes than NNCodec at equal model performance, generally compressing to less than half a bit per network weight and implicitly pruning insignificant weights.\nAdditionally, and in contrast to NNcodec, our method offers the same opportunities for inference speed-ups as OPTQ. By proving that file size and inference\ncost can be reduced simultaneously, we hope that our contribution shows a path\ntowards deploying large neural networks on end-user devices, alleviating privacy\nconcerns, regulatory constraints, and dependency on large service providers."},"_bibtex":{"value":"@misc{\nconzelmann2025ratedistortion,\ntitle={Rate/Distortion Constrained Model Quantization for Efficient Storage and Inference},\nauthor={Alexander Conzelmann and Robert Bamler},\nyear={2025},\nurl={https://openreview.net/forum?id=LnKDcqOfgy}\n}"},"title":{"value":"Rate/Distortion Constrained Model Quantization for Efficient Storage and Inference"},"pdf":{"value":"/pdf/590933963b6787647dabedc624b705fbd9b7ebf2.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"conzelmann|ratedistortion_constrained_model_quantization_for_efficient_storage_and_inference"},"authorids":{"value":["~Alexander_Conzelmann1","~Robert_Bamler1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Alexander Conzelmann","Robert Bamler"]}},"version":2},{"content":{"summary":{"value":"CHIMERA is a visual benchmark suite designed to evaluate vision-language models (VLMs) on diagram comprehension. It contains 6,000 training and 1,500 test diagrams. The authors identify three shortcut behaviors commonly exhibited by VLMs: (1) visual information memorization, (2) knowledge-recall shortcut, and (3) Clever-Hans shortcut. The latter two primarily arise from the models’ reliance on linguistic cues rather than genuine visual understanding. To address this, each CHIMERA entry integrates three modalities: visual, semantic, and textual. This enables a comprehensive assessment of model behavior. The benchmark defines four task levels: entity recognition, relation understanding, knowledge grounding, and visual reasoning. The authors evaluate across 15 VLMs revealing that much of their performance can be attributed to language bias rather than true multimodal reasoning. In particular, the Clever-Hans shortcut experiment exposes cases where models achieve high accuracy even when the visual input is omitted entirely."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"•\tThe bottom part of Figure 3 requires more clarity. Also at first glance, the Human Evaluation component is not clear.\n\n•\tIn Table 1, Annotator C’s assessments show high percentages for visual dependency – partially dependent and triple completeness – marginally insufficient. What is the reason for this? The difference with other annotators is not minor.\n\n•\tThe authors present model-wise performance in Table 5 of the Supplementary. I recommend that they include a concise visual summary of this information, or at least aggregate the results by model family within the main paper."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"•\tThe authors identify three prominent shortcut behaviors and base their analysis on these.\n\n•\tThe findings from the Clever-Hans shortcut experiment are interesting, demonstrating that models do not utilize information from images for visual questions.\n\n•\tThe construction of the test suite is well explained: the authors describe starting with Wiki Web2M and ultimately filtering 7,500 instances for CHIMERA through a semi-autonomous process.\n\n•\tThe paper is well-structured, clearly written, and includes appropriate figures and tables."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"•\tMost models achieve high scores (above 80%) across the majority of tasks. This raises the question of whether the primary purpose of the benchmark is simply to show that models rely more heavily on textual modality. Since the questions themselves do not appear to pose substantial challenges, the benchmark’s diagnostic value seems limited.\n\n•\tThe authors claim that it is surprising that models perform better on the visual modality than on the semantic modality (lines 334-338). However, given that most training data heavily represent visual and textual modalities, with comparatively limited exposure to semantic diagrams, such results are expected rather than surprising.\n\n•\tLLaMA3.2 and Gemini were used to construct the dataset, while the evaluation includes models from the LLaMA3.2 and Gemma3 families (also developed by Google). This overlap introduces a potential source of bias.\n\n•\tThe authors show in Figure 4 and Table 2 that there is a 6–8% gap between the textual modality and the visual or semantic modalities, concluding that models perform best on text (line 333). However, I believe this conclusion is influenced by a few weaker models. After recalculating Table 2 using the data from Table 5 in the Supplementary, excluding LLaMA3.2-11B, LLaVA1.6-7B, LLaVA1.6-13B, and BLIP3-4B, the observed gap becomes much smaller, suggesting that the overall trend may not be as pronounced as reported.\n\nThe recalculated average scores are: Visual - 90.7, Semantic - 89.3, and Linguistic - 93.3. As observed, the gaps become much less prominent, under 3% between the Visual and Textual modalities. This raises the question of whether the authors’ reported findings are broadly generalizable or primarily driven by a few underperforming models."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762924091357,"tcdate":1761873564080,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission13473/Reviewer_RvGE"],"signatures":["ICLR.cc/2026/Conference/Submission13473/Reviewer_RvGE"],"forum":"q3eB3PhtqD","number":2,"license":"CC BY 4.0","cdate":1761873564080,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission13473/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762924091357,"domain":"ICLR.cc/2026/Conference","replyto":"q3eB3PhtqD","id":"z3BItQWg7v","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["vision-language model","diagram understanding","multimodal reasoning","visual question answering","shortcut learning","knowledge grounding","benchmark dataset","multimodal evaluation"]},"primary_area":{"value":"interpretability and explainable AI"},"abstract":{"value":"Diagrams convey symbolic information in a visual format rather than a linear stream of words, making them especially challenging for AI models to process. While recent evaluations suggest that vision-language models (VLMs) perform well on diagram-related benchmarks, their reliance on knowledge, reasoning, or modality shortcuts raises concerns about whether they genuinely understand and reason over diagrams.\nTo address this gap, we introduce Chimera, a comprehensive test suite comprising 7,500 high-quality diagrams sourced from Wikipedia; each diagram is annotated with its symbolic content represented by semantic triples along with multi-level questions designed to assess four fundamental aspects of diagram comprehension: entity recognition, relation understanding, knowledge grounding, and visual reasoning.\nWe use Chimera to measure the presence of three types of shortcuts in visual question answering: \n(1) the visual-memorization shortcut, where VLMs rely on memorized visual patterns;\n(2) the knowledge-recall shortcut, where models leverage memorized factual knowledge instead of interpreting the diagram; and\n(3) the Clever-Hans shortcut, where models exploit superficial language patterns or priors without true comprehension. We evaluate 15 open-source VLMs from 7 model families on Chimera and find that their seemingly strong performance largely stems from shortcut behaviors: visual-memorization shortcuts have slight impact, knowledge-recall shortcuts play a moderate role, and Clever-Hans shortcuts contribute significantly.\nThese findings expose critical limitations in current VLMs and underscore the need for more robust evaluation protocols that benchmark genuine comprehension of complex visual inputs (e.g., diagrams) rather than question-answering shortcuts."},"_bibtex":{"value":"@misc{\nchi2026chimera,\ntitle={Chimera: Diagnosing Shortcut Learning in Visual-Language Understanding},\nauthor={Ziheng Chi and Yifan Hou and Chenxi Pang and Shaobo Cui and Mubashara Akhtar and Mrinmaya Sachan},\nyear={2026},\nurl={https://openreview.net/forum?id=q3eB3PhtqD}\n}"},"title":{"value":"Chimera: Diagnosing Shortcut Learning in Visual-Language Understanding"},"pdf":{"value":"/pdf/f92e46808334228be1f33c8c27812a84e924f889.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"chi|chimera_diagnosing_shortcut_learning_in_visuallanguage_understanding"},"authorids":{"value":["~Ziheng_Chi1","~Yifan_Hou1","~Chenxi_Pang1","~Shaobo_Cui1","~Mubashara_Akhtar1","~Mrinmaya_Sachan3"]},"authors":{"value":["Ziheng Chi","Yifan Hou","Chenxi Pang","Shaobo Cui","Mubashara Akhtar","Mrinmaya Sachan"]}},"version":2},{"content":{"summary":{"value":"This paper presents a diffusion-based video generation framework designed for generating synthetic video data for training robot policies. The idea is to use reference images of backgrounds and objects and generate synthetic versions of real demonstration datasets for better generalizability of the manipulation policies. While the idea isn't new, this paper focusses on multi-view video generation (rather than just one view) with cross-view geometric consistency while providing fine-grained control over object and background appearance. The framework is evaluated in two types of experiments: video generation realism and its effect on the manipulation policy."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1. How different can the generated scenes be from the original demonstrations before the policy performance degrades? For instance, can RoboTransfer insert novel object geometries or rearranged spatial layouts, or is it restricted to appearance-level (texture/color) variations?\n\n2. How does the model compare with recent video-diffusion or policy-learning baselines such as Cosmos-Transfer or VLA-based models (e.g., Pi0, OpenVLA)? Are there any quantitive performance metrics (besides just one qualitative example in Fig. 15)?\n\n3. Since the method relies predicted metric depth and surface normals, how sensitive is performance to inaccuracies in these estimates?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"The paper aims to bridge the sim-to-real gap and overcome the limited-data challenge for real-world robot learning. The proposed geometry-aware diffusion pipeline is indeed a good contribution. Specifically, the explicit disentanglement of geometry and appearance for robotic video generation allows for finer control generation and with the multi-view generation can be particularly useful for robot learning. Often, robot learning policies use multi-view data (e.g., an environment camera and a wrist camera) but single view video generation may lead to un-aligned multi-view policies. This paper directly addresses that challenge."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"While the overall idea is promising, the current version of the approach has several limitations.\n\n1. The framework does not generate new motions or actions data. As far as I understand, it simply replaces the original videos with the generated ones while keeping the robot trajectory the same. Thus, the \"synthetic\" demonstrations still have actual controls, and just synthetic videos. Therefore, it is unclear to me how far from the real demos can it generalize. The changes in the background are such that they don't really affect the actual trajectory and the novel objects inserted have the same shape with just different colors. This to be severely limits the generalization. \n\n2. On the empirical side, there were no comparisons with baselines in the main body of the paper. For video generation, the results only show ablations of the various losses constraint and on the policy learning side we see the same policy but with different levels of real/synthetic data. One would have expected to see comparisons with other baseline video generation and policy learning methods (including something like Pi0 or other VLAs that claim to be more inherently generalizable).\n\n3. The experiments with multi-view generation are limited to fixed camera views and it is not clear how this would generalize to a new setup. For example, would we have to retrain the network for a new view?"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762919104095,"tcdate":1762167925140,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission6846/Reviewer_G2VP"],"signatures":["ICLR.cc/2026/Conference/Submission6846/Reviewer_G2VP"],"forum":"WCY0l6z3Rm","number":4,"license":"CC BY 4.0","cdate":1762167925140,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission6846/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762919104095,"domain":"ICLR.cc/2026/Conference","replyto":"WCY0l6z3Rm","id":"RNMFN4STB8","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"TLDR":{"value":"RoboTransfer is a video synthesis framework for robotic manipulation that ensures multi-view consistency while enabling fine-grained, disentangled control."},"keywords":{"value":["Robotic","Video Generation","Imitation Learning","Data Augmentation"]},"supplementary_material":{"value":"/attachment/03e5750993071847358785f0e74d9ee5eef06a56.zip"},"primary_area":{"value":"applications to robotics, autonomy, planning"},"abstract":{"value":"The goal of general-purpose robotics is to create agents that can seamlessly adapt to and operate in diverse, unstructured human environments. Imitation learning has become a key paradigm for robotic manipulation, yet collecting large-scale and diverse demonstrations is prohibitively expensive. Simulators provide a cost-effective alternative, but the sim-to-real gap remains a major obstacle to scalability. We present RoboTransfer, a diffusion-based video generation framework for synthesizing robotic data. By leveraging cross-view feature interactions and globally consistent 3D geometry, RoboTransfer achieves multi-view geometric consistency while enabling fine-grained control over scene elements, including background editing and object replacement. Experiments show that RoboTransfer generates videos with improved geometric consistency and visual fidelity, and that policies trained on this data generalize better to novel, unseen scenarios. The code and datasets will be released upon acceptance."},"_bibtex":{"value":"@misc{\nliu.liu2025robotransfer,\ntitle={RoboTransfer: Geometry-Consistent Video Diffusion for Robotic Visual Policy Transfer},\nauthor={Liu.Liu and Xiaofeng Wang and Guosheng Zhao and Keyu Li and Wenkang Qin and Jiagang Zhu and Jiaxiong Qiu and Zheng Zhu and Guan Huang and Zhizhong Su},\nyear={2025},\nurl={https://openreview.net/forum?id=WCY0l6z3Rm}\n}"},"title":{"value":"RoboTransfer: Geometry-Consistent Video Diffusion for Robotic Visual Policy Transfer"},"pdf":{"value":"/pdf/553aee7ca3fb147a6558f072a6eadd177fbb8604.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"liuliu|robotransfer_geometryconsistent_video_diffusion_for_robotic_visual_policy_transfer"},"authorids":{"value":["~Liu.Liu19","~Xiaofeng_Wang5","~Guosheng_Zhao1","~Keyu_Li2","~Wenkang_Qin1","~Jiagang_Zhu3","~Jiaxiong_Qiu2","~Zheng_Zhu1","~Guan_Huang1","~Zhizhong_Su3"]},"authors":{"value":["Liu.Liu","Xiaofeng Wang","Guosheng Zhao","Keyu Li","Wenkang Qin","Jiagang Zhu","Jiaxiong Qiu","Zheng Zhu","Guan Huang","Zhizhong Su"]}},"version":2},{"content":{"venue":{"value":"IEEE Trans. Circuits Syst. Video Technol. 2022"},"pdf":{"value":"https://ieeexplore.ieee.org/iel7/76/9678106/09351972.pdf"},"venueid":{"value":"dblp.org/journals/TCSV/2022"},"paperhash":{"value":"zheng|stacked_multimodal_attention_network_for_contextaware_video_captioning"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Yi_Zheng:","https://dblp.org/search/pid/api?q=author:Yuejie_Zhang:","https://dblp.org/search/pid/api?q=author:Rui_Feng_0001:","~Tao_Zhang11","https://dblp.org/search/pid/api?q=author:Weiguo_Fan:"]},"html":{"value":"https://doi.org/10.1109/TCSVT.2021.3058626"},"_bibtex":{"value":"@article{DBLP:journals/tcsv/ZhengZFZF22,\n  author={Yi Zheng and Yuejie Zhang and Rui Feng and Tao Zhang and Weiguo Fan},\n  title={Stacked Multimodal Attention Network for Context-Aware Video Captioning},\n  year={2022},\n  cdate={1640995200000},\n  journal={IEEE Trans. Circuits Syst. Video Technol.},\n  volume={32},\n  number={1},\n  pages={31-42},\n  url={https://doi.org/10.1109/TCSVT.2021.3058626}\n}\n"},"abstract":{"value":"Recent neural models for video captioning usually employ an attention-based encoder-decoder framework. However, current approaches mainly attend to the motion features and object features of the video when generating the caption, but ignore the potential but useful historical information. Besides, exposure bias and vanishing gradients problems always exist in current caption generation models. In this paper, we propose a novel video captioning framework, named Stacked Multimodal Attention Network (SMAN). It adopts additional visual and textual historical information during caption generation as context features, employs a stacked architecture to process different features gradually, and utilizes the Reinforcement Learning method and coarse-to-fine training strategy to further improve the generated results. Both quantitative and qualitative experiments on the benchmark datasets of MSVD and MSR-VTT show the effectiveness and feasibility of our framework. The codes are available on https://github.com/zhengyi123456/SMAN."},"title":{"value":"Stacked Multimodal Attention Network for Context-Aware Video Captioning"},"authors":{"value":["Yi Zheng","Yuejie Zhang","Rui Feng","Tao Zhang","Weiguo Fan"]}},"tmdate":1744361072489,"pdate":1640995200000,"tcdate":1744361054700,"writers":["~"],"signatures":["~Tao_Zhang11"],"forum":"jW58Vi8ZFU","license":"CC BY-SA 4.0","number":390338,"cdate":1640995200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1744361072489,"domain":"DBLP.org","id":"jW58Vi8ZFU","version":2},{"content":{"summary":{"value":"This paper focuses on video question answering and proposes an instruction-tuning method called Agent-of-Thoughts Distillation (AoTD) for Video-LLMs. The authors first employ an agentic system to decompose complex video questions into sub-tasks which are iteratively solved by cutting-edge visual models to finally answer the video questions. The execution traces of this agentic system are then converted into chain-of-thoughts (CoTs), which are used for tuning Video-LLMs to generate answers with clear reasons and explainability. The authors conduct extensive experiments to show the effectiveness of AoTD in getting more accurate results and rationales on both multiple-choice video questions and open-ended video questions."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. What is the average length of the tested videos? Since the Video-LLMs receives a fixed number of frames, what if the key frames are not sampled for the model input because of the long video length? What will the rationale output be like in these cases?\n2. Will hallucination happen in the generated rationale? If so, how to avoid hallucination?\n3. In line 200, It seems that the CoTs of the agentic system follows a fixed order of tool calls: temporal grounding, object detection, and then question answering. Why using a fixed tool-call sequence instead of letting the LLM agent to decide the tool call sequence by its own? \n4. How will the tuned Video-LLMs answer the question like: “what is the main idea of the video?”"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The paper proposes a novel attempt to use CoTs to finetune video-LLMs for generating better answers with reasons, which is different from agent tuning where CoTs are traditionally used for finetuning LLMs for generating better planning and tool calls.\n2. The motivation of the paper is clear: using CoTs to finetune Video-LLMs in order to enhance the reasoning, spatial grounding and temporal grounding ability of the end-to-end models. The paper is well written. \n3. Extensive experiments and ablation study are conducted to illustrate the effectiveness of AoTD in boosting performance as well as endowing Video-LLMs with multi-step reasoning, temporal grounding and spatial grounding abilities."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The paper does not provide qualitative examples of the answers and rationales of the Video-LLMs tuned by AoTD.\n2. The generalization ability of AoTD on diverse question types remains to be discussed, since not all the video questions can be decomposed into the sub-tasks. For example, “explain the humor in this video”.\n3. Only perform object detection on video frames is not sufficient enough to accurately understand the video. Object tracking and object re-identification should be taken into consideration for temporal consistency."}},"nonreaders":[],"tmdate":1731428157101,"tcdate":1730944136782,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission13071/Reviewer_1h3R"],"signatures":["ICLR.cc/2025/Conference/Submission13071/Reviewer_1h3R"],"forum":"mMfDfJ8JFJ","number":4,"license":"CC BY 4.0","cdate":1730944136782,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission13071/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428157101,"domain":"ICLR.cc/2025/Conference","replyto":"mMfDfJ8JFJ","id":"JqjSDPmCpv","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video QA","Visual Understanding","MLLM","Agent-based Video Analysis"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"This paper tackles the problem of video question answering (VideoQA), \na task that often requires multi-step reasoning and a profound understanding of spatial-temporal dynamics. While large generative video-language models perform well on benchmarks, they often lack explainability and spatial-temporal grounding. \nIn this paper, we propose **A**gent-**o**f-**T**houghts **D**istillation (**AoTD**), a method that enhances generative models by incorporating automatically generated Chain-of-Thoughts (CoTs) into the instruction-tuning process. Specifically, we leverage an agent-based system to decompose complex questions into sub-tasks, and address them with specialized vision models, the intermediate results are then treated as reasoning chains. \nWe also introduce a verification mechanism using a large language model (LLM) to ensure the reliability of generated CoTs. Extensive experiments demonstrate that AoTD improves the performance on multiple-choice and open-ended benchmarks."},"_bibtex":{"value":"@misc{\nshi2024unlocking,\ntitle={Unlocking Video-{LLM} via Agent-of-Thoughts Distillation},\nauthor={Yudi Shi and Shangzhe Di and Qirui Chen and Weidi Xie},\nyear={2024},\nurl={https://openreview.net/forum?id=mMfDfJ8JFJ}\n}"},"title":{"value":"Unlocking Video-LLM via Agent-of-Thoughts Distillation"},"pdf":{"value":"/pdf/76ef8d449075a78ef1b9fccd577b171205566e35.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"shi|unlocking_videollm_via_agentofthoughts_distillation"},"authorids":{"value":["~Yudi_Shi1","~Shangzhe_Di1","~Qirui_Chen1","~Weidi_Xie3"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Yudi Shi","Shangzhe Di","Qirui Chen","Weidi Xie"]}},"version":2},{"content":{"summary":{"value":"This paper proposes Pyramid Attention Broadcast (PAB), a method for DiT-based video generation that leverages a pyramid-shaped broadcast design. This design is inspired by two key observations: (1) the attention output differences are minimal in the middle 70% of the diffusion steps, and (2) spatial, temporal, and cross-attention differences decrease in a hierarchical, pyramid-like manner. Based on these insights, PAB sets different broadcast ranges for each attention type. The method is further extended to distributed settings, supporting multi-GPU parallelism (e.g., 8-GPU, 16-GPU) to further reduce generation latency. Experiments on VBench and WebVid using three types of DiT models demonstrate that PAB effectively reduces latency and accelerates video generation."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"The broadcast sequence parallelism in this work relies on DSP, so it’s unclear how much of the latency reduction is directly attributable to PAB alone. In Table 4 or in the main text, it would be helpful to further clarify PAB’s independent contribution to latency improvements."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"The method’s effectiveness is demonstrated on multi-GPU setups, achieving real-time video generation (e.g., within 2 seconds) on an 8-card H100 configuration. This efficiency is also validated across multiple open-source Video DiT models."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. In Figure 4, the quantitative analysis of attention differences uses 30 inference steps based on Open-So ra, but the visualized attention differences cover 50 steps, creating inconsistency. Is this visualization based on Latte? Additionally, it’s unclear whether the attention feature difference plot is derived from an entire dataset analysis or a single video generation process.\n\n2. Some comparisons appear to be unfair. For example, in Figure 7, the original generation using a single device is compared to an 8-device setup, leading to a claimed 10.5x speedup. This comparison may be misleading and should be clarified.\n\n3. Key implementation details should be presented in the main paper. For instance, the most crucial setting, **PAB_246**, which defines the spatial, temporal, and cross-attention broadcast ranges, should be clearly explained rather than shown in supplementary materials.\n\n4. Broadcasting attention is not entirely new. For instance, TGATE, an image generation acceleration approach, also uses iterative caching and reusing of self-attention (SA) for acceleration. This reduces the novelty of the proposed approach.\n\n5. The qualitative results focus primarily on natural, static scenes. In real-world applications, generating videos with complex, dynamic actions, such as those involving people or animals, is more impactful. The qualitative evaluation scope is therefore too limited to showcase the method’s full potential in diverse scenarios."}},"nonreaders":[],"tmdate":1731428047839,"tcdate":1730633615641,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission4140/Reviewer_PKxB"],"signatures":["ICLR.cc/2025/Conference/Submission4140/Reviewer_PKxB"],"forum":"hDBrQ4DApF","number":3,"license":"CC BY 4.0","cdate":1730633615641,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission4140/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428047839,"domain":"ICLR.cc/2025/Conference","replyto":"hDBrQ4DApF","id":"pVitDPevzL","forumContent":{"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Diffusion Acceleration","DiT","Video Generation","Efficient","Real-Time","Parallelism","Sequence Parallelism"]},"supplementary_material":{"value":"/attachment/b75653ca8e723a551c7489bf441ab3253e9325f8.zip"},"primary_area":{"value":"generative models"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"We present Pyramid Attention Broadcast (PAB), a real-time, high quality and training-free approach for DiT-based video generation. Our method is founded on the observation that attention difference in the diffusion process exhibits a U-shaped pattern, indicating significant redundancy. We mitigate this by broadcasting attention outputs to subsequent steps in a pyramid style. It applies different broadcast strategies to each attention based on their variance for best efficiency. We further introduce broadcast sequence parallel for more efficient distributed inference. PAB demonstrates up to 10.5x speedup across three models compared to baselines, achieving real-time generation for up to 720p videos. We anticipate that our simple yet effective method will serve as a robust baseline and facilitate future research and application for video generation."},"_bibtex":{"value":"@inproceedings{\nzhao2025realtime,\ntitle={Real-Time Video Generation with Pyramid Attention Broadcast},\nauthor={Xuanlei Zhao and Xiaolong Jin and Kai Wang and Yang You},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=hDBrQ4DApF}\n}"},"title":{"value":"Real-Time Video Generation with Pyramid Attention Broadcast"},"pdf":{"value":"/pdf/0e3e0adc0fd4d73344f909e993ed31cad907d950.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"zhao|realtime_video_generation_with_pyramid_attention_broadcast"},"authorids":{"value":["~Xuanlei_Zhao1","~Xiaolong_Jin2","~Kai_Wang8","~Yang_You1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Xuanlei Zhao","Xiaolong Jin","Kai Wang","Yang You"]}},"version":2},{"content":{"venue":{"value":"CoRR 2026"},"pdf":{"value":"https://arxiv.org/pdf/2605.19397v1"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"huang|perceptionaware_video_semantic_communication"},"html":{"value":"https://doi.org/10.48550/arXiv.2605.19397"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2605-19397,\n  publtype={informal},\n  author={Yinhuan Huang and Zhijin Qin},\n  title={Perception-Aware Video Semantic Communication},\n  year={2026},\n  month={May},\n  cdate={1777593600000},\n  journal={CoRR},\n  volume={abs/2605.19397},\n  url={https://doi.org/10.48550/arXiv.2605.19397}\n}\n"},"abstract":{"value":"Ultra-high-resolution streaming and emerging immersive services are driving rapidly increasing wireless video traffic. However, perceptually pleasing video transmission over bandwidth-limited and latency-constrained wireless links remains challenging for conventional separated source-channel systems, which primarily target bit-level reliability and often suffer performance degradation under short-blocklength transmission. In addition, pixel-level distortion optimization does not necessarily align with human perception, while existing learned video codecs may incur high complexity and raise deployment issues. This paper proposes PVSC, a perception-aware video semantic communication framework for real-time wireless video transmission. PVSC eliminates explicit motion-vector transmission and exploits spatio-temporal feature coding to generate compact and channel-robust symbol streams. It also specifies side-information formatting, reference-buffer management, and lightweight rate control, enabling stable receiver-side reconstruction and bandwidth-adaptive inference with a single model. Extensive experiments demonstrate that PVSC achieves superior performance across diverse datasets, resolutions, GOP configurations, and channel conditions. Compared with the engineered ``VTM + 5G LDPC'' baseline, PVSC saves up to about 75% and 87% bandwidth at comparable LPIPS and DISTS, respectively, while enabling real-time inference on a single NVIDIA RTX 4090 GPU."},"title":{"value":"Perception-Aware Video Semantic Communication"},"authors":{"value":[{"fullname":"Yinhuan Huang","username":"~Yinhuan_Huang1"},{"fullname":"Zhijin Qin","username":"~Zhijin_Qin1"}]}},"tmdate":1784537718330,"pdate":1798675200000,"externalIds":["dblp:journals/corr/abs-2605-19397"],"tcdate":1784213787894,"writers":["~","OpenReview.net/Public_Article/DBLP.org","OpenReview.net/Support"],"signatures":["~Yinhuan_Huang1"],"forum":"xXMtqcLDHX","license":"CC BY-SA 4.0","number":60021,"cdate":1777593600000,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/DBLP.org/-/Record","OpenReview.net/Public_Article/-/Edit","OpenReview.net/Public_Article/-/Authorship_Claim"],"mdate":1784537718330,"domain":"OpenReview.net/Public_Article","id":"xXMtqcLDHX","version":2},{"content":{"summary":{"value":"This paper proposes a two-stage text-to-video generation framework, consisting of video content planning and grounded multi-scene video generation. The first module employs a large language model (LLM), such as GPT-4, to generate a video plan. The second module, trained with image-level layout annotations, generates a consistent multi-scene video given the video plan. The authors conduct various evaluations to demonstrate the effectiveness of their work."},"presentation":{"value":"2 fair"},"contribution":{"value":"3 good"},"soundness":{"value":"2 fair"},"strengths":{"value":"1. The proposed framework achieves high efficiency by not requiring video training data and maintaining good results with 87% of total parameters fixed.\n2. The authors develop several novel evaluation methods that provide solid comparisons between the proposed framework and previous works.\n3. The framework uses both high-level and low-level conditioning to enable fine-grained control over generated videos.\n4. Intuitively and effectively, the framework uses shared features for the same subject across different scenes to ensure multi-scene consistency."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"### A. Main paper\n1. This paper is somewhat too abstract throughout. The introduction is adequate, but I would expect to see more technical content in the following sections. For example, the loss function for the proposed image fine-tuning is not provided. Additionally, there is little evidence to support the correctness of the proposed methods beyond empirical results.\n2. The introduction to the datasets is limited. Some datasets provide fine-grained descriptions for each scene, while others only provide a single sentence for an entire video. Furthermore, the authors customize the Pororo-SV dataset by replacing character names with pronouns, but they do not justify this procedure. The lack of a clear explanation of the datasets makes it difficult to understand the task, such as the inputs and outputs for training and testing, and whether a large language model (LLM) is used.\n3. An ablation study should be conducted to demonstrate the effectiveness of the LLM. Additionally, if the LLM was used to refine prompts, these prompts should also be given to ModelScopeT2V to enable a fair comparison and provide readers with better insights.\n4. Since the consistency should be maintained regardless of the temporal distance between scenes, the authors should consider using the variance of CLIP features of all scenes instead of the average of similarities across adjacent scene pairs.\n5. The human evaluation does not have enough participants to provide reliable results.\n\n### B. Qualitative results\n1. The objects not exactly follow the bounding boxes.\n2. The \"pushing object\" video examples appear to show camera movement rather than object movement."},"confidence":{"value":"4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work."},"questions":{"value":"1. How do you replace the original animation characters in Pororo-SV with real-world entities? Do the edited videos look natural enough? Why do you use pronouns to replace character names, and wouldn't this make it difficult for ModelScopeT2V to guess the correct content, leading to an unfair comparison?\n2. What is the exact loss function used for finetuning?\n3. Why are some numbers not available in the results, such as FVD and FID for Coref-SV and Consistency for HiREST?"},"rating":{"value":"3: reject, not good enough"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699637047967,"tcdate":1699554204474,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission8416/Reviewer_pdWu"],"signatures":["ICLR.cc/2024/Conference/Submission8416/Reviewer_pdWu"],"forum":"5PkgaUwiY0","number":4,"license":"CC BY 4.0","cdate":1699554204474,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission8416/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699637047967,"domain":"ICLR.cc/2024/Conference","replyto":"5PkgaUwiY0","id":"atMBLroJ3f","forumContent":{"venue":{"value":"ICLR 2024 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Text-to-Video Generation","Large Language Models","Layout-Guided Video Generation","Temporal Consistency","Multi-Scene Video Generation","Layout Control"]},"supplementary_material":{"value":"/attachment/2436f0253ad2ca33bd87987c580d9851b9cb317b.zip"},"primary_area":{"value":"representation learning for computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Although recent text-to-video (T2V) generation methods have seen significant advancements, the majority of these works focus on producing short video clips of a single event with a single background (i.e., single-scene videos). Meanwhile, recent large language models (LLMs) have demonstrated their capability in generating layouts and programs to control downstream visual modules such as image generation models. This prompts an important question: can we leverage the knowledge embedded in these LLMs for temporally consistent long video generation? In this paper, we propose VideoDirectorGPT, a novel framework for consistent multi-scene video generation that uses the knowledge of LLMs for video content planning and grounded video generation. Specifically, given a single text prompt, we first ask our video planner LLM (GPT-4) to expand it into a ‘video plan’, which involves generating the scene descriptions, the entities with their respective layouts, the background for each scene, and consistency groupings of the entities and backgrounds. Next, guided by this output from the video planner, our video generator, named Layout2Vid, has explicit control over spatial layouts and can maintain temporal consistency of entities/backgrounds across multiple scenes, while being trained only with image-level annotations. Our experiments demonstrate that our proposed VideoDirectorGPT framework substantially improves layout and movement control in both single- and multi-scene video generation and can generate multi-scene videos with visual consistency across scenes, while achieving competitive performance with SOTAs in open-domain single-scene text-to-video generation. We also demonstrate that our framework can dynamically control the strength for layout guidance and can also generate videos with user-provided images. We hope our framework can inspire future work on integrating the planning ability of LLMs into consistent long video generation."},"_bibtex":{"value":"@misc{\nlin2024videodirectorgpt,\ntitle={VideoDirector{GPT}: Consistent Multi-Scene Video Generation via {LLM}-Guided Planning},\nauthor={Han Lin and Abhay Zala and Jaemin Cho and Mohit Bansal},\nyear={2024},\nurl={https://openreview.net/forum?id=5PkgaUwiY0}\n}"},"title":{"value":"VideoDirectorGPT: Consistent Multi-Scene Video Generation via LLM-Guided Planning"},"pdf":{"value":"/pdf/9d973f83e02170c22c014e1e66434974b4c189de.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Withdrawn_Submission"},"paperhash":{"value":"lin|videodirectorgpt_consistent_multiscene_video_generation_via_llmguided_planning"},"authorids":{"value":["~Han_Lin1","~Abhay_Zala1","~Jaemin_Cho1","~Mohit_Bansal2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Han Lin","Abhay Zala","Jaemin Cho","Mohit Bansal"]}},"version":2},{"content":{"summary":{"value":"The paper claims that the image encoder could capture rich spatial details while the video encoder could provide temporal contexts with sparse frames as input. Based on such observation, the authors propose to integrate both image encoder embeddings and video encoder embeddings to take both the benefits for video-language understanding. The performance on VCGBench, VCGBench-Diverse, MVBench and VideoMME shows the superiority of the trained model. The paper also proposes a new benchmark, called VCG-Diverse, posing its efforts towards question-answering of diversified video topics."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1. What is the instructions provided to the annotators to obtain good captions?\n2. Why not directly use the training data from plenty of existing works and see the effectiveness of the idea of integrating image encoder and video encoder?\n3. In terms of zero-shot question-answering, is Activitynet data excluded from the training set?\n4. Which results does Line 470-475 refer to? I don't see a cross-reference in the paper for this.\n5. Are results in tables 463 and 490 comparable? I understand from Lines 480-482 that these numbers might be based on different samples. Please correct me if I'm taking it wrong.\n6. By designing the new benchmarks, what are the insights and new conclusions for model design?"},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":1},"contribution":{"value":2},"strengths":{"value":"1. The idea of combining a high-resolution image encoder and a low-resolution sparse video encoder makes sense. It can somehow reduce the model training cost and inference latency, compared to full high-resolution dense video input.\n2. The paper introduces new training data curated by GPT-3.5, which might serve as new training data to feed the community.\n3. The paper introduces a new Video-QA benchmark using GPT-3.5 called VCGBench-Divese, which is based on human-curated descriptions.\n4. Experiments on various Video-QA benchmarks show the effectiveness of the trained model."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Related recent literature is missing. The paper has two efforts, one is towards better video understanding with LLM, e.g., [1][2][3], the other one is towards video-language benchmark and training data curation [4][5]. The paper lacks good literature reviews in both directions.\n2. Training data curation - necessity and fairness: To illustrate the effectiveness of the method and design. I think the experiments should be at least conducted on same data source and annotation approach. Curating new data  could be a good community contribution, however, the training data from existing works should be used to experiment to show the effectiveness of arch design.\n3. Inference data curation - video sources: The usage of HDVILA is doubtable. HDVILA was introduced as a pretraining dataset and an important source of training data. Using it for general Video-QA inference purposes may lead to leakage.\n4. Inference data curation - QA quality control: The VCGBench-Diverse QA is generated by GPT. How to control the quality? Is there any manual check involved?\n4. Experiments - Fairness: Missing some recent works. Please include the works for fair comparison.\n5. Experiments - Ablation: Adequate ablation experiments should be performed on the effectiveness of image encoder branch and video encoder branch on common video question answering benchmarks. In such way, the necessity of the dual branches can be validated.\n6. Format - the tables on Pages 9 and 10 don't have captions (excluding Table 4).\n7. The authors try to fit too much content into a single paper, which make the contribution of this paper ambiguous. For example, the integration of the image encoder and video encoder doesn't seem to be correlated to the data annotation process and new benchmark. I treasure their efforts, however, this might not be a good practice for me.\n\n[1] Llava-next-interleave: Tackling multi-image, video, and 3d in large multimodal models\n[2] Pllava: Parameter-free llava extension from images to videos for video dense captioning\n[3] Slowfast-llava: A strong training-free baseline for video large language models\n[4] ShareGPT4Video: Improving Video Understanding and Generation with Better Captions\n[5] Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos"}},"nonreaders":[],"tmdate":1731428662550,"tcdate":1730723727050,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission11214/Reviewer_8Pn4"],"signatures":["ICLR.cc/2025/Conference/Submission11214/Reviewer_8Pn4"],"forum":"YGWxpOI6Y0","number":5,"license":"CC BY 4.0","cdate":1730723727050,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission11214/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428662550,"domain":"ICLR.cc/2025/Conference","replyto":"YGWxpOI6Y0","id":"PUUBPaHKKe","forumContent":{"TLDR":{"value":"VideoGPT+: Spatiotemporal Aware Video Conversation Model"},"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["video-conversation-model","large multi-modal model","multi-modal","video-conversation","image-and-video","phi-3-min","vision-language","video-chatbot"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Building on the advances of language models, Large Multimodal Models (LMMs) have contributed significant improvements in video understanding. While the current video LMMs utilize advanced Large Language Models (LLMs), they rely on either image or video encoders to process visual inputs, each of which has its own limitations. Image encoders excel at capturing rich spatial details from frame sequences but lack explicit temporal context, which can be important in videos with intricate action sequences. On the other hand, video encoders provide temporal context but are often limited by computational constraints that lead to processing only sparse frames at lower resolutions, resulting in reduced contextual and spatial understanding. To this end, we introduce our model, which combines the complementary benefits of the image encoder (for detailed spatial understanding) and the video encoder (for global temporal context modeling). The model processes videos by dividing them into smaller segments and applies an adaptive pooling strategy on features extracted by both image and video encoders. Our architecture showcases improved performance across multiple video benchmarks, including VCGBench, MVBench and Zero-shot question-answering. Further, we develop 112K video-instruction set using a novel semi-automatic annotation pipeline which further improves the model performance. Additionally, to comprehensively evaluate video LMMs, we present our bench, covering 18 broad video categories such as lifestyle, sports, science, gaming, and surveillance videos. This benchmark with 4,354 question-answer pairs evaluates the generalization of existing LMMs on dense video captioning, spatial and temporal understanding, and complex reasoning, ensuring comprehensive assessment across diverse video types and dynamics. Our code, dataset, and pre-trained models will be publicly released."},"_bibtex":{"value":"@misc{\nmaaz2024videogpt,\ntitle={Video{GPT}+: Integrating Image and Video Encoders for Enhanced Video Understanding},\nauthor={Muhammad Maaz and Hanoona Abdul Rasheed and Salman Khan and Fahad Shahbaz Khan},\nyear={2024},\nurl={https://openreview.net/forum?id=YGWxpOI6Y0}\n}"},"title":{"value":"VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding"},"pdf":{"value":"/pdf/ec0f125e42fed1a2714c46de829865aaa0e50324.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"maaz|videogpt_integrating_image_and_video_encoders_for_enhanced_video_understanding"},"authorids":{"value":["~Muhammad_Maaz1","~Hanoona_Abdul_Rasheed1","~Salman_Khan4","~Fahad_Shahbaz_Khan1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Muhammad Maaz","Hanoona Abdul Rasheed","Salman Khan","Fahad Shahbaz Khan"]}},"version":2},{"content":{"venue":{"value":"CVPR 2026 Workshop VGBE"},"pdf":{"value":"/pdf/75d1575c5ca9fee46edd02233077a2dfc80cf738.pdf"},"keywords":{"value":["Diffusion Model","Reinforcement Learning","3D Vision"]},"submission_type":{"value":"Short Papers (up to 4 pages)"},"venueid":{"value":"thecvf.com/CVPR/2026/Workshop/VGBE"},"paperhash":{"value":"du|distilling_geometry_priors_for_3dconsistent_video_generation"},"authorids":{"value":["~Hongyang_Du2","~Junjie_Ye3","~Xiaoyan_Cong1","~Runhao_Li3","~Jingcheng_Ni2","~Aman_Agarwal4","~Zeqi_Zhou1","~Zekun_Li3","~Randall_Balestriero1","~Yue_Wang2"]},"abstract":{"value":"While recent video diffusion models (VDMs) produce visually impressive results, they fundamentally struggle to maintain 3D structural consistency, often resulting in object deformation or spatial drift. We hypothesize that these failures arise because standard denoising objectives lack explicit incentives for geometric coherence. To address this, we introduce VideoGPA (Video Geometric Preference Alignment), a data-efficient self-supervised framework that leverages a geometry foundation model to automatically derive dense preference signals that guide VDMs via Direct Preference Optimization (DPO). This approach effectively steers the generative distribution toward inherent 3D consistency without requiring human annotations. VideoGPA significantly enhances temporal stability, physical plausibility, and motion coherence using minimal preference pairs, consistently outperforming state-of-the-art baselines in extensive experiments."},"_bibtex":{"value":"@inproceedings{\ndu2026distilling,\ntitle={Distilling Geometry Priors for 3D-Consistent Video Generation},\nauthor={Hongyang Du and Junjie Ye and Xiaoyan Cong and Runhao Li and Jingcheng Ni and Aman Agarwal and Zeqi Zhou and Zekun Li and Randall Balestriero and Yue Wang},\nbooktitle={CVPR 2026 Workshop on Video Generative Models: Benchmarks and Evaluation},\nyear={2026},\nurl={https://openreview.net/forum?id=XLbjbCVtCZ}\n}"},"title":{"value":"Distilling Geometry Priors for 3D-Consistent Video Generation"},"authors":{"value":["Hongyang Du","Junjie Ye","Xiaoyan Cong","Runhao Li","Jingcheng Ni","Aman Agarwal","Zeqi Zhou","Zekun Li","Randall Balestriero","Yue Wang"]}},"tmdate":1774331923546,"pdate":1774331922583,"tcdate":1772225705235,"writers":["thecvf.com/CVPR/2026/Workshop/VGBE","thecvf.com/CVPR/2026/Workshop/VGBE/Submission3/Authors"],"signatures":["thecvf.com/CVPR/2026/Workshop/VGBE/Submission3/Authors"],"forum":"XLbjbCVtCZ","license":"CC BY 4.0","number":3,"cdate":1772225705235,"readers":["everyone"],"invitations":["thecvf.com/CVPR/2026/Workshop/VGBE/-/Submission","thecvf.com/CVPR/2026/Workshop/VGBE/-/Post_Submission","thecvf.com/CVPR/2026/Workshop/VGBE/-/Edit"],"mdate":1774331923546,"odate":1774331922583,"domain":"thecvf.com/CVPR/2026/Workshop/VGBE","id":"XLbjbCVtCZ","version":2},{"content":{"summary":{"value":"This work proposes a new wrist video generation model to alleviate the scarcity of wrist video data. It achieves this by proposing a reconstruction-generation pipeline. The reconstruction part aims to reconstruct a 4D point cloud from anchor videos and wrist view camera extrinsic parameters. The imperfect wrist view video can be rendered from these extrinsic parameters. The generation part focuses on generating a perfect wrist view video from these cues."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. The used training data is unclear. What did you use for the first stage of reconstruction and generation training? And for the second stage? Why was only 10k Droid data used? What about more Droid data?\n2. L150: S is missing an explanation.\n3. The subscript of y in lines 213 and 214 is different.\n4. The depth term of SPC is not very clear. Why is only the back term supervised? What is the rationale behind this loss design?\n5. How do you obtain correspondence points? What algorithm do you use?\n6. What does \"the i-th external view\" mean in L261?\n7. There is no explanation about \"cross-view fine-tuning.\" What does it mean? The training strategy seems novel; why don't you provide more details?\n8. What video generation model do you use? Is it trained from scratch in a newly designed network?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"The targeted problem is meaningful, and the solution is sound. This work proposes to upgrade VGGT with an extra head to predict wrist view camera extrinsics. To do so, this work introduces a correspondence supervision loss, referred to as spatial projection consistency loss. Afterward, this work injects the rendered imperfect wrist view video and anchor video into a video diffusion model to generate a perfect wrist video."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The presentation can be improved. I notice that the generated wrist view videos are repeated multiple times in your provided website (intro video and visualization part) as well as in the main manuscript. It is beneficial to showcase qualitative results as extensively as possible. Moreover, it is better to display the anchor video alongside the generated video. Furthermore,  the explanation is unclear: \n1. The used training data is unclear. What did you use for the first stage of reconstruction and generation training? And for the second stage? Why was only 10k Droid data used? What about more Droid data?\n2. L150: S is missing an explanation.\n3. The subscript of y in lines 213 and 214 is different.\n4. The depth term of SPC is not very clear. Why is only the back term supervised? What is the rationale behind this loss design?\n5. How do you obtain correspondence points? What algorithm do you use?\n6. What does \"the i-th external view\" mean in L261?\n7. There is no explanation about \"cross-view fine-tuning.\" What does it mean? The training strategy seems novel; why don't you provide more details?\n8. What video generation model do you use? Is it trained from scratch in a newly designed network?\n9. L426:  \"persis\"?"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762916865702,"tcdate":1761712947869,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission3607/Reviewer_fkUe"],"signatures":["ICLR.cc/2026/Conference/Submission3607/Reviewer_fkUe"],"forum":"Ilc2ybQWwH","number":2,"license":"CC BY 4.0","cdate":1761712947869,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission3607/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762916865702,"domain":"ICLR.cc/2026/Conference","replyto":"Ilc2ybQWwH","id":"HD9Y91XviV","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["World Model","Embodied AI"]},"supplementary_material":{"value":"/attachment/208350b4ad8b960ebe06d9603845075257990af3.zip"},"primary_area":{"value":"applications to robotics, autonomy, planning"},"abstract":{"value":"Wrist-view observations are crucial for VLA models as they capture fine-grained hand–object interactions that directly enhance manipulation performance.\nYet large-scale datasets rarely include such recordings, resulting in a substantial gap between abundant anchor views and scarce wrist views. Existing world models cannot bridge this gap, as they require a wrist-view first frame and thus fail to generate wrist-view videos from anchor views alone.\nAmid this gap, recent visual geometry models such as VGGT emerge with precisely the geometric and cross-view priors that make it possible to address such extreme viewpoint shifts.\nInspired by these insights, we propose WristWorld, the first 4D world model that generates wrist-view videos solely from anchor views.\nWristWorld operates in two stages: (i) Reconstruction, which extends VGGT and incorporates our proposed Spatial Projection Consistency (SPC) Loss to estimate geometrically consistent wrist-view poses and 4D point clouds; (ii) Generation, which employs our designed video generation model to synthesize temporally coherent wrist-view videos from the reconstructed perspective.\nExperiments on Droid, Calvin, and Franka Panda demonstrate state-of-the-art video generation with superior spatial consistency, while also improving VLA performance, raising the average task completion length on Calvin by 3.81% and closing 42.4% of the anchor-wrist view gap. See video results at anonymous page: https://wrist-world.github.io/"},"_bibtex":{"value":"@misc{\nqian2026wristworld,\ntitle={WristWorld: Generating Wrist-Views via 4D World Models for Robotic Manipulation},\nauthor={Zezhong Qian and Xiaowei Chi and Yuming Li and Shizun Wang and Sirui Han and Shanghang Zhang},\nyear={2026},\nurl={https://openreview.net/forum?id=Ilc2ybQWwH}\n}"},"title":{"value":"WristWorld: Generating Wrist-Views via 4D World Models for Robotic Manipulation"},"pdf":{"value":"/pdf/c751946463f7a953ff3d6e71707e9759371176f9.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"qian|wristworld_generating_wristviews_via_4d_world_models_for_robotic_manipulation"},"authorids":{"value":["~Zezhong_Qian3","~Xiaowei_Chi1","~Yuming_Li3","~Shizun_Wang1","~Sirui_Han1","~Shanghang_Zhang4"]},"authors":{"value":["Zezhong Qian","Xiaowei Chi","Yuming Li","Shizun Wang","Sirui Han","Shanghang Zhang"]}},"version":2},{"content":{"summary":{"value":"The authors proposed a novel self-supervised embedding approach AE2 to learn fine-grained view-invariant frame-wise video features from unpaired egocentric and exocentric videos. AE2 consists of two core designs: (1) an object-centric encoder that explicitly focuses on \nobjects; (2) a contrastive-based alignment objective that leverages temporally reversed frames as negative samples. Four datasets exhibit improvement in experiments, demonstrating a feasible framework for learning fine-grained view-invariant features."},"presentation":{"value":"3 good"},"contribution":{"value":"3 good"},"soundness":{"value":"3 good"},"strengths":{"value":"1. This paper presents a well-motivation idea and sound arguments in general. It makes sense to learn view-invariant video features form unpaired data.\n2. Ablation studies are conducted to evaluate the effectiveness of each design.\n3. The visualization of learned view-invariant video features for test videos is very impressive, which demonstrates effectively capture the progress of an action while remaining view-invariant."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. In Lines 225-226 “ In creating negative samples, we opt for reversing the frames in S rather than randomly shuffling them.” Although the authors explain the superiority of the reversing frames, authors need to provide further experiments comparing reversing frames with randomly shuffling them.\n2. Baseline's experimental results are so good that the baseline achieves SOTA results on most of the datasets. why is the baseline's experimental performance so good?\n3. Lack of analysis of experimental results in the ablation experiment section.\n4. In Training and Implementation Details, “the object-centre encoder network is optimized to minimize the loss above for 240 all pairs of video sequences in the training set, including ego-ego, exo-exo, and ego-exo pairs.”  The training process introduces data from the same viewpoint and paired data, but the network structure is designed specifically for unpaired data from different viewpoints. Can separate experiments be performed only on unpaired data from different viewpoints to demonstrate the validity of the proposed method?\n5. The language should be further improved by correcting some language errors, such as “ See Supp. for full implementation deatils ...” (Lines 232 in the paper)."},"confidence":{"value":"5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully."},"questions":{"value":"See Weaknesses"},"rating":{"value":"5: Borderline accept: Technically solid paper where reasons to accept outweigh reasons to reject, e.g., limited evaluation. Please use sparingly."},"code_of_conduct":{"value":"Yes"},"limitations":{"value":"None"}},"nonreaders":[],"tmdate":1702410817929,"tcdate":1688651677801,"writers":["NeurIPS.cc/2023/Conference","NeurIPS.cc/2023/Conference/Submission2068/Reviewer_CYjc"],"signatures":["NeurIPS.cc/2023/Conference/Submission2068/Reviewer_CYjc"],"forum":"Pj6X6GqNy8","number":3,"license":"CC BY 4.0","cdate":1688651677801,"mdate":1702410817929,"readers":["everyone"],"invitations":["NeurIPS.cc/2023/Conference/Submission2068/-/Official_Review","NeurIPS.cc/2023/Conference/-/Edit"],"domain":"NeurIPS.cc/2023/Conference","replyto":"Pj6X6GqNy8","id":"GdqXaD6ju9","forumContent":{"venue":{"value":"NeurIPS 2023 poster"},"keywords":{"value":["fine-grained video understanding","egocentric video","self-supervised learning","temporal alignment"]},"supplementary_material":{"value":"/attachment/9522a2e3745c925297c6377339f5499327e21558.zip"},"_bibtex":{"value":"@inproceedings{\nxue2023learning,\ntitle={Learning Fine-grained View-Invariant Representations from Unpaired Ego-Exo Videos via Temporal Alignment},\nauthor={Zihui Xue and Kristen Grauman},\nbooktitle={Thirty-seventh Conference on Neural Information Processing Systems},\nyear={2023},\nurl={https://openreview.net/forum?id=Pj6X6GqNy8}\n}"},"title":{"value":"Learning Fine-grained View-Invariant Representations from Unpaired Ego-Exo Videos via Temporal Alignment"},"paperhash":{"value":"xue|learning_finegrained_viewinvariant_representations_from_unpaired_egoexo_videos_via_temporal_alignment"},"TLDR":{"value":"a self-supervised embedding approach to learn fine-grained action features invariant to the egocentric and exocentric viewpoints"},"abstract":{"value":"The egocentric and exocentric viewpoints of a human activity look dramatically different, yet invariant representations to link them are essential for many potential applications in robotics and augmented reality.  Prior work is limited to learning view-invariant features from paired synchronized viewpoints.  We relax that strong data assumption and propose to learn fine-grained action features that are invariant to the viewpoints by aligning egocentric and exocentric videos in time, even when not captured simultaneously or in the same environment. To this end, we propose AE2, a self-supervised embedding approach with two key designs: (1) an object-centric encoder that explicitly focuses on regions corresponding to hands and active objects; (2) a contrastive-based alignment objective that leverages temporally reversed frames as negative samples. For evaluation, we establish a benchmark for fine-grained video understanding in the ego-exo context, comprising four datasets---including an ego tennis forehand dataset we collected, along with dense per-frame labels we annotated for each dataset. On the four datasets, our AE2 method strongly outperforms prior work in a variety of fine-grained downstream tasks, both in regular and cross-view settings."},"pdf":{"value":"/pdf/a88e226c00401c562932d4c6e21700aeee39d165.pdf"},"venueid":{"value":"NeurIPS.cc/2023/Conference"},"authorids":{"value":["~Zihui_Xue1","~Kristen_Grauman1"]},"authors":{"value":["Zihui Xue","Kristen Grauman"]}},"version":2},{"content":{"summary":{"value":"The paper introduces OmniVQA, a dataset for 360° Visual Question Answering (VQA) derived from the Stanford 2D–3D-S dataset, containing ~4.8k question–answer pairs focused on spatial reasoning in indoor environments. The authors also propose OmniVQABench, a benchmark for evaluating multimodal large language models on omnidirectional visual understanding, and a reinforcement fine-tuning approach (360-R1) using structured rewards (a combination of reasoning similarity, semantic accuracy, and answer formatting) to enhance performance of Qwen2.5-VL-7B."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1) How was “manual verification” of the generated QA pairs conducted, and what percentage of data received human review?\n2) Were dataset splits performed at the scene level to prevent leakage within Stanford 2D–3D-S?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":1},"strengths":{"value":"1) The motivation for developing a 360° VQA dataset is clear, as omnidirectional spatial reasoning is an underexplored area.\n2) The paper is well written, with transparent experimental details and ablations on reward components.\n3) The structured reward formulation for reasoning supervision is conceptually interesting and potentially generalizable."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1) Dataset annotation relies heavily on LLM-generated content (Qwen, DeepSeek, GPT-4o), with minimal clarity on human verification; this raises concerns about annotation quality and bias.\n2) Benchmark evaluation depends on DeepSeek-V3 as a judge, introducing circularity and bias in the reported scores.\n3) No human evaluation is provided to validate whether the observed gains correspond to genuine 360° reasoning improvements.\n4) The dataset is relatively small (≈5k QA pairs), restricting generalization.\n5) Baselines omit both state-of-the-art multimodal LLMs and geometry-aware 360° vision methods, weakening the empirical comparison.\n6) The reported improvements are potentially fragile, they may not be robust across evaluation judges, metrics, or alternative datasets, given the heavy dependence on a LLM-based scoring system."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762943086445,"tcdate":1761938660980,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission24458/Reviewer_W9W4"],"signatures":["ICLR.cc/2026/Conference/Submission24458/Reviewer_W9W4"],"forum":"VQV7SZ1wGy","number":3,"license":"CC BY 4.0","cdate":1761938660980,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission24458/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762943086445,"domain":"ICLR.cc/2026/Conference","replyto":"VQV7SZ1wGy","id":"KDXTED3w7p","forumContent":{"venue":{"value":"ICLR 2026 Conference Desk Rejected Submission"},"keywords":{"value":["Multi-Modal Large Models","Omnidirectional Vision"]},"supplementary_material":{"value":"/attachment/3c02b438db3fd52c35c8fe3885bc959173fcc743.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Omnidirectional images (ODIs), with their 360° field of view, provide unparalleled spatial awareness for immersive applications like augmented reality and embodied AI. However, the capability of existing multimodal large language models (MLLMs) to comprehend and reason about such panoramic scenes remains underexplored. This paper addresses this gap by introducing OmniVQA, the \\textbf{\\textit{first}} dataset and conducting the \\textbf{\\textit{first}} benchmark for omnidirectional visual question answering. Our evaluation of state-of-the-art MLLMs reveals significant limitations in handling omnidirectional visual question answering, highlighting persistent challenges in object localization, feature extraction, and hallucination suppression within panoramic contexts. These results underscore the disconnect between current MLLM capabilities and the demands of omnidirectional visual understanding, which calls for dedicated architectural or training innovations tailored to 360 ° imagery."},"_bibtex":{"value":"@misc{\nanonymous2026towards,\ntitle={Towards Omnidirectional Reasoning: A Dataset, Benchmark, and {GRPO}-based Method},\nauthor={Anonymous},\nyear={2026},\nurl={https://openreview.net/forum?id=VQV7SZ1wGy}\n}"},"title":{"value":"Towards Omnidirectional Reasoning: A Dataset, Benchmark, and GRPO-based Method"},"pdf":{"value":"/pdf/e9455f1c08a6853d44ff7324abafebdfaab5f78c.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Desk_Rejected_Submission"},"paperhash":{"value":"zhang|towards_omnidirectional_reasoning_a_dataset_benchmark_and_grpobased_method"},"authorids":{"value":["~Xinshen_ZHANG1","~Zhen_Ye2","~Xu_Zheng2"]},"authors":{"value":["Xinshen ZHANG","Zhen Ye","Xu Zheng"]}},"version":2},{"content":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["immersive content"]},"supplementary_material":{"value":"/attachment/5773686dc6b17b2632d3dc874878b8b3fbb061b7.zip"},"primary_area":{"value":"other topics in machine learning (i.e., none of the above)"},"abstract":{"value":"Generating a complete and explorable 360-degree visual world enables a wide range of downstream applications. While prior works have advanced the field, they remain constrained by either narrow field-of-view limitations, which hinder the synthesis of continuous and holistic scenes, or insufficient controllability that restricts free exploration by users or autonomous agents. To address this, we propose PanoWorld-X, a novel framework for high-fidelity and controllable panoramic video generation with diverse camera trajectories. \n First, we propose a novel pipeline for synthesizing panoramic video-trajectory dataset pairs in virtual 3D environments via Unreal Engine. This pipeline consists of four main steps and enables the collection of a large-scale dataset with rich scene diversity and accurate trajectory annotations.\n To achieve precise panoramic video generation, we identify that the bottleneck arises from the misalignment between the spherical geometry of panoramic data and the inductive priors of conventional video diffusion models. To address this, we leverage the spherical connectivity characteristics of panorama data, and propose a Sphere-Aware Diffusion Transformer that reprojects equirectangular features onto the spherical surface, thereby capturing geometric adjacency in the latent space. This design significantly improves both visual fidelity and spatiotemporal continuity.\n   Extensive experiments demonstrate that our PanoWorld-X achieves superior performance in various aspects, including motion range, control precision, and visual quality, underscoring its potential for real-world applications."},"_bibtex":{"value":"@misc{\nyin2026panoworldx,\ntitle={PanoWorld-X: Generating Explorable Panoramic Worlds via Sphere-Aware Video Diffusion},\nauthor={Yuyang Yin and Hao-Xiang Guo and Fangfu Liu and Mengyu Wang and Hanwen Liang and Eric Li and Yikai Wang and Xiaojie Jin and Yao Zhao and Yunchao Wei},\nyear={2026},\nurl={https://openreview.net/forum?id=iZyBEbq6jR}\n}"},"title":{"value":"PanoWorld-X: Generating Explorable Panoramic Worlds via Sphere-Aware Video Diffusion"},"pdf":{"value":"/pdf/8d2464ff9acb82ee4bdbbfda2c8b3031ed0755c7.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"yin|panoworldx_generating_explorable_panoramic_worlds_via_sphereaware_video_diffusion"},"authorids":{"value":["~Yuyang_Yin1","~Hao-Xiang_Guo1","~Fangfu_Liu2","~Mengyu_Wang5","~Hanwen_Liang2","~Eric_Li8","~Yikai_Wang2","~Xiaojie_Jin1","~Yao_Zhao1","~Yunchao_Wei1"]},"authors":{"value":["Yuyang Yin","Hao-Xiang Guo","Fangfu Liu","Mengyu Wang","Hanwen Liang","Eric Li","Yikai Wang","Xiaojie Jin","Yao Zhao","Yunchao Wei"]}},"tmdate":1770804968875,"tcdate":1758204586419,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission11906/Authors"],"signatures":["ICLR.cc/2026/Conference/Submission11906/Authors"],"forum":"iZyBEbq6jR","license":"CC BY 4.0","number":11906,"cdate":1758204586419,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/-/Submission","ICLR.cc/2026/Conference/-/Post_Submission","ICLR.cc/2026/Conference/Submission11906/-/Full_Submission","ICLR.cc/2026/Conference/Submission11906/-/Rebuttal_Revision","ICLR.cc/2026/Conference/-/Edit"],"mdate":1770804968875,"odate":1759896705795,"domain":"ICLR.cc/2026/Conference","id":"iZyBEbq6jR","version":2},{"content":{"submission_length":{"value":"Long submission (more than 12 pages of main content)"},"venue":{"value":"Under review for TMLR"},"abstract":{"value":"Deep learning models trained on datasets with spurious correlations can achieve high average accuracy whilst relying on shortcut features that do not generalise out of distribution. Whilst out-of-distribution testing can help in highlighting subgroup performance disparities arising from shortcut learning, it does not localise the regions within images that are associated with it. Existing research mostly utilises attribution maps from interpretability methods for understanding the spatial nature of spurious correlations. For example, conditional alignment methods separate task-relevant evidence from evidence tied to spurious correlations by comparing attribution maps from a task model, a sensitive attribute model, and a bias-reduced reference model. However, existing methods aggregate regional attribution rankings across images before computing dataset-level shortcut and task contribution maps, potentially masking recurring spatial shortcut patterns that occur in subsets of images. We address this limitation by computing per-image shortcut and task contribution maps and grouping their joint representations into recurring spatial patterns using K-means and non-negative matrix factorisation, and visualising the resulting shortcut groups through contribution maps and representative examples. Across CelebA, CheXpert, Waterbirds}, Camelyon17, and ISIC2019, and across ResNet and ViT models, the discovered shortcut groups reveal both shared and distinct spatial patterns of shortcut and task contribution, with varying subgroup composition and error rates, thereby enabling targeted inspection of image subsets with higher error rates. We perform input occlusion and internal test-time interventions to demonstrate that masking or suppression of task contribution regions substantially degrades the model classification performance and propose a combined shortcut suppression and task amplification feature intervention approach which generally reduces performance disparities."},"_bibtex":{"value":"@article{\nanonymous2026discovery,\ntitle={Discovery and Spatial Characterisation of Multiple Shortcut Groups for Auditing Vision Model Bias},\nauthor={Anonymous},\njournal={Submitted to Transactions on Machine Learning Research},\nyear={2026},\nurl={https://openreview.net/forum?id=Kvr5IFcz8t},\nnote={Under review}\n}"},"title":{"value":"Discovery and Spatial Characterisation of Multiple Shortcut Groups for Auditing Vision Model Bias"},"changes_since_last_submission":{"value":"Reformatted the manuscript using the current official TMLR template."},"pdf":{"value":"/pdf/82ec8ba96e6279b0304e3c1b4d2c1e0c2aa6f0a0.pdf"},"venueid":{"value":"TMLR/Under_Review"},"previous_TMLR_submission_url":{"value":"https://openreview.net/forum?id=dzjjLyKt35"},"assigned_action_editor":{"value":"~Chenyu_You1"}},"tmdate":1790453529976,"tcdate":1786705342493,"writers":["TMLR","TMLR/Paper11356/Authors"],"signatures":["TMLR/Paper11356/Authors"],"forum":"Kvr5IFcz8t","number":11356,"license":"CC BY 4.0","cdate":1786705342493,"readers":["everyone"],"invitations":["TMLR/-/Submission","TMLR/-/Edit","TMLR/-/Under_Review","TMLR/Paper11356/-/Revision"],"mdate":1790453529976,"odate":1787174845114,"domain":"TMLR","id":"Kvr5IFcz8t","version":2},{"content":{"research_area_keywords":{"value":"NLP Applications, Special Theme Track (Generalization of NLP Models)"},"languages_studied":{"value":"English"},"venue":{"value":"ACL ARR 2025 February Submission"},"_bibtex":{"value":"@inproceedings{\nanonymous2025shortcut,\ntitle={Shortcut Learning in Safety: The Impact of Keyword Bias in Safeguards},\nauthor={Anonymous},\nbooktitle={Submitted to ACL Rolling Review - February 2025},\nyear={2025},\nurl={https://openreview.net/forum?id=IOP5nuRx5S},\nnote={under review}\n}"},"title":{"value":"Shortcut Learning in Safety: The Impact of Keyword Bias in Safeguards"},"contribution_types":{"value":["Model analysis & interpretability"]},"abstract":{"value":"Safeguarding LLMs requires separating harmful prompts from safe ones.  We frame this reliance as a shortcut learning problem and conduct experiments revealing how existing models depend on specific keywords for classification rather than semantic understanding. Performance evaluations across six safety benchmarks show that models perform well when keyword distributions align but degrade on out-of-distribution prompts. Results from our counterfactual analysis demonstrate that current safeguard models are vulnerable to keyword distribution shifts due to shortcut learning. These findings highlight the importance of addressing shortcut learning to enhance the robustness of safeguard models."},"paper_type":{"value":"Short"},"pdf":{"value":"/pdf/e1ac6986d85a5163de293b6b4b0a585168748ee2.pdf"},"research_area":{"value":"NLP Applications"},"venueid":{"value":"aclweb.org/ACL/ARR/2025/February/Submission"}},"tmdate":1782961966315,"tcdate":1739548088123,"writers":["aclweb.org/ACL/ARR/2025/February","aclweb.org/ACL/ARR/2025/February/Submission2181/Authors"],"signatures":["aclweb.org/ACL/ARR/2025/February/Submission2181/Authors"],"forum":"IOP5nuRx5S","license":"CC BY 4.0","number":2181,"cdate":1739548088123,"readers":["everyone"],"invitations":["aclweb.org/ACL/ARR/2025/February/-/Submission","aclweb.org/ACL/ARR/2025/February/-/Edit","aclweb.org/ACL/ARR/2025/February/-/Post_Submission","aclweb.org/ACL/ARR/2025/February/Submission2181/-/Change_Reviewer_Nomination","aclweb.org/ACL/ARR/2025/February/Submission2181/-/Blind_Submission_License_Agreement","aclweb.org/ACL/ARR/2025/February/-/Preprint_Release_Submission","aclweb.org/ACL/ARR/2025/February/-/Preprint_Post_Submission"],"mdate":1782961966315,"odate":1746821355472,"domain":"aclweb.org/ACL/ARR/2025/February","id":"IOP5nuRx5S","version":2},{"content":{"summary":{"value":"This paper introduces a video VAE for video compression and generation tasks. It features 4 major components: a causal 3D residual block, spatio-temporal downsampling module, spatio-temporal attention module, and FILM encoder. The proposed VAE is tested with the autoencoding and generation tasks."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"(1) The title “joint image and video compression” is misleading. I believe the task is focused primarily on video generation, not video transmission or compression (i.e. the autoencoding task). If the task is video compression/autoencoding, the resulting rate-distortion performance of the proposed method should be compared with that of the prior works on end-to-end learned video compression. However, this was not done when the authors reported results on the video/image autoencoding task. Just because the resulting latents have smaller dimensionality does not suggest that they would have smaller bit rates.  \n\n(2) In Table 1, Video LDM (1 x 8 x 8) actually performs quite well particularly on DAVIS with large motion, as opposed to your better performing variant (4 x 8 x 8). Notably, Video LDM has relatively better TS performance. This leads to an impression that the proposed method cannot model well fast-motion sequences.\n\n(3) The model size and MAC/pixel are currently missing in Table 1. Also, except TS, all the quality metrics are for images, not video. Why not use VMAF?\n\n(4) I wonder if the competing methods adopt the GAN loss for training. This needs to be clarified in Table I. In addition, it is unclear how much the GAN loss contributes to the performance of the proposed method.\n\n(5) For the image autoencoding task, how is the proposed method used?\n\n(6) For the video generation task, it is unclear whether the proposed VAE is trained or fine-tuned together with the diffusion model. It appears to me that the proposed VAE is pre-trained to generate latents that follow a simple Gaussian distribution. \n\n(7) In addition, the proposed VAE can be a stand-alone approach to video generation. I wonder how it performs as compared to the diffusion-based modeling of the video latents.  \n\n(8) In Figure 4, I wonder whether the proposed method can generate well fast-motion sequences. \n\n(9) FILM encoder is shown to be effective in the ablation study. But, the proposed VAE does not appear to work well on the autoencoding task in Table I, as compared to Video LDM.  This is a bit confusing."},"rating":{"value":5},"details_of_ethics_concerns":{"value":"Not applicable."},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"The proposed method is evaluated with both the video/image autoencoding task and the video generation task. Moreover, extensive ablation studies were conducted to validate the effectiveness of the proposed method."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"(1) The title “joint image and video compression” is misleading. I believe the task is focused primarily on video generation, not video transmission or compression (i.e. the autoencoding task). If the task is video compression/autoencoding, the resulting rate-distortion performance of the proposed method should be compared with that of the prior works on end-to-end learned video compression. However, this was not done when the authors reported results on the video/image autoencoding task. Just because the resulting latents have smaller dimensionality does not suggest that they would have smaller bit rates.\n\n(2) In terms of the newly proposed components, I have the impression that their novelty is not very high.\n\n(3) The necessity of causality is unclear in the context of video generation. If the aim is to enable both image and video generation, the image generation task is not explored in the current writing."}},"nonreaders":[],"tmdate":1732770892805,"tcdate":1730622113480,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission2418/Reviewer_tD7c"],"signatures":["ICLR.cc/2025/Conference/Submission2418/Reviewer_tD7c"],"forum":"aRD1NqcXTC","number":3,"license":"CC BY 4.0","cdate":1730622113480,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission2418/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732770892805,"domain":"ICLR.cc/2025/Conference","replyto":"aRD1NqcXTC","id":"ia5jKJcLMi","forumContent":{"TLDR":{"value":"A causal video VAE for joint image and video tokenization"},"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Autoencoding","Generative Modelling","Causal Video VAE","FILM","Video Tokenization"]},"supplementary_material":{"value":"/attachment/64a77b89dfe3e76ce46a81375afe1fd4c32272aa.zip"},"primary_area":{"value":"generative models"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Generative modeling has seen significant advancements in image and video synthesis. However, the curse of dimensionality remains a significant obstacle, especially for video generation, given its inherently complex and high-dimensional nature. Many existing works rely on low-dimensional latent spaces from pretrained image autoencoders. However, this approach overlooks temporal redundancy in videos and often leads to temporally incoherent decoding. To address this issue, we propose a video compression network that reduces the dimensionality of visual data both spatially and temporally. Our model, based on a variational autoencoder, employs causal 3D convolution to handle images and videos jointly. The key contributions of our work include a scale-agnostic encoder for preserving video fidelity, a novel spatio-temporal down/upsampling block for robust long-sequence modeling, and a flow regularization loss for accurate motion decoding. \nOur approach outperforms competitors in video quality and compression rates across various datasets. Experimental analyses also highlight its potential as a robust autoencoder for video generation training."},"_bibtex":{"value":"@inproceedings{\nargaw2025highquality,\ntitle={High-Quality Joint Image and Video Tokenization with Causal {VAE}},\nauthor={Dawit Mureja Argaw and Xian Liu and Qinsheng Zhang and Joon Son Chung and Ming-Yu Liu and Fitsum Reda},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=aRD1NqcXTC}\n}"},"title":{"value":"High-Quality Joint Image and Video Tokenization with Causal VAE"},"pdf":{"value":"/pdf/55e58cfb0679e477e52cfc55ece2e1183259bcb5.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"argaw|highquality_joint_image_and_video_tokenization_with_causal_vae"},"authorids":{"value":["~Dawit_Mureja_Argaw1","~Xian_Liu1","~Qinsheng_Zhang1","~Joon_Son_Chung1","~Ming-Yu_Liu1","~Fitsum_Reda1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Dawit Mureja Argaw","Xian Liu","Qinsheng Zhang","Joon Son Chung","Ming-Yu Liu","Fitsum Reda"]}},"version":2},{"content":{"summary":{"value":"This paper presents CLAMR, a contextualized late-interaction retriever for multimodal video content retrieval, and makes three core contributions. Proposing a unified vision-language backbone that jointly encodes four modalities to enhance cross-modal contextualization, addressing the limitations of independent modality encoding in conventional methods. Introducing MULTIVENT 2.0++, a large-scale synthetic dataset with 371k modality-targeted queries, solving the scarcity of fine-grained modality-specific training data for multimodal retrieval."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"How does the modality-wise late-interaction mechanism (LIₘᵥ) specifically resolve conflicts when multiple modalities contain relevant information? The paper emphasizes single-modality targeting but lacks analysis of cross-modal corroboration scenarios.\nWhat are the root causes of the lower performance on vision-targeted queries? \nFor videos longer than 60 minutes, how does CLAMR’s frame sampling and retrieval efficiency degrade? Is there a strategy to optimize long-sequence processing?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"The joint encoding of multiple modalities and modality-wise late-interaction balance fine-grained token-level matching and computational efficiency, filling the gap of late-interaction’s underutilization in multimodal video retrieval.\nEvaluates on multiple benchmarks and conducts extensive ablations, ensuring the reliability of conclusions."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Focuses only on four modalities and does not explore other critical modalities in video content.\n Evaluations are primarily based on event-centric and general video datasets.\nWhile the joint encoding backbone aims to align modalities, the paper lacks analysis of alignment failures in temporally or semantically misaligned video content. \nThe model’s performance heavily relies on high-quality ASR transcripts and OCR text, making it vulnerable to low-resource scenarios where these modalities are noisy or unavailable. \nThe paper does not provide interpretability into how the modality-aware mechanism selects the “most relevant” modality for a given query."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1764360131099,"tcdate":1761954271465,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission15879/Reviewer_cXTo"],"signatures":["ICLR.cc/2026/Conference/Submission15879/Reviewer_cXTo"],"forum":"AXFuBS3ujj","number":4,"license":"CC BY 4.0","cdate":1761954271465,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission15879/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1764360131099,"domain":"ICLR.cc/2026/Conference","replyto":"AXFuBS3ujj","id":"bXIO0dA68A","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["multimodal retrieval","text-video retrieval","RAG"]},"supplementary_material":{"value":"/attachment/30f3f46b897718fb535393e45c6585c6c2f48bcf.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Online video content is richly multimodal: a single video might blend vision, speech, ambient audio, and on-screen text. Conventional retrieval systems typically treat these modalities as independent retrieval sources, which can lead to noisy and subpar results.\nIn this work, we explore multimodal video content retrieval, where relevance can be scored from a single modality or jointly across multiple modalities. Consequently, an effective retriever must dynamically determine which modality (or set of modalities) best address a given query. We introduce CLaMR, a multimodal, late-interaction retriever that jointly indexes four modalities: video frames, transcribed speech, on-screen text, and metadata. CLaMR jointly encodes all modalities within a unified multimodal backbone for improved contextualization and is trained to enhance dynamic modality selection via two key innovations. First, to overcome the lack of suitable training data, we introduce MultiVent 2.0++, a large-scale synthetic dataset built on MultiVent 2.0 (a collection of event-centric videos in various languages paired with English queries) with modality-targeted queries to teach modality selection. Next, we propose a modality-aware contrastive loss that trains the model on both a standard contrastive objective and an objective for learning correct modality usage. On the test sets of MultiVent 2.0++ and MSRVTT, we observe that conventional aggregation strategies, such as averaging similarities for baseline retrievers, often degrade performance by introducing noise from irrelevant modalities. In contrast, CLaMR consistently outperforms existing retrievers: on MultiVent 2.0++, CLaMR improves nDCG@10 by 25.6 points over the best-performing single-modality retriever and by 35.4 points over the best-performing multi-modality retriever. We illustrate the downstream utility of CLaMR with experiments on long-video QA, where it improves performance by 3.50% over LanguageBind on Video-MME and 1.42% over dense frame sampling on LongVideoBench."},"_bibtex":{"value":"@misc{\nwan2026clamr,\ntitle={{CL}a{MR}: Contextualized Late-Interaction for Multimodal Content Retrieval},\nauthor={David Wan and Han Wang and Elias Stengel-Eskin and Jaemin Cho and Mohit Bansal},\nyear={2026},\nurl={https://openreview.net/forum?id=AXFuBS3ujj}\n}"},"title":{"value":"CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval"},"pdf":{"value":"/pdf/64dc1ed7cf0393bc3ae6e960de0123892d1a795c.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"wan|clamr_contextualized_lateinteraction_for_multimodal_content_retrieval"},"authorids":{"value":["~David_Wan1","~Han_Wang9","~Elias_Stengel-Eskin1","~Jaemin_Cho1","~Mohit_Bansal2"]},"authors":{"value":["David Wan","Han Wang","Elias Stengel-Eskin","Jaemin Cho","Mohit Bansal"]}},"version":2},{"content":{"summary":{"value":"The paper proposes three action-conditioned video prediction training frameworks and shows empirical results of their efficacy on a robot video dataset."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"* How are discussions on diffusion-based video prediction (line 131 - 135) related to camera controls (line 103)? \n* What's the action space in RoAM? Is it continuous or discrete? Does the method scale to high-dimensional action space?"},"rating":{"value":5},"details_of_ethics_concerns":{"value":"N/A"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"* The paper adapts action-free VAE and flow-based formulation from video prediction literature to action-conditioned learning. \n* Empirical results show benefit of incorporating action information in training in terms of video prediction accuracy."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"* Missing related works. The paper revisits the literature on VAE-based and flow-matching-based video prediction models, without discussing the latter in the section of prior works. \n* The method restricts cameras to be static (line 148). This assumption does not hold in general for casual videos outside of the training data being used and limits the applicability of the method. \n* The assumption of causality between actions $a_t$ and observed framer $x_t$ is again specific to robot manipulation tasks with fixed cameras as in the RoAM dataset used in this paper. This assumption does not hold in generic videos. \n* The paper lacks a discussion of the potential benefit of incorporating actions into video modeling frameworks in general. Action-free video prediction training does not require action annotations and is much more scalable. If the downstream task is robot manipulation, an alternative approach is to train action-free video prediction models and extract robot actions via inverse dynamics [1]. More discussions would help strengthen the paper. If the goal is accurate video prediction itself, then how does the method compare to state-of-the-art video prediction architectures? \n* What's the relation of the proposed 3 distinct models? A much more extensive discussion on this would help clarify the motivation of developing three separate frameworks in the paper.  \n* The paper claims results on incorporating camera motion (line 014) but empirically only evaluate on datasets with fixed cameras. \n\n[1] Learning Universal Policies via Text-Guided Video Generation."}},"nonreaders":[],"tmdate":1731428773699,"tcdate":1730727016560,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission13808/Reviewer_XuYh"],"signatures":["ICLR.cc/2025/Conference/Submission13808/Reviewer_XuYh"],"forum":"VAvZ4oinpa","number":3,"license":"CC BY 4.0","cdate":1730727016560,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission13808/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428773699,"domain":"ICLR.cc/2025/Conference","replyto":"VAvZ4oinpa","id":"kjHdPflRPw","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"TLDR":{"value":"We propose variational models for learning action priors for video generation tasks in situations where the camera is also moving like in autonomous cars or robots."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Stochastic Video Generation","Variational Inference"]},"supplementary_material":{"value":"/attachment/ba85e8fd60e0de14ef7300be056e5cbf5f02a58a.pdf"},"primary_area":{"value":"probabilistic methods (Bayesian methods, variational inference, sampling, UQ, etc.)"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Long-term stochastic video generation remains challenging, especially with moving cameras. This scenario introduces complex interactions between camera movement and observed pixels, resulting in intricate spatio-temporal dynamics and partial observability issues. Current approaches often focus on pixel-level image reconstruction, neglecting explicit modeling of camera motion dynamics. Our proposed solution incorporates camera motion or action as an extended part of the observed image state, employing a multi-modal learning framework to simultaneously model both image and action. We introduce three models: (i) Video Generation with Learning Action Prior (VG-LeAP) that treats the image-action pair as an augmented state generated from a single latent stochastic process and uses variational inference to learn the image-action latent prior; (ii) Causal-LeAP, which establishes a causal relationship between action and the observed image frame, and learns a seperate action prior, conditioned on the observed image states along with the image prior; and (iii) RAFI, which integrates the augmented image-action state concept with a conditional flow matching framework, demonstrating that this action-conditioned image generation concept can be extended to other transformer-based architectures. Through comprehensive empirical studies on robotic video dataset, RoAM, we highlight the importance of multi-modal training in addressing partially observable video generation problems."},"_bibtex":{"value":"@misc{\nsarkar2025video,\ntitle={Video Generation with Learned Action Prior},\nauthor={Meenakshi Sarkar and Devansh Bhardwaj and Debasish Ghose},\nyear={2025},\nurl={https://openreview.net/forum?id=VAvZ4oinpa}\n}"},"title":{"value":"Video Generation with Learned Action Prior"},"pdf":{"value":"/pdf/7e0cbd10c17c4802fbb1a9110ab3deec2d204710.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"sarkar|video_generation_with_learned_action_prior"},"authorids":{"value":["~Meenakshi_Sarkar1","~Devansh_Bhardwaj1","~Debasish_Ghose1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Meenakshi Sarkar","Devansh Bhardwaj","Debasish Ghose"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a framework which integrates the idea from streaming dense video captioning and retrieval-augmented generation. Experiments on serveral benchmarks illustrate its effectiveness."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"As in weaknesses."},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"Extensive experiments illustrate that the proposed framework is an effective composition of the core ideas from streaming dense video captioning and retrieval-augmented caption generation."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Limited novelty: The primary contributions of this paper are heavily based on recent works$^{[1, 2]}$, with no significant improvements over previous methods. The use of retrieval-augmented generation is a common technique, and its application here adds little to the research community. The utilization of action information for video captioning has also been clained in previous works $^{[3]}$. \n\n2. Lack of in-depth analysis:  While this paper highlights the importance of information in verbs/actions compared to previous retrieval-augmented methods, it overlooks the potential drawbacks. Many actions can have similar visual representations, and relying on retrieved action information without correction methods may lead to errors. It would be beneficial to conduct a deeper analysis of the impact of incorporating action information, such as examining the accuracy of action recall and the distribution of actions in your dataset. And if the recall rate is high, indicating precise action perception by the model, what advantages does retrieval-augmentation offer by introducing more relevant information? Additionally, relying on actions seems to complicate the definition of learning objectives: “[BOS][s1][e1][caption text1][s2][e2][caption text2] ... [EOS]”. For a common phrase like “Pick up the apple and peel it,” what should the corresponding statement be?\n\n3. Poor presentation: Section 3 includes too many design details, presented mostly in text without considering the reader's experience. Consider enhancing this by incorporating illustrations and reducing the amount of text. It's also crucial to focus on simplicity and accuracy in your presentation.\n\n\n[1] Streaming Dense Video Captioning. CVPR 2024.\n\n[2] Do You Remember? Dense Video Captioning with Cross-Modal Memory Retrieval. CVPR 2024.\n\n[3] Syntax-Aware Action Targeting for Video Captioning. CVPR 2020."}},"nonreaders":[],"tmdate":1731427795028,"tcdate":1730713115345,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission3294/Reviewer_ML2Z"],"signatures":["ICLR.cc/2025/Conference/Submission3294/Reviewer_ML2Z"],"forum":"oO3oXJ19Pb","number":3,"license":"CC BY 4.0","cdate":1730713115345,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission3294/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427795028,"domain":"ICLR.cc/2025/Conference","replyto":"oO3oXJ19Pb","id":"SuqO2I0Xp6","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Dense video captioning","Online dense video captioning"]},"supplementary_material":{"value":"/attachment/22be9efeac9199c2db56442e763f88fbc635769f.pdf"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Dense video captioning requires solving the challenging tasks of temporally localizing events and generating descriptive captions within long video sequences. Existing methods often struggle to capture the evolving context within video streams and to produce accurate temporal alignment. To address this, we propose an online retrieval-augmented approach that processes video segments incrementally while dynamically retrieving relevant action phrases from a pre-constructed action-text corpus. This enriches the contextual information for both the video representation and the subsequent text decoder, improving the caption generation. Additionally, we present image-based simulated video pretraining, which mitigates the reliance on extensive video datasets by using image-level text-paired data aligned with the online video captioning format. Our experiments on the ViTT, YouCook2, and ActivityNet benchmarks demonstrate that our model significantly outperforms both existing global and online methods, validating its effectiveness."},"_bibtex":{"value":"@misc{\nkim2024actions,\ntitle={Actions Inspire Every Moment: Online Action-Augmented Dense Video Captioning},\nauthor={Dahun Kim and AJ Piergiovanni and Anelia Angelova},\nyear={2024},\nurl={https://openreview.net/forum?id=oO3oXJ19Pb}\n}"},"title":{"value":"Actions Inspire Every Moment: Online Action-Augmented Dense Video Captioning"},"pdf":{"value":"/pdf/ced2509c52424b1ede3d852681c4a02933dc58a7.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"kim|actions_inspire_every_moment_online_actionaugmented_dense_video_captioning"},"authorids":{"value":["~Dahun_Kim1","~AJ_Piergiovanni1","~Anelia_Angelova1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Dahun Kim","AJ Piergiovanni","Anelia Angelova"]}},"version":2},{"content":{"venue":{"value":"MobileHCI 2019"},"venueid":{"value":"dblp.org/conf/MHCI/2019"},"paperhash":{"value":"lai|a_shortcut_for_caret_positioning_on_touchscreen_phones"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Jianwei_Lai:","~Navid_Rajabi1","https://dblp.org/search/pid/api?q=author:Elahe_Javadi:"]},"html":{"value":"https://doi.org/10.1145/3338286.3340146"},"_bibtex":{"value":"@inproceedings{DBLP:conf/mhci/LaiRJ19,\n  author={Jianwei Lai and Navid Rajabi and Elahe Javadi},\n  title={A Shortcut for Caret Positioning on Touch-Screen Phones},\n  year={2019},\n  cdate={1546300800000},\n  pages={35:1-35:6},\n  url={https://doi.org/10.1145/3338286.3340146},\n  booktitle={MobileHCI},\n  crossref={conf/mhci/2019}\n}\n"},"abstract":{"value":"Moving a caret to the desired position in a text field is challenging and inefficient, especially when the user has only one hand available to hold and interact with a mobile phone. We propose a shortcut on keyboards that enables precise and efficient caret positioning in a text field on mobile devices. The proposed method uses a long press on a key to determine the desired position for the caret in the text field. A lab experiment was conducted to compare the shortcut with the traditional handle method and an existing method provided by the Google Keyboard. Even though participants were using the proposed shortcut for the first time, the technique achieved a higher task completion speed comparing to the Google method and was as quick as the traditional handle method. It also showed advantages in accuracy and user perceptions than the other two methods for one-handed interaction."},"title":{"value":"A Shortcut for Caret Positioning on Touch-Screen Phones"},"authors":{"value":["Jianwei Lai","Navid Rajabi","Elahe Javadi"]}},"tmdate":1718680232760,"pdate":1546300800000,"tcdate":1718680229673,"writers":["~"],"signatures":["~Navid_Rajabi1"],"forum":"X6maNlDAA4","license":"CC BY-SA 4.0","number":32263,"cdate":1546300800000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1718680232760,"domain":"DBLP.org","id":"X6maNlDAA4","version":2},{"content":{"summary":{"value":"This paper utilizes controllable video generation for learning to estimate 3D human poses robustly. RGB Video generation is supposed to help in generating diverse 2D human motion sequence through varying poses, scenes, and viewpoints. Besides using real 2D inputs, the idea behind using real-world detections is to design a robust and generalizable pose estimation model. The authors conduct experimental evaluation on various 3DHPE datasets to show the effectiveness of the proposed approach on real world and corrupt 2D inputs."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. What is the precise novelty beyond combining existing diffusion video generation models with pose augmentation?\n2. What is the core mechanism by which background variation improves generalization? The hypothesis is intuitive (“background matters”), but the paper lacks an analysis or visualization showing how added background diversity affects learned representations.\n3. What is the computational overhead of generating and filtering data? Since video diffusion generation is extremely costly, the scalability of this method to larger datasets or real-time adaptation is unclear. Additionally, for a generalizable model, training a video generation model on extensively diverse datasets seems to be crucial but at the same time unscalable."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The paper reframes the usual pose-only augmentation approach into a multi-modal data generation pipeline that explicitly models scene diversity in terms of pose, viewpoint, lighting, etc. which helps in building a generalizable 3D human pose estimation model.\n2. The idea of leveraging pose-guided video diffusion models is intuitive and a straightforward implementation should be easily achievable.\n3. Experiments include multiple datasets (H36M, PMR, 3DHP, 3DPW) and diverse metrics (MPJPE, P-MPJPE, velocity error). Moreover, the paper evaluates robustness under real-world corruptions (blur, compression, spatter), which strengthens the practical motivation.\n4. The effects of filtering ratios, pretraining strategies, and GT vs. detected 2D inputs are well-studied. Also, the results seem to demonstrate consistent improvements across nearly all configurations which suggests robustness of the method."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The technical contribution mainly lies in data generation and composition rather than in a novel model or algorithm which I feel is a major discussion point. The method builds on existing diffusion video generation models with minimal architectural innovation.\n2. The paper relies heavily on pretrained models like Animate Anyone (Hu et al.) and Latent Diffusion model (Rombach et al.) without domain-specific adaptation. It is unclear how much of the improvement comes from the inherent realism of these models rather than the proposed pipeline design.\n3. The method is validated primarily on H36M, PMR, and 3DHP which are all captured in a controlled environment. A demonstration on truly unseen in-the-wild datasets (e.g., MPII, COCO-Video, AMASS-based scenes) is missing and would support generalization claims.\n4. One of the concerns is that generating and filtering large-scale video data using diffusion models is resource-intensive. The paper lacks any discussion regarding this, e.g., training time, compute requirements, or efficiency trade-offs compared to simpler augmenters like PoseAug (Zhang et al.). No user or perceptual evaluation is provided for the realism or physical plausibility of generated videos."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762919291042,"tcdate":1761965702401,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission7124/Reviewer_WKNr"],"signatures":["ICLR.cc/2026/Conference/Submission7124/Reviewer_WKNr"],"forum":"Uh8NGna3VE","number":2,"license":"CC BY 4.0","cdate":1761965702401,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission7124/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762919291042,"domain":"ICLR.cc/2026/Conference","replyto":"Uh8NGna3VE","id":"uVxCgENd4K","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["3D Human Pose Estimation","Domain Generalization","Video Generation"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"n/a"},"_bibtex":{"value":"@misc{\nhu2025background,\ntitle={Background Matters: Robust 3D Human Pose Estimation via Controllable Video Generation},\nauthor={Xinhao Hu and Yiyi Zhang and Liqing Zhang and Jianfu Zhang},\nyear={2025},\nurl={https://openreview.net/forum?id=Uh8NGna3VE}\n}"},"title":{"value":"Background Matters: Robust 3D Human Pose Estimation via Controllable Video Generation"},"pdf":{"value":"/pdf/d1f4b0ea2f58c44f0da327365750d39e3772b6ff.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"hu|background_matters_robust_3d_human_pose_estimation_via_controllable_video_generation"},"authorids":{"value":["~Xinhao_Hu1","~Yiyi_Zhang1","~Liqing_Zhang2","~Jianfu_Zhang2"]},"authors":{"value":["Xinhao Hu","Yiyi Zhang","Liqing Zhang","Jianfu Zhang"]}},"version":2},{"content":{"venue":{"value":"ACCV (1) 2014"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-319-16865-4_38.pdf"},"venueid":{"value":"dblp.org/conf/ACCV/2014"},"paperhash":{"value":"bermudezcameo|minimal_solution_for_computing_pairs_of_lines_in_noncentral_cameras"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Jesus_Bermudez-Cameo:","~João_Pedro_Barreto1","https://dblp.org/search/pid/api?q=author:Gonzalo_López-Nicolás:","~José_Jesús_Guerrero1"]},"html":{"value":"https://doi.org/10.1007/978-3-319-16865-4_38"},"_bibtex":{"value":"@inproceedings{DBLP:conf/accv/Bermudez-CameoB14,\n  author={Jesus Bermudez-Cameo and João Pedro Barreto and Gonzalo López-Nicolás and José Jesús Guerrero},\n  title={Minimal Solution for Computing Pairs of Lines in Non-central Cameras},\n  year={2014},\n  cdate={1388534400000},\n  pages={585-597},\n  url={https://doi.org/10.1007/978-3-319-16865-4_38},\n  booktitle={ACCV (1)},\n  crossref={conf/accv/2014-1}\n}\n"},"abstract":{"value":"In non-central cameras, the complete geometry of a 3D line is mapped to each corresponding projection, therefore each line can be theoretically recovered from a single view. However, the solution of this problem is ill-conditioned due to the lack of effective baseline between rays. This limitation prevents from a practical implementation of the approach if lines are not close to the visual system. In this paper, we exploit additional geometric constraints to improve the results of line reconstruction from single images in non-central systems. In particular, we obtain the minimal solution for the case of a pair of intersecting orthogonal lines and for the case of a pair of parallel lines considering three rays from each line. The proposal has been evaluated with simulations and tested with real images."},"title":{"value":"Minimal Solution for Computing Pairs of Lines in Non-central Cameras"},"authors":{"value":["Jesus Bermudez-Cameo","João Pedro Barreto","Gonzalo López-Nicolás","José Jesús Guerrero"]}},"tmdate":1731488030096,"pdate":1388534400000,"tcdate":1729846786040,"writers":["~"],"signatures":["~Joao_Barreto1"],"forum":"CnqWEnuwXr","license":"CC BY-SA 4.0","number":160448,"cdate":1388534400000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1731488030096,"domain":"DBLP.org","id":"CnqWEnuwXr","version":2},{"content":{"venue":{"value":"ACCV (1) 2014"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-319-16865-4_38.pdf"},"venueid":{"value":"dblp.org/conf/ACCV/2014"},"paperhash":{"value":"bermudezcameo|minimal_solution_for_computing_pairs_of_lines_in_noncentral_cameras"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Jesus_Bermudez-Cameo:","~João_Pedro_Barreto1","https://dblp.org/search/pid/api?q=author:Gonzalo_López-Nicolás:","https://dblp.org/search/pid/api?q=author:José_Jesús_Guerrero:"]},"html":{"value":"https://doi.org/10.1007/978-3-319-16865-4_38"},"_bibtex":{"value":"@inproceedings{DBLP:conf/accv/Bermudez-CameoB14,\n  author={Jesus Bermudez-Cameo and João Pedro Barreto and Gonzalo López-Nicolás and José Jesús Guerrero},\n  title={Minimal Solution for Computing Pairs of Lines in Non-central Cameras},\n  year={2014},\n  cdate={1388534400000},\n  pages={585-597},\n  url={https://doi.org/10.1007/978-3-319-16865-4_38},\n  booktitle={ACCV (1)},\n  crossref={conf/accv/2014-1}\n}\n"},"abstract":{"value":"In non-central cameras, the complete geometry of a 3D line is mapped to each corresponding projection, therefore each line can be theoretically recovered from a single view. However, the solution of this problem is ill-conditioned due to the lack of effective baseline between rays. This limitation prevents from a practical implementation of the approach if lines are not close to the visual system. In this paper, we exploit additional geometric constraints to improve the results of line reconstruction from single images in non-central systems. In particular, we obtain the minimal solution for the case of a pair of intersecting orthogonal lines and for the case of a pair of parallel lines considering three rays from each line. The proposal has been evaluated with simulations and tested with real images."},"title":{"value":"Minimal Solution for Computing Pairs of Lines in Non-central Cameras"},"authors":{"value":["Jesus Bermudez-Cameo","João Pedro Barreto","Gonzalo López-Nicolás","José Jesús Guerrero"]}},"tmdate":1729846791831,"pdate":1388534400000,"tcdate":1729846781237,"writers":["~"],"signatures":["~Joao_Barreto1"],"forum":"vedZR4WDc7","license":"CC BY-SA 4.0","number":160425,"cdate":1388534400000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1729846791831,"domain":"DBLP.org","id":"vedZR4WDc7","version":2},{"content":{"venue":{"value":"ACCV (1) 2014"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-319-16865-4_38.pdf"},"venueid":{"value":"dblp.org/conf/ACCV/2014"},"paperhash":{"value":"bermudezcameo|minimal_solution_for_computing_pairs_of_lines_in_noncentral_cameras"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Jesus_Bermudez-Cameo:","~João_Pedro_Barreto1","https://dblp.org/search/pid/api?q=author:Gonzalo_López-Nicolás:","https://dblp.org/search/pid/api?q=author:José_Jesús_Guerrero:"]},"html":{"value":"https://doi.org/10.1007/978-3-319-16865-4_38"},"_bibtex":{"value":"@inproceedings{DBLP:conf/accv/Bermudez-CameoB14,\n  author={Jesus Bermudez-Cameo and João Pedro Barreto and Gonzalo López-Nicolás and José Jesús Guerrero},\n  title={Minimal Solution for Computing Pairs of Lines in Non-central Cameras},\n  year={2014},\n  cdate={1388534400000},\n  pages={585-597},\n  url={https://doi.org/10.1007/978-3-319-16865-4_38},\n  booktitle={ACCV (1)},\n  crossref={conf/accv/2014-1}\n}\n"},"abstract":{"value":"In non-central cameras, the complete geometry of a 3D line is mapped to each corresponding projection, therefore each line can be theoretically recovered from a single view. However, the solution of this problem is ill-conditioned due to the lack of effective baseline between rays. This limitation prevents from a practical implementation of the approach if lines are not close to the visual system. In this paper, we exploit additional geometric constraints to improve the results of line reconstruction from single images in non-central systems. In particular, we obtain the minimal solution for the case of a pair of intersecting orthogonal lines and for the case of a pair of parallel lines considering three rays from each line. The proposal has been evaluated with simulations and tested with real images."},"title":{"value":"Minimal Solution for Computing Pairs of Lines in Non-central Cameras"},"authors":{"value":["Jesus Bermudez-Cameo","João Pedro Barreto","Gonzalo López-Nicolás","José Jesús Guerrero"]}},"tmdate":1729846776618,"pdate":1388534400000,"tcdate":1729846770122,"writers":["~"],"signatures":["~Joao_Barreto1"],"forum":"cgkzjNZ2kU","license":"CC BY-SA 4.0","number":160405,"cdate":1388534400000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1729846776618,"domain":"DBLP.org","id":"cgkzjNZ2kU","version":2},{"content":{"summary":{"value":"This paper proposes RealisMotion, a novel framework for human video generation that achieves fine-grained control by explicitly decomposing the video content into four independent, composable elements: foreground subject, background video, human trajectory, and action patterns. The key idea is to perform motion editing, including trajectory and action, in a ground-aware 3D world coordinate system before fusing the elements via a specialized Video Diffusion Transformer, which is built upon the WAN-2.1 architecture. This approach successfully decouples geometry-sensitive control from appearance and temporal consistency. Experimental results indicate the effectiveness of the proposed method."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"How can the users obtain the control signals and use them to generate plausible results?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1.\tRealisMotion introduces an unparalleled level of explicit, independent control over four fundamental video elements, including subject, background, trajectory, and action.\n2.\tThe framework successfully integrates 3D motion priors with modern video diffusion priors, i.e., WAN-2.1-T2V. Experimental results verify the effectiveness of the proposed method."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1.\tThis reliance on large-scale internal data creates a major hurdle for reproducibility and makes it impossible for the academic community to verify the results or build upon the fine-tuned model. The performance gains might stem more from the scale and quality of this undisclosed training data than from the architectural novelty of RealisMotion itself.\n2.\tThe comparison against existing state-of-the-art methods (e.g., Animate Anyone, MotionCtrl, 3DTrajMaster) is inherently unfair if these methods are trained only on public datasets (e.g., public TikTok, standard motion datasets) while RealisMotion benefits from large-scale, potentially proprietary or industrial-scale internal video data.\n3.\tThe entire 3D motion pipeline is predicated on the accuracy of the GVHMR human mesh recovery method for initial pose, camera parameters, and the estimation of 3D points. It also relies on Depth Pro for depth and focal length estimation. Failures or noise in these heavy upstream models will inevitably compromise the quality and consistency of the 3D control signals, limiting the robustness of the system in real-world, highly challenging scenarios.\n4.\tThe inference process and the inference speed should also be included."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762917075224,"tcdate":1761786389603,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission3870/Reviewer_WgY7"],"signatures":["ICLR.cc/2026/Conference/Submission3870/Reviewer_WgY7"],"forum":"AvW39dAR8R","number":2,"license":"CC BY 4.0","cdate":1761786389603,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission3870/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762917075224,"domain":"ICLR.cc/2026/Conference","replyto":"AvW39dAR8R","id":"Dc3c8i9wrd","forumContent":{"TLDR":{"value":"We propose a decomposed human motion control and video generation framework that explicitly decouples motion from appearance, subject from background, and action from trajectory, enabling flexible mix-and-match composition of these elements."},"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Human Motion Control","Human Video Generation","Video Editing"]},"supplementary_material":{"value":"/attachment/0570422b04c804fa70007dac708e246f1383113e.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Generating human videos with realistic and controllable motions is a challenging task. While existing methods can generate visually compelling videos, they lack separate control over four key video elements: foreground subject, background video, human trajectory, and action patterns. In this paper, we propose a decomposed human motion control and video generation framework that explicitly decouples motion from appearance, subject from background, and action from trajectory, enabling flexible mix-and-match composition of these elements. Concretely, we first build a ground-aware 3D world coordinate system and perform motion editing directly in the 3D space. Trajectory control is implemented by unprojecting edited 2D trajectories into 3D with focal-length calibration and coordinate transformation, followed by speed alignment and orientation adjustment; actions are supplied by a motion bank or generated via text-to-motion methods. Then, based on modern text-to-video diffusion transformer models, we inject the subject as tokens for full attention, concatenate the background along the channel dimension, and add motion (trajectory and action) control signals by addition. Such a design opens up the possibility for us to generate realistic videos of anyone doing anything anywhere. Extensive experiments on benchmark datasets and real-world cases demonstrate that our method achieves state-of-the-art performance on both element-wise controllability and overall video quality. The source codes and project page are in the supplementary and at https://anonymous.4open.science/r/RealisMotion-anonymous-3870/ ."},"_bibtex":{"value":"@misc{\nliang2025realismotion,\ntitle={RealisMotion: Decomposed Human Motion Control and Video Generation in the World Space},\nauthor={Jingyun Liang and Jingkai Zhou and Shikai Li and Chenjie Cao and Lei Sun and Yichen Qian and Weihua Chen and Fan Wang},\nyear={2025},\nurl={https://openreview.net/forum?id=AvW39dAR8R}\n}"},"title":{"value":"RealisMotion: Decomposed Human Motion Control and Video Generation in the World Space"},"pdf":{"value":"/pdf/3e346b93880dc42b1820d10d7b026c7878f9e814.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"liang|realismotion_decomposed_human_motion_control_and_video_generation_in_the_world_space"},"authorids":{"value":["~Jingyun_Liang1","~Jingkai_Zhou2","~Shikai_Li1","~Chenjie_Cao1","~Lei_Sun5","~Yichen_Qian1","~Weihua_Chen1","~Fan_Wang6"]},"authors":{"value":["Jingyun Liang","Jingkai Zhou","Shikai Li","Chenjie Cao","Lei Sun","Yichen Qian","Weihua Chen","Fan Wang"]}},"version":2},{"content":{"summary":{"value":"This paper introduces VidMuse, an autoregressive model for generating music from video. The authors constructed a new large-scale dataset comprising 360k video-music pairs, which combines newly collected data with filtered existing music video data. The data construction and processing methods are described in detail. The model employs an autoregressive transformer decoder that generates music tokens directly, conditioned on long-term and short-term visual embeddings. Experimental results show that training on the new dataset and using LST fusion improves performance on both quantitative and qualitative metrics. The model is compared extensively with state-of-the-art models, and ablation studies are conducted. Demo videos are available on a webpage."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"- Did author consider visual tokens such as VQ-GAN tokens? Or combinations of different types of visual features?\n- What are GFLOPs for state-of-the-art models?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- One of the main contributions is the new constructed dataset. The dataset construction, filtering, preprocessing and usage are well-explained. The ethical statement section is critical and it addresses most of my concerns regarding a new music-video dataset.\n- The long-short-term visual feature module is well-motivated. The paper is generally well-written. It clearly describes each component of VidMuse and explains the module selection with accompanied experimental results.\n- Extensive demo videos are available in the anonymous webpage and supplementary material contains healthy additional details."},"flag_for_ethics_review":{"value":["Yes, Legal compliance (e.g., GDPR, copyright, terms of use)"]},"weaknesses":{"value":"- While music source separation can be used to remove the vocal soundtrack, the final generated music will not contain any vocal music. It is not a wrong choice but human vocal music is missing and obviously it is still playing a critical role in music. It just needs more investigation to generate both background music and reasonable vocal music. Besides, an evaluation/analysis for sound separation is needed. i.e., how well does the music sound separation (demucs) work for the collected data?\n- The audio in new collected video data actually contains more than music. Most movie trailers contain sound effects that are not necessary music. Unfortunately these sound effects will not be removed by music sound separation and it is unclear whether these sound effects will affect the quality.\n- The technical contribution regarding model component is relatively limited. The design of long-short term visual feature fusion is straightforward but in general I am okay with that for an application + dataset paper.\n- An analysis of music genre available in the dataset is missing. This is even more important than the video genre.\n- It seems like a pre-trained Encodec model is used in this work but why not fine-tune or even training a new Encodec model on this new dataset?\n- Although the effort and focus here is to generate music from video input only, it actually makes sense to incorporate additional text for better style control (like V2Meow). Since VidMuse leverages pre-trained MusicGen which is already a text-to-music model, why not maintaining its original capability while incorporating Long-short-term visual embeddings? This will potentially unlock more flexible applications.\n- As authors mentioned already, the ImageBind score is definitely not the best choice for video-music relevance. Even training a video-music contrastive learning model on top of this new dataset would be a better choice.\n- Some minor comments:\n  - Font size in Figure 2 is too small.\n  - Table 4 and Table 5 have bad overlapping\n  - Several videos (e.g., movie trailers) in the demo webpage seem like copyright protected. My impression is it might be ok for paper reviewing stage but please follow the correct guidance and use them carefully."}},"nonreaders":[],"tmdate":1731427364376,"tcdate":1730267586037,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission1045/Reviewer_LKtR"],"signatures":["ICLR.cc/2025/Conference/Submission1045/Reviewer_LKtR"],"forum":"6cGKi7FqJS","number":2,"license":"CC BY 4.0","cdate":1730267586037,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission1045/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427364376,"domain":"ICLR.cc/2025/Conference","replyto":"6cGKi7FqJS","id":"NyOSQyJpFR","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video-to-Music Generation","Transformer"]},"supplementary_material":{"value":"/attachment/9beeb773a63ce531e2fd1badaa250607f286a9c4.zip"},"primary_area":{"value":"generative models"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"In this work, we systematically study music generation conditioned solely on the video. First, we present a large-scale dataset by collecting 360K video-music pairs, including various genres such as movie trailers, advertisements, and documentaries. Furthermore, we propose VidMuse, a simple framework for generating music aligned with video inputs. VidMuse stands out by producing high-fidelity music that is both acoustically and semantically aligned with the video. By incorporating local and global visual cues, VidMuse enables the creation of coherent music tracks that consistently match the video content through Long-Short-Term modeling. Through extensive experiments, VidMuse outperforms existing models in terms of audio quality, diversity, and audio-visual alignment."},"_bibtex":{"value":"@misc{\ntian2024vidmuse,\ntitle={VidMuse: A Simple Video-to-Music Generation Framework with Long-Short-Term Modeling},\nauthor={Zeyue Tian and Zhaoyang Liu and Ruibin Yuan and Jiahao Pan and Qifeng Liu and Xu Tan and Qifeng Chen and Wei Xue and Yike Guo},\nyear={2024},\nurl={https://openreview.net/forum?id=6cGKi7FqJS}\n}"},"title":{"value":"VidMuse: A Simple Video-to-Music Generation Framework with Long-Short-Term Modeling"},"pdf":{"value":"/pdf/16417a40922983f089cc73c28dba6f7dbb0713da.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"tian|vidmuse_a_simple_videotomusic_generation_framework_with_longshortterm_modeling"},"authorids":{"value":["~Zeyue_Tian2","~Zhaoyang_Liu1","~Ruibin_Yuan1","~Jiahao_Pan1","~Qifeng_Liu1","~Xu_Tan1","~Qifeng_Chen1","~Wei_Xue5","~Yike_Guo1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Zeyue Tian","Zhaoyang Liu","Ruibin Yuan","Jiahao Pan","Qifeng Liu","Xu Tan","Qifeng Chen","Wei Xue","Yike Guo"]}},"version":2},{"content":{"summary":{"value":"The paper targets two frictions in Slot Attention (SA)–based object-centric learning (OCL): (1) cold-start queries on the first image/video frame, and (2) homogeneous aggregation transforms applied identically across all video frames despite their differing query conditions. The authors propose SmoothSA with two simple modifications: a tiny preheater that “preheats” first-frame queries using input features via self-distillation; and a differentiated recurrence policy using full three SA iterations on the first frame and a single iteration on non-first frames. Experiments on image (CLEVRtex, COCO, VOC) and video (YTVIS, VideoSAUR) OCL, plus downstream object recognition and VQA, show consistent gains over strong baselines such as SPOT and SlotContrast."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"The main questions are already reflected in the Weaknesses section above, concerning (1) the discrepancy between reported and original baseline performances, (2) the unclear interpretation of evaluation metrics across image and video datasets, and (3) the lack of clarity regarding training/testing segment lengths and sequence length generalization."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"Clear identification of two practical bottlenecks in SA pipelines (first-frame cold start; transform homogeneity across frames), with a minimal, easy-to-adopt solution. \n\nMethod simplicity: the preheater is a light Transformer decoder–style module trained via self-distillation inside the OCL model; the recurrence rule (3 iterations on the first frame, 1 on others) is plug-and-play. \n\nBroad empirical coverage: improvements reported on synthetic and real-world image datasets (ARI/ARI_fg/mIoU/mBO) and video datasets, with qualitative masks and downstream boosts in object recognition and VQA (GQA, CLEVRER). \n\nPositioning vs prior work is knowledgeable (e.g., BO-QSA, MetaSlot, SAVi/++, SlotContrast, STATM, SlotPi, RandSF.Q). The paper argues SmoothSA addresses issues orthogonal to prior query-initialization or query-prediction lines."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Unclear baseline performances:\nThe reported results for VideoSAUR and SlotContrast are substantially higher than those in the original papers, particularly in terms of FG-ARI. This discrepancy raises questions about the experimental comparability—whether the authors re-implemented these methods under different settings, used stronger backbones, or applied additional training tricks. A clear explanation is needed to ensure that the reported improvements are attributable to the proposed method rather than to differences in baseline implementations.\n\n2. Inconsistent evaluation metrics:\nThe paper evaluates SmoothSA on both image and video datasets but seems to use  same metrics for them. In image datasets, the image FG-ARI measures spatial clustering quality, while in video datasets the video FG-ARI captures both spatial segmentation and temporal consistency. However, the paper interprets these results uniformly, without clarifying the differences. This makes it difficult to fairly assess the improvements and understand whether the gains arise from better object segmentation, temporal stability, or both.\n\n3. Unclear experimental setup and sequence length generalization.\nThe paper does not specify the segment lengths used during training and testing. In video object-centric learning, training typically uses short clips while testing involves longer sequences—making sequence length generalization a central challenge. The absence of this information obscures how SmoothSA performs under distribution shifts in sequence length, and whether its recurrence design truly enhances temporal robustness."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762920992485,"tcdate":1762574881376,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission9379/Reviewer_FgBa"],"signatures":["ICLR.cc/2026/Conference/Submission9379/Reviewer_FgBa"],"forum":"dTcUXNfz2o","number":3,"license":"CC BY 4.0","cdate":1762574881376,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission9379/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762920992485,"domain":"ICLR.cc/2026/Conference","replyto":"dTcUXNfz2o","id":"ixMFGoE75m","forumContent":{"TLDR":{"value":"We address issues of query cold-start in slot attention iterations on the image or video's first frame and transform homogeneity in slot attention recurrences on the video frames, improving image and video OCL significantly."},"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Object-Centric Learning","Slot Attention","Object Discovery","Object Recognition","Dynamics Modeling"]},"supplementary_material":{"value":"/attachment/659acb9ee7e5401f5dbf0fa533df8a74eeeb4245.zip"},"primary_area":{"value":"unsupervised, self-supervised, semi-supervised, and supervised representation learning"},"abstract":{"value":"Slot Attention (SA) and its variants lie at the heart of mainstream Object-Centric Learning (OCL).\nObjects in an image can be aggregated into corresponding slot vectors, by \\textit{iteratively} refining cold-start query vectors, typically three times, via SA on image features.\nFor video, this aggregation is \\textit{recurrently} shared across frames, with queries cold-started on the first frame while transitioned from the previous frame's slots on non-first frames.\nHowever, cold-start queries lack sample-specific cues thus hindering precise aggregation on an image or a video's first frame;\nAlso, non-first frames' queries are already sample-specific thus requiring aggregation transforms different from the first frame.\nWe address these issues for the first time with our \\textit{SmoothSA}:\n(1) To smooth SA iterations on the image or video's first frame, we \\textit{preheat} the cold-start queries with rich information of input features, via a tiny module self-distilled inside OCL;\n(2) To smooth SA recurrences across all video frames, we \\textit{differentiate} the homogeneous transforms on the first and non-first frames, by using full and single iterations respectively.\nComprehensive experiments on object discovery, recognition and downstream benchmarks validate our method's effectiveness.\nFurther analyses illuminate how our method smooths SA iterations and recurrences.\nOur source code and training logs are provided in the supplement."},"_bibtex":{"value":"@misc{\nzhao2026smoothing,\ntitle={Smoothing Slot Attention Iterations and Recurrences},\nauthor={Rongzhen Zhao and Wenyan Yang and Juho Kannala and Joni Pajarinen},\nyear={2026},\nurl={https://openreview.net/forum?id=dTcUXNfz2o}\n}"},"title":{"value":"Smoothing Slot Attention Iterations and Recurrences"},"pdf":{"value":"/pdf/ba6cd9b279818a0e4cd8b9aafa6df8c0cf6b8c84.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"zhao|smoothing_slot_attention_iterations_and_recurrences"},"authorids":{"value":["~Rongzhen_Zhao2","~Wenyan_Yang1","~Juho_Kannala5","~Joni_Pajarinen2"]},"authors":{"value":["Rongzhen Zhao","Wenyan Yang","Juho Kannala","Joni Pajarinen"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a video generation pipeline based on MoVQ video decoding scheme. It consists of two stages: keyframes synthesis and video frame interpolation. It also compares two temporal conditioning approaches and different configurations of MoVQ-based models."},"presentation":{"value":"2 fair"},"contribution":{"value":"2 fair"},"soundness":{"value":"2 fair"},"strengths":{"value":"The contributions of this method are summarised as below: \n\n1. an end-to-end text-to-video latent diffusion pipeline that consists of key frames generation and frame interpolation\n\n2. separate temporal blocks for temporal modelling\n\n3. temporal output masking and data augmentations for robust VFI\n\n4. investigation of video decoders"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"My concern is the lack of novelty. The concert comments are below:\n\n1. This pipeline is not new. Various methods have tried to utilize text2image diffusion models for video generation. For example, the proposed temporal conditional scheme could be found at ``AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning\n''.\n\n2. Concatenating key frames along the channel dimension in video frame interpolation part is similar to \"IMAGEN VIDEO: HIGH DEFINITION VIDEO\nGENERATION WITH DIFFUSION MODELS\". \n\n3. The introduction of temporal layers is similar to \"Align your Latents: High-Resolution Video Synthesis with Latent Diffusion Models\". \n\n4. What does \"efficient\" in the title mean?"},"confidence":{"value":"5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully."},"questions":{"value":"See weakness."},"rating":{"value":"5: marginally below the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699637152865,"tcdate":1697371568247,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission9160/Reviewer_ubtc"],"signatures":["ICLR.cc/2024/Conference/Submission9160/Reviewer_ubtc"],"forum":"JBLgjRRuHG","number":1,"license":"CC BY 4.0","cdate":1697371568247,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission9160/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699637152865,"domain":"ICLR.cc/2024/Conference","replyto":"JBLgjRRuHG","id":"P1SgJ3d8n4","forumContent":{"venue":{"value":"ICLR 2024 Conference Withdrawn Submission"},"TLDR":{"value":"In this paper we propose a two-stage latent diffusion video generation architecture and a new MoVQ video decoding scheme. We also conduct experimental research to compare temporal blocks and temporal layers for architecture design."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["text-to-video","video generation","temporal consistency","frames interpolation","Inception Score","CLIPSIM","MoVQ video decoder"]},"supplementary_material":{"value":"/attachment/85e824fa7cd2f57c7ab32187b44fd2b90363c8a2.zip"},"primary_area":{"value":"generative models"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Multimedia generation approaches occupy a prominent place in artificial intelligence research. Text-to-image models achieved high quality results over the last years, however video synthesis methods recently started to develop. In this paper we present a new two-stage latent diffusion video generation architecture using a new MoVQ video decoding scheme. The first stage concerns keyframes synthesis, while the second one is devoted to interpolated frames generation. We compare two temporal conditioning approaches during evaluation and show the improvement of using temporal blocks over temporal layers in terms of IS and CLIPSIM metrics reflecting video generation quality aspects. We also evaluate different configurations of MoVQ-based video decoding scheme to achieve higher PSNR, SSIM, MSE and LPIPS scores. Finally, we compare our pipeline with existing solutions and achieve top-3 CLIPSIM metric score (0.2976)."},"_bibtex":{"value":"@misc{\nvladimir2024efficient,\ntitle={Efficient architectural aspects for text-to-video generation pipeline},\nauthor={Arkhipkin Sergeevich Vladimir and Zein Shaheen and Viacheslav Vasilev and Denis Valerievich Dimitrov and Andrey Kuznetsov},\nyear={2024},\nurl={https://openreview.net/forum?id=JBLgjRRuHG}\n}"},"title":{"value":"Efficient architectural aspects for text-to-video generation pipeline"},"pdf":{"value":"/pdf/de7dea6dc2c02c9d3dc83b1f407bdf5087d624b7.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Withdrawn_Submission"},"paperhash":{"value":"vladimir|efficient_architectural_aspects_for_texttovideo_generation_pipeline"},"authorids":{"value":["~Arkhipkin_Sergeevich_Vladimir1","~Zein_Shaheen1","~Viacheslav_Vasilev1","~Denis_Valerievich_Dimitrov1","~Andrey_Kuznetsov2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Arkhipkin Sergeevich Vladimir","Zein Shaheen","Viacheslav Vasilev","Denis Valerievich Dimitrov","Andrey Kuznetsov"]}},"version":2},{"content":{"summary":{"value":"The paper proposes AutoL2S, a training and inference framework that pairs long and short chain-of-thought (CoT) traces and introduces an <EASY> token so the model can decide when concise reasoning suffices. It formalizes a per-input decision rule that trades off predicted divergence between short and long answers against token cost, and extends the approach with AutoL2S-Plus, a length-aware fine-tuning stage that further calibrates expected reasoning length using AutoL2S as the reference policy. Empirically, across multiple reasoning benchmarks and two base LLMs, AutoL2S shortens outputs while preserving accuracy, with AutoL2S reporting up to about 57 percent reduction and AutoL2S-Plus reaching up to about 70 percent."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. Novelty and positioning. Which components are genuinely new relative to prior length-control and CoT-compression methods with gating or pruning? Please provide a table that maps each component in your method to the closest prior and clarifies the delta in assumptions and training signals. \n\n2. Core mechanism isolation. If you hold data, sampling, and training schedules fixed, how much of the gain comes from the <EASY> gate versus other steps? Please add ablations that toggle only the gate, only the long/short mixture, and only the length-aware finetuning.\n\n3. Failure modes. What happens when examples are mis-gated as “easy” but need long reasoning? Show targeted stress tests, qualitative error analyses, and accuracy drop conditioned on mis-gating.\n\n4. Theoretical guidance. Your analysis describes accuracy–cost tradeoffs but does not yield actionable prescriptions. Can you instantiate it to produce quantitative thresholds or schedules that match your best empirical settings?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"Originality lies in a clean recipe that pairs long and short CoT supervision with an explicit gating signal and a length-aware objective, giving a practical, model-agnostic way to control reasoning length. Quality shows as strong accuracy at substantially reduced tokens across several benchmarks, with clear training and inference procedures. Significance is high for latency and cost reduction in real deployments. Main weakness: novelty and contribution boundaries are not crisply isolated, with missing controlled ablations."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"While the paper presents a neat engineering recipe, the conceptual novelty feels limited: the core ideas of mixing long and short CoT traces, learning an “easy” gate, and adding a length-aware fine-tuning step closely echo prior work on CoT compression and length control. \n\nThe theoretical results are largely tautological reformulations of standard information-theoretic inequalities and risk trade-offs, and they do not yield actionable guidance for thresholds, divergence estimators, or training schedules that would change practice. \n\nEmpirically, the evaluation skews to math and physics with small backbones and teacher signals drawn from specific models, which raises concerns about robustness, domain generalization, and potential teacher or benchmark contamination; critical stress tests are missing, for example failure analysis when “easy” is mispredicted, calibration plots of the gate versus difficulty, distribution-shift tests, and cost-accuracy Pareto curves that include wall-clock and KV-cache metrics. \n\nSome reported results are hard to interpret or incomplete, e.g., masked AIME numbers and varying rejection-sampling settings without a principled selection rule, and comparisons occasionally lack strong recent baselines that optimize thinking length via other mechanisms."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762923245108,"tcdate":1762066308183,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission12320/Reviewer_gB3p"],"signatures":["ICLR.cc/2026/Conference/Submission12320/Reviewer_gB3p"],"forum":"LUttHOTlYz","number":4,"license":"CC BY 4.0","cdate":1762066308183,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission12320/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762923245108,"domain":"ICLR.cc/2026/Conference","replyto":"LUttHOTlYz","id":"ghVmfO3tVK","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Efficient Reasoning","LLM","Reasoning Models","Overthinking"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"The reasoning-capable large language models (LLMs) demonstrate strong performance in complex reasoning tasks but often suffer from overthinking issues after distillation, generating unnecessarily long chain-of-thought (CoT) reasoning paths for easy reasoning questions, thereby increasing inference cost and latency.\n  Recent work largely applies reinforcement learning to shorten reasoning paths in models that already possess reasoning capability. However, these approaches generalize poorly to non-reasoning LLMs, as they assume initial reasoning ability and rely on sparse, outcome-based rewards that make optimization unstable and limit effective learning.\n  In this paper, we propose Auto Long-Short Reasoning (AutoL2S), a dynamic and model-agnostic framework that enables LLMs to adaptively adjust reasoning length according to input complexity, while specifically targeting the stage of transferring non-reasoning LLMs into reasoning-capable but efficient ones via distillation.\n  AutoL2S introduces a learned mechanism in which LLMs are trained on data annotated with long and short CoT paths, together with a special \\<EASY\\> token that signals when long reasoning can be skipped.\n  During inference, the \\<EASY\\> token can indicate when the model can skip generating lengthy CoT reasoning.\n  Furthermore, we extend our framework with AutoL2S-Plus, which employs the AutoL2S as a reference model in a length-aware fine-tuning objective to calibrate expected reasoning length, enabling further efficiency gains without loss of accuracy.\n  We theoretically and empirically find that the joint training of long and short CoT paths not only enables dynamic reasoning but also helps the training of shorter CoT generation through knowledge transfer from longer CoT paths.\nAutoL2S reduces reasoning length by up to 70\\% without sacrificing performance, establishing it as an effective framework for scalable and efficient LLM reasoning."},"_bibtex":{"value":"@misc{\nluo2025autols,\ntitle={AutoL2S: Auto Long-Short Reasoning for Efficient Large Language Models},\nauthor={Feng Luo and Yu-Neng Chuang and Guanchu Wang and Hoang Anh Duy Le and Shaochen Zhong and Hongyi Liu and Jiayi Yuan and Yang Sui and Vladimir Braverman and Vipin Chaudhary and Xia Hu},\nyear={2025},\nurl={https://openreview.net/forum?id=LUttHOTlYz}\n}"},"title":{"value":"AutoL2S: Auto Long-Short Reasoning for Efficient Large Language Models"},"pdf":{"value":"/pdf/4f5ffa1d1a5809e7b2f9eaf4784bba9f55899f16.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"luo|autol2s_auto_longshort_reasoning_for_efficient_large_language_models"},"authorids":{"value":["~Feng_Luo4","~Yu-Neng_Chuang1","~Guanchu_Wang1","~Hoang_Anh_Duy_Le1","~Shaochen_Zhong1","~Hongyi_Liu5","~Jiayi_Yuan1","~Yang_Sui1","~Vladimir_Braverman1","~Vipin_Chaudhary2","~Xia_Hu4"]},"authors":{"value":["Feng Luo","Yu-Neng Chuang","Guanchu Wang","Hoang Anh Duy Le","Shaochen Zhong","Hongyi Liu","Jiayi Yuan","Yang Sui","Vladimir Braverman","Vipin Chaudhary","Xia Hu"]}},"version":2},{"content":{"summary":{"value":"The paper proposes ReWatch-R1, a novel approach to improve complex video reasoning in LVLMs by addressing the critical data bottleneck in existing methods. The authors introduce ReWatch, a large-scale dataset synthesized via a multi-stage agentic pipeline that includes temporally dense captions, high-difficulty multi-hop QA pairs, and video-grounded CoT traces generated through a Multi-Agent ReAct framework simulating human-like “re-watching.” They further develop an Observation & Reasoning (O&R) reward mechanism for RL, which jointly evaluates answer correctness and the factual grounding of intermediate reasoning steps."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1. Line 207: Why does the original set of 85K QA pairs yield over 170K multiple-choice QA pairs?  \n\n2. From the Chain-of-Thought (CoT) example shown in the bottom-right corner of Figure 2, is retrieving explicit timestamps truly necessary? Could the reasoning path be simplified—for instance, as follows:  \n> *\"... So, I’ll \\<action\\> retrieve segments focusing on the man with blonde, curly hair on a jet ski interacting with a passenger \\</action\\>. \\<observation\\> The man on the jet ski passes a sandwich to the passenger, who then takes a bite. \\</observation\\> This directly answers the question. The food item passed was a sandwich...\"*"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. Novel QA curation methods: ReWatch is carefully designed to enforce video dependency and multi-step reasoning through contrastive QA generation and rigorous filtering, effectively eliminating textual shortcuts and hallucination-prone supervision.\n\n2. Innovative O&R reward mechanism: By evaluating both final answers and the fidelity of intermediate observations and reasoning steps, the O&R reward explicitly discourages hallucination and promotes evidence-based reasoning."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Formula 4 merges captions across each time interval independently, which leads to a loss of referential consistency. For example, the caption for 0–10s might be “A man…”, while that for 10–20s is again “A man…”, even though it refers to the same individual as in the earlier segment. This inconsistency compromises the overall quality of the generated captions.\n\n2. The observation mechanism introduces inference latency. Despite yielding performance gains on certain video reasoning benchmarks, it provides only marginal improvements on most general video understanding benchmarks.\n\n3. The paper lacks comparisons with more recent and advanced baselines, such as VersaVid-R1 and GRPO-CARE.\n\n4. Quantitative analysis is missing: the paper does not present any concrete inference results or case studies from ReWatch-R1."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762927835103,"tcdate":1760585547208,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission18045/Reviewer_w6PG"],"signatures":["ICLR.cc/2026/Conference/Submission18045/Reviewer_w6PG"],"forum":"xindJJLSr1","number":1,"license":"CC BY 4.0","cdate":1760585547208,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission18045/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762927835103,"domain":"ICLR.cc/2026/Conference","replyto":"xindJJLSr1","id":"Aj5J9x9qxo","forumContent":{"TLDR":{"value":"We introduce an agent-based pipeline to synthesize a high-quality video reasoning dataset (ReWatch) and a novel reinforcement learning reward (O&R) to train LVLMs, achieving state-of-the-art performance."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Video Reasoning","Large Vision-Language Models (LVLMs)","Agentic Data Synthesis","Multi-Agent ReAct","Reinforcement Learning with Verifiable Reward (RLVR)","Chain-of-Thought (CoT)"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"While Reinforcement Learning with Verifiable Reward (RLVR) significantly advances image reasoning in Large Vision-Language Models (LVLMs), its application to complex video reasoning remains underdeveloped. This gap stems primarily from a critical data bottleneck: existing datasets lack the challenging, multi-hop questions and high-quality, video-grounded Chain-of-Thought (CoT) data necessary to effectively bootstrap RLVR. To address this, we introduce ReWatch, a large-scale dataset built to foster advanced video reasoning. We propose a novel multi-stage synthesis pipeline to synthesize its three components: ReWatch-Caption, ReWatch-QA, and ReWatch-CoT. A core innovation is our Multi-Agent ReAct framework for CoT synthesis, which simulates a human-like \"re-watching\" process to generate video-grounded reasoning traces by explicitly modeling information retrieval and verification. Building on this dataset, we develop ReWatch-R1 by post-training a strong baseline LVLM with Supervised Fine-Tuning (SFT) and our RLVR framework. This framework incorporates a novel Observation \\& Reasoning (O\\&R) reward mechanism that evaluates both the final answer's correctness and the reasoning's alignment with video content, directly penalizing hallucination. Our experiments show that ReWatch-R1 achieves state-of-the-art average performance on five challenging video reasoning benchmarks."},"_bibtex":{"value":"@inproceedings{\nzhang2026rewatchr,\ntitle={ReWatch-R1: Boosting Complex Video Reasoning in Large Vision-Language Models through Agentic Data Synthesis},\nauthor={Congzhi Zhang and Zhibin Wang and Yinchao Ma and Jiawei Peng and Yihan Wang and Qiang Zhou and Jun Song and Bo Zheng},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=xindJJLSr1}\n}"},"title":{"value":"ReWatch-R1: Boosting Complex Video Reasoning in Large Vision-Language Models through Agentic Data Synthesis"},"pdf":{"value":"/pdf/188764f9319a82af607db76b20fa18a29c9fcf00.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"zhang|rewatchr1_boosting_complex_video_reasoning_in_large_visionlanguage_models_through_agentic_data_synthesis"},"authorids":{"value":["~Congzhi_Zhang1","~Zhibin_Wang1","~Yinchao_Ma2","~Jiawei_Peng2","~Yihan_Wang29","~Qiang_Zhou8","~Jun_Song5","~Bo_Zheng5"]},"authors":{"value":["Congzhi Zhang","Zhibin Wang","Yinchao Ma","Jiawei Peng","Yihan Wang","Qiang Zhou","Jun Song","Bo Zheng"]}},"version":2},{"content":{"venue":{"value":"CoRR 2023"},"pdf":{"value":"http://arxiv.org/pdf/2310.12051v1"},"venueid":{"value":"dblp.org/journals/CORR/2023"},"paperhash":{"value":"williams|simpler_and_higher_lower_bounds_for_shortcut_sets"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Virginia_Vassilevska_Williams:","~Yinzhan_Xu2","https://dblp.org/search/pid/api?q=author:Zixuan_Xu:"]},"html":{"value":"https://doi.org/10.48550/arXiv.2310.12051"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2310-12051,\n  publtype={informal},\n  author={Virginia Vassilevska Williams and Yinzhan Xu and Zixuan Xu},\n  title={Simpler and Higher Lower Bounds for Shortcut Sets},\n  year={2023},\n  cdate={1672531200000},\n  journal={CoRR},\n  volume={abs/2310.12051},\n  url={https://doi.org/10.48550/arXiv.2310.12051}\n}\n"},"abstract":{"value":"We provide a variety of lower bounds for the well-known shortcut set problem: how much can one decrease the diameter of a directed graph on $n$ vertices and $m$ edges by adding $O(n)$ or $O(m)$ of shortcuts from the transitive closure of the graph. Our results are based on a vast simplification of the recent construction of Bodwin and Hoppenworth [FOCS 2023] which was used to show an $\\widetilde{\\Omega}(n^{1/4})$ lower bound for the $O(n)$-sized shortcut set problem. We highlight that our simplification completely removes the use of the convex sets by B\\'ar\\'any and Larman [Math. Ann. 1998] used in all previous lower bound constructions. Our simplification also removes the need for randomness and further removes some log factors. This allows us to generalize the construction to higher dimensions, which in turn can be used to show the following results. For $O(m)$-sized shortcut sets, we show an $\\Omega(n^{1/5})$ lower bound, improving on the previous best $\\Omega(n^{1/8})$ lower bound. For all $\\varepsilon > 0$, we show that there exists a $\\delta > 0$ such that there are $n$-vertex $O(n)$-edge graphs $G$ where adding any shortcut set of size $O(n^{2-\\varepsilon})$ keeps the diameter of $G$ at $\\Omega(n^\\delta)$. This improves the sparsity of the constructed graph compared to a known similar result by Hesse [SODA 2003]. We also consider the sourcewise setting for shortcut sets: given a graph $G=(V,E)$, a set $S\\subseteq V$, how much can we decrease the sourcewise diameter of $G$, $\\max_{(s, v) \\in S \\times V, \\text{dist}(s, v) < \\infty} \\text{dist}(s,v)$ by adding a set of edges $H$ from the transitive closure of $G$? We show that for any integer $d \\ge 2$, there exists a graph $G=(V, E)$ on $n$ vertices and $S \\subseteq V$ with $|S| = \\widetilde{\\Theta}(n^{3/(d+3)})$, such that when adding $O(n)$ or $O(m)$ shortcuts, the sourcewise diameter is $\\widetilde{\\Omega}(|S|^{1/3})$."},"title":{"value":"Simpler and Higher Lower Bounds for Shortcut Sets"},"authors":{"value":["Virginia Vassilevska Williams","Yinzhan Xu","Zixuan Xu"]}},"tmdate":1738178310715,"pdate":1672531200000,"tcdate":1738178294601,"writers":["~"],"signatures":["~Yinzhan_Xu2"],"forum":"whrZJgBX23","license":"CC BY-SA 4.0","number":287027,"cdate":1672531200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1738178310715,"domain":"DBLP.org","id":"whrZJgBX23","version":2},{"content":{"venue":{"value":"CoRR 2024"},"pdf":{"value":"https://arxiv.org/pdf/2411.06776v1"},"venueid":{"value":"dblp.org/journals/CORR/2024"},"paperhash":{"value":"dremin|machine_visionaware_quality_metrics_for_compressed_image_and_video_assessment"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Mikhail_Dremin:","~Konstantin_Kozhemyakov1","https://dblp.org/search/pid/api?q=author:Ivan_Molodetskikh:","https://dblp.org/search/pid/api?q=author:Malakhov_Kirill:","https://dblp.org/search/pid/api?q=author:Artur_Sagitov:","https://dblp.org/search/pid/api?q=author:Dmitriy_S._Vatolin:"]},"html":{"value":"https://doi.org/10.48550/arXiv.2411.06776"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2411-06776,\n  publtype={informal},\n  author={Mikhail Dremin and Konstantin Kozhemyakov and Ivan Molodetskikh and Malakhov Kirill and Artur Sagitov and Dmitriy S. Vatolin},\n  title={Machine vision-aware quality metrics for compressed image and video assessment},\n  year={2024},\n  cdate={1704067200000},\n  journal={CoRR},\n  volume={abs/2411.06776},\n  url={https://doi.org/10.48550/arXiv.2411.06776}\n}\n"},"abstract":{"value":"A main goal in developing video-compression algorithms is to enhance human-perceived visual quality while maintaining file size. But modern video-analysis efforts such as detection and recognition, which are integral to video surveillance and autonomous vehicles, involve so much data that they necessitate machine-vision processing with minimal human intervention. In such cases, the video codec must be optimized for machine vision. This paper explores the effects of compression on detection and recognition algorithms (objects, faces, and license plates) and introduces novel full-reference image/video-quality metrics for each task, tailored to machine vision. Experimental results indicate our proposed metrics correlate better with the machine-vision results for the respective tasks than do existing image/video-quality metrics."},"title":{"value":"Machine vision-aware quality metrics for compressed image and video assessment"},"authors":{"value":["Mikhail Dremin","Konstantin Kozhemyakov","Ivan Molodetskikh","Malakhov Kirill","Artur Sagitov","Dmitriy S. Vatolin"]}},"tmdate":1771835861799,"pdate":1735603200000,"externalIds":["dblp:journals/corr/abs-2411-06776"],"tcdate":1771835858288,"writers":["~"],"signatures":["~Konstantin_Kozhemyakov1"],"forum":"LsjqWjzknf","license":"CC BY-SA 4.0","number":826596,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1771835861799,"domain":"DBLP.org","id":"LsjqWjzknf","version":2},{"content":{"venue":{"value":"ICPR (23) 2024"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-031-78347-0_18.pdf"},"venueid":{"value":"dblp.org/conf/ICPR/2024"},"paperhash":{"value":"dremin|machine_visionaware_quality_metrics_for_compressed_image_and_video_assessment"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Mikhail_Dremin:","~Konstantin_Kozhemyakov1","https://dblp.org/search/pid/api?q=author:Ivan_Molodetskikh:","https://dblp.org/search/pid/api?q=author:Malakhov_Kirill:","https://dblp.org/search/pid/api?q=author:Sagitov_Artur:","https://dblp.org/search/pid/api?q=author:Dmitriy_S._Vatolin:"]},"html":{"value":"https://doi.org/10.1007/978-3-031-78347-0_18"},"_bibtex":{"value":"@inproceedings{DBLP:conf/icpr/DreminKMKAV24,\n  author={Mikhail Dremin and Konstantin Kozhemyakov and Ivan Molodetskikh and Malakhov Kirill and Sagitov Artur and Dmitriy S. Vatolin},\n  title={Machine Vision-Aware Quality Metrics for Compressed Image and Video Assessment},\n  year={2024},\n  cdate={1704067200000},\n  pages={266-282},\n  url={https://doi.org/10.1007/978-3-031-78347-0_18},\n  booktitle={ICPR (23)},\n  crossref={conf/icpr/2024-23}\n}\n"},"abstract":{"value":"A main goal in developing video-compression algorithms is to enhance human-perceived visual quality while maintaining file size. But modern video-analysis efforts such as detection and recognition, which are integral to video surveillance and autonomous vehicles, involve so much data that they necessitate machine-vision processing with minimal human intervention. In such cases, the video codec must be optimized for machine vision. This paper explores the effects of compression on detection and recognition algorithms (objects, faces, and license plates) and introduces novel full-reference image/video-quality metrics for each task, tailored to machine vision. Experimental results indicate our proposed metrics correlate better with the machine-vision results for the respective tasks than do existing image/video-quality metrics."},"title":{"value":"Machine Vision-Aware Quality Metrics for Compressed Image and Video Assessment"},"authors":{"value":["Mikhail Dremin","Konstantin Kozhemyakov","Ivan Molodetskikh","Malakhov Kirill","Sagitov Artur","Dmitriy S. Vatolin"]}},"tmdate":1771835861186,"pdate":1735603200000,"externalIds":["dblp:conf/icpr/DreminKMKAV24"],"tcdate":1771835858203,"writers":["~"],"signatures":["~Konstantin_Kozhemyakov1"],"forum":"EADUV9UFIw","license":"CC BY-SA 4.0","number":826595,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1771835861186,"domain":"DBLP.org","id":"EADUV9UFIw","version":2},{"content":{"summary":{"value":"This paper proposes TFMAudio, a text-to-audio generation framework that integrates Time-Frequency Mamba (TFM) and Energy-Aware Guidance (EAG). The model achieves linear-time long-form generation and maintains temporal and spectral consistency, producing up to 30-minute high-fidelity 44.1 kHz audio with strong semantic alignment."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"**1. On the Effectiveness of Energy-Aware Guidance (EAG)**\n\nThe ablation results suggest that EAG provides only marginal gains across objective metrics.\nCould the authors elaborate on specific scenarios or qualitative aspects where EAG meaningfully contributes to stability or audio fidelity?\n\n\n**2. On Length Controllability During Generation**\n\nCan TFMAudio support arbitrary-length generation (e.g., 13s or 27s) beyond fixed training configurations?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"**1. Effective Dual-Axis Modeling via Time-Frequency Mamba**\n\nThe proposed Time-Frequency Mamba performs 1D scans along both temporal and frequency axes, allowing the model to jointly capture temporal causality and spectral correlation — a capability that conventional Transformers struggle to achieve.\n\n**2. Linear-Time Complexity for Long Sequences**\n\nThe Mamba-based recurrent formulation enables linear computational complexity O(L) with respect to sequence length, offering a far more efficient alternative to the quadratic O(L²) cost of Transformers while maintaining long-range dependencies.\n\n**3. Stable and Scalable Audio Generation with Energy-Aware Guidance**\n\nThe introduced Energy-Aware Guidance (EAG) mitigates state drift by decomposing the flow-matching velocity field and adaptively damping unstable components, enabling reliable ultra-long (30-minute) 44.1 kHz audio generation with temporal consistency."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"**1. Marginal Impact of Energy-Aware Guidance (EAG)**\n\nAccording to the ablation results, the performance improvement from EAG is minimal, with only slight differences in objective metrics.\n\n**2. Lack of Flexible Length Control in Generation**\n\nWhile the paper demonstrates 10s and 30s generations, it is unclear whether the model allows arbitrary-length synthesis (e.g., 13s or 27s) or only supports pre-defined durations tied to training configurations."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762918253226,"tcdate":1761552129837,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission5772/Reviewer_xDov"],"signatures":["ICLR.cc/2026/Conference/Submission5772/Reviewer_xDov"],"forum":"WFgtMF6vk6","number":2,"license":"CC BY 4.0","cdate":1761552129837,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission5772/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762918253226,"domain":"ICLR.cc/2026/Conference","replyto":"WFgtMF6vk6","id":"v2ko9AABB5","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Text-to-Audio","Flow-Matching","Mamba","SSM","Diffusion"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Recent advancements in audio generation have been dominated by transformer-based diffusion models, which face challenges in extrapolating positional encodings and exhibit quadratic complexity in self-attention, limiting their consistency and efficiency for long-form generation.\nTo address these limitations, we propose TFMAudio, a novel latent audio generation model that integrates the strengths of Flow Matching and a custom-designed TFMamba backbone.\nTFMamba employs a dual-scan mechanism: TimeMamba captures long-range causal dependencies with linear complexity, while FrequencyMamba models spectral correlations such as harmonic structures. To enhance stability, we further introduce Energy-Aware Guidance (EAG), which mitigates state drift by adaptively regularizing classifier-free guidance. Experiments demonstrate that TFMAudio achieves state-of-the-art performance on text-to-audio benchmarks and exhibits robust extrapolation to ultra-long sequences. Remarkably, our model generates 30-minute high-fidelity audio while preserving temporal consistency and semantic alignment, significantly advancing the scalability and usability of text-to-audio models.\n  Demo:https://huggingface.co/spaces/tfmaudio/TFMAudio"},"_bibtex":{"value":"@misc{\ndai2026tfmaudio,\ntitle={{TFMA}udio: High-Fidelity Long-Form Text-to-Audio via Mamba-based Flow Matching},\nauthor={Hao Dai and Jagmohan Chauhan},\nyear={2026},\nurl={https://openreview.net/forum?id=WFgtMF6vk6}\n}"},"title":{"value":"TFMAudio: High-Fidelity Long-Form Text-to-Audio via Mamba-based Flow Matching"},"pdf":{"value":"/pdf/5833cc8c9f598257ed7359368750dd201f830de2.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"dai|tfmaudio_highfidelity_longform_texttoaudio_via_mambabased_flow_matching"},"authorids":{"value":["~Hao_Dai1","~Jagmohan_Chauhan1"]},"authors":{"value":["Hao Dai","Jagmohan Chauhan"]}},"version":2},{"content":{"summary":{"value":"This work proposes a video diffusion model to unifies the video generation and video recognition tasks. The proposed GenRec, during training, jointly optimize the generation objective and the classification objective. During inference, GenRec can deal with both the video generation conditioned on frames or classes, and video classification task. The experiments demonstrate that GenRec shows superior performance in both tasks."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"Pleaser refer to the weakness part. I would like to modify my rating according to the responses from the authors."},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"- The proposed GenRec unifies the video generation and video recognition tasks.\n- The experimental results validate the effectiveness of the proposed unified GenRec."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Considering that previous works [8,37,46] have already demonstrated that diffusion models have the capability of extracting sufficient features for understanding tasks, the combination of generation and recognition tasks to boost both tasks is too straight-forward, limiting the novelty of the paper.\n- It is unclear for the inference stage for video recognition. Like why the noised $\\tilde{z}$ is still incorporated for classification, which may lead to the recognition result undeterministic. During inference, is the multi-step denoising procedure sill required?\n- There is the lack of training/inference time cost comparison with previous video classification methods."},"limitations":{"value":"The authors have already claimed the limitations."}},"nonreaders":[],"tmdate":1730879276605,"tcdate":1720965138821,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission8924/Reviewer_caZh"],"signatures":["NeurIPS.cc/2024/Conference/Submission8924/Reviewer_caZh"],"forum":"YdfZP7qMzp","number":5,"license":"CC BY 4.0","cdate":1720965138821,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission8924/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730879276605,"domain":"NeurIPS.cc/2024/Conference","replyto":"YdfZP7qMzp","id":"YOuBBOCcws","forumContent":{"TLDR":{"value":"Unifying Video Generation and Recognition with Diffusion Models"},"venue":{"value":"NeurIPS 2024 poster"},"keywords":{"value":["video understanding","video generation","diffusion"]},"primary_area":{"value":"machine_vision"},"abstract":{"value":"Video diffusion models are able to generate high-quality videos by learning strong spatial-temporal priors on large-scale datasets. In this paper, we aim to investigate whether such priors derived from a generative process are suitable for video recognition, and eventually joint optimization of generation and recognition. Building upon Stable Video Diffusion, we introduce GenRec, the first unified framework trained with a random-frame conditioning process so as to learn generalized spatial-temporal representations. The resulting framework can naturally supports generation and recognition, and more importantly is robust even when visual inputs contain limited information. \nExtensive experiments demonstrate the efficacy of GenRec for both recognition and generation. In particular, GenRec achieves competitive recognition performance, offering 75.8% and 87.2% accuracy on SSV2 and K400, respectively. GenRec also performs the best on class-conditioned image-to-video generation, achieving 46.5 and 49.3 FVD scores on SSV2 and EK-100 datasets. Furthermore, GenRec demonstrates extraordinary robustness in scenarios that only limited frames can be observed. Code will be available at https://github.com/wengzejia1/GenRec."},"_bibtex":{"value":"@inproceedings{\nweng2024genrec,\ntitle={GenRec: Unifying Video Generation and Recognition with Diffusion Models},\nauthor={Zejia Weng and Xitong Yang and Zhen Xing and Zuxuan Wu and Yu-Gang Jiang},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=YdfZP7qMzp}\n}"},"title":{"value":"GenRec: Unifying Video Generation and Recognition with Diffusion Models"},"pdf":{"value":"/pdf/6e0d541821e88a536a6cb8b823d0813a8fe901bb.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"weng|genrec_unifying_video_generation_and_recognition_with_diffusion_models"},"authorids":{"value":["~Zejia_Weng1","~Xitong_Yang2","~Zhen_Xing2","~Zuxuan_Wu1","~Yu-Gang_Jiang1"]},"authors":{"value":["Zejia Weng","Xitong Yang","Zhen Xing","Zuxuan Wu","Yu-Gang Jiang"]}},"version":2},{"content":{"summary":{"value":"This paper proposes Task-Aware Mechanism (TAM), a hybrid Mixture-of-Experts (MoE) architecture applied to the Vision Tower (VT) of video-language models. Unlike prior works that mainly use MoE in LLMs for capacity scaling, TAM focuses on making the VT task-aware. A lightweight Inductor module (0.1B parameters) predicts the task type and dynamically adjusts video resolution and frame count based on the input query and video length. The authors also introduce a new dataset TA-116K for training the Inductor. The resulting model, TallVA-8B-A7B, achieves better performance across multiple video understanding benchmarks compared to state-of-the-art LVLMs with comparable LLM backbones."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"The manuscript needs more careful proofreading, e.g., “leverage video metadata to train a task-aware Hybrid-Gated MoE Vision Tower.” → “leverages,” singular."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. Well-written and easy to follow: The paper presents strong organization, with clear figures (e.g., Fig. 1 and Fig. 2) illustrating the pipeline and gating mechanisms.\n2. Comprehensive experiments: The evaluation spans diverse benchmarks (e.g., MVBench, EgoSchema, LongVideoBench, etc.), with consistent gains even when using smaller LLMs.\n3. Reproducibility: Implementation details, dataset composition, and ablation studies are described in depth, which enhances reproducibility."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Generalizability of the 8 task types: It is unclear whether the eight predefined video understanding categories comprehensively capture the diversity of real-world video tasks. If new task types emerge, will the Inductor or gating modules require retraining from scratch, or can they generalize via few-shot adaptation?\n\n2. Comparison with existing token merging or pruning works. Based on my understanding, the proposed gating mechanism can perform visual token compression before passing them to the LLMs, guided by the information from the input queries. However, some token merging or pruning methods (e.g., https://www.ecva.net/papers/eccv_2024/papers_ECCV/papers/02577.pdf and https://arxiv.org/abs/2405.16148) also reduce visual tokens, often based on the attention maps from previous layers or additional input cues, and may achieve a similar goal.\n\n3. Scalability beyond 7B-scale models: While experiments are comprehensive, all evaluations rely on 7B-scale LLMs. It would strengthen the paper to show whether the task-aware mechanism scales effectively to larger models or yields diminishing returns."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762918424730,"tcdate":1762075564547,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission6041/Reviewer_9C6K"],"signatures":["ICLR.cc/2026/Conference/Submission6041/Reviewer_9C6K"],"forum":"HpXGsaMPeB","number":2,"license":"CC BY 4.0","cdate":1762075564547,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission6041/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762918424730,"domain":"ICLR.cc/2026/Conference","replyto":"HpXGsaMPeB","id":"91zb6lEGk6","forumContent":{"TLDR":{"value":"We propose Task-Aware Mechanism that employs Hybrid Gating Strategy to endow MoE Vision Tower. TAM could intelligently determines the appropriate task category, the number of frames to sample, and the optimal resolution based on user's query."},"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Video Understanding; Multimodal Large Language Model; Large vision-language models; Mixture of Experts;"]},"supplementary_material":{"value":"/attachment/7c03a0cf7a3247df20a89b54fa9711ef11a64a67.zip"},"primary_area":{"value":"foundation or frontier models, including LLMs"},"abstract":{"value":"Does *Comprehending the main idea of a 2-hour movie* and *Counting the birds appearing in a 15-second clip* really warrant the same video processing pipeline? Recent successes of Mixture-of-Experts (MoE) architectures in language modeling have inspired explorations of MoE applications. However, existing MoE models mainly focus on Large Language Models (LLMs) while neglecting Vision Tower (VT) in multimodal models. MoE-LLMs are predominantly designed for capacity scaling, whereas VT contains three fundamentally distinct modules, indicating that directly copying MoE-LLM designs to VT is unlikely to be effective. Inspired by the emerging Task-Aware idea, we argue that MoE-VT architectures should embody the principle of *Right Tool for the Right Job*, providing suitable processing to different tasks. To address this, we propose Task-Aware Mechanism (TAM), a MoE-VT architecture that employs Hybrid Gating Strategy to endow VT with intrinsic Task-Aware ability. To equip the framework with task-aware capabilities, we further introduce a compact Inductor module with only 0.1B parameters, trained on our new dataset TA-116k. With the Inductor, TAM could dynamically determine the appropriate task category, the optimal resolution and number of frames to sample, based on the user query and the length of video. Leveraging TAM, we introduce the TallVA-8B-A7B model, which outperforms current SOTA methods across various benchmarks on comparable LLMs, demonstrating that TAM enables video understanding models to become more holistic on diverse tasks."},"_bibtex":{"value":"@misc{\nyin2026taskaware,\ntitle={Task-Aware Mechanism: Hybrid MoE Vision Tower Towards Holistic Video Understanding},\nauthor={Qishen Yin and Tanghui Jia and Peng Jin and Li Hao and Juntong Wu and Guanlin Lu and Beili Tang and Li Yuan},\nyear={2026},\nurl={https://openreview.net/forum?id=HpXGsaMPeB}\n}"},"title":{"value":"Task-Aware Mechanism: Hybrid MoE Vision Tower Towards Holistic Video Understanding"},"pdf":{"value":"/pdf/f8fc0d3d3303006691edd6fe4cfd78c5165e55a6.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"yin|taskaware_mechanism_hybrid_moe_vision_tower_towards_holistic_video_understanding"},"authorids":{"value":["~Qishen_Yin1","~Tanghui_Jia1","~Peng_Jin4","~Li_Hao7","~Juntong_Wu1","~Guanlin_Lu1","~Beili_Tang1","~Li_Yuan2"]},"authors":{"value":["Qishen Yin","Tanghui Jia","Peng Jin","Li Hao","Juntong Wu","Guanlin Lu","Beili Tang","Li Yuan"]}},"version":2},{"content":{"summary":{"value":"The paper presents a scalable multimodal question-answering (QA) synthesis pipeline designed to automatically generate domain-aware, reasoning-centric QA pairs. Building upon this pipeline, the authors construct the WeThink dataset, which comprises over 120K multimodal QA pairs accompanied by explicit reasoning paths. Furthermore, the paper introduces a hybrid reward mechanism integrated with reinforcement learning (RL) training, which substantially enhances the performance of multimodal large language models across multiple visual-language reasoning benchmarks. The work also demonstrates the potential of the automated data pipeline to continually improve model performance through the ongoing incorporation of diverse data."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"In Section 5.3, the authors note that increasing data diversity sometimes leads to performance degradation on certain benchmarks. Please provide a detailed analysis of the specific causes underlying this phenomenon, and discuss how this issue might be mitigated—for example, through improved data generation strategies or refined RL optimization approaches."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The scalable multimodal QA synthesis pipeline provides an efficient method for automatically generating domain-aware, reasoning-centric QA pairs. \n\n2. **The WeThink dataset is a valuable new resource** with over 120K multimodal QA pairs and explicit reasoning paths, catering to diverse domains and abilities.\n\n3. The proposed dataset have the potential to accelerate research in multimodal reasoning.\n\n4. The case in the Appendix C offer insightful examples of the model's capabilities.\n\n5. The paper is very well-written and easy to follow."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The paper's main strength lies in its engineering execution. However, it offers limited contributions in terms of novel, fundamental algorithms. The core methodology is largely an application or integration of existing techniques.\n\n2. The methodology is critically dependent on the output of powerful existing models, namely DeepSeek-R1 and Qwen2.5-VL-72B, for its data generation phase. This heavy reliance makes it difficult to assess the intrinsic capabilities of the method itself, distinct from the strengths of the models it leverages.\n\n3. The paper introduces a hybrid reward system that combines rule-based and model-based rewards. While the empirical results demonstrate its effectiveness, the design of such a composite reward function appears to be a relatively standard practice in RL for multimodal tasks, and thus offers limited novelty. \n\n4. The claim that \"our automated pipeline ensures continuous data diversity and scalable RL training\" is strong. While the dataset is diverse, demonstrating *continuous* enhancement and *scalability* through empirical evidence showing performance gains with progressively larger or more diverse datasets generated by the pipeline would be more convincing. The current experiments show improvements with the dataset, but the *continuous* and *scalable* nature of the pipeline needs more direct validation."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762920378286,"tcdate":1761790397136,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission8511/Reviewer_eVp6"],"signatures":["ICLR.cc/2026/Conference/Submission8511/Reviewer_eVp6"],"forum":"pe3Fx0awUx","number":2,"license":"CC BY 4.0","cdate":1761790397136,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission8511/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762920378286,"domain":"ICLR.cc/2026/Conference","replyto":"pe3Fx0awUx","id":"KqtcC41lXg","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["multimodal large language models","visual-language reasoning"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Building on the success of text-based reasoning models like DeepSeek-R1, extending these capabilities to multimodal reasoning holds great promise. While recent works have attempted to adapt DeepSeek-R1-style reinforcement learning (RL) training paradigms to multimodal large language models (MLLM), focusing on domain-specific tasks like math and visual perception, a critical question remains: How can we enhance visual-language reasoning through RL for different domains? To address this challenge, we make three key efforts: (1) A novel Scalable Multimodal QA Synthesis pipeline that autonomously generates domain-aware, reasoning-centric question-answer (QA) pairs directly from images across different domains. (2) The open-source WeThink dataset containing over 120K multimodal QA pairs with annotated reasoning paths, curated from 18 diverse dataset sources and covering various question domains. (3) A simple baseline incorporating a hybrid reward mechanism that combines rule-based verification with model-based assessment to optimize RL training efficiency across different task domains. Through comprehensive exploration of RL on our dataset, we demonstrate that the WeThink dataset significantly improves performance across diverse MLLM benchmarks. Furthermore, we highlight that our automated data pipeline can continuously increase data diversity, further boosting model performance. The anonymous code and dataset are available at \\url{https://anonymous.4open.science/r/WeThink-7C9A} and \\url{https://huggingface.co/datasets/WeThink/WeThink-Multimodal-Reasoning-120K}."},"_bibtex":{"value":"@misc{\nyang2025enhancing,\ntitle={Enhancing Vision-Language Reasoning via Reinforcement Learning with Scalable Multimodal {QA} Synthesis},\nauthor={Jie Yang and Feipeng Ma and Zitian Wang and Dacheng Yin and Kang Rong and Fengyun Rao and Ruimao Zhang},\nyear={2025},\nurl={https://openreview.net/forum?id=pe3Fx0awUx}\n}"},"title":{"value":"Enhancing Vision-Language Reasoning via Reinforcement Learning with Scalable Multimodal QA Synthesis"},"pdf":{"value":"/pdf/d2a563877d6642b94f6bceb9222f209d1ad14c39.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"yang|enhancing_visionlanguage_reasoning_via_reinforcement_learning_with_scalable_multimodal_qa_synthesis"},"authorids":{"value":["~Jie_Yang20","~Feipeng_Ma1","~Zitian_Wang2","~Dacheng_Yin1","~Kang_Rong2","~Fengyun_Rao2","~Ruimao_Zhang1"]},"authors":{"value":["Jie Yang","Feipeng Ma","Zitian Wang","Dacheng Yin","Kang Rong","Fengyun Rao","Ruimao Zhang"]}},"version":2},{"content":{"venue":{"value":"CoRR 2024"},"pdf":{"value":"http://arxiv.org/pdf/2403.19158v1"},"venueid":{"value":"dblp.org/journals/CORR/2024"},"paperhash":{"value":"ma|uncertaintyaware_deep_video_compression_with_ensembles"},"authorids":{"value":["~Wufei_Ma1","https://dblp.org/search/pid/api?q=author:Jiahao_Li:","https://dblp.org/search/pid/api?q=author:Bin_Li:","https://dblp.org/search/pid/api?q=author:Yan_Lu:"]},"html":{"value":"https://doi.org/10.48550/arXiv.2403.19158"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2403-19158,\n  publtype={informal},\n  author={Wufei Ma and Jiahao Li and Bin Li and Yan Lu},\n  title={Uncertainty-Aware Deep Video Compression with Ensembles},\n  year={2024},\n  cdate={1704067200000},\n  journal={CoRR},\n  volume={abs/2403.19158},\n  url={https://doi.org/10.48550/arXiv.2403.19158}\n}\n"},"abstract":{"value":"Deep learning-based video compression is a challenging task, and many previous state-of-the-art learning-based video codecs use optical flows to exploit the temporal correlation between successive frames and then compress the residual error. Although these two-stage models are end-to-end optimized, the epistemic uncertainty in the motion estimation and the aleatoric uncertainty from the quantization operation lead to errors in the intermediate representations and introduce artifacts in the reconstructed frames. This inherent flaw limits the potential for higher bit rate savings. To address this issue, we propose an uncertainty-aware video compression model that can effectively capture the predictive uncertainty with deep ensembles. Additionally, we introduce an ensemble-aware loss to encourage the diversity among ensemble members and investigate the benefits of incorporating adversarial training in the video compression task. Experimental results on 1080p sequences show that our model can effectively save bits by more than 20% compared to DVC Pro."},"title":{"value":"Uncertainty-Aware Deep Video Compression with Ensembles"},"authors":{"value":["Wufei Ma","Jiahao Li","Bin Li","Yan Lu"]}},"tmdate":1727463131610,"pdate":1704067200000,"tcdate":1727460079326,"writers":["~"],"signatures":["~Wufei_Ma1"],"forum":"QDkOXbp1i2","license":"CC BY-SA 4.0","number":92210,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1727463131610,"domain":"DBLP.org","id":"QDkOXbp1i2","version":2},{"content":{"venue":{"value":"IEEE Trans. Multim. 2024"},"pdf":{"value":"https://ieeexplore.ieee.org/iel7/6046/10384483/10461131.pdf"},"venueid":{"value":"dblp.org/journals/TMM/2024"},"paperhash":{"value":"ma|uncertaintyaware_deep_video_compression_with_ensembles"},"authorids":{"value":["~Wufei_Ma1","https://dblp.org/search/pid/api?q=author:Jiahao_Li:","https://dblp.org/search/pid/api?q=author:Bin_Li_0012:","https://dblp.org/search/pid/api?q=author:Yan_Lu_0001:"]},"html":{"value":"https://doi.org/10.1109/TMM.2024.3372352"},"_bibtex":{"value":"@article{DBLP:journals/tmm/MaLLL24,\n  author={Wufei Ma and Jiahao Li and Bin Li and Yan Lu},\n  title={Uncertainty-Aware Deep Video Compression With Ensembles},\n  year={2024},\n  cdate={1704067200000},\n  journal={IEEE Trans. Multim.},\n  volume={26},\n  pages={7863-7872},\n  url={https://doi.org/10.1109/TMM.2024.3372352}\n}\n"},"abstract":{"value":"Deep learning-based video compression is a challenging task, and many previous state-of-the-art learning-based video codecs use optical flows to exploit the temporal correlation between successive frames and then compress the residual error. Although these two-stage models are end-to-end optimized, the epistemic uncertainty in the motion estimation and the aleatoric uncertainty from the quantization operation lead to errors in the intermediate representations and introduce artifacts in the reconstructed frames. This inherent flaw limits the potential for higher bit rate savings. To address this issue, we propose an uncertainty-aware video compression model that can effectively capture the predictive uncertainty with deep ensembles. Additionally, we introduce an ensemble-aware loss to encourage the diversity among ensemble members and investigate the benefits of incorporating adversarial training in the video compression task. Experimental results on 1080p sequences show that our model can effectively save bits by more than 20% compared to DVC Pro."},"title":{"value":"Uncertainty-Aware Deep Video Compression With Ensembles"},"authors":{"value":["Wufei Ma","Jiahao Li","Bin Li","Yan Lu"]}},"tmdate":1727463126771,"pdate":1704067200000,"tcdate":1727460079289,"writers":["~"],"signatures":["~Wufei_Ma1"],"forum":"6MHLfjf9Fv","license":"CC BY-SA 4.0","number":92207,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1727463126771,"domain":"DBLP.org","id":"6MHLfjf9Fv","version":2},{"content":{"venue":{"value":"IEEE Trans. Circuits Syst. Video Technol. 1997"},"pdf":{"value":"https://ieeexplore.ieee.org/iel4/76/13348/00611175.pdf"},"venueid":{"value":"dblp.org/journals/TCSV/1997"},"paperhash":{"value":"heinzelman|networkdriven_motion_estimation_for_wireless_video_terminals"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Wendi_Rabiner_Heinzelman:","~Anantha_P._Chandrakasan1"]},"html":{"value":"https://doi.org/10.1109/76.611175"},"_bibtex":{"value":"@article{DBLP:journals/tcsv/HeinzelmanC97,\n  author={Wendi Rabiner Heinzelman and Anantha P. Chandrakasan},\n  title={Network-driven motion estimation for wireless video terminals},\n  year={1997},\n  cdate={852076800000},\n  journal={IEEE Trans. Circuits Syst. Video Technol.},\n  volume={7},\n  number={4},\n  pages={644-653},\n  url={https://doi.org/10.1109/76.611175}\n}\n"},"abstract":{"value":"This paper describes a new compression algorithm, termed network-driven motion estimation (NDME), which reduces the power dissipation of wireless video devices in a networked environment by exploiting the predictability of object motion. Since the location of an object in the current frame can often be predicted accurately from its location in previous frames, it is possible to optimally partition the motion estimation computation between the portable devices and high powered compute servers on the wired network. In network-driven motion estimation, a remote high-powered resource at the base-station (or on the wired network), predicts the motion vectors of the current frame from the motion vectors of the previous frames. The base-station sends these predicted motion vectors to a portable video encoder, where motion compensation proceeds as usual. Network-driven motion estimation adaptively adjusts the coding algorithm based on the amount of motion in the sequence, using motion prediction to code portions of the video sequence which contain a large amount of motion and conditional replenishment to code portions of the sequence which contain little scene motion. This algorithm achieves a reduction in the number of operations performed at the encoder for motion estimation by over two orders of magnitude while introducing minimal degradation to the decoded video compared with full search encoder-based motion estimation."},"title":{"value":"Network-driven motion estimation for wireless video terminals"},"authors":{"value":["Wendi Rabiner Heinzelman","Anantha P. Chandrakasan"]}},"tmdate":1769195660141,"pdate":883526400000,"externalIds":["dblp:journals/tcsv/HeinzelmanC97"],"tcdate":1769195412760,"writers":["~"],"signatures":["~Anantha_Chandrakasan1"],"forum":"MmNb474pBI","license":"CC BY-SA 4.0","number":793216,"cdate":852076800000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1769195660141,"domain":"DBLP.org","id":"MmNb474pBI","version":2},{"content":{"venue":{"value":"NeurIPS 2025 poster"},"keywords":{"value":["Reinforcement Learning with Human Feedback","Shortcut Learning","Reward Hacking"]},"primary_area":{"value":"social_and_economic_aspects_of_machine_learning"},"abstract":{"value":"In reinforcement learning from human feedback, preference-based reward models play a central role in aligning large language models to human-aligned behavior. However, recent studies show that these models are prone to reward hacking and often fail to generalize well due to over-optimization. They achieve high reward scores by exploiting shortcuts, that is, exploiting spurious features (e.g., response verbosity, agreeable tone, or sycophancy) that correlate with human preference labels in the training data rather than genuinely reflecting the intended objectives. In this paper, instead of probing these issues one at a time, we take a broader view of the reward hacking problem as shortcut behaviors and introduce a principled yet flexible approach to mitigate shortcut behaviors in preference-based reward learning. Inspired by the invariant theory in the kernel perspective, we propose Preference-based Reward Invariance for Shortcut Mitigation (PRISM), which learns group-invariant kernels with feature maps in a closed-form learning objective. Experimental results in several benchmarks show that our method consistently improves the accuracy of the reward model on diverse out-of-distribution tasks and reduces the dependency on shortcuts in downstream policy models, establishing a robust framework for preference-based alignment."},"_bibtex":{"value":"@inproceedings{\nye2025rectifying,\ntitle={Rectifying Shortcut Behaviors in Preference-based Reward Learning},\nauthor={Wenqian Ye and Guangtao Zheng and Aidong Zhang},\nbooktitle={The Thirty-ninth Annual Conference on Neural Information Processing Systems},\nyear={2025},\nurl={https://openreview.net/forum?id=m51t6RKfGH}\n}"},"title":{"value":"Rectifying Shortcut Behaviors in Preference-based Reward Learning"},"pdf":{"value":"/pdf/68f64407c11f7fe9069c3e63b8c90bfbf679caa6.pdf"},"venueid":{"value":"NeurIPS.cc/2025/Conference"},"paperhash":{"value":"ye|rectifying_shortcut_behaviors_in_preferencebased_reward_learning"},"authorids":{"value":["~Wenqian_Ye1","~Guangtao_Zheng1","~Aidong_Zhang2"]},"authors":{"value":["Wenqian Ye","Guangtao Zheng","Aidong Zhang"]}},"tmdate":1783627442984,"pdate":1758216610953,"tcdate":1745866183377,"writers":["NeurIPS.cc/2025/Conference","NeurIPS.cc/2025/Conference/Submission4636/Authors"],"signatures":["NeurIPS.cc/2025/Conference/Submission4636/Authors"],"forum":"m51t6RKfGH","license":"CC BY 4.0","number":4636,"cdate":1745866183377,"readers":["everyone"],"invitations":["NeurIPS.cc/2025/Conference/-/Submission","NeurIPS.cc/2025/Conference/-/Post_Submission","NeurIPS.cc/2025/Conference/Submission4636/-/Full_Submission","NeurIPS.cc/2025/Conference/Submission4636/-/Supplementary_Material","NeurIPS.cc/2025/Conference/-/Edit","NeurIPS.cc/2025/Conference/Submission4636/-/Camera_Ready_Revision"],"mdate":1783627442984,"odate":1761704765191,"domain":"NeurIPS.cc/2025/Conference","id":"m51t6RKfGH","version":2},{"content":{"venue":{"value":"SIGIR 2024"},"venueid":{"value":"dblp.org/conf/SIGIR/2024"},"paperhash":{"value":"zhao|comi_correct_and_mitigate_shortcut_learning_behavior_in_deep_neural_networks"},"authorids":{"value":["~Lili_Zhao3","~Qi_Liu3","~Linan_Yue1","~Wei_Chen54","https://dblp.org/search/pid/api?q=author:Liyi_Chen:","https://dblp.org/search/pid/api?q=author:Ruijun_Sun:","https://dblp.org/search/pid/api?q=author:Chao_Song:"]},"html":{"value":"https://doi.org/10.1145/3626772.3657729"},"_bibtex":{"value":"@inproceedings{DBLP:conf/sigir/Zhao0Y0CSS24,\n  author={Lili Zhao and Qi Liu and Linan Yue and Wei Chen and Liyi Chen and Ruijun Sun and Chao Song},\n  title={COMI: COrrect and MItigate Shortcut Learning Behavior in Deep Neural Networks},\n  year={2024},\n  cdate={1704067200000},\n  pages={218-228},\n  url={https://doi.org/10.1145/3626772.3657729},\n  booktitle={SIGIR},\n  crossref={conf/sigir/2024}\n}\n"},"abstract":{"value":"Deep Neural Networks (DNNs), despite their notable progress across information retrieval tasks, encounter the issues of shortcut learning and struggle with poor generalization due to their reliance on spurious correlations between features and labels. Current research mainly mitigates shortcut learning behavior using augmentation and distillation techniques, but these methods could be laborious and introduce unwarranted biases. To tackle these, in this paper, we propose COMI, a novel method to COrrect and MItigate shortcut learning behavior. Inspired by the ways students solve shortcuts in educational scenarios, we aim to reduce model's reliance on shortcuts and enhance its ability to extract underlying information integrated with standard Empirical Risk Minimization (ERM). Specifically, we first design Correct Habit (CoHa) strategy to retrieve the top m challenging samples for priority training, which encourages model to rely less on shortcuts in the early training. Then, to extract more meaningful underlying information, the information derived from ERM is separated into task-relevant and task-irrelevant information, the former serves as the primary basis for model predictions, while the latter is considered non-essential. However, within task-relevant information, certain potential shortcuts contribute to overconfident predictions. To mitigate this, we design Deep Mitigation (DeMi) network with shortcut margin loss to adaptively control the feature weights of shortcuts and eliminate their influence. Besides, to counteract unknown shortcut tokens issue in NLP, we adopt locally interpretable module-LIME to help recognize shortcut tokens. Finally, extensive experiments conducted on NLP and CV tasks demonstrate the effectiveness of COMI, which can perform well on both IID and OOD samples."},"title":{"value":"COMI: COrrect and MItigate Shortcut Learning Behavior in Deep Neural Networks"},"authors":{"value":["Lili Zhao","Qi Liu","Linan Yue","Wei Chen","Liyi Chen","Ruijun Sun","Chao Song"]}},"tmdate":1768973501557,"pdate":1704067200000,"tcdate":1722952466775,"writers":["~"],"signatures":["~Wei_Chen54"],"forum":"08q2ZdWTVE","license":"CC BY-SA 4.0","number":55452,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1768973501557,"domain":"DBLP.org","id":"08q2ZdWTVE","version":2},{"content":{"summary":{"value":"This paper introduces ViDiT-Q, a novel quantization method designed to address the unique challenges faced by diffusion transformers (DiTs) in text-to-image and video generation tasks. Large model sizes and multi-frame processing in video generation pose significant computational and memory costs, making efficient deployment on edge devices challenging. \nThe authors propose ViDiT-Q, a tailored quantization method for DiTs. This scheme effectively manages quantization errors by addressing specific challenges such as data distribution variations and channel imbalance. ViDiT-Q uses channel balancing to reduce color deviations and dynamic quantization to handle temporal variations in video sequences."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"**Please note that since I am not an expert in model quantization and do not have any background in this field, the weaknesses I provide may not be sufficient to reveal the shortcomings of the work.**\n\n1. While ViDiT-Q performs well at W8A8 and W4A8 quantization levels, there is a noticeable performance drop at lower activation bit-widths (such as W4A4 or W4A2). This indicates that the current mixed precision design has room for improvement, especially in fully leveraging the acceleration potential of 4-bit weights.\n\n2.ViDiT-Q introduces multiple quantization parameters (such as different $\\alpha$ values) to handle data variations across different timesteps. This complex parameter management increases the model's complexity.\n\n3. I believe the model can further introduce 8-bit Attention (SageAttention) to improve model efficiency, which has already been integrated into some video diffusion model libraries. I wonder if 8-bit attention mechanisms or some linear attention acceleration mechanisms can further improve your solution?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"ViDiT-Q reduces the incoherence of data distribution, thereby lowering quantization error, by combining scaling and rotation-based channel balancing methods. Specifically, the scaling method addresses the \"static\" channel imbalance at the initial denoising stage, while the rotation method handles the \"dynamic\" distribution variations over time.\n\nViDiT-Q uses channel balancing to reduce color deviations and dynamic quantization to handle temporal variations in video sequences.\n\nViDiT-Q is validated on various text-to-image and video generation models, demonstrating minimal degradation in visual quality and metrics even at W8A8 and W4A8 quantization levels.\n\nQualitative results show that ViDiT-Q maintains high image quality and text-image alignment, while naive PTQ methods produce highly blurred or noisy images."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"**Please note that since I am not an expert in model quantization and do not have any background in this field, the weaknesses I provide may not be sufficient to reveal the shortcomings of the work.**\n\n1. While ViDiT-Q performs well at W8A8 and W4A8 quantization levels, there is a noticeable performance drop at lower activation bit-widths (such as W4A4 or W4A2). This indicates that the current mixed precision design has room for improvement, especially in fully leveraging the acceleration potential of 4-bit weights.\n\n2.ViDiT-Q introduces multiple quantization parameters (such as different $\\alpha$ values) to handle data variations across different timesteps. This complex parameter management increases the model's complexity.\n\n3. I believe the model can further introduce 8-bit Attention (SageAttention) to improve model efficiency, which has already been integrated into some video diffusion model libraries. I wonder if 8-bit attention mechanisms or some linear attention acceleration mechanisms can further improve your solution?"}},"nonreaders":[],"tmdate":1731427485102,"tcdate":1730730731669,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission1762/Reviewer_2N42"],"signatures":["ICLR.cc/2025/Conference/Submission1762/Reviewer_2N42"],"forum":"E1N1oxd63b","number":3,"license":"CC BY 4.0","cdate":1730730731669,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission1762/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427485102,"domain":"ICLR.cc/2025/Conference","replyto":"E1N1oxd63b","id":"oqszlQVvSm","forumContent":{"venue":{"value":"ICLR 2025 Poster"},"TLDR":{"value":"We design the quantization scheme ViDiT-Q tailored for Video & Image Diffusion Transformer Quantization. It achieves W4A8 quantization with negligible performance loss,  which brings 2x memory and 1.5x end2end latency speedup."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["video generation","low-bit quantization","diffusion model"]},"supplementary_material":{"value":"/attachment/96fa401ac6d23e41e75c0080c445e99e64ed12da.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Diffusion transformers have demonstrated remarkable performance in visual generation tasks, such as generating realistic images or videos based on textual instructions. However, larger model sizes and multi-frame processing for video generation lead to increased computational and memory costs, posing challenges for practical deployment on edge devices. Post-Training Quantization (PTQ) is an effective method for reducing memory costs and computational complexity.\nWhen quantizing diffusion transformers, we find that existing quantization methods face challenges when applied to text-to-image and video tasks. To address these challenges, we begin by systematically analyzing the source of quantization error and conclude with the unique challenges posed by DiT quantization. Accordingly, we design an improved quantization scheme: ViDiT-Q (**V**ideo \\& **I**mage **Di**ffusion **T**ransformer **Q**uantization), tailored specifically for DiT models. We validate the effectiveness of ViDiT-Q across a variety of text-to-image and video models, achieving W8A8 and W4A8 with negligible degradation in visual quality and metrics. Additionally, we implement efficient GPU kernels to achieve practical 2-2.5x memory optimization and a 1.4-1.7x end-to-end latency speedup."},"_bibtex":{"value":"@inproceedings{\nzhao2025viditq,\ntitle={ViDiT-Q: Efficient and Accurate Quantization of Diffusion Transformers for Image and Video Generation},\nauthor={Tianchen Zhao and Tongcheng Fang and Haofeng Huang and Rui Wan and Widyadewi Soedarmadji and Enshu Liu and Shiyao Li and Zinan Lin and Guohao Dai and Shengen Yan and Huazhong Yang and Xuefei Ning and Yu Wang},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=E1N1oxd63b}\n}"},"title":{"value":"ViDiT-Q: Efficient and Accurate Quantization of Diffusion Transformers for Image and Video Generation"},"pdf":{"value":"/pdf/31d026067ce2d62cad5a383ceb3b48cfed171670.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"zhao|viditq_efficient_and_accurate_quantization_of_diffusion_transformers_for_image_and_video_generation"},"authorids":{"value":["~Tianchen_Zhao2","~Tongcheng_Fang1","~Haofeng_Huang3","~Rui_Wan2","~Widyadewi_Soedarmadji1","~Enshu_Liu1","~Shiyao_Li2","~Zinan_Lin1","~Guohao_Dai4","~Shengen_Yan1","~Huazhong_Yang2","~Xuefei_Ning1","~Yu_Wang3"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Tianchen Zhao","Tongcheng Fang","Haofeng Huang","Rui Wan","Widyadewi Soedarmadji","Enshu Liu","Shiyao Li","Zinan Lin","Guohao Dai","Shengen Yan","Huazhong Yang","Xuefei Ning","Yu Wang"]}},"version":2},{"content":{"summary":{"value":"This paper demonstrates current unconditional video generation models do not considering the subtle characteristics of real-world video and proposes a simple method using Gradient Reversal Layer (GRL) with lightweight CNN to disregard the implicitly encoded temporal information within each frame. The experiment results show that neglecting implicitly encoded temporal information does not affect generated video quality and can achieve better or comparable FVD score."},"presentation":{"value":"3 good"},"contribution":{"value":"2 fair"},"soundness":{"value":"2 fair"},"strengths":{"value":"This paper presents a very interesting perspective to estimate the realness of the generated video samples: the temporal locations of frames within random videos. This paper finds CNNs fail to classify the temporal locations from real-world video samples. But CNNs can precisely classify the temporal location of generated video samples. Based on this phenomenon, this paper proposes to use a lightweight CNN to disregard the implicitly encoded temporal information within each frame."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"I agree that the videos generated by the model should strive to be as similar as possible to real-world videos in various aspects. However, I have some doubts about your design using CNNs to classify the absolute positions of each frame in a 16-frame video. The positions of video frames should be relative rather than absolute. For example, after sampling many short videos of 16 frames each from a long video, the first frame of one short video may be the last frame of another video. This might make it difficult for CNNs to classify the position of every video frame in real-world datasets. However, for videos generated by video generation models, since they are trained on short videos (e.g., 16 frames) during their training phase, it is easy for the generation model to remember the relative positions between frames. This makes it easier for CNN classifiers to classify the positions of video frames. I believe that more training on real-world datasets may improve their classification performance.\n\nThis paper employs a Gradient Reversal Layer (GRL) to weaken the temporal information in each video frame. The authors use GRL in several places, such as, \"We integrate a Gradient Reversal Layer (GRL) along with an ImageNet pre-trained model,\" \"We adopt an adversarial training technique using GRL with a simple network,\" \"We propose a method consisting of GRL with the temporal classifier.\" These statements may have caused a lot of confusion for readers in understanding GRL. What exactly is GRL, and how does it function within the context of this article?\n\nIn terms of experiments, the authors do not provide a video demo to demonstrate the quality of its visual generation. I think in terms of video generation, the visual quality of the generated video results is far more important than the value of FVD.\n\nIn addition, there are some typos in the article:\nThe proposed method can be simply added to existing video generation methods in a plug-and-play manner. The full framework of the proposed method is shown in Fig. ??. -> In page 5\n\nit is negligible as the difference is only 5%p. -> In page 8"},"confidence":{"value":"5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully."},"questions":{"value":"see weakness"},"rating":{"value":"5: marginally below the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699635924908,"tcdate":1698947020576,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission14/Reviewer_LjvM"],"signatures":["ICLR.cc/2024/Conference/Submission14/Reviewer_LjvM"],"forum":"p6UwN2Rxhx","number":3,"license":"CC BY 4.0","cdate":1698947020576,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission14/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699635924908,"domain":"ICLR.cc/2024/Conference","replyto":"p6UwN2Rxhx","id":"1AGp4AV3b5","forumContent":{"venue":{"value":"Submitted to ICLR 2024"},"TLDR":{"value":"This paper uncovers how current video generation models inadvertently encode temporal information into frames, enabling accurate temporal classification by CNNs. We propose a method to eliminate this without compromising the FVD score."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Unconditional Video Generation; Video Generation"]},"primary_area":{"value":"generative models"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Unconditional video generation models seemed to generate realistic videos. However, in this paper, we delve into what could be the meaning of `realness' in the video generation models, taking into account that Convolutional Neural Networks (CNNs) are built with the inspiration from human visual neuroscience. Similar to human observers, we expected CNNs to struggle in classifying the temporal location of generated videos using a single frame due to the limited temporal information a single frame alone provides. However, our preliminary experiments unveil that current unconditional video generation models actually do inadvertently encode temporal location into each frame, enabling CNNs to correctly classify the temporal location of generated videos. To alleviate such a problem, we propose a method by adding the Gradient Reversal Layer (GRL) with lightweight CNN to the prior works to explicitly neglect this implicitly encoded temporal information. The experimental results, indeed, show that the implicit encoding of temporal information while training the unconditional video generator does negatively influence the FVD score. Moreover, experiments on diverse prior video generation models and datasets show that our approach can be used in a plug-and-play manner. Also, the results show the successful elimination of implicitly encoded temporal information without compromising the FVD score, highlighting the need to consider temporal classification accuracy as a supplementary metric in video generation models."},"_bibtex":{"value":"@misc{\nchoi2024unveiling,\ntitle={Unveiling Temporal Telltales: Are Unconditional Video Generation Models Implicitly Encoding Temporal Information?},\nauthor={Jaehyun Choi and Gyojin Han and Jiwan Hur and Jae Young Lee and Junmo Kim},\nyear={2024},\nurl={https://openreview.net/forum?id=p6UwN2Rxhx}\n}"},"title":{"value":"Unveiling Temporal Telltales: Are Unconditional Video Generation Models Implicitly Encoding Temporal Information?"},"pdf":{"value":"/pdf/8464b178825ff334fa72558bfd3475fb72bee392.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Rejected_Submission"},"paperhash":{"value":"choi|unveiling_temporal_telltales_are_unconditional_video_generation_models_implicitly_encoding_temporal_information"},"authorids":{"value":["~Jaehyun_Choi1","~Gyojin_Han1","~Jiwan_Hur1","~Jae_Young_Lee1","~Junmo_Kim1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Jaehyun Choi","Gyojin Han","Jiwan Hur","Jae Young Lee","Junmo Kim"]}},"version":2},{"content":{"summary":{"value":"This paper introduces a novel task and approach for one-shot medical video object segmentation using static image datasets. The authors address the challenge of limited annotated video data in medical imaging by proposing a framework that leverages readily available labeled static images to segment objects in medical videos with minimal annotation. The proposed method involves a two-stage process, training a one-shot segmentation model on images and temporal-aware test-time training via self-distillation. Experimental results demonstrate that the proposed method outperforms existing approaches in this low-data regime for video object segmentation on OS-I2V-Seg."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"Please see the weakness part 1,2,3,4,5,6."},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":1},"contribution":{"value":2},"strengths":{"value":"The paper is well-motivated. It presents a novel problem formulation by introducing one-shot medical video object segmentation using only static image datasets.\n\nThe proposed framework combines a memory mechanism with self-distillation during test-time training. This allows the model to leverage temporal information in videos without requiring additional annotated video data during training.\n\nExtensive experiments, including comparisons with state-of-the-art methods and ablation studies, demonstrate the effectiveness of the proposed approach."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. There are no details of the multi-scale feature enhancement module in the manuscript and in the figure.\n\n2. The process of usage of memory values is not revealed in Figure 1.\n\n3. Although ablation studies are presented, more detailed analysis on the impact of specific hyperparameters (e.g., memory bank size, selection of top-k affinities) could provide deeper insights into the method's performance and robustness.\n\n4. The results section is not well-organized or described. It would be better to provide a more detailed description and analysis and reorganize some content from the Appendix into the manuscript within the page limitation.\n\n5. It would be better to also collect other SOTA performance on HMC-QU, ASU-Mayo, and CAMUS to show the progress between the current work and the ultimate target based on temporal models. \n\n6. The possible reasons behind the dramatic performance decrease are not analyzed. For example, why does no TTT generally generate sub-optimal performance, but using some modules may slightly decrease the performance, and some make training collapse, yet using them all can generate the best performance? There is no deeper analysis of the performance gap and how and why some designs work."}},"nonreaders":[],"tmdate":1731429443331,"tcdate":1730605211911,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission13118/Reviewer_35mV"],"signatures":["ICLR.cc/2025/Conference/Submission13118/Reviewer_35mV"],"forum":"BhECSDSkAE","number":2,"license":"CC BY 4.0","cdate":1730605211911,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission13118/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731429443331,"domain":"ICLR.cc/2025/Conference","replyto":"BhECSDSkAE","id":"DzU0Flcj1z","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["medical video analysis","one-shot video object segmentation","test-time training","self-distillation"]},"primary_area":{"value":"unsupervised, self-supervised, semi-supervised, and supervised representation learning"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"This paper introduces a novel task and approach for one-shot medical video object segmentation using static image datasets. We address the critical challenge of limited annotated video data in medical imaging by proposing a framework that leverages readily available labeled static images to segment objects in medical videos with minimal annotation---specifically, a ground truth mask for only the first frame. Our method comprises training a one-shot segmentation model exclusively on images, followed by adapting it to medical videos through a test-time training strategy. This strategy incorporates a memory mechanism to utilize spatiotemporal context and employs self-distillation to maintain generalization capabilities. To facilitate research in this domain, we present OS-I2V-Seg, a comprehensive dataset comprising 28 categories in images and 4 categories in videos, totaling 68,416 image/frame-mask pairs. Extensive experiments demonstrate the efficacy of our approach in this extremely low-data regime for video object segmentation, establishing baseline performance on OS-I2V-Seg. The code and data will be made publicly available."},"_bibtex":{"value":"@misc{\nzheng2024temporalaware,\ntitle={Temporal-Aware Test-Time Training via Self-Distillation for One-Shot Image-to-Video Segmentation},\nauthor={Qicong Wang and Yilei Shi and Jingliang Hu and Xiao Xiang Zhu and Lichao Mou},\nyear={2024},\nurl={https://openreview.net/forum?id=BhECSDSkAE}\n}"},"title":{"value":"Temporal-Aware Test-Time Training via Self-Distillation for One-Shot Image-to-Video Segmentation"},"pdf":{"value":"/pdf/545f3a422f20d41a30e003f7fa3165c198d4e8d7.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"wang|temporalaware_testtime_training_via_selfdistillation_for_oneshot_imagetovideo_segmentation"},"authorids":{"value":["~Qicong_Wang2","~Yilei_Shi1","~Jingliang_Hu1","~Xiao_Xiang_Zhu1","~Lichao_Mou3"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Qicong Wang","Yilei Shi","Jingliang Hu","Xiao Xiang Zhu","Lichao Mou"]}},"version":2},{"content":{"summary":{"value":"This paper proposes Video-Qformer, a connection module that extracts spatiotemporal features from videos to enhance large language models (LLMs). The Video-Qformer architecture consists of an Attentive Pooling Module and a Q-Former with expert feed-forward networks (FFNs) specialized for spatial, temporal, and summarization queries. The model is evaluated on several benchmarks, including zero-shot VideoQA, Video-ChatGPT, video captioning, and video summarization, with ablation studies demonstrating the effectiveness of individual components."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"See Weaknesses"},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1) This paper introduces an Attentive Pooling Module that can decouple the spatiotemporal features of videos with cross-attention layers, which differs from the average pooling in Video-ChatGPT.\n2) Three FNN experts are introduced into the Q-former to handle the spatial, temporal, and summarization queries respectively.\n3) Ablation studies are conducted to verify each component of the Video-Qformer."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1) The introduction section (lines 30–50) spends excessive space on general background information that is only tangentially relevant to the main contribution—the design of an efficient video connection module for MLLMs. This could be streamlined for clarity.\n2) The paper does not adequately explain the motivation behind the Mixture-of-Experts (MoE) design or the rationale for decoupling spatiotemporal features through attentive pooling. Adding visualizations (e.g., attention maps) and deeper analysis of these choices would strengthen the work.\n3) This paper doesn't provide further insights or solve the main problems of current video MLLMs since it uses a setting of 8/16 frames for video understanding. If each frame is encoded as 32 tokens (just as what BLIP-2 did), 16 frames only constitute 512 tokens. Video representations with such a token length can be well accepted by almost all LLMs nowadays, there is no need to compress them further. Sometimes even an image-based MLLM will use 576 tokens to represent one single image (e.g., LLaVA-1.5).\n4) The experimental design raises concerns about fairness. The main comparison targets (Video-LLaMA and Video-ChatGPT) use much less video data to train their connectors, which complicates direct performance comparisons and raises doubts about the improvements of Video-Qformer."}},"nonreaders":[],"tmdate":1731428918429,"tcdate":1729411260593,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission11192/Reviewer_uaDg"],"signatures":["ICLR.cc/2025/Conference/Submission11192/Reviewer_uaDg"],"forum":"R6sIi9Kbxv","number":1,"license":"CC BY 4.0","cdate":1729411260593,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission11192/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428918429,"domain":"ICLR.cc/2025/Conference","replyto":"R6sIi9Kbxv","id":"iPCTUSWrfS","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Multimodal Large Language Model","Vision-Language Pretraining","Video Understanding"]},"primary_area":{"value":"foundation or frontier models, including LLMs"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Large language models (LLMs) have made remarkable strides in natural language processing tasks. However, effectively processing and understanding visual information remains a challenge for these models. To address this, multimodal large language models have been proposed, which integrate pre-trained visual encoders with LLMs. Although existing image-based approaches have shown success in aligning visual and textual modalities, extending these advancements to videos is challenging due to the richer visual and temporal information they contain. Current methods, including Video-ChatGPT and Video-LLaMA, have limitations in capturing inter-frame relationships and providing sufficient semantic context. To overcome these challenges, we propose Video Q-Former, a model that adaptively extracts spatiotemporal features from videos with a spatio-temporal querying transformer, enhancing the LLM’s comprehension of visual-language alignment. Extensive experiments demonstrate that our model achieves state-of-the-art performance across various datasets in zero-shot video question answering tasks."},"_bibtex":{"value":"@misc{\nnie2025video,\ntitle={Video Q-Former: Multimodal Large Language Model with Spatio-Temporal Querying Transformer Towards Video Understanding},\nauthor={Yuxiang Nie and Han Wang and Yanjie Wang and Can Huang and Liang Lin and Guanbin Li},\nyear={2025},\nurl={https://openreview.net/forum?id=R6sIi9Kbxv}\n}"},"title":{"value":"Video Q-Former: Multimodal Large Language Model with Spatio-Temporal Querying Transformer Towards Video Understanding"},"pdf":{"value":"/pdf/6a4b2bd8b1e48662f75e7fca3b2b64f4848d6d91.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"nie|video_qformer_multimodal_large_language_model_with_spatiotemporal_querying_transformer_towards_video_understanding"},"authorids":{"value":["~Yuxiang_Nie2","~Han_Wang18","~Yanjie_Wang2","~Can_Huang1","~Liang_Lin1","~Guanbin_Li2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Yuxiang Nie","Han Wang","Yanjie Wang","Can Huang","Liang Lin","Guanbin Li"]}},"version":2},{"content":{"summary":{"value":"This paper provides a theoretical framework analyzing shortcut learning in self-supervised learning (SSL) through the lens of eigenvalue decomposition of feature cross-correlation matrices. The authors introduce two concepts: extent bias (prioritizing features based on dimensional coverage) and amplitude bias (prioritizing features based on magnitude). Building on Simon et al. (2023)'s work on stepwise learning dynamics, they demonstrate that learning priority is fundamentally governed by eigenvalues of the feature cross-correlation matrix rather than semantic importance. The theoretical analysis is validated through toy models (linear networks, MLPs) and extended to semi-realistic datasets (Colored-MNIST, Modified Waterbirds)."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"What do the authors think the key takeaway from this work should be? \n\nIs the goal to provide a theoretical framework for future analysis or does this yield some clear empirical insights for mitigating or discovering spurious correlations in practice already?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"Rigorous theoretical framework: The eigenvalue decomposition analysis (Theorems 4.1, 4.2, 4.3, 5.1) provides precise mathematical characterization of when features are learned, with critical time points τⱼ ∝ 1/γⱼ clearly derived.\n\nClear experimental validation: Figure 1 demonstrates excellent alignment between theoretical predictions (dashed lines) and empirical results (solid lines) for loss, eigenvalues, and feature alignment evolution.\n\nComprehensive scope: Analysis extends beyond basic Barlow Twins to multiple SSL methods (SimCLR, VICReg - Section C), network architectures (linear, DLN, MLP - Sections B.1-B.3), and the redundancy reduction coefficient λ (Section 6.1, Figure 2).\n\nNovel formalization: Extent bias and amplitude bias provide useful conceptual frameworks. The connection between feature dimensionality (mₗ vs mₛ) and eigenvalue magnitude (γₗ = mₗ > γₛ = mₛ, Theorem 4.1) is elegantly established.\n\nSome empirical validation: Colored-MNIST experiments (Section 7.1, Figure 5) show the plateau at 70% accuracy directly validates the extent bias hypothesis in a controlled setting with varying object ratios."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Limited novelty of core insight: The observation that models learn high-dimensional/high-amplitude features first is well-established in the spectral bias literature (Rahaman et al. 2019, Tancik et al. 2020 - cited by authors) and also formulated as various other names (easy-to-learn, low-variance features etc.) in the literature. The main contribution is formalizing this specifically for SSL eigenvalue dynamics.\n\nLimited actionable insights: While the paper explains why shortcut learning occurs, it offers minimal guidance on how to mitigate it. The conclusion mentions \"designing mechanisms to encourage learning of generalizable features\" but provides no concrete methods.\nExperimental scope:\n\nModified Waterbirds experiments (Section 7.2) are interesting but only briefly described in appendix.\nNo experiments on standard SSL benchmarks (ImageNet pretraining + downstream tasks) to assess real-world impact.\nThe 70% plateau observation in Figure 5 is compelling but limited to artificially constructed spurious correlations."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762927271886,"tcdate":1761939895724,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission17357/Reviewer_xbxN"],"signatures":["ICLR.cc/2026/Conference/Submission17357/Reviewer_xbxN"],"forum":"P23YpnH3kq","number":3,"license":"CC BY 4.0","cdate":1761939895724,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission17357/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762927271886,"domain":"ICLR.cc/2026/Conference","replyto":"P23YpnH3kq","id":"I6nV8cGAvK","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["shortcut learning","self-supervised learning","stepwise learning","feature learning","learning dynamics"]},"supplementary_material":{"value":"/attachment/6569a8c2d6f8f8ef65ce686d8ebdf0850a5c8e20.zip"},"primary_area":{"value":"unsupervised, self-supervised, semi-supervised, and supervised representation learning"},"abstract":{"value":"Recent advances in self-supervised learning (SSL) have shown remarkable progress in representation learning. However, SSL models often exhibit shortcut learning phenomenon, where they exploit dataset-specific biases rather than learning generalizable features, sometimes leading to severe over-optimization on particular datasets. We present a theoretical framework that analyzes this shortcut learning phenomenon through the lens of $\\textit{extent bias}$ and $\\textit{amplitude bias}$. By investigating the relations among extent bias, amplitude bias, and learning priorities in SSL, we demonstrate that learning dynamics is fundamentally governed by the dimensional properties and amplitude of features rather than their semantic importance. Our analysis reveals how the eigenvalues of the feature cross-correlation matrix influence which features are learned earlier, providing insights into why models preferentially learn shortcut features over more generalizable features."},"_bibtex":{"value":"@misc{\nkim2026stepwise,\ntitle={Stepwise Feature Learning in Self-Supervised Learning},\nauthor={Juhwan Kim and Sungyoon Lee},\nyear={2026},\nurl={https://openreview.net/forum?id=P23YpnH3kq}\n}"},"title":{"value":"Stepwise Feature Learning in Self-Supervised Learning"},"pdf":{"value":"/pdf/22148320b5d0495551512bdf53979f497eb94e04.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"kim|stepwise_feature_learning_in_selfsupervised_learning"},"authorids":{"value":["~Juhwan_Kim1","~Sungyoon_Lee1"]},"authors":{"value":["Juhwan Kim","Sungyoon Lee"]}},"version":2},{"content":{"venue":{"value":"DAGM GCPR 2025"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-032-12840-9_8.pdf"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"viertola|video_object_segmentationaware_audio_generation"},"html":{"value":"https://doi.org/10.1007/978-3-032-12840-9_8"},"_bibtex":{"value":"@inproceedings{DBLP:conf/dagm/ViertolaIR25,\n  author={Ilpo Viertola and Vladimir Iashin and Esa Rahtu},\n  title={Video Object Segmentation-Aware Audio Generation},\n  year={2025},\n  cdate={1735689600000},\n  pages={106-122},\n  url={https://doi.org/10.1007/978-3-032-12840-9_8},\n  booktitle={DAGM GCPR},\n  crossref={conf/dagm/2025}\n}\n"},"abstract":{"value":"Existing multimodal audio generation models often lack precise user control, which limits their applicability in professional Foley workflows. In particular, these models focus on the entire video and do not provide precise methods for prioritizing a specific object within a scene, generating unnecessary background sounds, or focusing on the wrong objects. To address this gap, we introduce the novel task of video object segmentation-aware audio generation, which explicitly conditions sound synthesis on object-level segmentation maps. We present SAGANet, a new multimodal generative model that enables controllable audio generation by leveraging visual segmentation masks along with video and textual cues. Our model provides users with fine-grained and visually localized control over audio generation. To support this task and further research on segmentation-aware Foley, we propose Segmented Music Solos, a benchmark dataset of musical instrument performance videos with segmentation information. Our method demonstrates substantial improvements over current state-of-the-art methods and sets a new standard for controllable, high-fidelity Foley synthesis. Code, samples, and Segmented Music Solos are available at https://saganet.notion.site/."},"title":{"value":"Video Object Segmentation-Aware Audio Generation"},"authors":{"value":[{"fullname":"Ilpo Viertola","username":""},{"fullname":"Vladimir Iashin","username":"~Vladimir_Iashin1"},{"fullname":"Esa Rahtu","username":""}]}},"tmdate":1778360795005,"pdate":1767139200000,"externalIds":["dblp:conf/dagm/ViertolaIR25"],"tcdate":1778360791075,"writers":["~","OpenReview.net/Public_Article/DBLP.org","OpenReview.net/Support"],"signatures":["~Vladimir_Iashin1"],"forum":"I0IryoqDLz","license":"CC BY-SA 4.0","number":16872,"cdate":1735689600000,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/DBLP.org/-/Record","OpenReview.net/Public_Article/-/Edit"],"mdate":1778360795005,"domain":"OpenReview.net/Public_Article","id":"I0IryoqDLz","version":2},{"content":{"venue":{"value":"CoRR 2026"},"pdf":{"value":"https://arxiv.org/pdf/2608.28784v1"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"li|cleartextvideo_a_largescale_textcentric_video_dataset_bridging_video_restoration_and_scenetext_enhancement"},"html":{"value":"https://doi.org/10.48550/arXiv.2608.28784"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2608-28784,\n  publtype={informal},\n  author={Jinlong Li and Jiaming Ding and Dingfu Lu and Malcolm Hsiu and Chuang Ke and Kangning Yang and Bochen Guan and Lan Fu and Jie Cai and Huiming Sun and Zibo Meng},\n  title={ClearText-Video: A Large-Scale Text-Centric Video Dataset Bridging Video Restoration and Scene-Text Enhancement},\n  year={2026},\n  month={August},\n  cdate={1785542400000},\n  journal={CoRR},\n  volume={abs/2608.28784},\n  url={https://doi.org/10.48550/arXiv.2608.28784}\n}\n"},"abstract":{"value":"Multimodal Large Language Models (MLLMs) have recently made strong progress in visual--linguistic understanding. However, their performance on text-centric video reasoning remains highly sensitive to input quality. Real-world user-provided videos often contain motion blur, compression artifacts, noise, and low-resolution text, which impair reliable text reading and downstream reasoning. Whether MLLMs can robustly read and reason about real-world scene text under diverse quality conditions remains a fundamental open question. We introduce ClearText-Video (CTVid), a large-scale, scene-text-aware benchmark for studying text-centric video understanding under controlled quality variation. CTVid contains 4,639 real-world text-rich egocentric videos, 550K+ frames, 1.6M human-verified scene-text annotations, and 220K+ spatial/temporal question--answer pairs in Chinese and English. For each high-quality video, CTVid provides content-matched Degraded-Quality and Restored-Quality variants, supporting two task families: Text-Centric Video Restoration and Multi-Quality VideoQA. We evaluate 18 representative restoration methods and 16 state-of-the-art MLLMs on CTVid. The results show that visual enhancement does not guarantee textual fidelity or downstream reasoning gains: blur is more damaging than low resolution, restored videos can alter the textual evidence used by MLLMs, and OCR-only pipelines remain far below direct multimodal reasoning. CTVid exposes the gap between video restoration and text-grounded understanding, providing a rigorous foundation for restoration-aware, quality-robust text-centric video systems."},"title":{"value":"ClearText-Video: A Large-Scale Text-Centric Video Dataset Bridging Video Restoration and Scene-Text Enhancement"},"authors":{"value":[{"fullname":"Jinlong Li","username":""},{"fullname":"Jiaming Ding","username":""},{"fullname":"Dingfu Lu","username":""},{"fullname":"Malcolm Hsiu","username":""},{"fullname":"Chuang Ke","username":"~Chuang_Ke1"},{"fullname":"Kangning Yang","username":""},{"fullname":"Bochen Guan","username":""},{"fullname":"Lan Fu","username":""},{"fullname":"Jie Cai","username":""},{"fullname":"Huiming Sun","username":""},{"fullname":"Zibo Meng","username":""}]}},"tmdate":1791309429358,"pdate":1798675200000,"externalIds":["dblp:journals/corr/abs-2608-28784"],"tcdate":1791309396256,"writers":["~","OpenReview.net/Public_Article/DBLP.org","OpenReview.net/Support"],"signatures":["~Chuang_Ke1"],"forum":"FuddmGpvnv","license":"CC BY-SA 4.0","number":164450,"cdate":1785542400000,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/DBLP.org/-/Record","OpenReview.net/Public_Article/-/Edit"],"mdate":1791309429358,"domain":"OpenReview.net/Public_Article","id":"FuddmGpvnv","version":2},{"content":{"summary":{"value":"This work proposes a new privacy protection framework for face video. The framework includes three modules: face swapping, video prediction, and video steganography, which realizes the high visual quality of the cover video and the strong undetectability of the secret video.  \n\nI am sorry that this work should be rejected because it has many weaknesses."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1. Why does the framework need reversibility？\n\n2. Why not just hide the identity？\n\n3. What is the technical innovation of the work? Please explain how each part of the work (face swapping, video prediction, video steganography) is novel relative to existing techniques."},"rating":{"value":1},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. A new understandable framework to protect face privacy in video.\n\n2. The structure of the writing is clear"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Motivation of the Work is Ambiguous: The rationale behind the need for reversibility is unclear. Specifically, I do not see the significance of restoring the original video. The primary purpose of privacy protection is to eliminate sensitive information. In particular, \"irreversibility\" would more effectively enhance the strength of privacy protection. The authors need to clarify the scenarios in which reversibility is applicable.\n\n2. Implementation Approach is Unreasonable: The authors propose hiding a secret video within a cover video to protect facial privacy. However, I find this approach difficult to accept. This is because the secret video and the cover video share highly similar content (i.e., attributes other than identity), making it unnecessary to hide the secret video directly. In other words, since the only difference between the secret and cover videos is the identity, would it not be more feasible to simply hide the identity instead? \n\n3. Technical Innovation is Weak: The authors have integrated techniques such as face swapping, video prediction, and video steganography to construct the CausalVE framework, which makes it hard to identify any significant technical innovation in this work. The three contributions mentioned in the \"primary contributions\" section are all achievable by existing methods; this paper merely applies these techniques to facial privacy protection.\n\n4. Insufficient Experimental Evaluation: (1) Incorrect Comparison Trials: This work aims to protect facial privacy, and it should be compared with existing facial privacy protection methods rather than video steganography, like [1] or [2]. (2) Robustness Evaluation: Videos encoded in different formats may lose some information. Can this framework still achieve reversibility under such conditions?\n\n[1]The UU-Net_ Reversible Face De-Identification for Visual Surveillance Video Footage.\n[2]IdentityMask_Deep_Motion_Flow_Guided_Reversible_Face_Video_De-identification\n\n5. Practicality of the Work is Poor: The proposed framework incorporates various technologies and losses, making it difficult for users to understand. Additionally, the use of time-consuming techniques such as diffusion models results in high energy consumption and latency for CausalVE, complicating its integration into practical applications."}},"nonreaders":[],"tmdate":1733184560371,"tcdate":1729343789321,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission5154/Reviewer_Lhus"],"signatures":["ICLR.cc/2025/Conference/Submission5154/Reviewer_Lhus"],"forum":"waHmD2i1dv","number":2,"license":"CC BY 4.0","cdate":1729343789321,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission5154/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1733184560371,"domain":"ICLR.cc/2025/Conference","replyto":"waHmD2i1dv","id":"s8IarYo1Bm","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Bioprivacy","Diffusion model","Face swapping","Video Prediction","Reversible neural networks","Video Hiding"]},"primary_area":{"value":"alignment, fairness, safety, privacy, and societal considerations"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Advanced facial recognition technologies and recommender systems with inadequate privacy technologies and policies for facial interactions increase concerns about bioprivacy violations. With the proliferation of video and live-streaming websites, public-face video distribution and interactions pose greater privacy risks. Existing techniques typically address the risk of sensitive biometric information leakage through various privacy enhancement methods but pose a higher security risk by corrupting the information to be conveyed by the interaction data, or by leaving certain biometric features intact that allow an attacker to infer sensitive biometric information from them. To address these shortcomings, in this paper, we propose a neural network framework, CausalVE. We obtain cover images by adopting a diffusion model to achieve face swapping with face guidance and use the speech sequence features and spatiotemporal sequence features of the secret video for dynamic video inference and prediction to obtain a cover video with the same number of frames as the secret video. In addition, we hide the secret video by using reversible neural networks for video hiding so that the video can also disseminate secret data. Numerous experiments prove that our CausalVE has good security in public video dissemination and outperforms state-of-the-art methods from a qualitative, quantitative, and visual point of view."},"_bibtex":{"value":"@misc{\nhuang2024causalve,\ntitle={Causal{VE}: Face Video Privacy Encryption via Causal Video Prediction},\nauthor={Yubo Huang and Wenhao Feng and Xin Lai and Zixi Wang and Jingzehua Xu and Shuai Zhang and Hongjie He and Fan Chen},\nyear={2024},\nurl={https://openreview.net/forum?id=waHmD2i1dv}\n}"},"title":{"value":"CausalVE: Face Video Privacy Encryption via Causal Video Prediction"},"pdf":{"value":"/pdf/8f95871091b3fc7498487a989666fbe91c5968e8.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"huang|causalve_face_video_privacy_encryption_via_causal_video_prediction"},"authorids":{"value":["~Yubo_Huang2","~Wenhao_Feng2","~Xin_Lai4","~Zixi_Wang2","~Jingzehua_Xu1","~Shuai_Zhang6","~Hongjie_He1","~Fan_Chen10"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Yubo Huang","Wenhao Feng","Xin Lai","Zixi Wang","Jingzehua Xu","Shuai Zhang","Hongjie He","Fan Chen"]}},"version":2},{"content":{"venue":{"value":"CVPR 2026 Highlight"},"pdf":{"value":"/pdf/5fa07040eec626d49b96aba6bc8d95d6e9f0600e.pdf"},"venueid":{"value":"OpenReview.net/Archive"},"paperhash":{"value":"yang|videocof_unified_video_editing_with_temporal_reasoner"},"authorids":{"value":["~Xiangpeng_Yang1","~Ji_Xie2","~Yiyuan_Yang4","~Yan_Huang14","~Min_Xu5","~Qiang_Wu2"]},"html":{"value":"https://arxiv.org/pdf/2512.07469"},"abstract":{"value":"Existing video editing methods face a critical trade-off: expert models offer precision but rely on task-specific priors like masks, hindering unification; conversely, unified temporal in-context learning models are mask-free but lack explicit spatial cues, leading to weak instruction-to-region mapping and imprecise localization. To resolve this conflict, we propose VideoCoF, a novel Chain-of-Frames approach inspired by Chain-of-Thought reasoning. VideoCoF enforces a \"seeing, reasoning, then editing\" procedure by compelling the video diffusion model to first predict reasoning tokens (edit-region latents) before generating the target video tokens. This explicit reasoning step removes the need for user-provided masks while achieving precise instruction-to-region alignment and fine-grained video editing. Furthermore, we introduce a RoPE alignment strategy that leverages these reasoning tokens to ensure motion alignment and enable length extrapolation beyond the training duration. We demonstrate that with a minimal data cost of only 50k video pairs, VideoCoF achieves state-of-the-art performance on VideoCoF-Bench, validating the efficiency and effectiveness of our approach."},"title":{"value":"VideoCoF: Unified Video Editing with Temporal Reasoner"},"authors":{"value":["Xiangpeng Yang","Ji Xie","Yiyuan Yang","Yan Huang","Min Xu","Qiang Wu"]}},"tmdate":1777037998138,"pdate":1780416000000,"tcdate":1776846609710,"writers":["~Xiangpeng_Yang1","~Ji_Xie2","~Yiyuan_Yang4","~Yan_Huang14","~Min_Xu5","~Qiang_Wu2"],"signatures":["~Ji_Xie2"],"forum":"ZsGxyPA6fT","license":"CC BY 4.0","number":47273,"cdate":1776846609710,"readers":["everyone"],"invitations":["OpenReview.net/Archive/-/Direct_Upload"],"mdate":1777037998138,"domain":"OpenReview.net/Archive","id":"ZsGxyPA6fT","version":2},{"content":{"summary":{"value":"This paper proposed Frame-Voyager that learns to query informative frame combinations, based on the given textual queries in the task.\nAuthors introduced a new data collection and labeling pipeline, by ranking frame combinations using a pre-trained Video-LLM.\nExtensive experiments are conducted to support the effectiveness of the proposed Frame-Voyager."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Please see the weaknesses. And extra questions:\n\n(1) What is the trade-off between efficiency and effectiveness? Can you provide some metrics like running time/memory usage/flops to the proposed methods? As the model contains an extra keyframe localization stage, it is worthy showing those results to see the trade-off.\n\n(2) I notice that the proposed method uses VILA as a reference model to label rewards, and uses the same VILA series models for downstream tasks. Is this also a kind of self-rewarding strategy shown in the sevilla work? Or the proposed method/modules can zero transfer to other VLM like llava-ov/qwen-vl-2?\n\n(3) to answer my 1st weakness, I would suggest authors conduct extra experiments on Next-GQA to see the grounded QA results with grounded metrics.\n\n(4) Can you show some off-shelf llm-based localization tools like sevila localizer in Table 2? It would be interesting to see the comparison between a llm-based reasoning method with the proposed reward imitation learning method.\n\n(5)  Regarding the claim in Figure 3 and RQ3 (Lines 431-433), is this true? Related work (e.g., [1]) suggests that some observed issues may stem from limitations in the model side (VILA). Could the authors clarify this?\n\n[1] LONGVIDEOBENCH: A Benchmark for Long-context Interleaved Video-Language Understanding"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"This is a reasonable extension from previous keyframe selection work [1] where it relies more on single keyframe selection, and does not consider temporal relations/modeling among frames. The proposed pseudo-label scheme from a VLM + list of combinations of frames makes sense to me. \nAuthors also conduct extensive experiments across diverse popular benchmarks to show their effectiveness. \nOverall, I think the proposed Frame-Voyager makes a good contribution to the keyframe selection in video-language studies. \n\n[1] Self-chained image-language model for video localization and question answering. NeurIPS23"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. As the paper mentioned, this pseudo-label strategy is not scalable, and I think the frame combination part is a bit tricky. This is like creating artificial rewards according to video (like a sandbox in RL) to train the reward model. However,  the reward is not always reliable from a VLM even though we compute it according to GT answers. For example, as shown in some previous studies [2],  the model will generate correct answers when provided with wrong localized clips. It is true that we might give performance improvement when using such reward data created by a human-free pipeline, but it is more like an adapter / DPO style adaptation to the model rather than letting the model truly be grounded on query-related / informative frames. \n\n2. I think the weakness in 1 somehow affects the proposed method to require a relatively complex pre-training data construction/filtering/design, since the positive/negative signal is too sensitive according to video conditions in this framework from my view (e.g. frame blurring / redundancy).\n\n3. Also, the keyframe selection for video-language understanding is not a brand new topic, many related works in track try to propose different ways, from continuous space learning to discrete pipeline, to address question-aware moment detection. I would suggest to include those works [1,2,3,4] for a more comprehensive study. \n\n\n[1] ViLA: Efficient Video-Language Alignment for Video Question Answering. ECCV24.  \n[2] Can i trust your answer? visually grounded video question answering.  CVPR24.  \n[3] TimeCraft: Navigate Weakly-Supervised Temporal Grounded Video Question Answering via Bi-directional Reasoning. ECCV24.  \n[4] Self-Adaptive Sampling for Accurate Video Question Answering on Image Text Models. ACL24."}},"nonreaders":[],"tmdate":1733036034771,"tcdate":1730688436478,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission13344/Reviewer_pMFu"],"signatures":["ICLR.cc/2025/Conference/Submission13344/Reviewer_pMFu"],"forum":"LNL7zKvm7e","number":4,"license":"CC BY 4.0","cdate":1730688436478,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission13344/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1733036034771,"domain":"ICLR.cc/2025/Conference","replyto":"LNL7zKvm7e","id":"eBwdjhho5d","forumContent":{"venue":{"value":"ICLR 2025 Poster"},"TLDR":{"value":"Frame-Voyager queries optimal frame combinations for Video-LLMs based on task-specific queries, outperforming existing SOTA methods and achieving best results in Video Question Answering benchmarks as a plug-and-play solution."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video-LLM","Adaptive Frame Sampling"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Video Large Language Models (Video-LLMs) have made remarkable progress in video understanding tasks. However, they are constrained by the maximum length of input tokens, making it impractical to input entire videos. Existing frame selection approaches, such as uniform frame sampling and text-frame retrieval, fail to account for the information density variations in the videos or the complex instructions in the tasks, leading to sub-optimal performance. In this paper, we propose Frame-Voyager that learns to query informative frame combinations, based on the given textual queries in the task. To train Frame-Voyager, we introduce a new data collection and labeling pipeline, by ranking frame combinations using a pre-trained Video-LLM. Given a video of M frames, we traverse its T-frame combinations, feed them into a Video-LLM, and rank them based on Video-LLM's prediction losses. Using this ranking as supervision, we train Frame-Voyager to query the frame combinations with lower losses. In experiments, we evaluate Frame-Voyager on four Video Question Answering benchmarks by plugging it into two different Video-LLMs. The experimental results demonstrate that Frame-Voyager achieves impressive results in all settings, highlighting its potential as a plug-and-play solution for Video-LLMs."},"_bibtex":{"value":"@inproceedings{\nyu2025framevoyager,\ntitle={Frame-Voyager: Learning to Query Frames for Video Large Language Models},\nauthor={Sicheng Yu and CHENGKAI JIN and Huanyu Wang and Zhenghao Chen and Sheng Jin and ZHONGRONG ZUO and XU XIAOLEI and Zhenbang Sun and Bingni Zhang and Jiawei Wu and Hao Zhang and Qianru Sun},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=LNL7zKvm7e}\n}"},"title":{"value":"Frame-Voyager: Learning to Query Frames for Video Large Language Models"},"pdf":{"value":"/pdf/154af3dfa9ab115b4d4a01fa334c6ab45c7ad3af.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"yu|framevoyager_learning_to_query_frames_for_video_large_language_models"},"authorids":{"value":["~Sicheng_Yu3","~CHENGKAI_JIN2","~Huanyu_Wang2","~Zhenghao_Chen2","~Sheng_Jin3","~ZHONGRONG_ZUO1","~XU_XIAOLEI1","~Zhenbang_Sun1","~Bingni_Zhang1","~Jiawei_Wu9","~Hao_Zhang3","~Qianru_Sun2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Sicheng Yu","CHENGKAI JIN","Huanyu Wang","Zhenghao Chen","Sheng Jin","ZHONGRONG ZUO","XU XIAOLEI","Zhenbang Sun","Bingni Zhang","Jiawei Wu","Hao Zhang","Qianru Sun"]}},"version":2},{"content":{"summary":{"value":"The paper presents a zero-shot approach to subject-driven video generation with controlled motion, showing potential for reducing test-time fine-tuning overhead. The reference attention mechanism and mask-guided motion control provide innovations over previous methods by enhancing the fidelity of subject appearance within controlled bounding boxes."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"- The paper mentions data filtering steps, but additional details on the refinement of bounding boxes and control signals could provide readers with greater insight into handling imperfect data in training.\n\n- Given the limitations in complex motions and single-subject restriction, it's recommended that the paper discuss specific applications that could benefit from these features as they currently exist."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"- The DreamCustomizer framework proposed in this paper does not require fine-tuning during inference, and can directly customize the target subject and motion trajectory in a zero-shot situation. \n\n- This generation framework that does not require fine-tuning improves the efficiency of video generation and is conducive to a wide range of practical applications.\n\n- Users only need to provide the subject image and a set of bounding box sequences to generate a custom video, without the need for complex inference stage debugging."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- While DreamCustomizer claims tuning-free inference, the paper acknowledges the challenge of decoupling camera movement from object motion, leading to camera drift in certain contexts.\n\n-  DreamCustomizer is designed for single-subject customization and does not handle multi-subject videos, which is highlighted as a limitation but without proposed extensions.\n\n- The effectiveness of motion control relies heavily on the precision of bounding box annotations, making the system vulnerable to errors if box tracking is inconsistent.\n\n- Missing some important previous works:\n\n[1] MotionFollower: Editing video motion via lightweight score-guided diffusion. (2024).\n\n[2] FaceChain-ImagineID: Freely crafting high-fidelity diverse talking faces from disentangled audio. CVPR 2024\n\n[3] MotionEditor: Editing video motion via content-aware diffusion. CVPR 2024\n\n[4] Combo: Co-speech holistic 3D human motion generation and efficient customizable adaptation in harmony.  (2024)."}},"nonreaders":[],"tmdate":1731427221474,"tcdate":1730219880271,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission249/Reviewer_b9Q7"],"signatures":["ICLR.cc/2025/Conference/Submission249/Reviewer_b9Q7"],"forum":"TX0OsLcaWf","number":3,"license":"CC BY 4.0","cdate":1730219880271,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission249/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427221474,"domain":"ICLR.cc/2025/Conference","replyto":"TX0OsLcaWf","id":"39u65azPJS","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video Generation","Customized Generation"]},"supplementary_material":{"value":"/attachment/69fde2f78ddc277fe9af270ed5f24a639a020966.zip"},"primary_area":{"value":"generative models"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Recent advances in customized video generation have enabled users to create videos tailored to both specific subjects and motion trajectories. However, existing methods often require complicated test-time fine-tuning and struggle with balancing subject learning and motion control, limiting their real-world applications. In this paper, we present $\\textbf{DreamCustomizer}$, a zero-shot video customization framework capable of generating videos with a specific subject and motion trajectory, guided by a single image and a bounding box sequence, respectively, and without the need for test-time fine-tuning. Specifically, we introduce reference attention, which leverages the model’s inherent capabilities for subject learning, and devise a mask-guided motion module to achieve precise motion control by fully utilizing the robust motion signal of box masks derived from bounding boxes. While these two components achieve their intended functions, we empirically observe that motion control tends to dominate over subject learning. To address this, we propose two key designs: $\\textbf{1)}$ the masked reference attention, which integrates a blended latent mask modeling scheme into reference attention to enhance subject representations at the desired positions, and $\\textbf{2)}$ a reweighted diffusion loss, which differentiates the contributions of regions inside and outside the bounding boxes to ensure a balance between subject and motion control. Extensive experimental results on a newly curated dataset demonstrate that DreamCustomizer outperforms state-of-the-art methods in both subject customization and motion control. The dataset, code, and models will be made publicly available."},"_bibtex":{"value":"@misc{\nwei2025zeroshot,\ntitle={Zero-Shot Subject-Driven Video Customization with Precise Motion Control},\nauthor={Yujie Wei and Shiwei Zhang and Hangjie Yuan and Xiang Wang and Haonan Qiu and Rui Zhao and Yutong Feng and Feng Liu and Zhizhong Huang and Jiaxin Ye and Yingya Zhang and Hongming Shan},\nyear={2025},\nurl={https://openreview.net/forum?id=TX0OsLcaWf}\n}"},"title":{"value":"Zero-Shot Subject-Driven Video Customization with Precise Motion Control"},"pdf":{"value":"/pdf/e6763de00e1e6caf4b3ebd215b4f1a66f8e9adba.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"wei|zeroshot_subjectdriven_video_customization_with_precise_motion_control"},"authorids":{"value":["~Yujie_Wei1","~Shiwei_Zhang2","~Hangjie_Yuan1","~Xiang_Wang9","~Haonan_Qiu1","~Rui_Zhao12","~Yutong_Feng2","~Feng_Liu15","~Zhizhong_Huang1","~Jiaxin_Ye1","~Yingya_Zhang3","~Hongming_Shan1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Yujie Wei","Shiwei Zhang","Hangjie Yuan","Xiang Wang","Haonan Qiu","Rui Zhao","Yutong Feng","Feng Liu","Zhizhong Huang","Jiaxin Ye","Yingya Zhang","Hongming Shan"]}},"version":2},{"content":{"summary":{"value":"The authors propose the first Echocardiography generation model conditioned in ECG signals named ECHOPulse. The overall pipeline ECHOPulse consists of three stages. The first stage is video tokenizer training, which utilizes a VQ-VAE based model to train the tokenizer for ECHO video quantization. The second stage is to align the ECG and video tokens, and the ECG-FM is used to encode the ECG and masked token prediction is used for alignment. The final stage is video generation, which is autoregressive video token generation conditioned on ECG tokens."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. As mentioned in the weaknesses section, was there a specific reason the qualitative comparison for tokenization models (reconstructed video comparison) and video generation models (generated video comparison) were not conducted?\n2. Instead of using non-overlapping patches for video tokenization, would the VQ-VAE based tokenization improve performance using overlapping patches? Was this explored experimentally?"},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":3},"strengths":{"value":"1. The paper contains novel contribution in that it is the first paper to present Echo video generation conditioned on ECG signals, and also plan to release the ECHO videos paired with ECG dataset.\n\n2. The paper is well-written and easy to follow, with the pipeline figure covering all three steps used in EchoPulse.\n\n3. The experiments compare the tokenization performance and video generation performance, and compares to other related papers despite this paper being the first to explore ECG-guided ECHO video generation."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The video tokenization performance results in Table 1 compare the reconstruction metrics of MSE and MAE between EchoNet-Synthetic and ECHOPulse. However, there should also be qualitative result comparison to compare the reconstruction results between the two tokenization models, similar to qualitative video generation results in Figure 3. \n\n2. Similar to the video tokenization weakness above, the video generation experiments also do not contain any qualitative comparison between other models such as MoonShot, VideoComposer, and HeartBeat. The figures only display the qualitative results of ECHOPulse under different ECG conditions. Therefore, qualitative results should be displayed for each models. \n\n3. As mentioned in the Limitation and Future Works section, instead of using GAN loss for ECHOPulse, it will be better to see combination of diffusion models in this pipeline for ECHO video generation."}},"nonreaders":[],"tmdate":1733498891128,"tcdate":1730693455984,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission2496/Reviewer_A3Vb"],"signatures":["ICLR.cc/2025/Conference/Submission2496/Reviewer_A3Vb"],"forum":"i2r7LDjba3","number":5,"license":"CC BY 4.0","cdate":1730693455984,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission2496/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1733498891128,"domain":"ICLR.cc/2025/Conference","replyto":"i2r7LDjba3","id":"jspDXcoG9I","forumContent":{"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Medical video generation","ECHO synthesis","Multimodality","Wearable device","Medical foundation model"]},"supplementary_material":{"value":"/attachment/b36e4534abf15e7b6aaecc3733649ca6d9de5e5a.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Echocardiography (ECHO) is essential for cardiac assessments, but its video quality and interpretation heavily relies on manual expertise, leading to inconsistent results from clinical and portable devices. ECHO video generation offers a solution by improving automated monitoring through synthetic data and generating high-quality videos from routine health data. However, existing models often face high computational costs, slow inference, and rely on complex conditional prompts that require experts' annotations. To address these challenges, we propose ECHOPulse, an ECG-conditioned ECHO video generation model. ECHOPulse introduces two key advancements: (1) it accelerates ECHO video generation by leveraging VQ-VAE tokenization and masked visual token modeling for fast decoding, and (2) it conditions on readily accessible ECG signals, which are highly coherent with ECHO videos, bypassing complex conditional prompts. To the best of our knowledge, this is the first work to use time-series prompts like ECG signals for ECHO video generation. ECHOPulse not only enables controllable synthetic ECHO data generation but also provides updated cardiac function information for disease monitoring and prediction beyond ECG alone. Evaluations on three public and private datasets demonstrate state-of-the-art performance in ECHO video generation across both qualitative and quantitative measures. Additionally, ECHOPulse can be easily generalized to other modality generation tasks, such as cardiac MRI, fMRI, and 3D CT generation. We will make the synthetic ECHO dataset, along with the code and model, publicly available upon acceptance."},"_bibtex":{"value":"@inproceedings{\nli2025echopulse,\ntitle={{ECHOP}ulse: {ECG} Controlled Echocardio-gram Video Generation},\nauthor={Yiwei Li and Sekeun Kim and Zihao Wu and Hanqi Jiang and Yi Pan and Pengfei Jin and Sifan Song and Yucheng Shi and Xiaowei Yu and Tianze Yang and Tianming Liu and Quanzheng Li and Xiang Li},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=i2r7LDjba3}\n}"},"title":{"value":"ECHOPulse: ECG Controlled Echocardio-gram Video Generation"},"pdf":{"value":"/pdf/bb51df8d53a0dc570603952a85c0a1572d2a85ee.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"li|echopulse_ecg_controlled_echocardiogram_video_generation"},"authorids":{"value":["~Yiwei_Li2","~Sekeun_Kim1","~Zihao_Wu1","~Hanqi_Jiang2","~Yi_Pan6","~Pengfei_Jin1","~Sifan_Song1","~Yucheng_Shi2","~Xiaowei_Yu1","~Tianze_Yang2","~Tianming_Liu3","~Quanzheng_Li1","~Xiang_Li14"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Yiwei Li","Sekeun Kim","Zihao Wu","Hanqi Jiang","Yi Pan","Pengfei Jin","Sifan Song","Yucheng Shi","Xiaowei Yu","Tianze Yang","Tianming Liu","Quanzheng Li","Xiang Li"]}},"version":2},{"content":{"summary":{"value":"The paper proposes P2-DPO (Perceptual Processing Direct Preference Optimization), a self-supervised alignment method for Large Vision-Language Models (LVLMs) designed to mitigate hallucinations by focusing on Perceptual Processing rather than Perception failures.\nUnlike prior DPO or RLHF approaches that rely on human or synthetic preferences, P2-DPO generates on-policy, vision-aware preference pairs through two mechanisms: (1) Focus-and-Enhance pairs (for perceptual bottlenecks) and (2) Visual Robustness pairs (for degraded inputs) and uses them for Reinforcement Learning training. To enhance training efficacy, they further introduce a Calibration Loss to align visual evidence with text generation and a Dynamic Deficit-Weighting (DDW) scheme for adaptive balancing. Experiments on hallucination benchmarks (POPE, HallusionBench, AMBER, MMHal-Bench, TextVQA) show consistent improvements over DPO baselines and comparable results to human-feedback-based training."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"See weaknesses"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"The paper offers a compelling and underexplored perspective: separating Perceptual Processing failures (i.e., hallucinations despite correct attention) from Perception failures. This framing adds analytical depth and could motivate new diagnostic tools for LVLMs.\n\nThe idea of generating on-policy, vision-grounded preference pairs directly from the model’s own attention maps is interesting.\n\nThe experiments span several challenging hallucination benchmarks (POPE, HallusionBench, AMBER, MMHal-Bench, TextVQA) and include targeted validation (e.g., AFR, noise robustness), with clean tables and consistent results."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Although the method avoids human labels, generating multiple augmented views and attention-driven crops for each sample adds nontrivial computational overhead.  Could the authors provide results to demonstrate what is the runtime overhead (per training sample or epoch) compared to standard DPO?\n\nThe entire method depends on the quality of attention maps and cropping heuristics. If attention localization fails, the generated preference pairs may reinforce spurious regions or noise. How does the model behave when the attention maps are themselves incorrect—does this lead to reinforcing spurious crops? How would the approach handle multi-object scenes where attention maps have multiple disjoint foci?"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762920615555,"tcdate":1761571548837,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission8851/Reviewer_LaWn"],"signatures":["ICLR.cc/2026/Conference/Submission8851/Reviewer_LaWn"],"forum":"ekOwxTn65Y","number":1,"license":"CC BY 4.0","cdate":1761571548837,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission8851/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762920615555,"domain":"ICLR.cc/2026/Conference","replyto":"ekOwxTn65Y","id":"5cabkf8eBx","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["LLMs ; MLLMs ; Hallucination ; DPO"]},"primary_area":{"value":"foundation or frontier models, including LLMs"},"abstract":{"value":"Hallucination has recently garnered significant research attention in Large Vision-Language Models (LVLMs). Direct Preference Optimization (DPO) aims to learn directly from the corrected preferences provided by humans, thereby addressing the hallucination issue. Despite its success, this paradigm has yet to specifically target the perceptual bottleneck in attended regions or address insufficient Visual Robustness against image degradation. Furthermore, existing preference pairs are often vision-agnostic and their inherently off-policy nature limits their effectiveness in guiding model learning. To address these challenges, we propose Perceptual Processing Direct Preference Optimization (P$^2$-DPO), a novel training paradigm in which the model generates and learns from its own preference pairs, thereby directly addressing the identified visual bottlenecks while inherently avoiding the issues of vision-agnostic and off-policy data. It introduces: (1) an on-policy preference pairs construction method targeting Focus-and-Enhance perception and Visual Robustness, and (2) a well-designed Calibration Loss to precisely align visual signals with the causal generation of text. Experimental results demonstrate that with a comparable amount of training data and cost, P$^2$-DPO outperforms strong baselines that rely on costly human feedback on benchmarks. Furthermore, evaluations on Attention Region Fidelity (ARF) and image degradation scenarios validate the effectiveness of P$^2$-DPO in addressing perceptual bottleneck in attended regions and improving Visual Robustness against degraded inputs."},"_bibtex":{"value":"@inproceedings{\nzhang2026pdpo,\ntitle={P\\${\\textasciicircum}2\\$-{DPO}: Grounding Hallucination in Perceptual Processing via Calibration Direct Preference Optimization},\nauthor={ruipeng zhang and Zhihao Li and Haozhang Yuan and C.L.Philip Chen and Tong Zhang},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=ekOwxTn65Y}\n}"},"title":{"value":"P$^2$-DPO: Grounding Hallucination in Perceptual Processing via Calibration Direct Preference Optimization"},"pdf":{"value":"/pdf/14ba92ba4d75afc9ea61f6dca935bd5078b584ef.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"zhang|p^2dpo_grounding_hallucination_in_perceptual_processing_via_calibration_direct_preference_optimization"},"authorids":{"value":["~Ruipeng_Zhang4","~Zhihao_Li21","~Haozhang_Yuan1","~C.L.Philip_Chen1","~Tong_Zhang14"]},"authors":{"value":["Ruipeng Zhang","Zhihao Li","Haozhang Yuan","C.L.Philip Chen","Tong Zhang"]}},"version":2},{"content":{"venue":{"value":"MICCAI (1) 2024"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-031-72378-0_48.pdf"},"venueid":{"value":"dblp.org/conf/MICCAI/2024"},"paperhash":{"value":"wu|gazedirected_vision_gnn_for_mitigating_shortcut_learning_in_medical_image"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Shaoxuan_Wu:","https://dblp.org/search/pid/api?q=author:Xiao_Zhang_0028:","https://dblp.org/search/pid/api?q=author:Bin_Wang:","https://dblp.org/search/pid/api?q=author:Zhuo_Jin:","https://dblp.org/search/pid/api?q=author:Hansheng_Li:","~Jun_Feng2"]},"html":{"value":"https://doi.org/10.1007/978-3-031-72378-0_48"},"_bibtex":{"value":"@inproceedings{DBLP:conf/miccai/WuZWJLF24,\n  author={Shaoxuan Wu and Xiao Zhang and Bin Wang and Zhuo Jin and Hansheng Li and Jun Feng},\n  title={Gaze-Directed Vision GNN for Mitigating Shortcut Learning in Medical Image},\n  year={2024},\n  cdate={1704067200000},\n  pages={514-524},\n  url={https://doi.org/10.1007/978-3-031-72378-0_48},\n  booktitle={MICCAI (1)},\n  crossref={conf/miccai/2024-1}\n}\n"},"abstract":{"value":"Deep neural networks have demonstrated remarkable performance in medical image analysis. However, its susceptibility to spurious correlations due to shortcut learning raises concerns about network interpretability and reliability. Furthermore, shortcut learning is exacerbated in medical contexts where disease indicators are often subtle and sparse. In this paper, we propose a novel gaze-directed Vision GNN (called GD-ViG) to leverage the visual patterns of radiologists from gaze as expert knowledge, directing the network toward disease-relevant regions, and thereby mitigating shortcut learning. GD-ViG consists of a gaze map generator (GMG) and a gaze-directed classifier (GDC). Combining the global modelling ability of GNNs with the locality of CNNs, GMG generates the gaze map based on radiologists’ visual patterns. Notably, it eliminates the need for real gaze data during inference, enhancing the network’s practical applicability. Utilizing gaze as the expert knowledge, the GDC directs the construction of graph structures by incorporating both feature distances and gaze distances, enabling the network to focus on disease-relevant foregrounds. Thereby avoiding shortcut learning and improving the network’s interpretability. The experiments on two public medical image datasets demonstrate that GD-ViG outperforms the state-of-the-art methods, and effectively mitigates shortcut learning. Our code is available at https://github.com/SX-SS/GD-ViG."},"title":{"value":"Gaze-Directed Vision GNN for Mitigating Shortcut Learning in Medical Image"},"authors":{"value":["Shaoxuan Wu","Xiao Zhang","Bin Wang","Zhuo Jin","Hansheng Li","Jun Feng"]}},"tmdate":1769258228390,"pdate":1735603200000,"externalIds":["dblp:conf/miccai/WuZWJLF24"],"tcdate":1769258211715,"writers":["~"],"signatures":["~Jun_Feng2"],"forum":"k1R7ZTFjVE","license":"CC BY-SA 4.0","number":794077,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1769258228390,"domain":"DBLP.org","id":"k1R7ZTFjVE","version":2},{"content":{"venue":{"value":"CoRR 2024"},"pdf":{"value":"https://arxiv.org/pdf/2406.14050v2"},"venueid":{"value":"dblp.org/journals/CORR/2024"},"paperhash":{"value":"wu|gazedirected_vision_gnn_for_mitigating_shortcut_learning_in_medical_image"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Shaoxuan_Wu:","https://dblp.org/search/pid/api?q=author:Xiao_Zhang_0028:","https://dblp.org/search/pid/api?q=author:Bin_Wang:","https://dblp.org/search/pid/api?q=author:Zhuo_Jin:","https://dblp.org/search/pid/api?q=author:Hansheng_Li:","~Jun_Feng2"]},"html":{"value":"https://doi.org/10.48550/arXiv.2406.14050"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2406-14050,\n  publtype={informal},\n  author={Shaoxuan Wu and Xiao Zhang and Bin Wang and Zhuo Jin and Hansheng Li and Jun Feng},\n  title={Gaze-directed Vision GNN for Mitigating Shortcut Learning in Medical Image},\n  year={2024},\n  cdate={1704067200000},\n  journal={CoRR},\n  volume={abs/2406.14050},\n  url={https://doi.org/10.48550/arXiv.2406.14050}\n}\n"},"abstract":{"value":"Deep neural networks have demonstrated remarkable performance in medical image analysis. However, its susceptibility to spurious correlations due to shortcut learning raises concerns about network interpretability and reliability. Furthermore, shortcut learning is exacerbated in medical contexts where disease indicators are often subtle and sparse. In this paper, we propose a novel gaze-directed Vision GNN (called GD-ViG) to leverage the visual patterns of radiologists from gaze as expert knowledge, directing the network toward disease-relevant regions, and thereby mitigating shortcut learning. GD-ViG consists of a gaze map generator (GMG) and a gaze-directed classifier (GDC). Combining the global modelling ability of GNNs with the locality of CNNs, GMG generates the gaze map based on radiologists' visual patterns. Notably, it eliminates the need for real gaze data during inference, enhancing the network's practical applicability. Utilizing gaze as the expert knowledge, the GDC directs the construction of graph structures by incorporating both feature distances and gaze distances, enabling the network to focus on disease-relevant foregrounds. Thereby avoiding shortcut learning and improving the network's interpretability. The experiments on two public medical image datasets demonstrate that GD-ViG outperforms the state-of-the-art methods, and effectively mitigates shortcut learning. Our code is available at https://github.com/SX-SS/GD-ViG."},"title":{"value":"Gaze-directed Vision GNN for Mitigating Shortcut Learning in Medical Image"},"authors":{"value":["Shaoxuan Wu","Xiao Zhang","Bin Wang","Zhuo Jin","Hansheng Li","Jun Feng"]}},"tmdate":1769258216819,"pdate":1735603200000,"externalIds":["dblp:journals/corr/abs-2406-14050"],"tcdate":1769258209922,"writers":["~"],"signatures":["~Jun_Feng2"],"forum":"dDMZyaovQQ","license":"CC BY-SA 4.0","number":794066,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1769258216819,"domain":"DBLP.org","id":"dDMZyaovQQ","version":2},{"content":{"summary":{"value":"- The paper presents TeaserGen, a two-stage system for creating teasers from long documentaries, addressing challenges like audiovisual alignment, smooth transitions, and factual accuracy. To support this, the authors developed DocumentaryNet, a dataset of 1,269 documentaries with teasers, including multimodal elements like video, narration, and sound effects. \n\n- TeaserGen first generates teaser narration from the documentary’s transcript using a large language model, creating an engaging summary. It then pairs visuals with narration through either a pre-trained contrastive language-vision model or a deep sequential model to match visuals accurately.\n\n- Results show that TeaserGen outperforms baseline models in maintaining coherence and alignment, offering a streamlined approach to automated teaser generation. DocumentaryNet and TeaserGen together provide valuable tools for advancing multimodal content modeling in documentary summarization."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"- Could you clarify the notation used in rows 210 and 211? Specifically, is  $S$  intended to represent a sequence of language tokens or audio waveforms? Consistent notation throughout the paper would enhance understanding and prevent confusion.\n\n- Have you considered evaluating alternative architectures or conducting ablation studies to assess the robustness of relying solely on the pretrained VTGHL model? Exploring different models could strengthen the validation of your approach.\n\n- What is the rationale behind setting the minimum clip length to three seconds and the overlap between clips to one second? Providing theoretical or empirical justification for these specific constraints would help in understanding their impact on the results.\n\n- Given that a frame rate of 1 FPS may be insufficient for capturing dynamic video content, have you experimented with higher frame rates? How does increasing the frame rate affect the model's ability to capture motion and improve prediction accuracy?\n\n- How does your model address the issue of repetitive scenes arising from multiple frames sharing identical sentence embeddings? Have you explored methods to incorporate fine-grained temporal information or annotations to enhance diversity and temporal coherence?\n\n- Could you elaborate on how the VTGHLS threshold of 0.64 was determined? Additionally, have you investigated how varying this threshold influences the results, perhaps through an ablation study?\n\n-  Given that the TeaserGen-LR model uses only three transformer layers, have you tested deeper architectures to see if they capture complex patterns more effectively? Would increasing the number of layers improve performance?\n\n- Have you considered using standardized quantitative evaluation metrics like ROUGE, BLEU, BERTScore, or perplexity for assessing the generated narration text? Including these metrics could enhance the reproducibility and comparability of your study.\n\n- Relying solely on L2 distance as the loss function might not fully capture perceptual similarities. Have you experimented with alternative loss functions, such as perceptual loss, SSIM loss and so on to potentially achieve more nuanced image generation?\n\n- Considering that the models were trained for only 15 epochs on a relatively small test set of 49 documentaries, have you explored training for more epochs or using a larger dataset? How might this affect model performance and generalizability?  To test the generalizability of your approach, have you considered evaluating the model on additional datasets beyond the 49 documentaries? How does the model perform on different genres or types of video content?\n\n- Could you provide details on the computational resources required by the diffusion prior model? Understanding the computational overhead would help assess the practicality and scalability of your approach in real-world applications.\n\n- Higher scene change rates may enhance visual diversity, have you evaluated their impact on the narrative coherence of the teasers? How do you balance diversity with maintaining a coherent and engaging storyline?\n\n- You employ different CLIP models (CLIP-ViT-B/32 and CLIP-ViT-L/14) for different components of your system. Have you considered the potential inconsistencies this might introduce? Would using the same CLIP model throughout improve the alignment between textual and visual data?\n\n-How does the frame extraction rate influence the model's performance in terms of capturing essential visual information? Have you analyzed the trade-offs between computational efficiency and the richness of visual features at different frame rates?"},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"- Tackles a unique problem in automated teaser generation for documentaries with TeaserGen, a creative, narration-centered two-stage approach that combines large language models with language-vision models for cohesive narration and visual alignment, showing effective and innovative use of existing technologies.\n\n- Provides solid empirical support with comparisons to baseline models across objective (e.g., F1 score, CLIPScore) and subjective metrics, as well as the introduction of DocumentaryNet, a multimodal dataset with documentary-teaser pairs that enriches resources available for this research area.\n\n- The work has significant potential impact by addressing a real-world gap in video summarization for documentary-style content, with applications in multimedia and educational fields, and establishes a foundation for further multimodal research, likely to stimulate new directions in long-form video modeling."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The paper presents an innovative approach to video teaser generation using pretrained language-vision models, but several issues need to be addressed to enhance its clarity and robustness. In rows 210 and 211, there is a notation inconsistency where $S$ is defined as a sequence of language tokens, yet later each $S_i$ is referred to as a waveform (audio signal). This inconsistency creates confusion, and it's crucial for the notation to consistently represent either language tokens or audio waveforms throughout the paper to avoid misunderstandings.\n\n- In Section 4.2.1, the method relies heavily on a single pretrained VTGHS model without sufficient ablation studies or comparisons with alternative architectures, which weakens the validation of the approach. The constraints imposed such as a minimum clip length of three seconds and a one-second overlap between clips, appear arbitrary and lack theoretical or empirical justification, raising questions about their effectiveness and impact on the results. Additionally, using a frame rate of only one frame per second (1 FPS), as mentioned in rows 453 to 455, is inadequate for videos that change dynamically. This low frame rate hinders the model's ability to effectively capture motion and make accurate predictions.\n\n- In Section 4.2.2, the model extracts features at a low frame rate of 1 FPS, causing multiple frames to share identical sentence embeddings. This coarse temporal resolution fails to capture dynamic changes within the video, leading to overly repetitive scenes. The absence of fine-grained temporal annotations in the dataset prevents the model from effectively distinguishing and assigning unique embeddings to semantically similar frames occurring at different times. As a result, the approach struggles to maintain diversity and temporal coherence in the generated visual content.\n\n- Regarding threshold selection, in row 320, the paper states, `We estimate a VTGHLS of 0.64 for the ground truth teasers`, but there is insufficient explanation about the criteria used to select this threshold value. The paper does not analyze how varying this threshold affects the results. Since there are ablation studies on changing the matching score function, as mentioned in Section 5.6, it is necessary to explore this threshold value in more detail to understand its impact on the overall pipeline.\n\n- Table 4 may present an unfair comparison because the TeaserGen models utilize advanced decoding techniques like beam search, which can enhance performance, while the baseline models do not use these techniques. For a fair comparison that accurately reflects each model's true capabilities, all models should employ similar decoding methods.\n\n- A significant limitation in the methodology is the heavy reliance on subjective metrics without incorporating standardized quantitative measures, as mentioned in rows 468 to 470. This over-reliance weakens the study's reproducibility and generalizability. While subjective listening tests provide valuable insights, the absence of automatic evaluation metrics commonly used in natural language processing and computer vision (such as ROUGE, BLEU, BERTScore, or perplexity for the generated narration tex) makes it difficult to objectively compare the results with other approaches or validate the findings across different contexts.\n\n- The TeaserGen-LR model uses a limited architecture with only three transformer layers, which may not capture complex patterns as effectively as the diffusion prior's more extensive 12-block backbone. Relying solely on L2 distance as the loss function might not fully capture perceptual similarities, leading to less nuanced image generation. Additionally, training the model for only 15 epochs on a small test set of 49 documentaries raises concerns about underfitting and limits the generalizability of the results.\n\n- The study does not discuss the potential computational overhead introduced by incorporating the diffusion prior, which could affect the practicality and scalability of the approach. While higher scene change rates may increase visual diversity, they might also compromise the narrative coherence of the generated teasers. Addressing these issues would enhance the study's robustness and applicability."}},"nonreaders":[],"tmdate":1731427419520,"tcdate":1730343446707,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission1391/Reviewer_VbpS"],"signatures":["ICLR.cc/2025/Conference/Submission1391/Reviewer_VbpS"],"forum":"G1n50BMqzm","number":2,"license":"CC BY 4.0","cdate":1730343446707,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission1391/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427419520,"domain":"ICLR.cc/2025/Conference","replyto":"G1n50BMqzm","id":"ZRA3bg93yW","forumContent":{"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Teaser Generation","Multimodal Learning","Vision-Language Model"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Teasers are an effective tool for promoting content in entertainment, commercial and educational fields. However, creating an effective teaser for long videos is challenging for it requires long-range multimodal modeling capability for the input videos, while necessitating maintaining audiovisual alignments, managing scene transitions and preserving factual accuracy for the output teasers. Due to the lack of a publicly-available dataset, progress along this research direction has been hindered. In this work, we present DocumentaryNet, a collection of 1,269 documentaries paired with their teasers, featuring multimodal data streams of video, speech, music, sound effects and narrations. With DocumentaryNet, we propose a new two-stage system for generating teasers from long documentaries. The proposed TeaserGen system first generates the teaser narration from the transcribed narration from the documentary using a pretrained large language model, and then selects the most relevant visual content to accompany the generated narration through language-vision models. For narration-video matching, we explore two approaches: a pretraining-based model using pretrained contrastive language-vision models and a deep sequential model that learns the mapping between the narrations and visuals. Our experimental results show that the pretraining-based approach is more effective at identifying relevant visual content than directly trained deep autoregressive models."},"_bibtex":{"value":"@inproceedings{\nxu2025teasergen,\ntitle={TeaserGen: Generating Teasers for Long Documentaries},\nauthor={Weihan Xu and Paul Pu Liang and Haven Kim and Julian McAuley and Taylor Berg-Kirkpatrick and Hao-Wen Dong},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=G1n50BMqzm}\n}"},"title":{"value":"TeaserGen: Generating Teasers for Long Documentaries"},"pdf":{"value":"/pdf/44b76094009a1157e83bc6a2869184a2b23497ba.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"xu|teasergen_generating_teasers_for_long_documentaries"},"authorids":{"value":["~Weihan_Xu1","~Paul_Pu_Liang1","~Haven_Kim1","~Julian_McAuley1","~Taylor_Berg-Kirkpatrick1","~Hao-Wen_Dong1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Weihan Xu","Paul Pu Liang","Haven Kim","Julian McAuley","Taylor Berg-Kirkpatrick","Hao-Wen Dong"]}},"version":2},{"content":{"summary":{"value":"This paper introduces a new framework that directly models raw 3D motion sequences to animate arbitrary characters from reference images. Unlike previous methods that rely on 2D pose renderings, MTVCraft encodes 3D joint trajectories into compact 4D motion tokens using a 4D Motion Tokenizer (4DMoT), which preserves spatial-temporal dynamics and eliminates the need for pixel-level pose alignment. These tokens are then integrated into a Motion-aware Video Diffusion Transformer (MV-DiT) featuring 4D motion attention and 4D positional encodings, enabling expressive, disentangled control of motion and appearance. Implemented on both CogVideoX-5B and Wan-2.1-14B backbones, MTVCraft achieves state-of-the-art results, with superior zero-shot generalization to unseen characters, styles, and non-human subjects."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"The codebook size of 8,192 with a 3072-dimensional embedding is relatively large compared to those commonly used in motion generation models. Would decreasing its dimension or size help reduce unused codes or improve resolution?"},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"A key strength of integrating a motion-generation pipeline into video generation lies in its ability to provide explicit temporal and structural control over motion, resulting in more coherent and realistic dynamics than end-to-end pixel-based video synthesis. By introducing intermediate motion representations, such as in motion generation domain [1,2], the framework captures fine-grained spatial-temporal cues that general video models often overlook. This separation of motion from appearance allows the generator to synthesize consistent, expressive, and physically plausible movements while preserving identity and style, effectively bridging human motion understanding with visual generation. Moreover, as demonstrated in previous literature [3], coupling motion modeling within video diffusion enables strong generalization to arbitrary characters. Also due to the modular paradigm, the method scales naturally across different backbones, whether smaller transformer-based models like CogVideoX or Wan-2.1, since the motion tokens are from unified SMPL.\n\n[1] Momask : MoMask: Generative Masked Modeling of 3D Human Motions, cvpr 2024\n[2] SALAD : salad skeleton-aware latent diffusion for text-driven motion generation and editing, cvpr 2025\n[3] AnyMoLe : AnyMoLe: Any Character Motion In-betweening Leveraging Video Diffusion Models, cvpr 2025"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"A potential weakness of this paradigm is that the 4D motion compression via 4DMoT is conceptually straightforward and not architecturally novel. The encoder-decoder with vector quantization closely follows standard VQVAE formulations, and while it effectively transforms SMPL joint trajectories into compact motion tokens, it does not introduce fundamentally new techniques in motion encoding or representation learning. However, despite this structural simplicity, the usage and integration of such motion tokenization within a large-scale video generation framework remains a meaningful contribution."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762916944310,"tcdate":1761032174558,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission3719/Reviewer_1LKS"],"signatures":["ICLR.cc/2026/Conference/Submission3719/Reviewer_1LKS"],"forum":"m7AQM9H6wa","number":1,"license":"CC BY 4.0","cdate":1761032174558,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission3719/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762916944310,"domain":"ICLR.cc/2026/Conference","replyto":"m7AQM9H6wa","id":"42gHeOXrYl","forumContent":{"TLDR":{"value":"We propose MTVCraft, a novel paradigm for animating arbitrary characters with 4D motion tokens."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Character Animation","Motion Tokenization","Video Generation"]},"supplementary_material":{"value":"/attachment/fe58496e08aed1320a5392bc95d7fd45a5c94bbf.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Character image animation has rapidly advanced with the rise of digital humans. However, existing methods rely largely on 2D-rendered pose images for motion guidance, which limits generalization and discards essential 4D information for open-world animation. To address this, we propose MTVCraft (Motion Tokenization Video Crafter), the first framework that directly models raw 3D motion sequences (i.e., 4D motion) for character image animation. Specifically, we introduce 4DMoT (4D motion tokenizer) to quantize 3D motion sequences into 4D motion tokens. Compared to 2D-rendered pose images, 4D motion tokens offer more robust spatial-temporal cues and avoid strict pixel-level alignment between pose images and the character, enabling more flexible and disentangled control. Next, we introduce MV-DiT (Motion-aware Video DiT). By designing unique motion attention with 4D positional encodings, MV-DiT can effectively leverage motion tokens as 4D compact yet expressive context for character image animation in the complex 4D world. We implement MTVCraft on both CogVideoX-5B (small scale) and Wan-2.1-14B (large scale), demonstrating that our framework is easily scalable and can be applied to models of varying sizes. Experiments on the TikTok and Fashion benchmarks demonstrate our state-of-the-art performance. Moreover, powered by robust motion tokens, MTVCraft showcases unparalleled zero-shot generalization. It can animate arbitrary characters in full-body and half-body forms, and even non-human objects across diverse styles and scenarios. Hence, it marks a significant step forward in this field and opens a new direction for pose-guided video generation. Our project page is available at https://github.com/DINGYANB/MTVCrafter. A scaled version has been commercially deployed and is available at https://telestudio.teleagi.cn/generatevideo/creativeWorkshop."},"_bibtex":{"value":"@inproceedings{\nding2026mtvcraft,\ntitle={{MTVC}raft: Tokenizing 4D Motion for Arbitrary Character Animation},\nauthor={Yanbo Ding and Xirui Hu and Guo Zhi Zhi and Yan Zhang and Xinrui Wang and Zhixiang He and Chi Zhang and Yali Wang and Xuelong Li},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=m7AQM9H6wa}\n}"},"title":{"value":"MTVCraft: Tokenizing 4D Motion for Arbitrary Character Animation"},"pdf":{"value":"/pdf/4e63364f06783cd9ad648b2791f901df70dad464.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"ding|mtvcraft_tokenizing_4d_motion_for_arbitrary_character_animation"},"authorids":{"value":["~Yanbo_Ding2","~Xirui_Hu1","~Guo_Zhi_Zhi1","~Yan_Zhang62","~Xinrui_Wang9","~Zhixiang_He1","~Chi_Zhang14","~Yali_Wang1","~Xuelong_Li2"]},"authors":{"value":["Yanbo Ding","Xirui Hu","Guo Zhi Zhi","Yan Zhang","Xinrui Wang","Zhixiang He","Chi Zhang","Yali Wang","Xuelong Li"]}},"version":2},{"content":{"venue":{"value":"ACL workshop'25"},"pdf":{"value":"/pdf/e31f1fd9588bd60c05393d5d00293a1bb20653a0.pdf"},"venueid":{"value":"OpenReview.net/Archive"},"paperhash":{"value":"tasawong|shortcut_learning_in_safety_the_impact_of_keyword_bias_in_safeguards"},"authorids":{"value":["~Panuthep_Tasawong1","~Napat_Laosaengpha1","~Wuttikorn_Ponwitayarat1","~Sitiporn_Sae_Lim1","~Potsawee_Manakul1","~Samuel_Cahyawijaya1","~Can_Udomcharoenchaikit1","~Peerat_Limkonchotiwat1","~Ekapol_Chuangsuwanich1","~Sarana_Nutanong1"]},"html":{"value":"https://aclanthology.org/2025.llmsec-1.14/"},"abstract":{"value":"This paper investigates the problem of shortcut learning in safety guardrails for large language models (LLMs). It reveals that current safeguard models often rely excessively on superficial cues, such as specific keywords that are spuriously correlated with training labels, rather than genuinely understanding the input’s semantics or intent. As a result, their performance degrades significantly when there is a shift in keyword distribution. The paper also examines the impact of reducing shortcut reliance, showing that merely minimizing shortcut influence is insufficient. To build robust safeguard models, it is equally crucial to promote the use of intended features."},"title":{"value":"Shortcut Learning in Safety: The Impact of Keyword Bias in Safeguards"},"authors":{"value":["Panuthep Tasawong","Napat Laosaengpha","Wuttikorn Ponwitayarat","Sitiporn Sae Lim","Potsawee Manakul","Samuel Cahyawijaya","Can Udomcharoenchaikit","Peerat Limkonchotiwat","Ekapol Chuangsuwanich","Sarana Nutanong"]}},"tmdate":1778035740107,"pdate":1753981200000,"tcdate":1778035740107,"writers":["~Panuthep_Tasawong1","~Napat_Laosaengpha1","~Wuttikorn_Ponwitayarat1","~Sitiporn_Sae_Lim1","~Potsawee_Manakul1","~Samuel_Cahyawijaya1","~Can_Udomcharoenchaikit1","~Peerat_Limkonchotiwat1","~Ekapol_Chuangsuwanich1","~Sarana_Nutanong1"],"signatures":["~Napat_Laosaengpha1"],"forum":"gsG9wjPNa2","license":"CC BY 4.0","number":48711,"cdate":1778035740107,"readers":["everyone"],"invitations":["OpenReview.net/Archive/-/Direct_Upload"],"mdate":1778035740107,"domain":"OpenReview.net/Archive","id":"gsG9wjPNa2","version":2},{"content":{"venue":{"value":"CoRR 2025"},"venueid":{"value":"dblp.org/journals/CORR/2025"},"paperhash":{"value":"guo|shortft_diffusion_model_alignment_via_shortcutbased_finetuning"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Xiefan_Guo:","~Miaomiao_Cui2","https://dblp.org/search/pid/api?q=author:Liefeng_Bo:","https://dblp.org/search/pid/api?q=author:Di_Huang_0001:"]},"html":{"value":"https://doi.org/10.48550/arXiv.2507.22604"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2507-22604,\n  publtype={informal},\n  author={Xiefan Guo and Miaomiao Cui and Liefeng Bo and Di Huang},\n  title={ShortFT: Diffusion Model Alignment via Shortcut-based Fine-Tuning},\n  year={2025},\n  month={July},\n  cdate={1751328000000},\n  journal={CoRR},\n  volume={abs/2507.22604},\n  url={https://doi.org/10.48550/arXiv.2507.22604}\n}\n"},"abstract":{"value":"Backpropagation-based approaches aim to align diffusion models with reward functions through end-to-end backpropagation of the reward gradient within the denoising chain, offering a promising perspective. However, due to the computational costs and the risk of gradient explosion associated with the lengthy denoising chain, existing approaches struggle to achieve complete gradient backpropagation, leading to suboptimal results. In this paper, we introduce Shortcut-based Fine-Tuning (ShortFT), an efficient fine-tuning strategy that utilizes the shorter denoising chain. More specifically, we employ the recently researched trajectory-preserving few-step diffusion model, which enables a shortcut over the original denoising chain, and construct a shortcut-based denoising chain of shorter length. The optimization on this chain notably enhances the efficiency and effectiveness of fine-tuning the foundational model. Our method has been rigorously tested and can be effectively applied to various reward functions, significantly improving alignment performance and surpassing state-of-the-art alternatives."},"title":{"value":"ShortFT: Diffusion Model Alignment via Shortcut-based Fine-Tuning"},"authors":{"value":["Xiefan Guo","Miaomiao Cui","Liefeng Bo","Di Huang"]}},"tmdate":1763000782270,"pdate":1735689600000,"externalIds":["dblp:journals/corr/abs-2507-22604"],"tcdate":1763000777961,"writers":["~"],"signatures":["~miaomiao_cui1"],"forum":"bmNx3Ya54o","license":"CC BY-SA 4.0","number":699104,"cdate":1751328000000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1763000782270,"domain":"DBLP.org","id":"bmNx3Ya54o","version":2},{"content":{"venue":{"value":"Signal Image Video Process. 2026"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/s11760-026-05297-3.pdf"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"jiao|livenet_a_unified_lightweight_network_for_lowlight_image_and_video_enhancement_with_temporal_consistency"},"html":{"value":"https://doi.org/10.1007/s11760-026-05297-3"},"_bibtex":{"value":"@article{DBLP:journals/sivp/JiaoHLJHP26,\n  author={Minghai Jiao and Haoyang He and Yixian Liu and Wenyan Jiang and Jiangang Hu and Yuhuai Peng},\n  title={LIVE-Net: A unified lightweight network for low-light image and video enhancement with temporal consistency},\n  year={2026},\n  month={May},\n  cdate={1777593600000},\n  journal={Signal Image Video Process.},\n  volume={20},\n  number={6},\n  pages={346},\n  url={https://doi.org/10.1007/s11760-026-05297-3}\n}\n"},"abstract":{"value":"Under low-light conditions, images and videos suffer compounded degradation, including loss of dark-region details, color shifts, and noise amplification. Existing methods are typically designed separately for images or videos, making it challenging to balance enhancement quality, temporal consistency, and efficiency in a single framework. This study proposes LIVE-Net, a lightweight down-up sampling backbone that unifies image and video low-light enhancement. Its core components include: (1) the Local Saliency-Aware PoolFormer block, which integrates SimAM for adaptive attention to dark-region outliers while achieving linear-complexity global aggregation via a parameter-free pooling token mixer; (2) the Selective Kernel Fusion module and the Grouped Shuffle Attention Fusion module, enabling content-adaptive detail-semantic trade-offs and robust cross-stage aggregation with minimal cost; (3) Depthwise Convolution Gated Recurrent Unit with a Temporal Difference Consistency loss for lightweight temporal modeling and flicker suppression without optical flow. With only 0.13M parameters and 1.48G FLOPs, LIVE-Net achieves 23.18 dB/0.868 and 26.37 dB/0.943 on LOL-v2 real and synthetic datasets, remains competitive on LOL-v1 and SID, and improves temporal consistency on DID and SDSD while maintaining single-frame quality, demonstrating its unified, robust, and real-time capabilities."},"title":{"value":"LIVE-Net: A unified lightweight network for low-light image and video enhancement with temporal consistency"},"authors":{"value":[{"fullname":"Minghai Jiao","username":""},{"fullname":"Haoyang He","username":""},{"fullname":"Yixian Liu","username":"~Yixian_Liu3"},{"fullname":"Wenyan Jiang","username":""},{"fullname":"Jiangang Hu","username":""},{"fullname":"Yuhuai Peng","username":""}]}},"tmdate":1784973978555,"pdate":1798675200000,"externalIds":["dblp:journals/sivp/JiaoHLJHP26"],"tcdate":1784973974207,"writers":["~","OpenReview.net/Public_Article/DBLP.org","OpenReview.net/Support"],"signatures":["~Yixian_Liu3"],"forum":"tjskH9J8A8","license":"CC BY-SA 4.0","number":104260,"cdate":1777593600000,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/DBLP.org/-/Record","OpenReview.net/Public_Article/-/Edit"],"mdate":1784973978555,"domain":"OpenReview.net/Public_Article","id":"tjskH9J8A8","version":2},{"content":{"summary":{"value":"The paper introduces a novel approach to video understanding by leveraging prompt guidance to enhance the performance of video language models (VLMs). The methodology focuses on reducing video redundancy and extracting key content to enhance the performance of VLMs. It uses CLIP-based visual-prompt alignment to extract relevant visual information and compresses visual sequences using convolution-style pooling. Direct Preference Optimization (DPO) was also used to improve its performance. In experiments, PPLLaVA demonstrated good results on both long and short video benchmarks, and achieves over an 80% compression rate."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. The method seems similar to the involution kernel[1]. What are the  differences between the two? \n2. Have any ablation studies been conducted using regular 3D convolution for pooling instead of prompt-guided methods? \n3. In Table 5, why is the context length for average pooling set to 576 instead of 1024?\n4. Can PPLLaVA  extend to multimodal **generative** models?\n\n[1] Involution: Inverting the Inherence of Convolution for Visual Recognition"},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The paper introduces methods for video understanding by leveraging prompt guidance through the use of CLIP-based visual-prompt alignment and convolution-style pooling.\n2. PPLLaVA shows versatility by performing well on both long and short video benchmarks, demonstrating robust performance for varied video sequence understanding. This adaptability is crucial for handling diverse video lengths and complexities, making it a robust solution for varied video sequence understanding5.\n3. PPLLaVA uses Direct Preference Optimization (DPO) to reduce hallucinations in video-based dialogue and also applies CLIP Context Extension to expand text encoding capacity."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. PPLLaVA does not show leading results compared to some training-free compression methods like SLOWFAST-LLAVA[1], on benchmarks such as MSVD and MSRVTT.\n2. The use of Direct Preference Optimization (DPO) and Proximal Policy Optimization (PPO) lacks innovation. Also, LLaVA-Next-Video also achieve great accuracy improvements using above methods, but this paper does not highlight any unique advantages of these methods within PPLLaVA. \n\n[1] SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models"}},"nonreaders":[],"tmdate":1731427503293,"tcdate":1730656501463,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission1865/Reviewer_Qkbp"],"signatures":["ICLR.cc/2025/Conference/Submission1865/Reviewer_Qkbp"],"forum":"qUZY7ymDPr","number":1,"license":"CC BY 4.0","cdate":1730656501463,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission1865/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427503293,"domain":"ICLR.cc/2025/Conference","replyto":"qUZY7ymDPr","id":"DbXjUMw9uV","forumContent":{"TLDR":{"value":"visual token pooling with instruction guidance, 1/8 throughput, better video&image performance."},"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video LLM","Prompt-guided Pooling","PPLLaVA"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"The past year has witnessed the significant advancement of video-based large language models. However, the challenge of developing a unified model for both short and long video understanding remains unresolved. Most existing video LLMs cannot handle hour-long videos, while methods custom for long videos tend to be ineffective for shorter videos and images. In this paper, we identify the key issue as the redundant content in videos. To address this, we propose a novel pooling strategy that simultaneously achieves token compression and instruction-aware visual feature aggregation. Our model is termed Prompt-guided Pooling LLaVA, or PPLLaVA for short. Specifically, PPLLaVA consists of three core components: the CLIP-based visual-prompt alignment that extracts visual information relevant to the user's instructions, the prompt-guided pooling that compresses the visual sequence to arbitrary scales using convolution-style pooling, and the clip context extension designed for lengthy prompt common in visual dialogue. Moreover, our codebase also integrates the most advanced video Direct Preference Optimization (DPO) and visual interleave training. Extensive experiments have validated the performance of our model. With superior throughput, PPLLaVA achieves better results on image benchmarks as a video LLM, while achieving state-of-the-art performance across various video benchmarks, excelling in tasks ranging from caption generation to multiple-choice questions, and handling video lengths from seconds to hours. The codes are promised to be made public."},"_bibtex":{"value":"@misc{\nliu2025ppllava,\ntitle={{PPLL}a{VA}: Varied Video Sequence Understanding With Prompt Guidance},\nauthor={Ruyang Liu and Chen Li and Haoran Tang and Yixiao Ge and Ying Shan and Haibo Lu and Jiankun Yang},\nyear={2025},\nurl={https://openreview.net/forum?id=qUZY7ymDPr}\n}"},"title":{"value":"PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance"},"pdf":{"value":"/pdf/bb6fc0273484f207f47d34813d971987179b27a5.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"liu|ppllava_varied_video_sequence_understanding_with_prompt_guidance"},"authorids":{"value":["~Ruyang_Liu1","~Chen_Li34","~Haoran_Tang4","~Yixiao_Ge2","~Ying_Shan2","~Haibo_Lu1","~Jiankun_Yang1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Ruyang Liu","Chen Li","Haoran Tang","Yixiao Ge","Ying Shan","Haibo Lu","Jiankun Yang"]}},"version":2},{"content":{"summary":{"value":"This paper proposed SyncFlow to simultaneously generate audio and video from text, where the dual-diffusion transformer structure and modality adapter are designed to jointly model audio and video with temporal synchronization. The empirical evaluations show that the proposed method generates better aligned video and audio."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. Did you compare the proposed method with MM-Diffusion and TAVGBench discussed in the part of the related work? \n2. The text caption used in the proposed method is obtained from video. It is better for T2AV to consider more suitable text captions by describing both audio and video. \n3. The ImageBind score seems to only consider the semantic alignment. Did you try other metrics to evaluate the temporal alignment performance of the proposed model?\n4. Frob Table 2, why the ground truth of the ImageBind score is not optimal like IB Gen-V & GT-T?"},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The paper proposes flow-matching-based models for text-to-audio-video joint generation by designing dual diffusion transformers to handle each modality. \n2. The modality-decoupled multi-stage training strategy is proposed to efficiently make use of data and improve the quality and temporal alignment between generated audio and video. \n3. The paper is generally well-written."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Although the modality adapter is designed to align video with audio, there are no further explanations why just considering unidirectional (video-to-audio) fusion and how the adapter achieves the temporal alignment beyond semantic alignment. \n2. It is a little confusing why text captions are only injected into the video branch rather than both video and audio branches.\n3. Some experiment details are unclear and missing. For example, what are the learning rate and optimizer for (multi-stage) training? Whether the Video VAE and Audio VAE are jointly trained, individually trained, or from frozen pre-trained ones. What are the used text-to-audio (T2A-1 and T2A-1) and video-to-audio (V2A-1 and V2A-2) models? \n4. It is better to present subjective performance by providing the demo or video files."}},"nonreaders":[],"tmdate":1731428915470,"tcdate":1730592590329,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission7555/Reviewer_mJv8"],"signatures":["ICLR.cc/2025/Conference/Submission7555/Reviewer_mJv8"],"forum":"J2EmNMLoxv","number":2,"license":"CC BY 4.0","cdate":1730592590329,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission7555/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428915470,"domain":"ICLR.cc/2025/Conference","replyto":"J2EmNMLoxv","id":"bvs30N7IZy","forumContent":{"TLDR":{"value":"SyncFlow generates synchronized video and audio from text using a dual-diffusion-transformer, with efficient multi-stage training strategy, and strong zero-shot capabilities on V2A and generalization new video resolution."},"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Text-to-Audio-Video-Joint Generation","Flow matching","Diffusion Transformer"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Video and audio are closely correlated modalities that humans naturally perceive together. While recent advancements have enabled the generation of audio or video from text, producing both modalities simultaneously still typically relies on either a cascaded process or multi-modal contrastive encoders. These approaches, however, often lead to suboptimal results due to inherent information losses during inference and conditioning. In this paper, we introduce SyncFlow, a system that is capable of simultaneously generating temporally synchronized audio and video from text. The core of SyncFlow is the proposed dual-diffusion-transformer (d-DiT) architecture, which enables joint video and audio modelling with proper information fusion. To efficiently manage the computational cost of joint audio and video modelling, SyncFlow utilizes a multi-stage training strategy that separates video and audio learning before joint fine-tuning. Our empirical evaluations demonstrate that SyncFlow produces audio and video outputs that are more correlated than baseline methods with significantly enhanced audio quality and audio-visual correspondence. Moreover, we demonstrate strong zero-shot capabilities of SyncFlow, including zero-shot video-to-audio generation and adaptation to novel video resolutions without further training."},"_bibtex":{"value":"@misc{\nliu2025syncflow,\ntitle={SyncFlow: Temporally Aligned Joint Audio-Video Generation from Text},\nauthor={Haohe Liu and Gael Le Lan and Xinhao Mei and Zhaoheng Ni and Anurag Kumar and Varun K. Nagaraja and Wenwu Wang and Mark D Plumbley and Yangyang Shi and Vikas Chandra},\nyear={2025},\nurl={https://openreview.net/forum?id=J2EmNMLoxv}\n}"},"title":{"value":"SyncFlow: Temporally Aligned Joint Audio-Video Generation from Text"},"pdf":{"value":"/pdf/a79408a33fe7d3c044ea05b5acc2491b01e943b3.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"liu|syncflow_temporally_aligned_joint_audiovideo_generation_from_text"},"authorids":{"value":["~Haohe_Liu2","~Gael_Le_Lan1","~Xinhao_Mei1","~Zhaoheng_Ni1","~Anurag_Kumar1","~Varun_K._Nagaraja1","~Wenwu_Wang1","~Mark_D_Plumbley1","~Yangyang_Shi1","~Vikas_Chandra2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Haohe Liu","Gael Le Lan","Xinhao Mei","Zhaoheng Ni","Anurag Kumar","Varun K. Nagaraja","Wenwu Wang","Mark D Plumbley","Yangyang Shi","Vikas Chandra"]}},"version":2},{"content":{"summary":{"value":"This paper proposed a new text-to-video generation method based on latent diffusion model (LDM) on text-to-image generation. The contribution of the proposed method has three folds: 1) A modified LDM UNet with temporal modeling modules for text-to-video generation. 2) A new training dataset for the text-to-video generation extended from a few image and video datasets. 3) A set of strategies for tackling auto-regressive long video generation including FNR, PNS, and DSG. The experimental results on UCF101 show that the proposed method outperforms the compared baselines."},"presentation":{"value":"4 excellent"},"contribution":{"value":"2 fair"},"soundness":{"value":"2 fair"},"strengths":{"value":"- [writing] This paper is well-written. The content is easy to follow and all figures and tables are straight-forward.\n- [method] The way of converting existing video and image datasets into text-to-video training datasets is insightful for other video generation works since high-quality diverse video datasets are expensive to obtain.\n- [method] The auto-regressive long video generation strategy based on noise reshuffle and reuse is interesting.\n- [experiment] The experimental results on UCF101 show that the proposed method can outperform the compared strong baselines in video quality metrics."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- [method] The text-to-video diffusion model based on LDM has been widely studied by Align-Your-Latent,  MagicVideo, and Latent-Shift, where the proposed temporal UNet shares a similar idea of those existing methods that use temporal convolution/attention, resulting in a lack of novelty in the architecture design.\n- [method] For the noise reuse part, PNS and DSG are to directly share part of the noise of the already generated frames, which makes sense to me. However, FNR reverses the noise of the previously generated frames, which I think would contradict the temporal flow of the video. For example, for video frames f_n-2, f_n-1, f_n, f_n+1, f_n+2, letting f_n-2 and f_n-1 be already generated conditioned frames, in FNR, the conditioned frames for generating f_n, f_n+1, f_n+2 would be f_n-1, f_n-2 rather than f_n-2, f_n-1, where the temporal order is reversed. I think the authors should give more explanation here as to why the reversed order is better.\n- [experiment] Only experimental results on UCF101 are provided, while other baselines compared the zero-shot MSRVTT performance, which I think this paper should also follow.\n- [experiment] No user study comparison is presented."},"confidence":{"value":"5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully."},"questions":{"value":"- In the comparison experiments on UCF101, most existing works only train the model on the webvid10M dataset, I wonder what is the training dataset used to produce the results in Table 1.\n- I found the video generated by the proposed method is very flurry, especially the background. Do you have any explanation for it?"},"rating":{"value":"5: marginally below the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636062955,"tcdate":1698721985548,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission1356/Reviewer_btaf"],"signatures":["ICLR.cc/2024/Conference/Submission1356/Reviewer_btaf"],"forum":"qGZTrj48Lb","number":2,"license":"CC BY 4.0","cdate":1698721985548,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission1356/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636062955,"domain":"ICLR.cc/2024/Conference","replyto":"qGZTrj48Lb","id":"2Po2eRkgEY","forumContent":{"TLDR":{"value":"A text-to-video diffusion model is proposed for iteratively generating longer videos."},"venue":{"value":"ICLR 2024 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["diffusion model","video generation","artificial intelligence"]},"primary_area":{"value":"generative models"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Inspired by the remarkable success of Latent Diffusion Models (LDMs) for image synthesis, we study LDM for text-to-video generation, which is a formidable challenge due to the computational and memory constraints during both model training and inference. A single LDM is usually only capable of generating a very limited number of video frames. Some existing works focus on separate prediction models for generating more video frames, which suffer from additional training cost and frame-level jittering, however. In this paper, we propose a framework called \"Reuse and Diffuse\" dubbed *VidRD* to produce more frames following the frames already generated by an LDM. Conditioned on an initial video clip with a small number of frames, additional frames are iteratively generated by reusing the original latent features and following the previous diffusion process. Besides, for the autoencoder used for translation between pixel space and latent space, we inject temporal layers into its decoder and fine-tune these layers for higher temporal consistency. We also propose a set of strategies for composing video-text data that involve diverse content from multiple existing datasets including video datasets for action recognition and image-text datasets. Extensive experiments show that our method achieves good results in both quantitative and qualitative evaluations. Our project page is available at [https://anonymous0x233.github.io/ReuseAndDiffuse/](https://anonymous0x233.github.io/ReuseAndDiffuse/)."},"_bibtex":{"value":"@misc{\ngu2024reuse,\ntitle={Reuse and Diffuse: Iterative Denoising for Text-to-Video Generation},\nauthor={Jiaxi Gu and Shicong Wang and Haoyu Zhao and Tianyi Lu and Xing Zhang and Zuxuan Wu and Songcen Xu and Wei Zhang and Yu-Gang Jiang and Hang Xu},\nyear={2024},\nurl={https://openreview.net/forum?id=qGZTrj48Lb}\n}"},"title":{"value":"Reuse and Diffuse: Iterative Denoising for Text-to-Video Generation"},"pdf":{"value":"/pdf/3e4a69e35dc3f76fa87869adade1d326bb1fe7f3.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Withdrawn_Submission"},"paperhash":{"value":"gu|reuse_and_diffuse_iterative_denoising_for_texttovideo_generation"},"authorids":{"value":["~Jiaxi_Gu1","~Shicong_Wang1","~Haoyu_Zhao2","~Tianyi_Lu1","~Xing_Zhang5","~Zuxuan_Wu1","~Songcen_Xu1","~Wei_Zhang45","~Yu-Gang_Jiang1","~Hang_Xu1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Jiaxi Gu","Shicong Wang","Haoyu Zhao","Tianyi Lu","Xing Zhang","Zuxuan Wu","Songcen Xu","Wei Zhang","Yu-Gang Jiang","Hang Xu"]}},"version":2},{"content":{"summary":{"value":"This paper introduces a new video editing task called generative video compositing and proposes the first feasible solution, GenCompositor. Based on the Diffusion Transformer architecture, this method can automatically composite dynamic elements from a foreground video into a background video according to user-specified trajectories and dimensions."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"Please refer to the weaknesses."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. This work is the first to introduce the \"generative video compositing\" task, which aims to automatically composite foreground video elements into a background video using generative models while supporting user control over attributes such as trajectory and scale. \n2. A complete architecture is proposed, including a background preservation branch, DiT fusion blocks, and Extended Rotary Position Embedding (ERoPE), aiming to address the three main challenges: background consistency, foreground injection, and user control."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Appendix F shows that extending RoPE along height, width, or temporal dimensions yields nearly identical training loss curves, leading authors to conclude these three directions are equivalent. However, this dimension-agnostic behavior seems contradictory to the inherent spatial-temporal characteristics of videos. If ERoPE truly models spatial layout relationships, spatial extensions (height/width) should outperform temporal extension since they directly encode spatial proximity. Conversely, if it captures motion dynamics, temporal extension should be more effective given the causal and directional nature of video frames. The complete equivalence across all three dimensions suggests that ERoPE may simply function by expanding the position encoding to avoid feature conflicts, rather than genuinely modeling spatial-temporal interactions between foreground and background. Could the authors clarify this dimension-independence phenomenon and explain whether simpler alternatives (such as learnable position offsets for different video sources) might achieve comparable results?\n\n2. The practical usage requires SAM2 for foreground segmentation and optical flow algorithms for trajectory tracking, and these preprocessing steps may introduce errors and increase system complexity, yet the paper does not discuss how these errors propagate and accumulate to affect the final results.\n\n3. The paper only compares its method with two video harmonization approaches and two trajectory-controlled generation methods, but fails to include comparisons with recent state-of-the-art methods in related tasks such as video object insertion, which share similarities with generative video compositing."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762915818305,"tcdate":1761911529861,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission1570/Reviewer_sd3W"],"signatures":["ICLR.cc/2026/Conference/Submission1570/Reviewer_sd3W"],"forum":"ynim5u2N4i","number":4,"license":"CC BY 4.0","cdate":1761911529861,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission1570/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762915818305,"domain":"ICLR.cc/2026/Conference","replyto":"ynim5u2N4i","id":"SNy99k7IGg","forumContent":{"TLDR":{"value":"GenCompositor is capable of effortlessly compositing different videos guided by user-specified trajectories and scales."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Diffusion Models","Video Editing","Video Compositing"]},"supplementary_material":{"value":"/attachment/27f18c510c51d9af8b342e392cb17ce1d09641cc.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Video compositing combines live-action footage to create video production, serving as a crucial technique in video creation and film production. Traditional pipelines require intensive labor efforts and expert collaboration, resulting in lengthy production cycles and high manpower costs. To address this issue, we automate this process with generative models, called generative video compositing. This new task strives to adaptively inject identity and motion information of foreground video to the target video in an interactive manner, allowing users to customize the size, motion trajectory, and other attributes of the dynamic elements added in final video. Specifically, we designed a novel Diffusion Transformer (DiT) pipeline based on its intrinsic properties. To maintain consistency of the target video before and after editing, we revised a light-weight DiT-based background preservation branch with masked token injection. As to inherit dynamic elements from other sources, a DiT fusion block is proposed using full self-attention, along with a simple yet effective foreground augmentation for training. Besides, for fusing background and foreground videos with different layouts based on user control, we developed a novel position embedding, named Extended Rotary Position Embedding (ERoPE). Finally, we curated a dataset comprising 61K sets of videos for our new task, called VideoComp. This data includes complete dynamic elements and high-quality target videos. Experiments demonstrate that our method effectively realizes generative video compositing, outperforming existing possible solutions in fidelity and consistency. Project is available at https://gencompositor.github.io/"},"_bibtex":{"value":"@inproceedings{\nyang2026gencompositor,\ntitle={GenCompositor: Generative Video Compositing with Diffusion Transformer},\nauthor={Shuzhou Yang and Xiaoyu Li and Xiaodong Cun and Guangzhi Wang and Lingen Li and Ying Shan and Jian Zhang},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=ynim5u2N4i}\n}"},"title":{"value":"GenCompositor: Generative Video Compositing with Diffusion Transformer"},"pdf":{"value":"/pdf/cfd1ebc349fc91f9a0fce0b4ac6a434601465415.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"yang|gencompositor_generative_video_compositing_with_diffusion_transformer"},"authorids":{"value":["~Shuzhou_Yang1","~Xiaoyu_Li2","~Xiaodong_Cun1","~Guangzhi_Wang1","~Lingen_Li1","~Ying_Shan2","~Jian_Zhang22"]},"authors":{"value":["Shuzhou Yang","Xiaoyu Li","Xiaodong Cun","Guangzhi Wang","Lingen Li","Ying Shan","Jian Zhang"]}},"version":2},{"content":{"summary":{"value":"This paper proposes TS-Attn, a training-free attention mechanism that improves multi-event video generation by dynamically separating and modulating cross-attention between motion regions and multi-event textual conditions. By introducing motion region extraction and event-aware attention modulation, the method reduces temporal misalignment and cross-event coupling, achieving better temporal coherence and event accuracy. TS-Attn can be plugged into existing diffusion-based video models without retraining, yielding substantial performance gains on StoryEval-Bench with only ~2% extra inference cost."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"See weakness. Overall, I lean towards borderline for the current version and am happy to update my rating if my questions are well answered."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- Well-motivated and intuitive idea that directly targets temporal attention entanglement in multi-event video generation.\n\n- Training-free and plug-and-play design makes it broadly applicable across existing diffusion models.\n\n- Extensive experiments and ablations demonstrate consistent improvements and robustness across architectures and benchmarks.\n\n- Clear presentation and visualizations that effectively explain both the mechanism and empirical benefits."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"**Lack of comparison with prior methods**\nThe paper omits several highly relevant works that also manipulate cross-attention maps to achieve fine-grained event grounding without retraining, such as DreamRunner [1], VideoTetris [2], and TALC [3].\nThese methods similarly align textual tokens with corresponding visual regions through attention reweighting, making them conceptually close to TS-Attn. However, the authors neither cite nor compare with them. Including these approaches as baselines or at least discussing their differences would strengthen the paper’s positioning and contribution clarity.\n\n**Unclear generalization to multi-subject scenarios**\nThe proposed motion-region extraction appears to assume a single dominant subject, computing masks for the entire video latents.\nIn cases involving subject transitions (e.g., a person leaves and a cat enters), this design may fail to isolate subject-specific motion regions, leading to incorrect or conflicting event-to-visual grounding.\nClarifying how TS-Attn handles such cases, or showing examples involving multiple subjects, would improve the completeness of the work.\n\n---\n[1] DreamRunner: Fine-Grained Compositional Story-to-Video Generation with Retrieval-Augmented Motion Adaptation, 2024.\n\n[2] VideoTetris: Towards Compositional Text-to-Video Generation, 2024.\n\n[3] TALC: Time-Aligned Captions for Multi-Scene Text-to-Video Generation, 2024."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762915735586,"tcdate":1762005184904,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission1323/Reviewer_Qohm"],"signatures":["ICLR.cc/2026/Conference/Submission1323/Reviewer_Qohm"],"forum":"QixNhagZ9t","number":3,"license":"CC BY 4.0","cdate":1762005184904,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission1323/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762915735586,"domain":"ICLR.cc/2026/Conference","replyto":"QixNhagZ9t","id":"KfMclvwXca","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Video generation","Diffusion model"]},"supplementary_material":{"value":"/attachment/12f78047d142448e02e683183762ae698be23030.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Generating high-quality videos from complex temporal descriptions, which refer to prompts containing multiple sequential actions, remains a significant challenge. Existing methods are constrained by an inherent trade-off: using multiple short prompts fed sequentially into the model improves action fidelity but compromises temporal consistency, while a single complex prompt preserves consistency at the cost of prompt following capability. We attribute this problem to two primary causes: temporal misalignment between video content and the prompt, and conflicting attention coupling between motion-related visual objects and their associated text conditions. To address these challenges, we propose a novel, training-free attention mechanism, Temporal-wise Separable Attention (TS-Attn), which dynamically rearranges attention distribution to ensure temporal awareness and global coherence in multi-event scenarios. TS-Attn can be seamlessly integrated into various pre-trained text-to-video models, boosting StoryEval-Bench scores by 33.5% and 16.4% on Wan2.1-T2V-14B and Wan2.2-T2V-A14B with only a 2% increase in inference time. It also supports plug-and-play usage across models for multi-event image-to-video generation. The source code and video demos are available in the supplementary materials."},"_bibtex":{"value":"@inproceedings{\nzhang2026tsattn,\ntitle={{TS}-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation},\nauthor={Hongyu Zhang and Yufan Deng and Zilin Pan and Peng-Tao Jiang and Bo Li and Qibin Hou and Zhen Dong and Zhiyang Dou and Daquan Zhou},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=QixNhagZ9t}\n}"},"title":{"value":"TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation"},"pdf":{"value":"/pdf/6b996f3ab25840409fd41e7a0f4dc1a8af081b11.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"zhang|tsattn_temporalwise_separable_attention_for_multievent_video_generation"},"authorids":{"value":["~Hongyu_Zhang6","~Yufan_Deng4","~Zilin_Pan1","~Peng-Tao_Jiang1","~Bo_Li20","~Qibin_Hou1","~Zhen_Dong3","~Zhiyang_Dou1","~Daquan_Zhou1"]},"authors":{"value":["Hongyu Zhang","Yufan Deng","Zilin Pan","Peng-Tao Jiang","Bo Li","Qibin Hou","Zhen Dong","Zhiyang Dou","Daquan Zhou"]}},"version":2},{"content":{"venue":{"value":"IEEE Access 2018"},"pdf":{"value":"https://ieeexplore.ieee.org/iel7/6287639/8274985/08424025.pdf"},"venueid":{"value":"dblp.org/journals/ACCESS/2018"},"paperhash":{"value":"aliyu|interferenceaware_multipath_video_streaming_in_vehicular_environments"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Ahmed_Aliyu:","https://dblp.org/search/pid/api?q=author:Abdul_Hanan_Abdullah:","https://dblp.org/search/pid/api?q=author:Nauman_Aslam:","https://dblp.org/search/pid/api?q=author:Ayman_Altameem:","https://dblp.org/search/pid/api?q=author:Raja_Zahilah_Radzi:","https://dblp.org/search/pid/api?q=author:Rupak_Kharel:","~Mufti_Mahmud1","https://dblp.org/search/pid/api?q=author:Shiv_Prakash:","https://dblp.org/search/pid/api?q=author:Mohammed_Joda_Usman:"]},"html":{"value":"https://doi.org/10.1109/ACCESS.2018.2854784"},"_bibtex":{"value":"@article{DBLP:journals/access/AliyuAAARKMPU18,\n  author={Ahmed Aliyu and Abdul Hanan Abdullah and Nauman Aslam and Ayman Altameem and Raja Zahilah Radzi and Rupak Kharel and Mufti Mahmud and Shiv Prakash and Mohammed Joda Usman},\n  title={Interference-Aware Multipath Video Streaming in Vehicular Environments},\n  year={2018},\n  cdate={1514764800000},\n  journal={IEEE Access},\n  volume={6},\n  pages={47610-47626},\n  url={https://doi.org/10.1109/ACCESS.2018.2854784}\n}\n"},"abstract":{"value":"The multipath transmission is one of the suitable transmission methods for high data rate oriented communication such as video streaming. Each video packets are split into smaller frames for parallel transmission via different paths. One path may interfere with another path due to these parallel transmissions. The multipath oriented interference is due to the route coupling which is one of the major challenges in vehicular traffic environments. The route coupling increases channel contention resulting in video packet collision. In this context, this paper proposes an Interference-aware Multipath Video Streaming (I-MVS) framework focusing on link and node disjoint optimal paths. Specifically, a multipath vehicular network model is derived. The model is utilized to develop interference-aware video streaming method considering angular driving statistics of vehicles. The quality of video streaming links is measured based on packet error rate considering non-circular transmission range oriented shadowing effects. Algorithms are developed as a complete operational I-MVS framework. The comparative performance evaluation attests the benefit of the proposed framework considering various video streaming related metrics."},"title":{"value":"Interference-Aware Multipath Video Streaming in Vehicular Environments"},"authors":{"value":["Ahmed Aliyu","Abdul Hanan Abdullah","Nauman Aslam","Ayman Altameem","Raja Zahilah Radzi","Rupak Kharel","Mufti Mahmud","Shiv Prakash","Mohammed Joda Usman"]}},"tmdate":1746242953745,"pdate":1514764800000,"tcdate":1746242865396,"writers":["~"],"signatures":["~Mufti_Mahmud1"],"forum":"bq83cr3Of7","license":"CC BY-SA 4.0","number":409564,"cdate":1514764800000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1746242953745,"domain":"DBLP.org","id":"bq83cr3Of7","version":2},{"content":{"summary":{"value":"This paper focuses on analyzing the strengths and limitations of existing video diffusion models for different video understanding tasks. A probing framework is proposed to extract the video feature representations from existing pre-trained models and learn lightweight task heads for different downstream video understanding tasks.  Comprehensive experiments across six models and four downstream video tasks are conducted and some findings are demonstrated."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"please see the weaknesses section."},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1) The investigation of different characteristics of image-/video-based diffusions is interesting and meaningful for the video-understanding community;\n2) The proposed probing framework is technically sound and the experimental results are comprehensive.\n3) The paper is well-written, making it easy to follow and understand."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1) Although many findings are demonstrated by the comprehensive results, **the reviewer finds these findings are well-known consensus in the video understanding community**. For example, modeling the human motion or object dynamics is important for analyzing the videos, it is straightforward that most video-based diffusions contribute their models to this problem. Therefore, for instance, the findings \"video diffusion models demonstrate exceptional proficiency in capturing motion patterns and temporal dynamics\" that the paper try to demonstrate, **cannot provide new insights to the community**;\n\n2) The used video action recognition datasets, i.e., UCF and HMDB, actually do not require modeling too much temporal/dynamic semantics for recognition. To make the demonstration more convincing, the reviewer thinks experiments should be conducted on the action datasets that require effective temporal dynamics modeling (e.g., something v2, Kinetics).\n\nminor issues:\n1) the symbol zT in line #231 should be corrected;\n2) the sentence in line #289~290 should be revised;"}},"nonreaders":[],"tmdate":1731427291033,"tcdate":1730472721599,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission633/Reviewer_uAKH"],"signatures":["ICLR.cc/2025/Conference/Submission633/Reviewer_uAKH"],"forum":"SIZhZrU41O","number":2,"license":"CC BY 4.0","cdate":1730472721599,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission633/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427291033,"domain":"ICLR.cc/2025/Conference","replyto":"SIZhZrU41O","id":"SdnBoOiA4X","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Diffusion Models","Video Understanding","Representation Learning"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Diffusion models have demonstrated significant progress in visual perception tasks due to their ability to capture fine-grained, object-centric features through large-scale vision-language pretraining. While their success in image-based tasks is well-established, extending this capability to the domain of video understanding remains a key challenge.  In this work, we explore the potential of diffusion models for video understanding by analyzing the feature representations learned by both image- and video-based diffusion models, alongside non-generative, self-supervised approaches. We propose a unified probing framework to evaluate six models across four core video understanding tasks: action recognition, object discovery, scene understanding, and label propagation. Our findings reveal that video diffusion models consistently rank among the top performers, particularly excelling at modeling temporal dynamics and scene structure. This observation not only sets them apart from image-based diffusion models but also opens a new direction for advancing video understanding, offering a fresh alternative to traditional discriminative pre-training objectives. Interestingly, we demonstrate that higher generation performance does not always correlate with improved performance in downstream tasks, highlighting the importance of careful representation selection. Overall, our results suggest that video diffusion models hold substantial promise for video understanding by effectively capturing both spatial and temporal information, positioning them as strong competitors in this evolving domain."},"_bibtex":{"value":"@misc{\nbao2024video,\ntitle={Video Diffusion Models Learn the Structure of the Dynamic World},\nauthor={Zhipeng Bao and Anurag Bagchi and Yu-Xiong Wang and Pavel Tokmakov and Martial Hebert},\nyear={2024},\nurl={https://openreview.net/forum?id=SIZhZrU41O}\n}"},"title":{"value":"Video Diffusion Models Learn the Structure of the Dynamic World"},"pdf":{"value":"/pdf/fdbdf04b193cc1eedaf473e0f0fb5b83b1b813ab.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"bao|video_diffusion_models_learn_the_structure_of_the_dynamic_world"},"authorids":{"value":["~Zhipeng_Bao1","~Anurag_Bagchi1","~Yu-Xiong_Wang1","~Pavel_Tokmakov2","~Martial_Hebert1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Zhipeng Bao","Anurag Bagchi","Yu-Xiong Wang","Pavel Tokmakov","Martial Hebert"]}},"version":2},{"content":{"venue":{"value":"ICDCN 2014"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-642-45249-9_28.pdf"},"venueid":{"value":"dblp.org/conf/ICDCN/2014"},"paperhash":{"value":"ahmedin|exploiting_scalable_video_coding_for_content_aware_downlink_video_delivery_over_lte"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Ahmed_Ahmedin:","https://dblp.org/search/pid/api?q=author:Kartik_Pandit:","~Dipak_Ghosal1","https://dblp.org/search/pid/api?q=author:Amitabha_Amitava_Ghosh:"]},"html":{"value":"https://doi.org/10.1007/978-3-642-45249-9_28"},"_bibtex":{"value":"@inproceedings{DBLP:conf/icdcn/AhmedinPGG14,\n  author={Ahmed Ahmedin and Kartik Pandit and Dipak Ghosal and Amitabha Amitava Ghosh},\n  title={Exploiting Scalable Video Coding for Content Aware Downlink Video Delivery over LTE},\n  year={2014},\n  cdate={1388534400000},\n  pages={423-437},\n  url={https://doi.org/10.1007/978-3-642-45249-9_28},\n  booktitle={ICDCN},\n  crossref={conf/icdcn/2014}\n}\n"},"abstract":{"value":"We propose a content aware scheduler to allocate resources for video delivery on the downlink of a Long Term Evolution (LTE) network. We consider multiple users subscribe to a video streaming service, and request videos encoded in H.264 Scalable Video Coding format. The scheduler maximizes the average video quality across all users by assigning resource blocks based on their device capabilities, link qualities, and available resources. We measure video quality using two full reference metrics: peak signal-to-noise ratio (PSNR) and structural similarity (SSIM) index. We formulate the video delivery problem first as an integer linear program (ILP), and then reduce it to the multiple choice knapsack problem (MCKP). To solve the MCKP, we propose two fast heuristics with reduced processing overhead at the eNodeB, and a fully polynomial-time approximate scheme (FPTAS) using dynamic programming and profit-scaling. Our evaluation results indicate that the heuristics are within a factor of \\(\\frac{1}{2}\\), and the FPTAS is very close to the optimal obtained from an ILP solver. We also propose a signaling mechanism to implement the content aware scheduler in existing LTE systems, and evaluate the impact of signaling delay on video distortion using both indoor and outdoor measurements collected from AT&T and T-Mobile networks."},"title":{"value":"Exploiting Scalable Video Coding for Content Aware Downlink Video Delivery over LTE"},"authors":{"value":["Ahmed Ahmedin","Kartik Pandit","Dipak Ghosal","Amitabha Amitava Ghosh"]}},"tmdate":1769371455664,"pdate":1419984000000,"externalIds":["dblp:conf/icdcn/AhmedinPGG14"],"tcdate":1769371424817,"writers":["~"],"signatures":["~Dipak_Ghosal1"],"forum":"f6D8PqHy33","license":"CC BY-SA 4.0","number":801015,"cdate":1388534400000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1769371455664,"domain":"DBLP.org","id":"f6D8PqHy33","version":2},{"content":{"TLDR":{"value":"Extract mixture of interpretable model from a Blackbox and provide FOL explanation to mitigate shortcut learning."},"venue":{"value":"SCIS 2023 Poster"},"pdf":{"value":"/pdf/54a221e229465c0fe965e2ecf2a7592ea3b8d618.pdf"},"keywords":{"value":["Explainable AI","Posthoc Explanation","Interpretability","Shortcuts"]},"venueid":{"value":"ICML.cc/2023/Workshop/SCIS"},"abstract":{"value":"We use concept-based interpretable models to mitigate shortcut learning. Existing methods lack interpretability.\nBeginning with a Blackbox, we iteratively carve out a mixture of interpretable experts (MoIE) and a residual network. Each expert explains a subset of data using First Order Logic (FOL). While explaining a sample, the FOL from biased BB-derived MoIE detects the shortcut effectively. Finetuning the BB with Metadata Normalization (MDN) eliminates the shortcut. The FOLs from the finetuned-BB-derived MoIE verify the elimination of the shortcut. Our experiments show that MoIE does not hurt the accuracy of the original BB and eliminates shortcuts effectively."},"title":{"value":"Tackling Shortcut Learning in Deep Neural Networks: An Iterative Approach with Interpretable Models"}},"tmdate":1690513095311,"pdate":1687228756359,"tcdate":1684794307342,"writers":["ICML.cc/2023/Workshop/SCIS","ICML.cc/2023/Workshop/SCIS/Submission20/Authors"],"signatures":["ICML.cc/2023/Workshop/SCIS/Submission20/Authors"],"forum":"m5vnLHfNy7","number":20,"cdate":1684794307342,"mdate":1690513095311,"readers":["everyone"],"invitations":["ICML.cc/2023/Workshop/SCIS/-/Submission","ICML.cc/2023/Workshop/SCIS/-/Post_Submission","ICML.cc/2023/Workshop/SCIS/-/Edit","ICML.cc/2023/Workshop/SCIS/Submission20/-/Post-Decision_Revision"],"domain":"ICML.cc/2023/Workshop/SCIS","id":"m5vnLHfNy7","version":2},{"content":{"summary":{"value":"This paper introduces an innovative set of methods designed to enhance video generation. These include the 3D-MBQ-VAE, which improves video compression, a novel text-to-video generation framework, a spectral transformer-based denoising network for superior video generation, and a Sketch Guided Video Inpainting task that utilizes Low-Rank Adaptation (LoRA) for efficient fine-tuning."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"1.Based on my understanding, the videos in the WebVid10M dataset all contain watermarks. I'm curious as to how the videos generated by the model, which was trained on this dataset, do not have any watermarks. Can the authors explain this?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1.The authors propose a denoising network named Spectral Transformer that processes video latents in the frequency domain using the Fourier Transform, capturing global context and long-range dependencies effectively.\n\n2.The paper is the first to address the downstream task of sketch-guided video inpainting, using LORA adaptors for parameter-efficient fine-tuning of the denoising network."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1.3D-MBQ-VAE Concerns:\n\n(1)W.A.L.T[1] has demonstrated that using a 3D casual VAE can allow for joint training with both images and videos in text-to-video models, significantly improving performance compared to training solely with videos. Can the authors discuss whether the proposed 3D-MBQ-VAE can also support such joint image and video training? If feasible, it would be beneficial to see an experiment demonstrating this capability and comparing it with that of W.A.L.T.\n\n(2) In the comparative experiments listed in Table 1, could the authors provide comparative results of MAGVIT-v2[2] in regard to Video Compression Metrics or Video Reconstruction tasks?\n\n (3) Codebook (Vocabulary) Size Influence: As far as I understand, the size of the codebook could considerably affect the 3D-MBQ-VAE's reconstruction ability. However, the paper doesn't seem to mention any information relating to the codebook size. Could the authors provide additional details on the impact of varying codebook sizes on video reconstruction?\n\n2.Experiment Setting for Text-to-Video Model: In Table 3, the authors used WebVID10M as the training data and also computed FVD and other metrics on WebVID10M for testing. However, the comparative methods did not use WebVID10M in their training data, which suggests that other text-to-video methods computed FVD and other metrics in a zero-shot manner. In addition, specifics like the sampling schedule, sampling steps, and classifier-free guidance scale used by the authors' proposed method and the comparison methods were not disclosed. To ensure a fair comparison, could the authors evaluate all models (including theirs) in a zero-shot manner on a common test set, like FVD of UCF-101? Also, providing a table or an appendix detailing all the hyperparameters and settings for each method would greatly enhance reproducibility and fairness in comparison.\n\n------------\n\n[1].Gupta, Agrim, et al. \"Photorealistic video generation with diffusion models.\" arXiv preprint arXiv:2312.06662 (2023).\n\n[2].Yu, Lijun, et al. \"Language Model Beats Diffusion--Tokenizer is Key to Visual Generation.\" arXiv preprint arXiv:2310.05737 (2023)."}},"nonreaders":[],"tmdate":1732523876378,"tcdate":1730536793151,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission1775/Reviewer_UAh7"],"signatures":["ICLR.cc/2025/Conference/Submission1775/Reviewer_UAh7"],"forum":"bW9fGYo44s","number":2,"license":"CC BY 4.0","cdate":1730536793151,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission1775/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732523876378,"domain":"ICLR.cc/2025/Conference","replyto":"bW9fGYo44s","id":"xhDvvppNNT","forumContent":{"TLDR":{"value":"High Quality text to video generation with discrete diffusion"},"venue":{"value":"ICLR 2025 Spotlight"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["text2video","VQ-Diffusion","video Inpainting","Large scale pretraining"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"The spatio-temporal complexity of video data presents significant challenges in tasks such as compression, generation, and inpainting. We present four key contributions to address the challenges of spatiotemporal video processing. First, we introduce the 3D Mobile Inverted Vector-Quantization Variational Autoencoder (3D-MBQ-VAE), which combines Variational Autoencoders (VAEs) with masked modeling to enhance spatiotemporal video compression. The model achieves superior temporal consistency and state-of-the-art (SOTA) reconstruction quality by employing a novel training strategy with full frame masking. Second, we present MotionAura, a text-to-video generation framework that utilizes vector-quantized diffusion models to discretize the latent space and capture complex motion dynamics, producing temporally coherent videos aligned with text prompts. Third, we propose a spectral transformer-based denoising network that processes video data in the frequency domain using the Fourier Transform. This method effectively captures global context and long-range dependencies for high-quality video generation and denoising. Lastly, we introduce a downstream task of Sketch Guided Video Inpainting. This task leverages Low-Rank Adaptation (LoRA) for parameter-efficient fine-tuning. Our models achieve SOTA performance on a range of benchmarks.  Our work offers robust frameworks for spatiotemporal modeling and user-driven video content manipulation."},"_bibtex":{"value":"@inproceedings{\nsusladkar2025motionaura,\ntitle={MotionAura: Generating High-Quality and Motion Consistent Videos using Discrete Diffusion},\nauthor={Onkar Kishor Susladkar and Jishu Sen Gupta and Chirag Sehgal and Sparsh Mittal and Rekha Singhal},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=bW9fGYo44s}\n}"},"title":{"value":"MotionAura: Generating High-Quality and Motion Consistent Videos using Discrete Diffusion"},"pdf":{"value":"/pdf/45b7da6fce6fc272027212b6c17632d43a4c1530.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"susladkar|motionaura_generating_highquality_and_motion_consistent_videos_using_discrete_diffusion"},"authorids":{"value":["~Onkar_Kishor_Susladkar1","~Jishu_Sen_Gupta1","~Chirag_Sehgal1","~Sparsh_Mittal1","~Rekha_Singhal1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Onkar Kishor Susladkar","Jishu Sen Gupta","Chirag Sehgal","Sparsh Mittal","Rekha Singhal"]}},"version":2},{"content":{"summary":{"value":"This paper proposes using a pre-trained video diffusion model to generate multi-view images for reconstruction. Additionally, a 3D-aware video-to-video refiner is trained to enhance the multi-view images to higher resolutions. Both models are trained in two stages on the Objaverse dataset using the A100 GPU for three days."},"suitability":{"value":3},"strengths":{"value":"1. Experiments validate the effectiveness of the architecture, including the 3D-aware Multi-view Refinement stage and the interpolation view number 𝑀 in 3D reconstruction.\n\n2. Hi3D can generate diverse and plausible instances, each with distinct geometric structures or textures, using a video diffusion model."},"confidence":{"value":3},"rating":{"value":2},"limitations":{"value":"1. Lack of novelty in this approach, as many similar works, such as VFusion3D, V3D, SV3D,  have already been conducted using video models for 3D generation.\n\n2. The pipeline is overly redundant, requiring three stages: Multi-view Generation, 3D-aware Multi-view Refinement, and 3D Mesh Extraction, which includes Gaussian splatting and SDF-based reconstruction. The inference efficiency is too low.\n\n3. There is a lack of experimentation, making the results unconvincing. Additionally, there is no comparison with other video model-based methods, such as VFusion3D, and V3D."}},"nonreaders":[],"tmdate":1721657404323,"tcdate":1716262280125,"writers":["acmmm.org/ACMMM/2024/Conference","acmmm.org/ACMMM/2024/Conference/Submission5194/Reviewer_ZhWA"],"signatures":["acmmm.org/ACMMM/2024/Conference/Submission5194/Reviewer_ZhWA"],"forum":"VEHNTupyIU","number":2,"license":"CC BY 4.0","cdate":1716262280125,"readers":["everyone"],"invitations":["acmmm.org/ACMMM/2024/Conference/Submission5194/-/Official_Review","acmmm.org/ACMMM/2024/Conference/-/Edit","acmmm.org/ACMMM/2024/Conference/Submission5194/Official_Review2/-/Final_Rating"],"mdate":1721657404323,"domain":"acmmm.org/ACMMM/2024/Conference","replyto":"VEHNTupyIU","id":"4KaD376Usm","forumContent":{"venue":{"value":"MM2024 Poster"},"supplementary_material":{"value":"/attachment/ecfc245e2ed05acceb6ca817a4b7704a31f61091.zip"},"abstract":{"value":"Despite having tremendous progress in image-to-3D generation, existing methods still struggle to produce multi-view consistent images with high-resolution textures in detail, especially in the paradigm of 2D diffusion that lacks 3D awareness. In this work, we present High-resolution Image-to-3D model (Hi3D), a new video diffusion based paradigm that redefines a single image to multi-view images as 3D-aware sequential image generation (i.e., orbital video generation). This methodology delves into the underlying temporal consistency knowledge in video diffusion model that generalizes well to geometry consistency across multiple views in 3D generation. Technically, Hi3D first empowers the pre-trained video diffusion model with 3D-aware prior (camera pose condition), yielding multi-view images with low-resolution texture details. A 3D-aware video-to-video refiner is learnt to further scale up the multi-view images with high-resolution texture details. Such high-resolution multi-view images are further augmented with novel views through 3D Gaussian Splatting, which are finally leveraged to obtain high-fidelity meshes via 3D reconstruction. Extensive experiments on both novel view synthesis and single view reconstruction demonstrate that our Hi3D manages to produce superior multi-view consistency images with highly-detailed textures."},"relevance_to_conference":{"value":"This work aims to tackle the challenging multimedia content creation task (image-to-3D). The created 3D content can be seamlessly combined with other forms of multimedia data like audio, video, and text. Such multimodal processing allows for the creation of richer and more interactive experiences."},"_bibtex":{"value":"@inproceedings{\nyang2024hid,\ntitle={Hi3D: Pursuing High-Resolution Image-to-3D Generation with Video Diffusion Models},\nauthor={Haibo Yang and Yang Chen and Yingwei Pan and Ting Yao and Zhineng Chen and Chong-Wah Ngo and Tao Mei},\nbooktitle={ACM Multimedia 2024},\nyear={2024},\nurl={https://openreview.net/forum?id=VEHNTupyIU}\n}"},"title":{"value":"Hi3D: Pursuing High-Resolution Image-to-3D Generation with Video Diffusion Models"},"pdf":{"value":"/pdf/ee8b1ac5d9ff251982b30845c541dda807383cb8.pdf"},"venueid":{"value":"acmmm.org/ACMMM/2024/Conference"},"paperhash":{"value":"yang|hi3d_pursuing_highresolution_imageto3d_generation_with_video_diffusion_models"},"primary_subject_area":{"value":"[Generation] Generative Multimedia"},"authorids":{"value":["~Haibo_Yang3","~Yang_Chen20","~Yingwei_Pan1","~Ting_Yao1","~Zhineng_Chen1","~Chong-Wah_Ngo4","~Tao_Mei3"]},"authors":{"value":["Haibo Yang","Yang Chen","Yingwei Pan","Ting Yao","Zhineng Chen","Chong-Wah Ngo","Tao Mei"]}},"version":2},{"content":{"venue":{"value":"CVPR 2026"},"abstract":{"value":"Generating high-fidelity audio that is both semantically meaningful and temporally synchronized with silent videos remains a challenging problem in video-to-audio generation. Existing approaches often fail to capture fine-grained temporal correspondence between visual events and audio dynamics, leading to unrealistic or desynchronized outputs. To address these limitations, we propose VisioSonic, a Video-Aligned Sound generation framework that unifies flow-matching diffusion and preference-guided alignment. VisioSonic introduces a multimodal conditioning module that jointly leverages video frames and textual cues to provide semantic and frame-level temporal guidance. A co-attention diffusion transformer efficiently fuses visual and audio representations, enabling content-aware sound synthesis with minimal computation costs. To further enhance alignment beyond supervised training, we introduce Semantic-Temporal Alignment Ranked Direct Preference Optimization (STAR-DPO), a novel preference-learning paradigm that automatically generates audio candidates, ranks them based on both semantic and temporal alignment, and subsequently fine-tunes the diffusion model using the derived preference pairs. Extensive experiments on various benchmarks demonstrate that VisioSonic achieves state-of-the-art audio-video synchronization and audio fidelity while using the fewest trainable parameters among competing approaches. Project page: https://kaiw7.github.io/VisioSonic/"},"_bibtex":{"value":"@inproceedings{\nwang2026hear,\ntitle={Hear What You See: Video-to-Audio Generation with Diffusion Transformer and Semantic-Temporal Alignment-Ranked Direct Preference Optimization},\nauthor={Kai Wang and Tao Zhou and Jiayi Lei and Jing Wang and Jinman Zhao and Weiguo Pian and Yuan Cheng and Yapeng Tian and Peng Gao and Bin Fu and Yihao Liu and Dimitrios Hatzinakos and Yuewen Cao},\nbooktitle={Conference on Computer Vision and Pattern Recognition 2026},\nyear={2026},\nurl={https://openreview.net/forum?id=TWJQCh0awj}\n}"},"title":{"value":"Hear What You See: Video-to-Audio Generation with Diffusion Transformer and Semantic-Temporal Alignment-Ranked Direct Preference Optimization"},"pdf":{"value":"https://openaccess.thecvf.com/content/CVPR2026/papers/Wang_Hear_What_You_See_Video-to-Audio_Generation_with_Diffusion_Transformer_and_CVPR_2026_paper.pdf"},"venueid":{"value":"thecvf.com/CVPR/2026/Conference"},"paperhash":{"value":"wang|hear_what_you_see_videotoaudio_generation_with_diffusion_transformer_and_semantictemporal_alignmentranked_direct_preference_optimization"},"authorids":{"value":["~Kai_Wang39","~Tao_Zhou17","~jiayi_lei1","~Jing_Wang80","~Jinman_Zhao2","~Weiguo_Pian1","~Yuan_Cheng1","~Yapeng_Tian1","~Peng_Gao3","~Bin_Fu1","~Yihao_Liu1","~Dimitrios_Hatzinakos1","~Yuewen_Cao1"]},"authors":{"value":["Kai Wang","Tao Zhou","Jiayi Lei","Jing Wang","Jinman Zhao","Weiguo Pian","Yuan Cheng","Yapeng Tian","Peng Gao","Bin Fu","Yihao Liu","Dimitrios Hatzinakos","Yuewen Cao"]}},"tmdate":1789656742483,"pdate":1789656439011,"tcdate":1765222739660,"writers":["thecvf.com/CVPR/2026/Conference","thecvf.com/CVPR/2026/Conference/Submission41517/Authors"],"signatures":["thecvf.com/CVPR/2026/Conference/Submission41517/Authors"],"forum":"TWJQCh0awj","license":"CC BY 4.0","number":41517,"cdate":1765222739660,"readers":["everyone"],"invitations":["thecvf.com/CVPR/2026/Conference/-/Submission","thecvf.com/CVPR/2026/Conference/Submission41517/-/Full_Submission","thecvf.com/CVPR/2026/Conference/-/Post_Submission","thecvf.com/CVPR/2026/Conference/Submission41517/-/Supplementary_Material","thecvf.com/CVPR/2026/Conference/-/Edit","thecvf.com/CVPR/2026/Conference/-/Compute_Flag"],"mdate":1789656742483,"odate":1789656439011,"domain":"thecvf.com/CVPR/2026/Conference","id":"TWJQCh0awj","version":2},{"content":{"venue":{"value":"IET Computer Vision"},"pdf":{"value":"https://onlinelibrary.wiley.com/doi/pdf/10.1049/cvi2.12100"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"yang|multistage_attention_network_for_videobased_person_reidentification"},"html":{"value":"https://doi.org/10.1049/cvi2.12100"},"abstract":{"value":"Video-based person re-identification (Re-ID) has received increasing attention in video surveillance analysis in recent years. To extract relevant information of the target, many existing methods utilise the attention mechanism in the residual block of the ResNet. However, these methods only focus on the residual block and ignore the output of the shortcut part, which also contains rich information about the person. To solve this problem, a different aspect of network design is investigated: the insert position of the attention module. To simultaneously explore the discriminative information in both the residual block and the shortcut, a novel multi-stage attention method is proposed by inserting the attention mechanism between stages of ResNet. Using this method can effectively extract the rich discriminative features of the target to better distinguish different pedestrians and improve the feature extraction capabilities of the model. Extensive experiments are conducted on four popular video-based person Re-ID datasets to demonstrate the effectiveness of the authors’ proposed method and display its superiority with the existing video-based person Re-ID methods."},"title":{"value":"Multi‐stage attention network for video‐based person re‐identification"},"authors":{"value":[{"fullname":"Fan Yang"},{"fullname":"Wei Li","username":"~Wei_Li29"},{"fullname":"Binbin Liang"},{"fullname":"Songchen Han"},{"fullname":"Xuan Zhu"}]}},"tmdate":1789092504884,"pdate":1659312000000,"externalIds":["doi:10.1049/cvi2.12100"],"tcdate":1762407212047,"writers":["~","OpenReview.net/Public_Article/ORCID.org","OpenReview.net/Support"],"signatures":["~Wei_Li29"],"forum":"DXcEenf55o","license":"CC BY-SA 4.0","number":7425,"cdate":1673837242003,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/ORCID.org/-/Record","OpenReview.net/Public_Article/-/Edit","OpenReview.net/-/Edit"],"mdate":1789092504884,"domain":"OpenReview.net/Public_Article","id":"DXcEenf55o","version":2},{"content":{"venueid":{"value":"OpenReview.net/Archive"},"paperhash":{"value":"yang|multistage_attention_network_for_videobased_person_reidentification"},"authorids":{"value":["~Fan_Yang47","~Wei_Li29","~Binbin_Liang1","~Songchen_Han1","~Xuan_Zhu5"]},"html":{"value":"https://ietresearch.onlinelibrary.wiley.com/doi/10.1049/cvi2.12100"},"abstract":{"value":"Video-based person re-identification (Re-ID) has received increasing attention in video surveillance analysis in recent years. To extract relevant information of the target, many existing methods utilise the attention mechanism in the residual block of the ResNet. However, these methods only focus on the residual block and ignore the output of the shortcut part, which also contains rich information about the person. To solve this problem, a different aspect of network design is investigated: the insert position of the attention module. To simultaneously explore the discriminative information in both the residual block and the shortcut, a novel multi-stage attention method is proposed by inserting the attention mechanism between stages of ResNet. Using this method can effectively extract the rich discriminative features of the target to better distinguish different pedestrians and improve the feature extraction capabilities of the model. Extensive experiments are conducted on four popular video-based person Re-ID datasets to demonstrate the effectiveness of the authors’ proposed method and display its superiority with the existing video-based person Re-ID methods."},"title":{"value":"Multi-stage attention network for video-based person re-identification"},"authors":{"value":["Fan Yang","Wei Li","Binbin Liang","Songchen Han","Xuan Zhu"]}},"tmdate":1741137204085,"pdate":1648569600000,"tcdate":1741137173169,"writers":["~Fan_Yang47","~Wei_Li29","~Binbin_Liang1","~Songchen_Han1","~Xuan_Zhu5"],"signatures":["~Fan_Yang47"],"forum":"3IkxHlrgbr","license":"CC BY-NC-ND 4.0","number":33713,"cdate":1741137173169,"readers":["everyone"],"invitations":["OpenReview.net/Archive/-/Direct_Upload"],"mdate":1741137204085,"domain":"OpenReview.net/Archive","id":"3IkxHlrgbr","version":2},{"content":{"venue":{"value":"CoRR 2025"},"pdf":{"value":"https://arxiv.org/pdf/2512.18888v1"},"venueid":{"value":"dblp.org/journals/CORR/2025"},"paperhash":{"value":"achara|localising_shortcut_learning_in_pixel_space_via_ordinal_scoring_correlations_for_attribution_representations_oscar"},"authorids":{"value":["~Akshit_Achara1","https://dblp.org/search/pid/api?q=author:Peter_Triantafillou:","https://dblp.org/search/pid/api?q=author:Esther_Puyol-Antón:","https://dblp.org/search/pid/api?q=author:Alexander_Hammers:","~Andrew_P._King1"]},"html":{"value":"https://doi.org/10.48550/arXiv.2512.18888"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2512-18888,\n  publtype={informal},\n  author={Akshit Achara and Peter Triantafillou and Esther Puyol-Antón and Alexander Hammers and Andrew P. King},\n  title={Localising Shortcut Learning in Pixel Space via Ordinal Scoring Correlations for Attribution Representations (OSCAR)},\n  year={2025},\n  month={December},\n  cdate={1764547200000},\n  journal={CoRR},\n  volume={abs/2512.18888},\n  url={https://doi.org/10.48550/arXiv.2512.18888}\n}\n"},"abstract":{"value":"Deep neural networks often exploit shortcuts. These are spurious cues which are associated with output labels in the training data but are unrelated to task semantics. When the shortcut features are associated with sensitive attributes, shortcut learning can lead to biased model performance. Existing methods for localising and understanding shortcut learning are mostly based upon qualitative, image-level inspection and assume cues are human-visible, limiting their use in domains such as medical imaging. We introduce OSCAR (Ordinal Scoring Correlations for Attribution Representations), a model-agnostic framework for quantifying shortcut learning and localising shortcut features. OSCAR converts image-level task attribution maps into dataset-level rank profiles of image regions and compares them across three models: a balanced baseline model (BA), a test model (TS), and a sensitive attribute predictor (SA). By computing pairwise, partial, and deviation-based correlations on these rank profiles, we produce a set of quantitative metrics that characterise the degree of shortcut reliance for TS, together with a ranking of image-level regions that contribute most to it. Experiments on CelebA, CheXpert, and ADNI show that our correlations are (i) stable across seeds and partitions, (ii) sensitive to the level of association between shortcut features and output labels in the training data, and (iii) able to distinguish localised from diffuse shortcut features. As an illustration of the utility of our method, we show how worst-group performance disparities can be reduced using a simple test-time attenuation approach based on the identified shortcut regions. OSCAR provides a lightweight, pixel-space audit that yields statistical decision rules and spatial maps, enabling users to test, localise, and mitigate shortcut reliance. The code is available at https://github.com/acharaakshit/oscar"},"title":{"value":"Localising Shortcut Learning in Pixel Space via Ordinal Scoring Correlations for Attribution Representations (OSCAR)"},"authors":{"value":["Akshit Achara","Peter Triantafillou","Esther Puyol-Antón","Alexander Hammers","Andrew P. King"]}},"tmdate":1772172818672,"pdate":1767139200000,"externalIds":["dblp:journals/corr/abs-2512-18888"],"tcdate":1769657262381,"writers":["~"],"signatures":["~Akshit_Achara1"],"forum":"4NBjIHrWvO","license":"CC BY-SA 4.0","number":814318,"cdate":1764547200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1772172818672,"domain":"DBLP.org","id":"4NBjIHrWvO","version":2},{"content":{"venue":{"value":"IEEE Trans. Circuits Syst. Video Technol. 2025"},"pdf":{"value":"https://ieeexplore.ieee.org/iel7/76/10949577/09984695.pdf"},"venueid":{"value":"dblp.org/journals/TCSV/2025"},"paperhash":{"value":"amirpour|deepstream_video_streaming_enhancements_using_compressed_deep_neural_networks"},"authorids":{"value":["","","~Christian_Timmerer1"]},"html":{"value":"https://doi.org/10.1109/TCSVT.2022.3229079"},"_bibtex":{"value":"@article{DBLP:journals/tcsv/AmirpourGT25,\n  author={Hadi Amirpour and M. Ghanbari and Christian Timmerer},\n  title={DeepStream: Video Streaming Enhancements Using Compressed Deep Neural Networks},\n  year={2025},\n  month={April},\n  cdate={1743465600000},\n  journal={IEEE Trans. Circuits Syst. Video Technol.},\n  volume={35},\n  number={4},\n  pages={3786-3797},\n  url={https://doi.org/10.1109/TCSVT.2022.3229079}\n}\n"},"abstract":{"value":"In HTTP Adaptive Streaming (HAS), each video is divided into smaller segments, and each segment is encoded at multiple pre-defined bitrates to construct a bitrate ladder. To optimize bitrate ladders, per-title encoding approaches encode each segment at various bitrates and resolutions to determine the convex hull. From the convex hull, an optimized bitrate ladder is constructed, resulting in an increased Quality of Experience (QoE) for end-users. With the ever-increasing efficiency of deep learning-based video enhancement approaches, they are more and more employed at the client-side to increase the QoE, specifically when GPU capabilities are available. Therefore, scalable approaches are needed to support end-user devices with both CPU and GPU capabilities (denoted as CPU-only and GPU-available end-users, respectively) as a new dimension of a bitrate ladder. To address this need, we propose DeepStream, a scalable content-aware per-title encoding approach to support both CPU-only and GPU-available end-users. (i) To support backward compatibility, DeepStream constructs a bitrate ladder based on any existing per-title encoding approach. Therefore, the video content will be provided for legacy end-user devices with CPU-only capabilities as a base layer (BL). (ii) For high-end end-user devices with GPU capabilities, an enhancement layer (EL) is added on top of the base layer comprising lightweight video super-resolution deep neural networks (DNNs) for each bitrate-resolution pair of the bitrate ladder. A content-aware video super-resolution approach leads to higher video quality, however, at the cost of bitrate overhead. To reduce the bitrate overhead for streaming content-aware video super-resolution DNNs, DeepCABAC, context-adaptive binary arithmetic coding for DNN compression, is used. Furthermore, the similarity among (i) segments within a scene and (ii) frames within a segment are used to reduce the training costs of DNNs. Experimental results show bitrate savings of 34% and 36% to maintain the same PSNR and VMAF, respectively, for GPU-available end-users, while the CPU-only users get the desired video content as usual."},"title":{"value":"DeepStream: Video Streaming Enhancements Using Compressed Deep Neural Networks"},"authors":{"value":["Hadi Amirpour","M. Ghanbari","Christian Timmerer"]}},"tmdate":1772203658399,"pdate":1767139200000,"externalIds":["dblp:journals/tcsv/AmirpourGT25"],"tcdate":1772203655213,"writers":["~"],"signatures":["~Christian_Timmerer1"],"forum":"ieJVWzh9bS","license":"CC BY-SA 4.0","number":833528,"cdate":1743465600000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1772203658399,"domain":"DBLP.org","id":"ieJVWzh9bS","version":2},{"content":{"summary":{"value":"This paper proposes NeuroClips, a framework that decodes high-fidelity and smooth video from fMRI. NeuroClips uses a semantics reconstructor for video keyframes to ensure semantic accuracy and consistency, and a perception reconstructor for capturing low-level perceptual details, ensuring video smoothness."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"- Why is there no model or loss function for movement(motion)?\n- Are you willing to make the code publicly available?"},"rating":{"value":7},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"NeuroClips is the framework to decouple high-level semantics and low-level perception flows for fMRI-to-video reconstruction, achieving high-fidelity and smooth video outputs.\nThe framework addresses the temporal resolution gap between fMRI and video data, ensuring smooth and consistent video outputs through innovative modules like Inception Extension and Temporal Upsampling."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The video captures the semantic meaning well but fails to accurately follow the ground truth movement."},"limitations":{"value":"As mentioned above, motion reconstruction has not been resolved."}},"nonreaders":[],"tmdate":1730878637882,"tcdate":1720801557861,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission351/Reviewer_85K8"],"signatures":["NeurIPS.cc/2024/Conference/Submission351/Reviewer_85K8"],"forum":"8qu52Fl1Dt","number":3,"license":"CC BY 4.0","cdate":1720801557861,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission351/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730878637882,"domain":"NeurIPS.cc/2024/Conference","replyto":"8qu52Fl1Dt","id":"pGUO6cv3yW","forumContent":{"venue":{"value":"NeurIPS 2024 oral"},"TLDR":{"value":"This paper proposed NeuroClips, a new state-of-the-art fMRI-to-video reconstruction framework, achieving smooth high-fidelity video reconstruction of up to 6s at 8FPS."},"keywords":{"value":["fMRI visual decoding; fMRI-to-video Reconstruction"]},"primary_area":{"value":"neuroscience_and_cognitive_science"},"flagged_for_ethics_review":{"value":true},"abstract":{"value":"Reconstruction of static visual stimuli from non-invasion brain activity fMRI achieves great success, owning to advanced deep learning models such as CLIP and Stable Diffusion. However, the research on fMRI-to-video reconstruction remains limited since decoding the spatiotemporal perception of continuous visual experiences is formidably challenging. We contend that the key to addressing these challenges lies in accurately decoding both high-level semantics and low-level perception flows, as perceived by the brain in response to video stimuli. To the end, we propose NeuroClips, an innovative framework to decode high-fidelity and smooth video from fMRI. NeuroClips utilizes a semantics reconstructor to reconstruct video keyframes, guiding semantic accuracy and consistency, and employs a perception reconstructor to capture low-level perceptual details, ensuring video smoothness. During inference, it adopts a pre-trained T2V diffusion model injected with both keyframes and low-level perception flows for video reconstruction. Evaluated on a publicly available fMRI-video dataset, NeuroClips achieves smooth high-fidelity video reconstruction of up to 6s at 8FPS, gaining significant improvements over state-of-the-art models in various metrics, e.g., a 128% improvement in SSIM and an 81% improvement in spatiotemporal metrics. Our project is available at https://github.com/gongzix/NeuroClips."},"_bibtex":{"value":"@inproceedings{\ngong2024neuroclips,\ntitle={NeuroClips: Towards High-fidelity and Smooth f{MRI}-to-Video Reconstruction},\nauthor={Zixuan Gong and Guangyin Bao and Qi Zhang and Zhongwei Wan and Duoqian Miao and Shoujin Wang and Lei Zhu and Changwei Wang and Rongtao Xu and Liang Hu and Ke Liu and Yu Zhang},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=8qu52Fl1Dt}\n}"},"title":{"value":"NeuroClips: Towards High-fidelity and Smooth fMRI-to-Video Reconstruction"},"pdf":{"value":"/pdf/258f5ea41fed74143053a220d1c9971bc970b99a.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"gong|neuroclips_towards_highfidelity_and_smooth_fmritovideo_reconstruction"},"authorids":{"value":["~Zixuan_Gong2","~Guangyin_Bao1","~Qi_Zhang25","~Zhongwei_Wan1","~Duoqian_Miao1","~Shoujin_Wang1","~Lei_Zhu8","~Changwei_Wang2","~Rongtao_Xu1","~Liang_Hu1","~Ke_Liu11","~Yu_Zhang60"]},"authors":{"value":["Zixuan Gong","Guangyin Bao","Qi Zhang","Zhongwei Wan","Duoqian Miao","Shoujin Wang","Lei Zhu","Changwei Wang","Rongtao Xu","Liang Hu","Ke Liu","Yu Zhang"]}},"version":2},{"content":{"summary":{"value":"The paper proposes Generative View Stitching (GVS), a training-free sampling method for camera-guided long-video generation that synthesizes all frames in parallel rather than autoregressively. GVS targets a key failure mode of autoregressive (AR) video diffusion—collisions and collapse when the predefined camera path would “walk through” previously generated content because the model cannot see the future. GVS divides the sequence into overlapping diffusion windows, jointly denoises target chunks with their past and future neighbors using any Diffusion Forcing (DF) video model. Two additions make this work for video: Omni Guidance (a guidance rule that strengthens conditioning on past & future within DF) and a cyclic conditioning mechanism for loop closure (long-range consistency)."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. How does GVS conceptually differ from prior camera-controllable or pose-guided video diffusion methods?\n\n2. Can you provide more details about the DF backbone (architecture, conditioning format, training data)?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. Address the future-conditioning problem in camera-guided video and adapts diffusion stitching to video without retraining, leveraging DF’s token-wise noise masking. \n\n2. The Omni Guidance formulation modifies the joint score toward a conditional score using DF’s null-conditioning, and cyclic conditioning explicitly addresses loop closure, both are new in this context.\n\n3. Enables stable, collision-free, long camera-guided rollouts, including looped or topologically tricky paths (e.g., “Impossible Staircase”), using off-the-shelf DF video models, which lowers the barrier to long-video generation without retraining."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Limited discussion of related work in camera control.\nWhile the paper introduces a strong motivation around camera-guided video generation, it provides little discussion of prior camera control or camera trajectory-conditioned video synthesis methods. Recent works such as VideoComposer [1], CameraCtrl [2] have explored related spatial or pose-guided control mechanisms.\n\n2. Lack of comparison with camera-control-based baselines. No quantitative or qualitative comparisons are made with existing camera-controllable video diffusion methods, making it hard to judge relative controllability and spatial consistency.\n\n3. Missing ablation study on Diffusion Forcing (DF) effectiveness.  An ablation contrasting DF-free and DF-enabled variants would clarify how much DF drives the gains.\n\n4. Insufficient model description and unclear generality across base models.\nThe paper under-specifies the DF backbone, leaving ambiguity about architecture and conditioning. It is also unclear whether results hold on stronger base models like Wan 2.1 [3].\n \n5. Restricted applicability to Diffusion Forcing models.\nThe proposed method is tightly coupled to the Diffusion Forcing architecture, and cannot be directly applied to other video diffusion architectures.\n\n[1] \"VideoComposer: Compositional Video Synthesis with Motion Controllability\". \n[2] \"CameraCtrl: Enabling Camera Control for Text-to-Video Generation\". \n[3] \"Wan: Open and Advanced Large-Scale Video Generative Models\"."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762921298866,"tcdate":1761905302070,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission9815/Reviewer_rey1"],"signatures":["ICLR.cc/2026/Conference/Submission9815/Reviewer_rey1"],"forum":"fpQpQbFPCU","number":2,"license":"CC BY 4.0","cdate":1761905302070,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission9815/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762921298866,"domain":"ICLR.cc/2026/Conference","replyto":"fpQpQbFPCU","id":"xU97mYlnLr","forumContent":{"TLDR":{"value":"A non-autoregressive alternative to length extrapolation of video diffusion models, which enables collision-free video generation for predefined camera trajectories."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Long Video Generation","Camera-guided Video Generation","Video Diffusion Models"]},"primary_area":{"value":"generative models"},"abstract":{"value":"Autoregressive video diffusion models are capable of long rollouts that are stable and consistent with history, but they are unable to guide the current generation with conditioning from the future. In camera-guided video generation with a predefined camera trajectory, this limitation leads to collisions with the generated scene, after which autoregression quickly collapses. To address this, we propose Generative View Stitching (GVS), which samples the entire sequence in parallel such that the generated scene is faithful to every part of the predefined camera trajectory. Our main contribution is a sampling algorithm that extends prior work on diffusion stitching for robot planning to video generation. While such stitching methods usually require a specially trained model, GVS is compatible with any off-the-shelf video model trained with Diffusion Forcing, a prevalent sequence diffusion framework that we show already provides the affordances necessary for stitching. We then introduce Omni Guidance, a technique that enhances the temporal consistency in stitching by conditioning on both the past and future, and that enables our proposed loop-closing mechanism for delivering long-range coherence. Overall, GVS achieves camera-guided video generation that is stable, collision-free, frame-to-frame consistent, and closes loops for a variety of predefined camera paths, including Oscar Reutersvärd’s Impossible Staircase."},"_bibtex":{"value":"@inproceedings{\nsong2026generative,\ntitle={Generative View Stitching},\nauthor={Chonghyuk Song and Michal Stary and Boyuan Chen and George Kopanas and Vincent Sitzmann},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=fpQpQbFPCU}\n}"},"title":{"value":"Generative View Stitching"},"pdf":{"value":"/pdf/af2e5f0d6f9fea66d6edeb4740472dd782567544.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"song|generative_view_stitching"},"authorids":{"value":["~Chonghyuk_Song1","~Michal_Stary1","~Boyuan_Chen2","~George_Kopanas1","~Vincent_Sitzmann1"]},"authors":{"value":["Chonghyuk Song","Michal Stary","Boyuan Chen","George Kopanas","Vincent Sitzmann"]}},"version":2},{"content":{"summary":{"value":"The paper presents StreamV2V, a diffusion model for real-time, streaming video-to-video (V2V) translation using user prompts. It achieves efficiency by maintaining a continually updated feature bank of past frames, which it references to enhance incoming frames without needing fine-tuning. StreamV2V seamlessly integrates with image diffusion models and operates at 20 FPS on an A100 GPU, making it significantly faster than competing models. User studies and quantitative metrics validate StreamV2V's strong temporal coherence and adaptability."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. how does the proposed framework still outperform other video-to-video translation techniques for variations of diffusion models? Do different models with the proposed framework still outperform other video-to-video translation techniques?\n\n2. what does the running-time comparison look like on CPU Inferences rather than GPU inferences?"},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. this paper addresses the challenge of real-time video-to-video translation via the streamV2V framework with backward-looking and training-free techniques. Since streamV2V is training-free and fine-tuning-free compared to other video-to-video techniques, the proposed approach is feasible to be deployed to practical applications.\n\n2. the streamV2V framework enhances coherence between adjacent translated video frames through the extended self-attention layer and a proposed feature fusion approach for intermediate features of the current and the past frames. Due to the extended self-attention layer and the feature fusion approach, streamV2V effectively reuses the intermediate features from the past video frames to generate coherent translated video frameworks between adjacent frames.\n\n3. the streamV2V stores projected keys, projected values, and outputs of intermediate features of the past video frames through a feature bank with a dynamic merging mechanism. With the dynamic feature bank, the streamV2V framework trivially reuses historical features to align the outputs of intermediate blocks of the current and the past frames."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. the proposed streamV2V framework fails to perform video-to-video translation for high-quality video generation in some scenarios while a prompt is provided.\n\n2. the proposed streamV2V framework fails to perform video-to-video translation when input video includes rapid motion movements.\n\n3. to test the generalization of the proposed framework, the manuscript misses experimentation comparisons for variations of difusion models."}},"nonreaders":[],"tmdate":1732618736737,"tcdate":1730609542613,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission6311/Reviewer_DwFq"],"signatures":["ICLR.cc/2025/Conference/Submission6311/Reviewer_DwFq"],"forum":"AMkf7h7HER","number":3,"license":"CC BY 4.0","cdate":1730609542613,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission6311/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732618736737,"domain":"ICLR.cc/2025/Conference","replyto":"AMkf7h7HER","id":"8usJwQMfIb","forumContent":{"TLDR":{"value":"We present StreamV2V to support real-time video-to-video translation for streaming input."},"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Streaming video translation","diffusion models","feature banks"]},"supplementary_material":{"value":"/attachment/6636cc26116c9f1106fdd3c3b05da0232d939b8c.zip"},"primary_area":{"value":"generative models"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"This paper introduces StreamV2V, a diffusion model that achieves real-time streaming video-to-video (V2V) translation with user prompts. \n  Unlike prior V2V methods using batches to process limited frames, we opt to process frames in a streaming fashion, to support unlimited frames.\n  At the heart of StreamV2V lies a backward-looking principle that relates the present to the past. \n  This is realized by maintaining a feature bank, which archives information from past frames.\n  For incoming frames, StreamV2V extends self-attention to include banked keys and values, and directly fuses similar past features into the output.\n  The feature bank is continually updated by merging stored and new features, making it compact yet informative.\n  StreamV2V stands out for its adaptability and efficiency, seamlessly integrating with image diffusion models without fine-tuning.\n  It can run 20 FPS on one A100 GPU, being 15$\\times$, 46$\\times$, 108$\\times$, and 158$\\times$ faster than FlowVid, CoDeF, Rerender, and TokenFlow, respectively. \n  Quantitative metrics and user studies confirm StreamV2V's exceptional ability to maintain temporal consistency."},"_bibtex":{"value":"@inproceedings{\nliang2025looking,\ntitle={Looking Backward: Streaming Video-to-Video Translation with Feature Banks},\nauthor={Feng Liang and Akio Kodaira and Chenfeng Xu and Masayoshi Tomizuka and Kurt Keutzer and Diana Marculescu},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=AMkf7h7HER}\n}"},"title":{"value":"Looking Backward: Streaming Video-to-Video Translation with Feature Banks"},"pdf":{"value":"/pdf/65906303ad329767664a6e8fca6ed65a144ae4eb.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"liang|looking_backward_streaming_videotovideo_translation_with_feature_banks"},"authorids":{"value":["~Feng_Liang3","~Akio_Kodaira1","~Chenfeng_Xu1","~Masayoshi_Tomizuka2","~Kurt_Keutzer1","~Diana_Marculescu4"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Feng Liang","Akio Kodaira","Chenfeng Xu","Masayoshi Tomizuka","Kurt Keutzer","Diana Marculescu"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a method to enhance the multi-view and temporal consistency in the video generation.  To this end,  the method design a reconstruct-then-synthesis approach: given an generated video conditioned on an image, it first refines its depth estimation to obtain a 3D Gaussian representation of the background scene and objects.  Then, the 3D representations can be rendered in different viewpoints, and then the rendering results are fed into the video generation model for inpainting."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"How is the quality of generated videos related to the video depth? Besides, please also report the timing statistics of the proposed method."},"rating":{"value":4},"details_of_ethics_concerns":{"value":"NA"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1.  This paper tackles a challenging problem: removing the structure flickering in the generated video under large viewpoint change. \n2.  The proposed method combines the video depth estimation,  3D Gaussian representation to reconstruct 3D structure from the initial video after removing the noise and outliers. This reconstructed 3D structure can help to increase the multi-view and temporal consistency  of generated videos."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1.  While the experimental results are impressive,  the advantage of this method over STOA camera trajectory control method for video generation are not comprehensive.  I would like to see more comparisons with video generation methods,  for instance,  Animateanything: Consistent and Controllable Animation for Video Generation CVPR 2025.  \n\n2. The proposed method is a post-processing method for video generation that heavily relies on 3D reconstruction. Therefore, the generalizability of this method relies on the generalizability of video depth estimation and so on.   More tests on various kind of scenes  are necessary to verify the claimed contribution in this introduction, since it is difficult to obtain accurate and consistent depth for dynamic objects with some materials, such as glass and highly  specular materials."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762918403621,"tcdate":1761984609238,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission5999/Reviewer_mZU9"],"signatures":["ICLR.cc/2026/Conference/Submission5999/Reviewer_mZU9"],"forum":"guUaZN0kyC","number":3,"license":"CC BY 4.0","cdate":1761984609238,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission5999/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762918403621,"domain":"ICLR.cc/2026/Conference","replyto":"guUaZN0kyC","id":"udJTHdDLpE","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Scene Generation","Video Generation"]},"supplementary_material":{"value":"/attachment/b14a16ad6bae67e7e5b500c0d788de074931df51.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"We present WorldCrafter, a novel framework that enables interactive dynamic scene generation from a single image by leveraging geometry-aware and temporal modeling. Existing methods often suffer from texture distortion, structural inaccuracies, and temporal flickering under large viewpoint changes. These issues mainly caused by explicit pixel-wise reprojection strategies. To address these challenges, WorldCrafter introduces two complementary modules: 1) Geometry-aware Video Depth Refinement, which enhances structural fidelity by refining depth with multi-frame geometric priors and semantic cues; and  2) Object-consistent Temporal Modeling, which disentangles video frames into object-level layers to improve coherence between static backgrounds and dynamic foregrounds. These components form a unified rendering-inpainting framework for photorealistic and camera-controllable dynamic scene generation. Experiments demonstrate that WorldCrafter produces geometrically accurate and temporally coherent results across diverse scenes and camera trajectories."},"_bibtex":{"value":"@misc{\ndong2025worldcrafter,\ntitle={WorldCrafter: Dynamic Scene Generation from a Single Image with Geometric and Temporal Consistency},\nauthor={Haoye Dong and Gim Hee Lee},\nyear={2025},\nurl={https://openreview.net/forum?id=guUaZN0kyC}\n}"},"title":{"value":"WorldCrafter: Dynamic Scene Generation from a Single Image with Geometric and Temporal Consistency"},"pdf":{"value":"/pdf/bda6aad4d682f68ba69bd1ef445627742737bfa4.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"dong|worldcrafter_dynamic_scene_generation_from_a_single_image_with_geometric_and_temporal_consistency"},"authorids":{"value":["~Haoye_Dong1","~Gim_Hee_Lee1"]},"authors":{"value":["Haoye Dong","Gim Hee Lee"]}},"version":2},{"content":{"venue":{"value":"ECCV (74) 2026"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-032-37029-7_12.pdf"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"li|cleartextvideo_a_largescale_textcentric_video_dataset_bridging_video_restoration_and_scenetext_enhancement"},"html":{"value":"https://doi.org/10.1007/978-3-032-37029-7_12"},"_bibtex":{"value":"@inproceedings{DBLP:conf/eccv/LiDLHKYGFCSM26,\n  author={Jinlong Li and Jiaming Ding and Dingfu Lu and Malcolm Hsiu and Chuang Ke and Kangning Yang and Bochen Guan and Lan Fu and Jie Cai and Huiming Sun and Zibo Meng},\n  title={ClearText-Video: A Large-Scale Text-Centric Video Dataset Bridging Video Restoration and Scene-Text Enhancement},\n  year={2026},\n  cdate={1767225600000},\n  pages={188-207},\n  url={https://doi.org/10.1007/978-3-032-37029-7_12},\n  booktitle={ECCV (74)},\n  crossref={conf/eccv/2026-74}\n}\n"},"abstract":{"value":"Multimodal Large Language Models (MLLMs) have recently demonstrated strong progress in visual–linguistic understanding, yet their performance on text-centric video reasoning remains highly sensitive to input quality. Real-world user-provided videos frequently contain motion blur, compression artifacts, noise, and low-resolution text, substantially impairing reliable text reading and downstream reasoning. Whether MLLMs can robustly read and reason over in-the-wild text under diverse quality conditions remains an unanswered fundamental question. We introduce ClearText-Video (CTVid), a large-scale, scene-text-aware benchmark for studying text-centric video understanding under controlled quality variation. CTVid contains 4,639 real-world text-rich egocentric videos, 550K+ frames, 1.6M human-verified scene-text annotations, and 220K+ spatial/temporal question–answer pairs in Chinese and English. For each high-quality video, CTVid provides content-matched Degraded-Quality and Restored-Quality variants, supporting two task families: Text-Centric Video Restoration and Multi-Video Quality VideoQA. We evaluate 18 representative restoration methods and 16 state-of-the-art MLLMs on CTVid. The results show that visual enhancement does not guarantee textual fidelity or downstream reasoning gains: blur is more damaging than low resolution, restored videos can alter the textual evidence used by MLLMs, and OCR-only pipelines remain far below direct multimodal reasoning. CTVid exposes the gap between video restoration and text-grounded understanding, providing a rigorous foundation for restoration-aware, quality-robust text-centric video systems. Benchmark: https://github.com/jinlong17/CTVid-Bench."},"title":{"value":"ClearText-Video: A Large-Scale Text-Centric Video Dataset Bridging Video Restoration and Scene-Text Enhancement"},"authors":{"value":[{"fullname":"Jinlong Li","username":""},{"fullname":"Jiaming Ding","username":""},{"fullname":"Dingfu Lu","username":""},{"fullname":"Malcolm Hsiu","username":""},{"fullname":"Chuang Ke","username":"~Chuang_Ke1"},{"fullname":"Kangning Yang","username":""},{"fullname":"Bochen Guan","username":""},{"fullname":"Lan Fu","username":""},{"fullname":"Jie Cai","username":""},{"fullname":"Huiming Sun","username":""},{"fullname":"Zibo Meng","username":""}]}},"tmdate":1791309430032,"pdate":1798675200000,"externalIds":["dblp:conf/eccv/LiDLHKYGFCSM26"],"tcdate":1791309396224,"writers":["~","OpenReview.net/Public_Article/DBLP.org","OpenReview.net/Support"],"signatures":["~Chuang_Ke1"],"forum":"oaIegpTq0I","license":"CC BY-SA 4.0","number":164449,"cdate":1767225600000,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/DBLP.org/-/Record","OpenReview.net/Public_Article/-/Edit"],"mdate":1791309430032,"domain":"OpenReview.net/Public_Article","id":"oaIegpTq0I","version":2},{"content":{"correctness":{"value":"Yes"},"summary_and_contributions":{"value":"The paper introduces a novel benchmark for Video Corpus Moment Retrieval (VCMR) that focuses on fine-grained video understanding. The authors argue that existing VCMR benchmarks are limited to coarse-grained understanding, which hampers the precise localization of video moments when given fine-grained queries. To address this limitation, the paper proposes a more challenging VCMR scenario that requires models to retrieve the best-matched video moment from a corpus with other partially matched candidates.\nThe paper's main contributions are as follows:\n1. It proposes a challenging scenario requiring precise video moment localization in response to fine-grained queries amidst similar candidates.\n2. It introduces an automated pipeline using LLMs and LMMs to create detailed video annotations with a noise evaluator to enhance accuracy. And it develops high-quality datasets (Charades-FIG, DiDeMo-FIG, ActivityNet-FIG) to facilitate fine-grained VCMR.\n4. It demonstrates the inadequacy of existing models when evaluated on the new datasets, underlining the importance of fine-grained understanding in VCMR."},"confidence":{"value":4},"documentation":{"value":"Yes."},"rating":{"value":4},"title":{"value":"The paper introduces a novel benchmark for Video Corpus Moment Retrieval (VCMR) that focuses on fine-grained video understanding."},"ethics":{"value":"No."},"clarity":{"value":"No."},"review":{"value":"Pros\n1. The paper introduces an automated pipeline in detail, which uses LLMs and LMMs to create detailed video annotations. In addition, to alleviate the hallucination of LLMs, it proposes a noise evaluator to enhance accuracy.\n2. Based on the above pipeline, the paper introduces a novel and more challenging VCMR benchmark concentrate more on fine-grained video features, which is expected to be valuable to the community.\n\nCons\n1. The paper does not provide evaluation results by enough ablation studies. The author only give the final result and ignore to reveal how much role each component in the pipeline plays in the final evaluation result. e.g., Statics Enhanced Captioning, Dynamics Enhanced Captioning, Fine-Granularity Aware Noise Evaluator……\n2. The article lacks some qualitative analysis or visualization of the test samples, which allows readers to more intuitively feel how this solution focuses on the fine-grained features of the video.\n3. There are some problems with formatting. Tables 3 and 4 on page 9 should follow Table 2 instead of being inserted in the middle of Chapter 5: conclusion."},"strengths":{"value":"1. The paper introduces an automated pipeline in detail, which uses LLMs and LMMs to create detailed video annotations. In addition, to alleviate the hallucination of LLMs, it proposes a noise evaluator to enhance accuracy.\n2. Based on the above pipeline, the paper introduces a novel and more challenging VCMR benchmark concentrate more on fine-grained video features, which is expected to be valuable to the community."},"flag_for_ethics_review":{"value":"2: No, there are no or only very minor ethics concerns"},"relation_to_prior_work":{"value":"Yes."},"opportunities_for_improvement":{"value":"1. Can you supply more experiment results to support the validity of your paper？Especially the ablation studies of each component.\n2. It would be beneficial to generate visualization for the test samples to understand the fine-grained capability of the benchmark. This information may be included in the Appendix if necessary."},"additional_feedback":{"value":"Please refer to the weaknesses and questions part."},"limitations":{"value":"Same as the weakness."}},"nonreaders":[],"tmdate":1731500648196,"tcdate":1721833531504,"writers":["NeurIPS.cc/2024/Datasets_and_Benchmarks_Track","NeurIPS.cc/2024/Datasets_and_Benchmarks_Track/Submission246/Reviewer_G9CT"],"signatures":["NeurIPS.cc/2024/Datasets_and_Benchmarks_Track/Submission246/Reviewer_G9CT"],"forum":"ZMn2SPUgkU","number":2,"license":"CC BY 4.0","cdate":1721833531504,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Datasets_and_Benchmarks_Track/Submission246/-/Official_Review","NeurIPS.cc/2024/Datasets_and_Benchmarks_Track/-/Edit"],"mdate":1731500648196,"domain":"NeurIPS.cc/2024/Datasets_and_Benchmarks_Track","replyto":"ZMn2SPUgkU","id":"zkoVtgwzlZ","forumContent":{"TLDR":{"value":"We propose VERIFIED, an automatic video-text annotation pipeline to generate captions with reliable fine-grained statics and dynamics, and create a fine-grained VCMR benchmark."},"venue":{"value":"NeurIPS 2024 Track Datasets and Benchmarks Poster"},"keywords":{"value":["Fine-Grained Video Understanding","Video Corpus Moment Retrieval","Dynamics"]},"supplementary_material":{"value":"/attachment/d010312e58e11dd55f60515e860632a76a93a8df.pdf"},"abstract":{"value":"Existing Video Corpus Moment Retrieval (VCMR) is limited to coarse-grained understanding that hinders precise video moment localization when given fine-grained queries. In this paper, we propose a more challenging fine-grained VCMR benchmark requiring methods to localize the best-matched moment from the corpus with other partially matched candidates. To improve the dataset construction efficiency and guarantee high-quality data annotations, we propose VERIFIED, an automatic \\underline{V}id\\underline{E}o-text annotation pipeline to generate captions with \\underline{R}el\\underline{I}able \\underline{FI}n\\underline{E}-grained statics and \\underline{D}ynamics. Specifically, we resort to large language models (LLM) and large multimodal models (LMM) with our proposed Statics and Dynamics Enhanced Captioning modules to generate diverse fine-grained captions for each video. To filter out the inaccurate annotations caused by the LLM hallucination, we propose a Fine-Granularity Aware Noise Evaluator where we fine-tune a video foundation model with disturbed hard-negatives augmented contrastive and matching losses. With VERIFIED, we construct a more challenging fine-grained VCMR benchmark containing Charades-FIG, DiDeMo-FIG, and ActivityNet-FIG which demonstrate a high level of annotation quality. We evaluate several state-of-the-art VCMR models on the proposed dataset, revealing that there is still significant scope for fine-grained video understanding in VCMR."},"_bibtex":{"value":"@inproceedings{\nchen2024verified,\ntitle={{VERIFIED}: A Video Corpus Moment Retrieval Benchmark for Fine-Grained Video Understanding},\nauthor={Houlun Chen and Xin Wang and Hong Chen and Zeyang Zhang and Wei Feng and Bin Huang and Jia Jia and Wenwu Zhu},\nbooktitle={The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track},\nyear={2024},\nurl={https://openreview.net/forum?id=ZMn2SPUgkU}\n}"},"title":{"value":"VERIFIED: A Video Corpus Moment Retrieval Benchmark for Fine-Grained Video Understanding"},"pdf":{"value":"/pdf/93ca5dd2e78eb5960639e34d307fa0133e3bbd42.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Datasets_and_Benchmarks_Track"},"paperhash":{"value":"chen|verified_a_video_corpus_moment_retrieval_benchmark_for_finegrained_video_understanding"},"authorids":{"value":["~Houlun_Chen1","~Xin_Wang17","~Hong_Chen9","~Zeyang_Zhang1","~Wei_Feng11","~Bin_Huang4","~Jia_Jia1","~Wenwu_Zhu1"]},"authors":{"value":["Houlun Chen","Xin Wang","Hong Chen","Zeyang Zhang","Wei Feng","Bin Huang","Jia Jia","Wenwu Zhu"]}},"version":2},{"content":{"venue":{"value":"ICML 2026 regular"},"keywords":{"value":["Video Generation Model","3D Reconstruction","Reinforcement Learning","Self-Supervised Learning"]},"supplementary_material":{"value":"/attachment/bd16c70b69b06456c840644d6f2ad0e7f0cf5188.zip"},"_bibtex":{"value":"@inproceedings{\ndu2026videogpa,\ntitle={Video{GPA}: Distilling Geometry Priors for 3D-Consistent Video Generation},\nauthor={Hongyang Du and Junjie Ye and Xiaoyan Cong and Runhao Li and Jingcheng Ni and Aman Agarwal and Zeqi Zhou and Zekun Li and Randall Balestriero and Yue Wang},\nbooktitle={Forty-third International Conference on Machine Learning},\nyear={2026},\nurl={https://openreview.net/forum?id=neygndmdoS}\n}"},"title":{"value":"VideoGPA: Distilling Geometry Priors for 3D-Consistent Video Generation"},"paperhash":{"value":"du|videogpa_distilling_geometry_priors_for_3dconsistent_video_generation"},"originally_submitted_PDF":{"value":"/pdf/17c3b325771ba32f283ab17ff20ce349b884829c.pdf"},"TLDR":{"value":"VideoGPA leverages geometric priors to align video diffusion models, significantly improving 3D consistency."},"primary_area":{"value":"deep_learning->generative_models_and_autoencoders"},"abstract":{"value":"While recent video diffusion models (VDMs) produce visually impressive results, they fundamentally struggle to maintain 3D structural consistency, often resulting in object deformation or spatial drift. We hypothesize that these failures arise because standard denoising objectives lack explicit incentives for geometric coherence. To address this, we introduce VideoGPA (Video Geometric Preference Alignment), a data-efficient self-supervised framework that leverages a geometry foundation model to automatically derive dense preference signals that guide VDMs via Direct Preference Optimization (DPO). This approach effectively steers the generative distribution toward inherent 3D consistency without requiring human annotations. VideoGPA significantly enhances temporal stability, geometric plausibility, and motion coherence using minimal preference pairs, consistently outperforming state-of-the-art baselines in extensive experiments."},"link_to_code":{"value":"https://hongyang-du.github.io/VideoGPA-Website/"},"pdf":{"value":"/pdf/6959df1a6955fcefb233d586ca70a1513418d6b1.pdf"},"lay_summary":{"value":"AI video generators can create strikingly realistic footage, but they often fail in subtle ways: objects warp as the camera moves, backgrounds drift, and scenes lose their physical coherence over time. These failures happen not because the models are too small or undertrained, but because they are never explicitly taught to respect the laws of 3D space — they learn to make videos look plausible frame by frame, without any guarantee that the underlying geometry is consistent.\n\nWe introduce VideoGPA, a lightweight method that teaches existing video generators to \"think in 3D.\" The key idea is to use a separate geometry model — one that can reconstruct a 3D scene from a video — as an automatic quality inspector. If a generated video is geometrically correct, the inspector should be able to reconstruct it accurately; if not, the reconstruction error is high. We use this error as a free signal to rank generated videos from best to worst, and then train the video model to prefer the geometrically better ones. No human labeling is required.\n\nThe result is a video generator that produces more stable, physically plausible videos — with fewer floating objects, fewer warping backgrounds, and better motion coherence — using only a small amount of additional training. Improvements also transfer to dynamic scenes the model was never explicitly trained on, suggesting that learning 3D consistency teaches a general physical intuition rather than a narrow trick."},"venueid":{"value":"ICML.cc/2026/Conference"},"authorids":{"value":["~Hongyang_Du2","~Junjie_Ye3","~Xiaoyan_Cong1","~Runhao_Li3","~Jingcheng_Ni2","~Aman_Agarwal4","~Zeqi_Zhou1","~Zekun_Li3","~Randall_Balestriero1","~Yue_Wang2"]},"authors":{"value":["Hongyang Du","Junjie Ye","Xiaoyan Cong","Runhao Li","Jingcheng Ni","Aman Agarwal","Zeqi Zhou","Zekun Li","Randall Balestriero","Yue Wang"]}},"tmdate":1790067310671,"pdate":1777576264593,"tcdate":1768680211436,"writers":["ICML.cc/2026/Conference","ICML.cc/2026/Conference/Submission6720/Authors"],"signatures":["ICML.cc/2026/Conference/Submission6720/Authors"],"forum":"neygndmdoS","license":"CC BY 4.0","number":6720,"cdate":1768680211436,"readers":["everyone"],"invitations":["ICML.cc/2026/Conference/-/Submission","ICML.cc/2026/Conference/-/Edit","ICML.cc/2026/Conference/-/Post_Submission","ICML.cc/2026/Conference/Submission6720/-/Reciprocal_Reviewing_Correction","ICML.cc/2026/Conference/Submission6720/-/Full_Submission","ICML.cc/2026/Conference/Submission6720/-/Camera_Ready_Revision"],"mdate":1790067310671,"odate":1782341922926,"domain":"ICML.cc/2026/Conference","id":"neygndmdoS","version":2},{"content":{"venue":{"value":"CoRR 2026"},"pdf":{"value":"https://arxiv.org/pdf/2601.23286v2"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"du|videogpa_distilling_geometry_priors_for_3dconsistent_video_generation"},"html":{"value":"https://doi.org/10.48550/arXiv.2601.23286"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2601-23286,\n  publtype={informal},\n  author={Hongyang Du and Junjie Ye and Xiaoyan Cong and Runhao Li and Jingcheng Ni and Aman Agarwal and Zeqi Zhou and Zekun Li and Randall Balestriero and Yue Wang},\n  title={VideoGPA: Distilling Geometry Priors for 3D-Consistent Video Generation},\n  year={2026},\n  month={January},\n  cdate={1767225600000},\n  journal={CoRR},\n  volume={abs/2601.23286},\n  url={https://doi.org/10.48550/arXiv.2601.23286}\n}\n"},"abstract":{"value":"While recent video diffusion models (VDMs) produce visually impressive results, they fundamentally struggle to maintain 3D structural consistency, often resulting in object deformation or spatial drift. We hypothesize that these failures arise because standard denoising objectives lack explicit incentives for geometric coherence. To address this, we introduce VideoGPA (Video Geometric Preference Alignment), a data-efficient self-supervised framework that leverages a geometry foundation model to automatically derive dense preference signals that guide VDMs via Direct Preference Optimization (DPO). This approach effectively steers the generative distribution toward inherent 3D consistency without requiring human annotations. VideoGPA significantly enhances temporal stability, physical plausibility, and motion coherence using minimal preference pairs, consistently outperforming state-of-the-art baselines in extensive experiments."},"title":{"value":"VideoGPA: Distilling Geometry Priors for 3D-Consistent Video Generation"},"authors":{"value":[{"fullname":"Hongyang Du","username":""},{"fullname":"Junjie Ye","username":""},{"fullname":"Xiaoyan Cong","username":"~Xiaoyan_Cong1"},{"fullname":"Runhao Li","username":""},{"fullname":"Jingcheng Ni","username":""},{"fullname":"Aman Agarwal","username":""},{"fullname":"Zeqi Zhou","username":""},{"fullname":"Zekun Li","username":""},{"fullname":"Randall Balestriero","username":""},{"fullname":"Yue Wang","username":""}]}},"tmdate":1778348320927,"pdate":1798675200000,"externalIds":["dblp:journals/corr/abs-2601-23286"],"tcdate":1778348303825,"writers":["~","OpenReview.net/Public_Article/DBLP.org","OpenReview.net/Support"],"signatures":["~Xiaoyan_Cong1"],"forum":"duJE8qE4uA","license":"CC BY-SA 4.0","number":16811,"cdate":1767225600000,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/DBLP.org/-/Record","OpenReview.net/Public_Article/-/Edit"],"mdate":1778348320927,"domain":"OpenReview.net/Public_Article","id":"duJE8qE4uA","version":2},{"content":{"summary":{"value":"This paper proposes \"Video Policy,\" a framework for learning robot policies by leveraging pre-trained video generation models. The method uses a dual U-Net architecture: a fine-tuned Stable Video Diffusion (SVD) model generates future video frames, and a smaller action decoder, conditioned on the video model's internal features, generates corresponding actions. The central claim is that a two-stage training process—where the video model is frozen before training the action decoder—is superior to end-to-end joint training. The authors present results on the RoboCasa and Libero10 benchmarks, claiming state-of-the-art performance and high sample efficiency with as few as 50 demonstrations. \n\nHowever, I think the overall novelty of the paper is largely diluted by many related works that guide robot policy via video generation, and the main finding is not original as well."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"**On Architectural Choices:** Why is a modular, causal `Video -> Action` framework fundamentally better than more integrated approaches like UVA's joint latent space or UWM's unified transformer? \n\n**On System Reliability:** Given that policy failures are directly caused by failures in video generation, the current system appears fundamentally brittle. What concrete mechanisms could be implemented to assess the physical plausibility of generated videos and enable the policy to recover from generative failures?\n\n**On Practical Deployment:** An inference time of 9 seconds is prohibitive for real-world robotics. Beyond general optimism about future diffusion model acceleration, what specific architectural or algorithmic changes to *your proposed method* do you see as the most viable path to achieving real-time control?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"**Strong reported performance on benchmarks**: The paper reports high success rates on the RoboCasa (63% average) and Libero10 (94% average) benchmarks, outperforming several baselines, including some large-scale Vision-Language-Action (VLA) models, while using significantly less demonstration data. These results provide initial evidence for the potential of leveraging powerful video priors.  \n\n**Good video generation quality**: The authors give a good demonstration of the capability of the video generation model, which works well in most cases.\n\n**Good representation**: The paper delivery is clear, and the details are available for understanding the proposed method."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"**Overstated Novelty and Lack of Meaningful Comparison to Concurrent Work:** The paper frames the use of video models for policy learning as a novel contribution. This is a significant overstatement. The 2024-2025 period has seen a proliferation of work on this exact topic, including but not limited to concurrent models like **UVA**[1] , **UWM**[2] , and **VPP**[3] , which explore unified video-action architectures. The paper fails to properly situate itself within this crowded landscape, and its core novelty is limited to the specific finding about two-stage training, which is not so surprising, as action finetuning from a  pretrained video backbone seems a natural idea.\n\n**Prohibitive Computational Cost:** The reported inference time of **9 seconds for a single action sequence on an A100 GPU** is a severe limitation that is understated by the authors. This latency makes the system completely unsuitable for real-time, closed-loop control, which is a prerequisite for robotics in any dynamic environment. Dismissing this as a general problem that future hardware will solve is insufficient; it is a fundamental constraint on the proposed method's viability.  \n\n**Over dependent on the Video Generation Results** The policy's complete dependence on the video generator is a critical design flaw. The authors admit that real-world failures are caused by \"unrealistic video predictions\"—where the model imagines a physically impossible future. This creates a brittle, open-loop system. The paper proposes no mechanism to detect or recover from these generative failures, making the approach unreliable for practical applications.  \n\n**Reference**:\n\n[1] Li, S., Gao, Y., Sadigh, D., & Song, S. (2025). Unified Video Action Model. *ArXiv, abs/2503.00200*.\n\n[2] Zhu, C., Yu, R., Feng, S., Burchfiel, B., Shah, P., & Gupta, A. (2025). Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets. *ArXiv, abs/2504.02792*.\n\n[3] Hu, Y., Guo, Y., Wang, P., Chen, X., Wang, Y., Zhang, J., Sreenath, K., Lu, C., & Chen, J. (2024). Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations. *ArXiv, abs/2412.14803*."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762920944355,"tcdate":1761999785157,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission9310/Reviewer_rAra"],"signatures":["ICLR.cc/2026/Conference/Submission9310/Reviewer_rAra"],"forum":"cWczH8ontO","number":3,"license":"CC BY 4.0","cdate":1761999785157,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission9310/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762920944355,"domain":"ICLR.cc/2026/Conference","replyto":"cWczH8ontO","id":"kS39pdLgGa","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Behavior Cloning","Video Generation","Representation Learning"]},"supplementary_material":{"value":"/attachment/524fa64b2ccd106dcea68fc0fcfa7b106bff0d0c.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Despite tremendous progress in dexterous manipulation, current visuomotor policies remain fundamentally limited by two challenges: they struggle to generalize under perceptual or behavioral distribution shifts, and their performance is constrained by the size of human demonstration data. In this paper, we use video generation as a proxy for robot policy learning to address both limitations simultaneously.\nWe propose Video Policy, a modular framework that combines video and action generation that can be trained end-to-end. Our results demonstrate that learning to generate videos of robot behavior allows for the extraction of policies with minimal demonstration data, significantly improving robustness and sample efficiency. Our method shows strong generalization to unseen objects, backgrounds, and tasks, both in simulation and the real world. We further highlight that task success is closely tied to the generated video, with action-free video data providing critical benefits for generalizing to novel tasks. By leveraging large-scale video generative models, we achieve superior performance compared to recent VLAs and video-action models, paving the way for more scalable and data-efficient robot policy learning."},"_bibtex":{"value":"@misc{\nliang2025video,\ntitle={Video Generators are Robot Policies},\nauthor={Junbang Liang and Pavel Tokmakov and Ruoshi Liu and Sruthi Sudhakar and Paarth Shah and Rares Andrei Ambrus and Carl Vondrick},\nyear={2025},\nurl={https://openreview.net/forum?id=cWczH8ontO}\n}"},"title":{"value":"Video Generators are Robot Policies"},"pdf":{"value":"/pdf/cdf5affcaceddafbd0c7100e9ca873e764edfd23.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"liang|video_generators_are_robot_policies"},"authorids":{"value":["~Junbang_Liang2","~Pavel_Tokmakov2","~Ruoshi_Liu2","~Sruthi_Sudhakar1","~Paarth_Shah2","~Rares_Andrei_Ambrus1","~Carl_Vondrick2"]},"authors":{"value":["Junbang Liang","Pavel Tokmakov","Ruoshi Liu","Sruthi Sudhakar","Paarth Shah","Rares Andrei Ambrus","Carl Vondrick"]}},"version":2},{"content":{"summary":{"value":"In this paper, the authors propose the first trajectory-aware VLM, TraceVLM, to handle the correspondence between the human attention trajectories and the linguistic representation of a video. To finetune the QwenVL-2.5, a dataset construction pipeline is created to generate high-quality reasoning-enabled training data. The experiments demonstrate the effectiveness of the model on both trajectory-aware tasks and related benchmarks."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Please refer to the weakness."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The authors propose TraceVLM, the first VLM to achieve trajectory-aware video understanding/reasoning.\n2. To facilitate the model training, the authors construct a data annotation pipeline to obtain the training data based on existing trajectory video data. And the experiments demonstrate the effectiveness of the dataset and training strategy.\n3. The Geometric Simplification is efficient in reducing redundancy and noise.\n4. Not only in trajectory-aware tasks, but also in other visual benchmarks, TraceVLM has comparable or superior performance compared with other VLMs."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Major:\n1. In the introduction, line 42, the author claimed, “they often focus their attention on the primary regions of an image while neglecting surrounding contextual information, and may even be distracted by irrelevant areas”, and in Figure 1, the example depicts a man in a car holding a phone. The question is: (a) how to differentiate important context information and unimportant context information. Actually, in Figure 1, even humans will ignore the background because of the higher exposure and irrelevance of the main part of the image, and this is an exact example of NOT being distracted by irrelevant areas. (b) Is there any example showing “distracted by irrelevant areas”?\n2. In terms of the motivation, can the authors add some discussion of the real-world application of such a technique? What is the point of developing trajectory-driven localized understanding? \n3. In Section 3.2.3, line 223, what is ϵ_base? And how is w_i determined for each segment?\n4. In Section 3.2.5, what is the size of the segmentation codebook?\n5. In the data generation pipeline, how to ensure the correctness of the reasoning generation? Does the predefined global-->object-->paragraph --> local reasoning chain (Figure 4) work for every case? Maybe when the question changes, the reasoning process will vary a lot, since the focus and the scene structure or the intrinsic logic will be totally different.\n6. In the visualization of 4.5, it would be better if these cases also include the original Qwen2.5VL-7B, so that the attention map difference can be observed.\n\nMinor:\n1. Line 45, add a reference for this: “Research on human visual attention trajectories also plays an important role in domains such as virtual reality and autonomous driving.”\n2. In Figure 4, Step 1, the text is unclear due to the color (light yellow/green part)\n3. In Figure 4, Step 2-1, the Q1 and Q2 look incomplete."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762915836920,"tcdate":1762247093046,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission1632/Reviewer_Q5oH"],"signatures":["ICLR.cc/2026/Conference/Submission1632/Reviewer_Q5oH"],"forum":"8tZEohkGpB","number":4,"license":"CC BY 4.0","cdate":1762247093046,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission1632/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762915836920,"domain":"ICLR.cc/2026/Conference","replyto":"8tZEohkGpB","id":"bMOK6RlWf3","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Trajectory; Interactivate; Large Vision Language Model"]},"supplementary_material":{"value":"/attachment/6d1a4a3f3d197037b3fddfd68af9933a64f376f8.zip"},"primary_area":{"value":"foundation or frontier models, including LLMs"},"abstract":{"value":"Recent Large Vision Language Models (LVLMs) demonstrate remarkable capabilities in image understanding and natural language generation. However, current approaches focus predominantly on global image understanding, struggling to simulate human visual attention trajectories and explain associations between generated descriptions and specific image regions. To address this challenge, we propose TraceVLM, a unified vision-language model that integrates trajectory-aware spatial understanding within an end-to-end framework. TraceVLM employs a Trajectory-aware Visual Perception (TVP) module for deep bidirectional fusion of visual features and trajectory information. We utilize a geometric simplification algorithm to extract semantic keypoints from raw trajectories and propose a three-stage training pipeline where trajectory information guides description generation and region localization. We further extend TraceVLM to attention trajectory-guided segmentation and video scene understanding tasks, enabling cross-frame trajectory tracking and temporal attention analysis. Based on large vision-language model reasoning capabilities, we construct the Reasoning-based Interactive Localized Narratives (RILN) dataset to enhance logical reasoning and interpretability. Extensive experiments on trajectory-guided captioning, text-guided trajectory prediction, image understanding, and segmentation demonstrate that TraceVLM achieves state-of-the-art performance, establishing a foundation for intuitive human-computer spatial interaction and interpretable visual understanding."},"_bibtex":{"value":"@misc{\nyang2026connecting,\ntitle={Connecting Where You Look With What You Understand: Trajectory-Driven Localized Understanding for Interactive Vision-Language Models},\nauthor={Fan Yang and Yousong Zhu and Shurong Zheng and Yufei Zhan and Hongyin Zhao and Xin Li and Chaoyang Zhao and Zhaowen Li and Yaowei Wang and Jinqiao Wang},\nyear={2026},\nurl={https://openreview.net/forum?id=8tZEohkGpB}\n}"},"title":{"value":"Connecting Where You Look With What You Understand: Trajectory-Driven Localized Understanding for Interactive Vision-Language Models"},"pdf":{"value":"/pdf/91874da74cdb210b5097bd1550deb3def1533b2a.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"yang|connecting_where_you_look_with_what_you_understand_trajectorydriven_localized_understanding_for_interactive_visionlanguage_models"},"authorids":{"value":["~Fan_Yang22","~Yousong_Zhu1","~Shurong_Zheng1","~Yufei_Zhan1","~Hongyin_Zhao3","~Xin_Li24","~Chaoyang_Zhao1","~Zhaowen_Li1","~Yaowei_Wang1","~Jinqiao_Wang1"]},"authors":{"value":["Fan Yang","Yousong Zhu","Shurong Zheng","Yufei Zhan","Hongyin Zhao","Xin Li","Chaoyang Zhao","Zhaowen Li","Yaowei Wang","Jinqiao Wang"]}},"version":2},{"content":{"summary":{"value":"In this work, the authors propose a LongViTU benchmark for Long-Form Video Understanding. Basically, they leverage Ego4D as data source, and develop a three-stage pipeline for QA annotation and revision. First, it builds up a hierarchical video tree to describe videos in different temporal scales. Second, they apply a sliding window approach to any subtree, and generate QA of subtree by GPT4. Third, they use GPT-4 to make a thorough revision of the generated question-answering pairs."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Please see weakness section."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1 Topic is good. Long-form video understanding is a challenging but important problem. Developping a benchmark for instruct tuning and evaluation is critical in this problem.\n\n2 Experiments are sufficient. The experimental studies are interesting to show the challenges and potentials of this benchmark."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1 This benchmark is based on EGO4D. Hence, the annotation would be similar to EgoTaskQA. As shown in Table 1, the difference is the increasing scale of data set and the newly-added timestep annotations. Is such timestep annotation important or not? Are there any expermental results to show its impact on your benchmark ?\n\n2 The hierarchical video tree style design is similar to [MoVQA: A Benchmark of Versatile Question-Answering for Long-Form Movie Understanding, arXiv:2312.04817]. \n\n3 The paper writing should be refined. The structure is OK, while the content is not quite easy to read."}},"nonreaders":[],"tmdate":1731428926446,"tcdate":1730451400373,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission7604/Reviewer_sDiL"],"signatures":["ICLR.cc/2025/Conference/Submission7604/Reviewer_sDiL"],"forum":"4j9plQoOH1","number":1,"license":"CC BY 4.0","cdate":1730451400373,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission7604/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428926446,"domain":"ICLR.cc/2025/Conference","replyto":"4j9plQoOH1","id":"D4x0ql92bi","forumContent":{"TLDR":{"value":"We propose a large-scale instruction-tuning dataset for long-form video understanding."},"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["vision language models","instruction-tuning","long-form video understanding"]},"primary_area":{"value":"datasets and benchmarks"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"This paper presents LongViTU, a large-scale (~121k QA pairs, ~900h videos), automatically generated dataset for long-form video understanding. Our key idea is inspired by the success of Large Language Models (LLMs) and Multimodal Language Models (MLMs) that are fueled by machine-generated instruction-following data (*e.g.*, InstructGPT, LLaVA). We developed a *systematic* approach to produce massive question-answeringing pairs tailored to virtually unbounded long videos by organizing them into a ***hierarchical tree***, incorporating ***self-revision*** mechanisms to guarantee high quality. We curate LongViTU for each QA pair: 1) involves a long context (average *certificate length* of 4.6 minutes); 2) requires rich knowledge and condensed reasoning (commonsense, causality, planning, *etc.*); 3) explicit labels the timestamps of relevant events throughout the entire video. Furthermore, LongViTU provides a benchmark to facilitate future research in instruction-following for long-form videos. Our experiments first reveal the performance gap between open-source video MLMs and their commercial counterparts (*e.g.*, Gemini-1.5-Pro) on this benchmark. Supervised Fine-Tuning (SFT) on open-source models led to Video-LLaVA achieving the best performance, with a GPT-4 score of $50.7$, closely following $52.3$ by the leading closed-source model Gemini-1.5-Pro, underscoring the substantial challenge posed by our benchmark. Further SFT on LongViTU with Video-LLaVA resulted in improvements of $30.7$% on the In-Distribution (ID) benchmark EgoSchema; $12.9$% and $0.6$% on the Out-of-Distribution (OOD) benchmarks WorldQA and VideoMME, respectively. These outcomes demonstrate the effectiveness and robust OOD generalizability of our proposed instruction-tuning scheme for long-form video understanding. The dataset, SFT models, and code are publicly available on the anonymous page [LongViTU](https://longvitu.github.io)."},"_bibtex":{"value":"@misc{\nwu2024longvitu,\ntitle={LongVi{TU}: Instruction Tuning for Long-Form Video Understanding},\nauthor={Rujie Wu and Xiaojian Ma and Hai Ci and Yue Fan and Yuxuan Wang and Haozhe Zhao and Qing Li and Yizhou Wang},\nyear={2024},\nurl={https://openreview.net/forum?id=4j9plQoOH1}\n}"},"title":{"value":"LongViTU: Instruction Tuning for Long-Form Video Understanding"},"pdf":{"value":"/pdf/e663a2eb9e041444826a666f95acc8764c6e736b.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"wu|longvitu_instruction_tuning_for_longform_video_understanding"},"authorids":{"value":["~Rujie_Wu2","~Xiaojian_Ma1","~Hai_Ci1","~Yue_Fan2","~Yuxuan_Wang4","~Haozhe_Zhao1","~Qing_Li1","~Yizhou_Wang1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Rujie Wu","Xiaojian Ma","Hai Ci","Yue Fan","Yuxuan Wang","Haozhe Zhao","Qing Li","Yizhou Wang"]}},"version":2},{"content":{"summary":{"value":"This paper introduces MF², a new benchmark designed to evaluate narrative comprehension in long-form movies (average 88 minutes, 53 movies in total). \n26 humans watched movies and generated paired statements consisting of an accurate factual description (fact) and a minimally modified false statement (fib). \nFor each pair, they specified the required scene granularity (single/multi/global) for answering and selected one or more comprehension dimension from {event/entity understanding, temporal perception, emotion understanding, causal reasoning}. \nIn total, 868 pairs were constructed. \nThis paper provide the result on evaluated performance using both publicly available recent VLM models and proprietary models such as GPT-4o and Gemini 2.5 Pro, alongside human assessments. \nAlso, the results of ablation studies are provided to examine the impact of input modality in (video,  subtitle)."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"Please see weaknesses."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"- Benchmark datasets are valuable assets for our community. The authors commit to releasing the full dataset, code, and movies, ensuring reproducibility and supporting future research.\n- Movies are long-form video that mostly self-contained stories. MF² focuses on holistic narrative understanding on full-length movies, requiring models to reason about the story’s core elements, unlike previous benchmarks that emphasize “needle-in-a-haystack” details. This work seems to pioneer a new area to deal with holistic understanding and reasoning across the entire narrative in movie-level videos.\n- MF² is expected high-quality, human-annotated benchmark dataset constructed with intensive human labors including full-time watching movies.\n- The multi-faceted analysis—across models, modalities, reasoning types, and comprehension dimensions—is thorough and convincing. The inclusion of human baselines is particularly valuable. It is good information as a baseline for potential users to choose the proposed benchmark dataset."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- While minimal-edit fibs are effective and annotators filters ambiguous cases, some may be too obvious or, conversely, too subtle, potentially confusing both humans and models.\n- Considering the previous benchmarks, this reviewer do not intend to challenge the current configuration. While the authors directly compares the values from the models and humans presented in Table 3, humans would need to watch the video and subtitles without sound for fair comparison. \n- For multi-scene cases, the range of multiples is not provided. I think it would be helpful to show its distribution, if the values are collected.\n- “global” claims seems to be new and substantial problems to address the overall understanding of the movie. However, it seems there are only 59 cases. Considering the total number of movies, 53, there are only 1.11 claims per a movie. As a main contribution, it looks too small.\n- (Minor) In Figures 3-5, the graph visibility could be improved. Rather than simply distinguishing by color, adding another method would make it more visible in grayscale printout."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762928336085,"tcdate":1762271358828,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission18623/Reviewer_BnV2"],"signatures":["ICLR.cc/2026/Conference/Submission18623/Reviewer_BnV2"],"forum":"cC3TW1s309","number":4,"license":"CC BY 4.0","cdate":1762271358828,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission18623/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762928336085,"domain":"ICLR.cc/2026/Conference","replyto":"cC3TW1s309","id":"Zz8dRFsCmF","forumContent":{"TLDR":{"value":"A benchmark for long video understanding"},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["vision-language models","long video understanding","memory consolidation"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Despite recent progress in vision-language models (VLMs), holistic understanding of long-form video content remains a significant challenge, partly due to limitations in current benchmarks. Many focus on peripheral, \"needle-in-a-haystack\"  details, encouraging context-insensitive retrieval over deep comprehension. Others rely on large-scale, semi-automatically generated questions (often produced by language models themselves) that are easier for models to answer but fail to reflect genuine understanding. In this paper, we introduce $\\textbf{MF$^2$}$, a new benchmark for evaluating whether models can comprehend, consolidate, and recall key narrative information---requiring integration of both visual and linguistic modalities---from full-length movies ($\\textbf{50-170 minutes long}$). MF$^2$ includes over 50 full-length, $\\textbf{open-licensed}$ movies, each paired with manually constructed sets of claim pairs---one true (fact) and one plausible but false (fib), totalling over 850 pairs. These claims target core narrative elements such as $\\textbf{character motivations}$ and $\\textbf{emotions}$, $\\textbf{causal chains}$, and $\\textbf{event order}$, and refer to $\\textbf{memorable moments}$ that humans can recall without rewatching the movie. Instead of multiple-choice formats, we adopt a binary claim evaluation protocol: for each pair, models must correctly identify both the true and false claims. This reduces biases like answer ordering and enables a more precise assessment of reasoning. Our experiments demonstrate that both open-weight and closed state-of-the-art models fall well short of human performance, underscoring the relative ease of the task for humans and their superior ability to retain and reason over critical narrative information---an ability current VLMs lack."},"_bibtex":{"value":"@misc{\nzaranis2026movie,\ntitle={Movie Facts and Fibs ({MF}\\${\\textasciicircum}2\\$): A Benchmark for Long Movie Understanding},\nauthor={Emmanouil Zaranis and Ant{\\'o}nio Farinhas and Saul Santos and Beatriz Canaverde and Miguel Moura Ramos and Aditya K Surikuchi and Andr{\\'e} Guilherme Nunes Viveiros and Baohao Liao and Elena Bueno-Benito and Nithin Sivakumaran and Pavlo Vasylenko and Shoubin Yu and Sonal Sannigrahi and Wafaa Mohammed and Ben Peters and Danae Sanchez Villegas and Elias Stengel-Eskin and Giuseppe Attanasio and Jaehong Yoon and Stella Frank and Alessandro Suglia and Chrysoula Zerva and Desmond Elliott and Mariella Dimiccoli and Mohit Bansal and Oswald Lanz and Raffaella Bernardi and Raquel Fern{\\'a}ndez and Sandro Pezzelle and Vlad Niculae and Andre Martins},\nyear={2026},\nurl={https://openreview.net/forum?id=cC3TW1s309}\n}"},"title":{"value":"Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding"},"pdf":{"value":"/pdf/5bdf44097eb82ba53623a01715adc367cc8b0428.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"zaranis|movie_facts_and_fibs_mf^2_a_benchmark_for_long_movie_understanding"},"authorids":{"value":["~Emmanouil_Zaranis1","~António_Farinhas1","~Saul_Santos1","~Beatriz_Canaverde1","~Miguel_Moura_Ramos1","~Aditya_K_Surikuchi1","~André_Guilherme_Nunes_Viveiros1","~Baohao_Liao1","~Elena_Bueno-Benito1","~Nithin_Sivakumaran1","~Pavlo_Vasylenko1","~Shoubin_Yu1","~Sonal_Sannigrahi1","~Wafaa_Mohammed1","~Ben_Peters1","~Danae_Sanchez_Villegas1","~Elias_Stengel-Eskin1","~Giuseppe_Attanasio1","~Jaehong_Yoon1","~Stella_Frank2","~Alessandro_Suglia1","~Chrysoula_Zerva1","~Desmond_Elliott1","~Mariella_Dimiccoli1","~Mohit_Bansal2","~Oswald_Lanz3","~Raffaella_Bernardi1","~Raquel_Fernández1","~Sandro_Pezzelle1","~Vlad_Niculae2","~Andre_Martins1"]},"authors":{"value":["Emmanouil Zaranis","António Farinhas","Saul Santos","Beatriz Canaverde","Miguel Moura Ramos","Aditya K Surikuchi","André Guilherme Nunes Viveiros","Baohao Liao","Elena Bueno-Benito","Nithin Sivakumaran","Pavlo Vasylenko","Shoubin Yu","Sonal Sannigrahi","Wafaa Mohammed","Ben Peters","Danae Sanchez Villegas","Elias Stengel-Eskin","Giuseppe Attanasio","Jaehong Yoon","Stella Frank","Alessandro Suglia","Chrysoula Zerva","Desmond Elliott","Mariella Dimiccoli","Mohit Bansal","Oswald Lanz","Raffaella Bernardi","Raquel Fernández","Sandro Pezzelle","Vlad Niculae","Andre Martins"]}},"version":2},{"content":{"summary":{"value":"In this work, the authors propose a training-free framework, StreamChat, for streaming video understanding. StreamChat uses a hierarchical memory system to efficiently process long video sequences for real-time, multi-turn dialogue. With parallel scheduling, it improves speed and reduces latency. \n\nThe authors also introduce a new benchmark, StreamBench, which tests video understanding across diverse scenarios, and results show StreamChat outperforms current models in accuracy and response time."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Q1: Will the example in the prompt impact the models' prediction? For example, your example in the prompt is \"{'pred': 'A'}\". Does it increase the probability of predicting \"A\"?\n\nQ2: The StreamBench benchmark supports only QA tasks now. Are there any plans to extend it to include additional streaming video understanding tasks, such as dense streaming video captioning or temporal/spatiotemporal grounding?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":1},"strengths":{"value":"S1: The paper introduces a novel hierarchical memory mechanism that compresses video representations over long sequences, enabling efficient video feature retrieval in real-time multi-turn dialogue contexts.\n\nS2: I believe the introduction of StreamBench as a benchmark for streaming video understanding is a significant contribution that can fill a critical gap in the field. By offering a standardized evaluation framework, StreamBench enables more rigorous comparisons across models, driving advancements in streaming video understanding.\n\nS3: The proposed StreamChat achieves the best performance on StreamBench, surpassing other state-of-the-art methods, which is impressive."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"W1: StreamChat was compared only with several open-source models. Including proprietary models like GPT-4o and Gemini-1.5 in the StreamBench evaluation would provide a more comprehensive comparison. \n\nW2: Human performance is absent in the comparison, which would help illustrate the gap between current models and human capabilities if included.\n\nW3: The authors compare their new benchmark only with older datasets or benchmarks, such as MSVD, MSRVTT, and ActivityNet. It would strengthen the evaluation to include comparisons with more recent benchmarks for video understanding, such as Seed-Bench [1], Video-Bench [2], MVBench [3], and LVBench [4].\n\nW4: The paper lacks an analysis of factors that impact performance on StreamBench, such as the input frame sequence length or the language model size used by the models.\n\n[1] Li, Bohao, et al. \"Seed-bench: Benchmarking multimodal llms with generative comprehension.\" arXiv preprint arXiv:2307.16125 (2023).\n\n[2] Ning, Munan, et al. \"Video-bench: A comprehensive benchmark and toolkit for evaluating video-based large language models.\" arXiv preprint arXiv:2311.16103 (2023).\n\n[3] Li, Kunchang, et al. \"Mvbench: A comprehensive multi-modal video understanding benchmark.\" Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2024.\n\n[4] Wang, Weihan, et al. \"LVBench: An Extreme Long Video Understanding Benchmark.\" arXiv preprint arXiv:2406.08035 (2024)."}},"nonreaders":[],"tmdate":1731427310241,"tcdate":1730666438586,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission743/Reviewer_Bg29"],"signatures":["ICLR.cc/2025/Conference/Submission743/Reviewer_Bg29"],"forum":"JbPb6RieNC","number":2,"license":"CC BY 4.0","cdate":1730666438586,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission743/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427310241,"domain":"ICLR.cc/2025/Conference","replyto":"JbPb6RieNC","id":"N82zHf8ifZ","forumContent":{"venue":{"value":"ICLR 2025 Poster"},"TLDR":{"value":"A Vertasile Approach and Benchmark for Streaming Video Understanding and Multi-round Interaction"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Streaming Video Understanding; Video MLLM; Hierarchical Memory System"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Recent advances in Large Language Models (LLMs) have enabled the development of Video-LLMs, advancing multimodal learning by bridging video data with language tasks. However, current video understanding models struggle with processing long video sequences, supporting multi-turn dialogues, and adapting to real-world dynamic scenarios. To address these issues, we propose StreamChat, a training-free framework for streaming video reasoning and conversational interaction. StreamChat leverages a novel hierarchical memory system to efficiently process and compress video features over extended sequences, enabling real-time, multi-turn dialogue. Our framework incorporates a parallel system scheduling strategy that enhances processing speed and reduces latency, ensuring robust performance in real-world applications. Furthermore, we introduce StreamBench, a versatile benchmark that evaluates streaming video understanding across diverse media types and interactive scenarios, including multi-turn interactions and complex reasoning tasks.  Extensive evaluations on StreamBench and other public benchmarks demonstrate that  StreamChat significantly outperforms existing\nstate-of-the-art models in terms of accuracy and response times, confirming its effectiveness for streaming video understanding. Code is available at StreamChat."},"_bibtex":{"value":"@inproceedings{\nxiong2025streaming,\ntitle={Streaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge},\nauthor={Haomiao Xiong and Zongxin Yang and Jiazuo Yu and Yunzhi Zhuge and Lu Zhang and Jiawen Zhu and Huchuan Lu},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=JbPb6RieNC}\n}"},"title":{"value":"Streaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge"},"pdf":{"value":"/pdf/6b4412176320c83956b2b47bd354a372b72935a8.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"xiong|streaming_video_understanding_and_multiround_interaction_with_memoryenhanced_knowledge"},"authorids":{"value":["~Haomiao_Xiong1","~Zongxin_Yang1","~Jiazuo_Yu1","~Yunzhi_Zhuge1","~Lu_Zhang7","~Jiawen_Zhu1","~Huchuan_Lu1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Haomiao Xiong","Zongxin Yang","Jiazuo Yu","Yunzhi Zhuge","Lu Zhang","Jiawen Zhu","Huchuan Lu"]}},"version":2},{"content":{"venue":{"value":"Multim. Tools Appl. 2023"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/s11042-023-14774-7.pdf"},"venueid":{"value":"dblp.org/journals/MTA/2023"},"paperhash":{"value":"nishimura|stateaware_video_procedural_captioning"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Taichi_Nishimura:","~Atsushi_Hashimoto1","https://dblp.org/search/pid/api?q=author:Yoshitaka_Ushiku:","~Hirotaka_Kameko1","https://dblp.org/search/pid/api?q=author:Shinsuke_Mori:"]},"html":{"value":"https://doi.org/10.1007/s11042-023-14774-7"},"_bibtex":{"value":"@article{DBLP:journals/mta/NishimuraHUKM23,\n  author={Taichi Nishimura and Atsushi Hashimoto and Yoshitaka Ushiku and Hirotaka Kameko and Shinsuke Mori},\n  title={State-aware video procedural captioning},\n  year={2023},\n  month={October},\n  cdate={1696118400000},\n  journal={Multim. Tools Appl.},\n  volume={82},\n  number={24},\n  pages={37273-37301},\n  url={https://doi.org/10.1007/s11042-023-14774-7}\n}\n"},"abstract":{"value":"Video procedural captioning (VPC), which generates procedural text from instructional videos, is an essential task for scene understanding and real-world applications. The main challenge of VPC is to describe how to manipulate materials accurately. This paper focuses on this challenge by designing a new VPC task, generating a procedural text from the clip sequence of an instructional video and material set. In this task, the state of materials is sequentially changed by manipulations, yielding their state-aware visual representations (e.g., eggs are transformed into cracked, stirred, then fried forms). The essential difficulty is to convert such visual representations into textual representations; that is, a model should track the material states after manipulations to better associate the cross-modal relations. To achieve this, we propose a novel VPC method, which modifies an existing textual simulator for tracking material states as a visual simulator and incorporates it into a video captioning model. Our experimental results show the effectiveness of the proposed method, which outperforms state-of-the-art video captioning models. We further analyze the learned embedding of materials to demonstrate that the simulators capture their state transition."},"title":{"value":"State-aware video procedural captioning"},"authors":{"value":["Taichi Nishimura","Atsushi Hashimoto","Yoshitaka Ushiku","Hirotaka Kameko","Shinsuke Mori"]}},"tmdate":1767594329634,"pdate":1672531200000,"tcdate":1731282817449,"writers":["~"],"signatures":["~Atsushi_Hashimoto1"],"forum":"SO0o7qzIdX","license":"CC BY-SA 4.0","number":178013,"cdate":1696118400000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1767594329634,"domain":"DBLP.org","id":"SO0o7qzIdX","version":2},{"content":{"summary":{"value":"This paper proposes VideoAgent, a model designed to improve video generation quality based on feedback from a VLM. Specifically, VideoAgent trains both a video generation model and a video refinement model, which share parameters and are based on diffusion processes. The video generation model functions as a conventional text-to-video model, while the video refinement model takes as input the ground-truth video, the predicted video, and feedback from the VLM to produce a refined video. The final output video can be mapped to an action trajectory. When the environment is interactive and can confirm task success, VideoAgent performs rollouts within the environment to gather additional data for training (but very slow now). This method is evaluated at the policy level in the Meta-World and iTHOR benchmarks, and for video quality on the Bridge dataset."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"When I read this paper, I encountered some areas of confusion, and I hope my feedback can help the authors improve the clarity of the writing. If there are any inaccuracies in my understanding, I welcome any corrections.\n\n1. The motivation for introducing the consistency model should be clearly articulated. I didn’t see a compelling reason to use the consistency model, as it seems possible to achieve similar results by using the training objective of DDPM and employing DDIM for inference.\n\n2. Lines 119-120 should describe the loss for the consistency model rather than for DDPM. In DDPM, the loss should be calculated from ($x^{t+1}$ to $x^t$), which also makes the \"vanilla objective for video diffusion\" mentioned in line 197 confusing.\n\n3. In line 224, it seems that video generation should be a single step if using a consistency model. Why are multiple steps required, and if they are, what distinguishes this approach from DDIM?\n\n4. The motivation in Section 3.1 could be more clearly defined. The overall objective of the proposed method is to (1) generate video and (2) refine the video based on feedback until it is accepted by the VLM, using the same parameters. However, this goal is not clearly explained at the beginning of Section 3.1, which requires readers to infer the motivation after understanding the method.\n\n5. Line 164 states, \"the model can learn to preserve the realistic part of the video while refining the hallucinatory part,\" which could be misleading. Feedback is applied to the entire video, not frame-level feedback. This statement may somewhat overstate the model’s capability.\n\n6. It is unclear how the parameters are shared between the video generation model and the video refinement model.\n\n7. It is not shown in the method how the video maps to action. Indicating where this is discussed in the paper would be helpful."},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. The idea of VideoAgent is novel; I haven’t seen an iterative approach to improving video generation quality before.\n2. Obtaining feedback on videos from VLM is feasible."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The writing section can be strengthened. The distinction between the consistency model and DDPM is vague in this paper, yet it is crucial. Given that I believe DDPM + DDIM can achieve the same results (referring to the implementation of video generation models and video refinement models), it’s even more necessary to explain the necessity and motivation for using the consistency model. Specific issues can be referenced in the questions section.\n\n2. The experimental section of the paper is relatively weak. The baseline is only AVDC, a typical video generation model without a corresponding video refinement model. Possible baselines could include VLP[1] and a naive text-to-video diffusion model (under similar training and inference FLOP). Additionally, results from traditional baselines based on BC and RL should also be included for reference.\n\n3. More realistic robotic manipulation benchmarks should be considered, such as Simper[2].\n\n4. It would be better to conduct experiments on a real robot; currently, the focus is only on testing quality using real-robot videos. Success rate evaluations are missing, which may necessitate the use of a real robot. If a real robot is not available, providing results from Simper would be acceptable.\n\n5. Even after refinement, the video generation quality of VideoAgent on the Bridge dataset is not high compared to modern classic diffusion models, and the resolution is very low.\n\n6. Fundamentally, generating a video and then generating actions is not widely regarded as a correct approach; the improvements made in this paper are still limited by the fundamental flaws of video planning. The video generation is too slow to be suitable as a policy. If the paper could demonstrate the importance of video generation for policy, it would greatly strengthen its contributions.\n\n7. The paper seems to avoid the discussion of video generation speed. In reality, how many actions can be generated per second with this method? Or what is the average time required to generate a single action? The resolution should also be reported.\n\n[1] Video Language Planning.\n[2] Evaluating Real-World Robot Manipulation Policies in Simulation"}},"nonreaders":[],"tmdate":1732522825513,"tcdate":1730658318717,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission12522/Reviewer_b49X"],"signatures":["ICLR.cc/2025/Conference/Submission12522/Reviewer_b49X"],"forum":"JaRihIHbZm","number":3,"license":"CC BY 4.0","cdate":1730658318717,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission12522/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732522825513,"domain":"ICLR.cc/2025/Conference","replyto":"JaRihIHbZm","id":"nQQWdFX1SD","forumContent":{"TLDR":{"value":"We propose VideoAgent to self improve video generation by refining video plans using external feedback, significantly reducing hallucinations and enhancing task success in simulated robotic manipulation."},"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["sequential decision making","video generation","self improvement"]},"supplementary_material":{"value":"/attachment/334107ddd0aa8c18f860cb4a9e36883922a1eda5.zip"},"primary_area":{"value":"reinforcement learning"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Video generation has been used to generate visual plans for controlling robotic systems. Given an image observation and a language instruction, previous work has generated video plans which are then converted to robot controls to be executed. However, a major bottleneck in leveraging video generation for control lies in the quality of the generated videos, which often suffer from hallucinatory content and unrealistic physics, resulting in low task success when control actions are extracted from the generated videos. While scaling up dataset and model size provides a partial solution, integrating external feedback is both natural and essential for grounding video generation in the real world. With this observation, we propose VideoAgent for self-improving generated video plans based on external feedback. Instead of directly executing the generated video plan, VideoAgent first refines the generated video plans using a novel procedure which we call self-conditioning consistency, utilizing feedback from a pretrained vision-language model (VLM). As the refined video plan is being executed, VideoAgent collects additional data from the environment to further improve video plan generation. Experiments in simulated robotic manipulation from MetaWorld and iTHOR show that VideoAgent drastically reduces hallucination, thereby boosting success rate of downstream manipulation tasks. We further illustrate that VideoAgent can effectively refine real-robot videos, providing an early indicator that robotics can be an effective tool in grounding video generation in the physical world."},"_bibtex":{"value":"@misc{\nsoni2025videoagent,\ntitle={VideoAgent: Self-Improving Video Generation},\nauthor={Achint Soni and Sreyas Venkataraman and Abhranil Chandra and Sebastian Fischmeister and Percy Liang and Bo Dai and Sherry Yang},\nyear={2025},\nurl={https://openreview.net/forum?id=JaRihIHbZm}\n}"},"title":{"value":"VideoAgent: Self-Improving Video Generation"},"pdf":{"value":"/pdf/5fa2e6bbbf4455305abf2b22f2478584bf778202.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"soni|videoagent_selfimproving_video_generation"},"authorids":{"value":["~Achint_Soni1","~Sreyas_Venkataraman1","~Abhranil_Chandra1","~Sebastian_Fischmeister1","~Percy_Liang1","~Bo_Dai1","~Sherry_Yang1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Achint Soni","Sreyas Venkataraman","Abhranil Chandra","Sebastian Fischmeister","Percy Liang","Bo Dai","Sherry Yang"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a novel approach for text-based video-to-video editing that eliminates the need for resource-heavy finetuning for each video and model. The authors introduce a synthetic paired video dataset tailored for video-to-video transfer tasks, taking inspiration from image transfer methods such as Instruct Pix2Pix. This method translates the Prompt-to-Prompt model to videos and efficiently generates paired samples, each consisting of an input video and its edited counterpart. They also propose Long Video Sampling Correction (LVSC), ensuring consistent long videos across batches. Their method outperforms existing methods like Tune-A-Video in terms of text-based video-to-video editing, paving new avenues for exploration and deployment."},"presentation":{"value":"3 good"},"contribution":{"value":"3 good"},"soundness":{"value":"3 good"},"strengths":{"value":"1. The proposed method eliminates the need for per-video-per-model finetuning, potentially saving significant computational resources. The creation of a synthetic paired video dataset tailored for video-to-video transfer tasks is a novel approach that could prove beneficial for training models in this domain. The introduction of LVSC addresses the challenge of maintaining consistency in long videos across batches, a notable improvement over existing methods.\n2. Sufficient and comprehensive experiments (both quantitatively and qualitatively) on the comparisons are given with prior arts and ablations of key designs. The given method gives notable quantitative improvements and its visual results faithfully follow the given instructions compared with other approaches from the supp."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The generated video shows visual appeals in the given content with different styles while presenting severe jitter in the newly added content.\n2. The performance of the proposed method relies on the synthetic paired video dataset. Though it gives reasonable sampling strategies on the generated data, if this dataset doesn't closely match real-world scenarios, it may limit the model's utility.\n3. The paper does not talk about potential failure cases or limitations of their approach in the main paper, which could help us better understand the proposed system."},"confidence":{"value":"4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work."},"questions":{"value":"1. It would be better to quantify the difference between the generated video dataset and some reference one, e.g., computing FID between them. It may also be helpful to validate the effectiveness of the given sampling criteria.\n2. How compatible is this approach with different types of video content, various editing instructions, and text-to-video generation methods?"},"rating":{"value":"6: marginally above the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1700737363345,"tcdate":1698786656594,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission873/Reviewer_BaKN"],"signatures":["ICLR.cc/2024/Conference/Submission873/Reviewer_BaKN"],"forum":"IoKRezZMxF","number":2,"license":"CC BY 4.0","cdate":1698786656594,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission873/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1700737363345,"domain":"ICLR.cc/2024/Conference","replyto":"IoKRezZMxF","id":"BXU4Afgzrm","forumContent":{"venue":{"value":"ICLR 2024 poster"},"TLDR":{"value":"We've developed a synthetic dataset to train a text-based video editing model, eliminating the need for per-video fine-tuning, and introduced a method for seamless long video editing."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Computer Vision","Video Editing","Diffusion Model"]},"supplementary_material":{"value":"/attachment/9cab55bbe2e33a78087b54a1213ffdd7a40cefee.zip"},"primary_area":{"value":"generative models"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"We introduce a novel and efficient approach for text-based video-to-video editing that eliminates the need for resource-intensive per-video-per-model finetuning. At the core of our approach is a synthetic paired video dataset tailored for video-to-video transfer tasks. Inspired by Instruct Pix2Pix's image transfer via editing instruction, we adapt this paradigm to the video domain. Extending the Prompt-to-Prompt to videos, we efficiently generate paired samples, each with an input video and its edited counterpart. Alongside this, we introduce the Long Video Sampling Correction during sampling, ensuring consistent long videos across batches. Our method surpasses current methods like Tune-A-Video, heralding substantial progress in text-based video-to-video editing and suggesting exciting avenues for further exploration and deployment."},"_bibtex":{"value":"@inproceedings{\ncheng2024consistent,\ntitle={Consistent Video-to-Video Transfer Using Synthetic Dataset},\nauthor={Jiaxin Cheng and Tianjun Xiao and Tong He},\nbooktitle={The Twelfth International Conference on Learning Representations},\nyear={2024},\nurl={https://openreview.net/forum?id=IoKRezZMxF}\n}"},"title":{"value":"Consistent Video-to-Video Transfer Using Synthetic Dataset"},"pdf":{"value":"/pdf/e91267e5675cf150b685ada8cd0303644c45a25f.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference"},"paperhash":{"value":"cheng|consistent_videotovideo_transfer_using_synthetic_dataset"},"authorids":{"value":["~Jiaxin_Cheng1","~Tianjun_Xiao1","~Tong_He5"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Jiaxin Cheng","Tianjun Xiao","Tong He"]}},"version":2},{"content":{"venue":{"value":"ICLR 2027 Conference Submission"},"TLDR":{"value":"We show that current video embedding models poorly capture motion dynamics, introduce a benchmark for motion-sensitive video retrieval, and propose a motion-aware embedding model to help close this gap."},"keywords":{"value":["Motion-Aware Video Retreival","Motion Dynamics","Video Retrieval"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"We identify a fundamental limitation in current video retrieval models: they rely heavily on static appearance cues and inadequately capture the motion dynamics that distinguish visually similar actions. To systematically study this limitation, we introduce MAVR, a motion-centric benchmark and training dataset built largely from newly collected videos for fine-grained text-to-video and video-to-text retrieval. MAVR combines motion-defined actions, scene-matched hard negatives, and a controlled caption scheme that separates action information from scene and viewpoint cues. Despite performing strongly on a broader multimodal embedding benchmark, existing models perform poorly on MAVR, revealing a substantial gap in motion-sensitive retrieval. As a first step toward closing this gap, we introduce MAVE, a motion-aware embedding architecture that augments a multimodal embedding model with temporally informed tokens from a dedicated video pathway. When trained on MAVR, MAVE substantially outperforms state-of-the-art retrieval baselines on actions represented during training, demonstrating the benefit of explicit temporal modeling. However, these gains do not consistently generalize to unseen actions underscoring that robust motion-aware retrieval remains an open challenge."},"_bibtex":{"value":"@inproceedings{\nanonymous2026beyond,\ntitle={Beyond Static Cues: Motion-Aware Video Retrieval},\nauthor={Anonymous},\nbooktitle={Submitted to The Fifteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=K0LzQ1216m},\nnote={under review}\n}"},"title":{"value":"Beyond Static Cues: Motion-Aware Video Retrieval"},"pdf":{"value":"/pdf/8c90f6ca364fb8721c987d724df8267dbb6d9f55.pdf"},"venueid":{"value":"ICLR.cc/2027/Conference/Submission"}},"tmdate":1791221025816,"tcdate":1787942403172,"writers":["ICLR.cc/2027/Conference","ICLR.cc/2027/Conference/Submission5973/Authors"],"signatures":["ICLR.cc/2027/Conference/Submission5973/Authors"],"forum":"K0LzQ1216m","license":"CC BY 4.0","number":5973,"cdate":1787942403172,"readers":["everyone"],"invitations":["ICLR.cc/2027/Conference/-/Submission","ICLR.cc/2027/Conference/-/Edit","ICLR.cc/2027/Conference/-/Submission_Change_Before_Bidding","ICLR.cc/2027/Conference/Submission5973/-/Full_Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Reviewing"],"mdate":1791221025816,"odate":1791031328286,"domain":"ICLR.cc/2027/Conference","id":"K0LzQ1216m","version":2},{"content":{"venue":{"value":"CVPR 2026"},"abstract":{"value":"We propose Context-aware Video-text Alignment (CVA), a novel framework to address a significant challenge in video temporal grounding: achieving temporally sensitive video-text alignment that remains robust to irrelevant background context. %CVA integrates complementary data-centric and architectural innovations to enhance contextual robustness. Our framework is built on three key components. First, we propose Query-aware Context Diversification (QCD), a new data augmentation strategy that ensures only semantically unrelated content is mixed in. It builds a video-text similarity-based pool of replacement clips to simulate diverse contexts while preventing the false negative caused by query-agnostic mixing. Second, we introduce the Context-invariant Boundary Discrimination (CBD) loss, a contrastive loss that enforces semantic consistency at challenging temporal boundaries, making their representations robust to contextual shifts and hard negatives. Third, we introduce the Context-enhanced Transformer Encoder (CTE), a hierarchical architecture that combines windowed self-attention and bidirectional cross-attention with learnable queries to capture multi-scale temporal context. Through the synergy of these data-centric and architectural enhancements, CVA achieves state-of-the-art performance on major VTG benchmarks, including QVHighlights and Charades-STA. Notably, our method achieves a significant improvement of approximately 5 points in Recall@1 (R1) scores over state-of-the-art methods, highlighting its effectiveness in mitigating false negatives."},"_bibtex":{"value":"@inproceedings{\nmoon2026cva,\ntitle={{CVA}: Context-aware Video-text Alignment for Video Temporal Grounding},\nauthor={Sungho Moon and Seunghun Lee and Jiwan Seo and Sunghoon Im},\nbooktitle={Conference on Computer Vision and Pattern Recognition 2026},\nyear={2026},\nurl={https://openreview.net/forum?id=yYAofV4EMb}\n}"},"title":{"value":"CVA: Context-aware Video-text Alignment for Video Temporal Grounding"},"pdf":{"value":"https://openaccess.thecvf.com/content/CVPR2026/papers/Moon_CVA_Context-aware_Video-text_Alignment_for_Video_Temporal_Grounding_CVPR_2026_paper.pdf"},"venueid":{"value":"thecvf.com/CVPR/2026/Conference"},"paperhash":{"value":"moon|cva_contextaware_videotext_alignment_for_video_temporal_grounding"},"authorids":{"value":["~Sungho_Moon1","~Seunghun_Lee3","~Jiwan_Seo2","~Sunghoon_Im1"]},"authors":{"value":["Sungho Moon","Seunghun Lee","Jiwan Seo","Sunghoon Im"]}},"tmdate":1789656567658,"pdate":1789656437779,"tcdate":1765221906689,"writers":["thecvf.com/CVPR/2026/Conference","thecvf.com/CVPR/2026/Conference/Submission38005/Authors"],"signatures":["thecvf.com/CVPR/2026/Conference/Submission38005/Authors"],"forum":"yYAofV4EMb","license":"CC BY 4.0","number":38005,"cdate":1765221906689,"readers":["everyone"],"invitations":["thecvf.com/CVPR/2026/Conference/-/Submission","thecvf.com/CVPR/2026/Conference/Submission38005/-/Full_Submission","thecvf.com/CVPR/2026/Conference/-/Post_Submission","thecvf.com/CVPR/2026/Conference/Submission38005/-/Supplementary_Material","thecvf.com/CVPR/2026/Conference/-/Edit","thecvf.com/CVPR/2026/Conference/-/Compute_Flag"],"mdate":1789656567658,"odate":1789656437779,"domain":"thecvf.com/CVPR/2026/Conference","id":"yYAofV4EMb","version":2},{"content":{"venue":{"value":"MMM (5) 2025"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-981-96-2074-6_36.pdf"},"venueid":{"value":"dblp.org/conf/MMM/2025"},"paperhash":{"value":"cheng|interactive_video_search_with_multimodal_llm_video_captioning"},"authorids":{"value":["~Yu-Tong_Cheng1","https://dblp.org/search/pid/api?q=author:Jiaxin_Wu_0001:","~Zhixin_Ma1","~Jiangshan_He1","https://dblp.org/search/pid/api?q=author:Xiao-Yong_Wei:","https://dblp.org/search/pid/api?q=author:Chong-Wah_Ngo:"]},"html":{"value":"https://doi.org/10.1007/978-981-96-2074-6_36"},"_bibtex":{"value":"@inproceedings{DBLP:conf/mmm/ChengWMHWN25,\n  author={Yu-Tong Cheng and Jiaxin Wu and Zhixin Ma and Jiangshan He and Xiao-Yong Wei and Chong-Wah Ngo},\n  title={Interactive Video Search with Multi-modal LLM Video Captioning},\n  year={2025},\n  cdate={1735689600000},\n  pages={302-309},\n  url={https://doi.org/10.1007/978-981-96-2074-6_36},\n  booktitle={MMM (5)},\n  crossref={conf/mmm/2025-5}\n}\n"},"abstract":{"value":"Cross-modal representation learning is essential for interactive text-to-video search tasks. However, the representation learning is limited by the size and quality of video-caption pairs. To improve the search accuracy, we propose to enlarge the size of available video-caption pairs by leveraging multi-model LLM on video captioning. Specifically, we use LLM to generate video captions for a large video collection (i.e., WebVid dataset) and use the generated video-caption pairs to pre-train a text-to-video search model. Additionally, we use LLM to generate fine-grained captions for test video collections to enable text-to-caption retrieval. Furthermore, we build a semantic overview of the retrieved rank list based on the detailed captions in our interactive video retrieval system which act as hints for user to refine their query. Experimental results show that the generated captions are effective in improving the search accuracy of both AVS and T-KIS tasks on the TRECVid datasets."},"title":{"value":"Interactive Video Search with Multi-modal LLM Video Captioning"},"authors":{"value":["Yu-Tong Cheng","Jiaxin Wu","Zhixin Ma","Jiangshan He","Xiao-Yong Wei","Chong-Wah Ngo"]}},"tmdate":1779787004530,"pdate":1735689600000,"tcdate":1740455326830,"writers":["~"],"signatures":["~Zhixin_Ma1"],"forum":"t1VhwdRfiP","license":"CC BY-SA 4.0","number":340524,"cdate":1735689600000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1779787004530,"domain":"DBLP.org","id":"t1VhwdRfiP","version":2},{"content":{"summary":{"value":"In the context of matrix sensing with deep matrix factorizations, the paper analyzes the inductive bias of interpolators with minimal Hessian trace, which is a well-known measure of sharpness. Specifically, under a Restricted Isometry Property (RIP) assumption on the linear measurements, it establishes that the minimal Hessian trace with which the deep matrix factorization can express an interpolator is approximately equal to the latter’s nuclear norm. In turn, this yields guarantees on the recovery of the ground truth matrix, as well as a separation result from the recovery obtained by the minimal Frobenius norm interpolator.\n\nAs an additional contribution, develops a closed form expression for the minimal Hessian trace of an interpolator when there is only a single linear measurement."},"presentation":{"value":"4 excellent"},"contribution":{"value":"3 good"},"soundness":{"value":"4 excellent"},"strengths":{"value":"1. Well-written and easy to follow. The motivation, problem at hand, and main results are clearly described.\n\n2. Helpful discussions are provided for each of the technical results, in particular regarding their implications and relation to existing work.\n\n3. The technical contributions establish a new connection between flatness, in terms of minimal Hessian trace, and generalization for deep matrix factorizations (previous work of Ding et al. 2022 studies only depth two factorizations). Despite conventional wisdom and empirical evidence suggesting that flatness may lead to good generalization, formally showing that this is the case has proven challenging. This attests to the significance of the current paper’s results.\nPersonally, I found it interesting that depth does not help in the analyzed setting, in the sense that minimizing the Hessian trace amounts to approximately minimizing the nuclear norm, similar to the depth L = 2 case. This may suggest that deep matrix factorization with RIP measurements is an unsatisfactory setting for uncovering the benefits of depth in deep learning and its (possible) relation with flatness.\n"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The points below pertain to the scope of the technical contributions, suggesting places for improvement. I do not believe that points (2) and (3) significantly harm the quality of the current paper and are perhaps considerations for future work. For (1), it seems that it may be straightforward to incorporate, in which case I believe it can increase the generality of the current work.\n\n1. Only interpolators with exactly minimal Hessian trace are considered. Since in practice a more realistic view is that it is only approximately minimized, if possible, it is worth (even if in an appendix) extending the generalization results to interpolators whose Hessian trace is approximately minimal.\n\n2. The current paper characterizes the regularizer induced by minimizing the Hessian trace only for interpolators. Since using explicit regularization can lead to solutions that don't entirely minimize the loss to zero, I believe it is important to understand the inductive bias of minimizing the Hessian trace for non-interpolating linear mappings as well. While more complex, it may be more realistic.\n\n3. The current paper only treats the question of what interpolating the data with minimal Hessian trace implies, and not the complementary question of whether it, or other measures of sharpness more generally, are implicitly minimized in deep matrix factorization by standard optimizers.\n\nAdditional (minor) comments and typos:\n- Typo in line 75: “regularzier” should be “regularizer”.\n- Typo in line 101: I believe “observe” should have been “observed”.\n- Typo in equation after line 166: I believe there is a missing gradient symbol in the middle expression and a factor of two in both the middle and right expressions.\n\n"},"confidence":{"value":"4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work."},"questions":{"value":"Have you considered other notions of sharpness besides the trace of the Hessian, e.g. its the maximal eigenvalue? In particular, can minimizing the maximal eigenvalue of the Hessian ensure generalization in the considered setting, or is it too weak and one must look at the trace of the Hessian?"},"rating":{"value":"7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations."},"code_of_conduct":{"value":"Yes"},"limitations":{"value":"The authors have adequately addressed limitations of the work."}},"nonreaders":[],"tmdate":1702411454098,"tcdate":1688421182106,"writers":["NeurIPS.cc/2023/Conference","NeurIPS.cc/2023/Conference/Submission13921/Reviewer_ecS4"],"signatures":["NeurIPS.cc/2023/Conference/Submission13921/Reviewer_ecS4"],"forum":"2hQ7MBQApp","number":2,"license":"CC BY 4.0","cdate":1688421182106,"mdate":1702411454098,"readers":["everyone"],"invitations":["NeurIPS.cc/2023/Conference/Submission13921/-/Official_Review","NeurIPS.cc/2023/Conference/-/Edit"],"domain":"NeurIPS.cc/2023/Conference","replyto":"2hQ7MBQApp","id":"TMLFkMiiS3","forumContent":{"venue":{"value":"NeurIPS 2023 poster"},"keywords":{"value":["Sharpness minimization","Deep learning","Matrix factorization","Deep linear networks","Implicit bias","SGD","Trace of Hessian regularizer"]},"supplementary_material":{"value":"/attachment/0f0ed4619c0edc162df334fe30e3e932f1b85e2d.pdf"},"_bibtex":{"value":"@inproceedings{\ngatmiry2023what,\ntitle={What is the Inductive Bias of Flatness Regularization? A Study of Deep Matrix Factorization Models},\nauthor={Khashayar Gatmiry and Zhiyuan Li and Tengyu Ma and Sashank J. Reddi and Stefanie Jegelka and Ching-Yao Chuang},\nbooktitle={Thirty-seventh Conference on Neural Information Processing Systems},\nyear={2023},\nurl={https://openreview.net/forum?id=2hQ7MBQApp}\n}"},"title":{"value":"What is the Inductive Bias of Flatness Regularization? A Study of Deep Matrix Factorization Models"},"paperhash":{"value":"gatmiry|what_is_the_inductive_bias_of_flatness_regularization_a_study_of_deep_matrix_factorization_models"},"abstract":{"value":"Recent works on over-parameterized neural networks have shown that  the stochasticity in optimizers has the implicit regularization effect of minimizing the sharpness of the loss function (in particular, the trace of its Hessian) over the family zero-loss solutions. More explicit forms of flatness regularization also empirically improve the generalization performance. However, it remains unclear why and when flatness regularization leads to better generalization. \nThis work takes the first step towards understanding the inductive bias of the minimum trace of the Hessian solutions in an important setting: learning deep linear networks from linear measurements, also known as \\emph{deep matrix factorization}. We show that with the standard Restricted Isometry Property (RIP) on the measurements, minimizing the trace of Hessian is approximately equivalent to minimizing the Schatten 1-norm of the corresponding end-to-end matrix parameters (i.e., the product of all layer matrices), which in turn leads to better generalization."},"pdf":{"value":"/pdf/4043ffb8be238fac592a296a07c31643f83ad421.pdf"},"venueid":{"value":"NeurIPS.cc/2023/Conference"},"authorids":{"value":["~Khashayar_Gatmiry1","~Zhiyuan_Li2","~Tengyu_Ma1","~Sashank_J._Reddi1","~Stefanie_Jegelka3","~Ching-Yao_Chuang1"]},"authors":{"value":["Khashayar Gatmiry","Zhiyuan Li","Tengyu Ma","Sashank J. Reddi","Stefanie Jegelka","Ching-Yao Chuang"]}},"version":2},{"content":{"summary":{"value":"This paper examines the task of adapting a Video-LLM, pre-trained on short videos, to handle long videos without additional training. It identifies key challenges related to the fixed video encoder, modality projector, and the limited context window size of the LLM. To address these issues, the authors propose a video token rearrangement technique and an extension method for the LLM's context window to accommodate a greater number of visual tokens. Additionally, they introduce a KV-cache compression mechanism to minimize memory usage during inference. These innovations enable the proposed INTP-Video-LLaVA model to process videos with up to 32 frames."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"My primary concern is the scope of the experimental section. It is recommended that the authors expand their experimental analysis to encompass a variety of Video-LLMs and conduct ablation studies for each design decision. This would offer a more thorough understanding of the findings for the readers. Please find more details in Weaknesses."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. **Innovation and Relevance:** The paper introduces a novel training-free approach to extend the input video length compatibility for Video-LLMs. This contribution is significant and addresses a timely issue in the field. \n\n2. **Clarity and Coherence:** The paper is commendably well-structured and it is easy to follow.\n\n3. **Insightful Analysis:** The paper offers a thorough explanation of the challenges associated with video encoder limitations, LLM context window size constraints, and KV-cache management during inference in current Video-LLMs. This analysis is particularly valuable for the community, providing insights that can guide future research and development."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The applicability of the proposed training-free techniques across a range of Video-LLMs is a critical aspect to assess. The paper, however, presents experimental results solely for Video-LLaVA. It is recommended that the authors expand their experiments to include additional Video-LLMs to demonstrate the broader applicability of the techniques.\n\n2. The paper lacks certain crucial experiments. An ablation study examining the impact of the video token rearrangement techniques and the efficacy of every design choice of the RoPE interpolation methods is recommended. Such studies would provide a more comprehensive understanding of the contributions of these specific aspects to the overall performance."}},"nonreaders":[],"tmdate":1731427986204,"tcdate":1730722317706,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission3931/Reviewer_4Edy"],"signatures":["ICLR.cc/2025/Conference/Submission3931/Reviewer_4Edy"],"forum":"QrTvFCa4nX","number":4,"license":"CC BY 4.0","cdate":1730722317706,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission3931/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427986204,"domain":"ICLR.cc/2025/Conference","replyto":"QrTvFCa4nX","id":"7AQ6hh6VFo","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"TLDR":{"value":"Obtain a Video-LLM which can process more frames in a totally training-free manner."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video Understanding"]},"primary_area":{"value":"foundation or frontier models, including LLMs"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Advancements in Large Language Models (LLMs) inspire various strategies for integrating video modalities. \nA key approach is Video-LLMs, which incorporate an optimizable interface linking sophisticated video encoders to LLMs. \nHowever, due to computation and data limitations, these Video-LLMs are typically pre-trained to process only short videos, limiting their broader application for understanding longer video content. Additionally, fine-tuning Video-LLMs to handle longer videos is cost-prohibitive.\nConsequently, it becomes essential to explore the interpolation of Video-LLMs under a completely training-free setting. In this paper, we first identify the primary challenges in interpolating Video-LLMs: (1) the video encoder and modality alignment projector are fixed, preventing the integration of additional frames into Video-LLMs, and (2) the LLM backbone is limited in its content length capabilities, which complicates the processing of an increased number of video tokens.\nTo address these challenges, we propose a specific INTerPolation method for Video-LLMs (INTP-Video-LLMs). We introduce an alternative video token rearrangement technique that circumvents limitations imposed by the fixed video encoder and alignment projector. Furthermore, we introduce a training-free LLM context window extension method to enable Video-LLMs to understand a correspondingly increased number of visual tokens."},"_bibtex":{"value":"@misc{\nshang2025interpolating,\ntitle={Interpolating Video-{LLM}s:  Toward Longer-sequence {LMM}s in a Training-free Manner},\nauthor={Yuzhang Shang and Bingxin Xu and Weitai Kang and Mu Cai and Yuheng Li and Zehao Wen and Zhen Dong and Kurt Keutzer and Yong Jae Lee and Yan Yan},\nyear={2025},\nurl={https://openreview.net/forum?id=QrTvFCa4nX}\n}"},"title":{"value":"Interpolating Video-LLMs:  Toward Longer-sequence LMMs in a Training-free Manner"},"pdf":{"value":"/pdf/610d4d569fba56bc840ff0a6d6108535129e04d6.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"shang|interpolating_videollms_toward_longersequence_lmms_in_a_trainingfree_manner"},"authorids":{"value":["~Yuzhang_Shang1","~Bingxin_Xu1","~Weitai_Kang1","~Mu_Cai1","~Yuheng_Li1","~Zehao_Wen1","~Zhen_Dong3","~Kurt_Keutzer1","~Yong_Jae_Lee2","~Yan_Yan6"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Yuzhang Shang","Bingxin Xu","Weitai Kang","Mu Cai","Yuheng Li","Zehao Wen","Zhen Dong","Kurt Keutzer","Yong Jae Lee","Yan Yan"]}},"version":2},{"content":{"venue":{"value":"IEEE Trans. Circuits Syst. Video Technol. 2025"},"pdf":{"value":"https://ieeexplore.ieee.org/iel8/76/10885778/10706923.pdf"},"venueid":{"value":"dblp.org/journals/TCSV/2025"},"paperhash":{"value":"yu|prompting_videolanguage_foundation_models_with_domainspecific_finegrained_heuristics_for_video_question_answering"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Ting_Yu_0002:","https://dblp.org/search/pid/api?q=author:Kunhao_Fu:","~Shuhui_Wang1","https://dblp.org/search/pid/api?q=author:Qingming_Huang:","https://dblp.org/search/pid/api?q=author:Jun_Yu_0002:"]},"html":{"value":"https://doi.org/10.1109/TCSVT.2024.3475510"},"_bibtex":{"value":"@article{DBLP:journals/tcsv/YuFWHY25,\n  author={Ting Yu and Kunhao Fu and Shuhui Wang and Qingming Huang and Jun Yu},\n  title={Prompting Video-Language Foundation Models With Domain-Specific Fine-Grained Heuristics for Video Question Answering},\n  year={2025},\n  month={February},\n  cdate={1738368000000},\n  journal={IEEE Trans. Circuits Syst. Video Technol.},\n  volume={35},\n  number={2},\n  pages={1615-1630},\n  url={https://doi.org/10.1109/TCSVT.2024.3475510}\n}\n"},"abstract":{"value":"Video Question Answering (VideoQA) represents a crucial intersection between video understanding and language processing, requiring both discriminative unimodal comprehension and sophisticated cross-modal interaction for accurate inference. Despite advancements in multi-modal pre-trained models and video-language foundation models, these systems often struggle with domain-specific VideoQA due to their generalized pre-training objectives. Addressing this gap necessitates bridging the divide between broad cross-modal knowledge and the specific inference demands of VideoQA tasks. To this end, we introduce HeurVidQA, a framework that leverages domain-specific entity-action heuristics to refine pre-trained video-language foundation models. Our approach treats these models as implicit knowledge engines, employing domain-specific entity-action prompters to direct the model’s focus toward precise cues that enhance reasoning. By delivering fine-grained heuristics, we improve the model’s ability to identify and interpret key entities and actions, thereby enhancing its reasoning capabilities. Extensive evaluations across multiple VideoQA datasets demonstrate that our method significantly outperforms existing models, underscoring the importance of integrating domain-specific knowledge into video-language models for more accurate and context-aware VideoQA."},"title":{"value":"Prompting Video-Language Foundation Models With Domain-Specific Fine-Grained Heuristics for Video Question Answering"},"authors":{"value":["Ting Yu","Kunhao Fu","Shuhui Wang","Qingming Huang","Jun Yu"]}},"tmdate":1767802542407,"pdate":1735689600000,"externalIds":["dblp:journals/tcsv/YuFWHY25"],"tcdate":1767802539897,"writers":["~"],"signatures":["~Shuhui_Wang1"],"forum":"3JY3SxbMek","license":"CC BY-SA 4.0","number":730836,"cdate":1738368000000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1767802542407,"domain":"DBLP.org","id":"3JY3SxbMek","version":2},{"content":{"venue":{"value":"PCM (2) 2015"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-319-24078-7_40.pdf"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"xing|synthesisaware_regionbased_3d_video_coding"},"html":{"value":"https://doi.org/10.1007/978-3-319-24078-7_40"},"_bibtex":{"value":"@inproceedings{DBLP:conf/pcm/XingWJW15,\n  author={Zhiwei Xing and Anhong Wang and Jian Jin and Yingchun Wu},\n  title={Synthesis-Aware Region-Based 3D Video Coding},\n  year={2015},\n  cdate={1420070400000},\n  pages={400-409},\n  url={https://doi.org/10.1007/978-3-319-24078-7_40},\n  booktitle={PCM (2)},\n  crossref={conf/pcm/2015-2}\n}\n"},"abstract":{"value":"In a depth-image-based rendering (DIBR)-based 3D video system, the original 3D video is commonly compressed from the point of the video itself, paying little attention to its contribution to the synthesized virtual view. This paper first proposes a method to divide the original video into different regions according to its contribution to the synthesized view and then proposes to measure the regions using compressive sensing with different measurement rates, leading to a synthesis-aware region-based 3D video coding approach. Experimental results show that our approach can achieve better synthesized quality under the same equivalent measurement rate. Our approach is suitable for the applications when the virtual view is more important than the original views."},"title":{"value":"Synthesis-Aware Region-Based 3D Video Coding"},"authors":{"value":[{"fullname":"Zhiwei Xing","username":""},{"fullname":"Anhong Wang","username":"~Anhong_Wang1"},{"fullname":"Jian Jin","username":""},{"fullname":"Yingchun Wu","username":""}]}},"tmdate":1786454759547,"pdate":1451520000000,"externalIds":["dblp:conf/pcm/XingWJW15"],"tcdate":1786454743965,"writers":["~","OpenReview.net/Public_Article/DBLP.org","OpenReview.net/Support"],"signatures":["~Anhong_Wang1"],"forum":"g5a6okSLr2","license":"CC BY-SA 4.0","number":129695,"cdate":1420070400000,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/DBLP.org/-/Record","OpenReview.net/Public_Article/-/Edit"],"mdate":1786454759547,"domain":"OpenReview.net/Public_Article","id":"g5a6okSLr2","version":2},{"content":{"summary":{"value":"This paper focuses on long video editing with training-free diffusion models.  The authors propose a useful cross-window attention mechanism to ensure the consistency and length of the video. They also leverage DDIM for accurate control and a video frame interpolation model to mitigate the frame-level flickering issue. The authors presented rich and excellent experimental results, and provided their code and products in supplementary materials."},"soundness":{"value":"3 good"},"confidence":{"value":"5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully."},"questions":{"value":"Cross-window attention is proposed to improve inter-window consistency, then how do the authors ensure consistency within the window? \nAs shown in Figure 5 (a), fully crossframe attention-based ControlVideo not only maintains the loyalty of the background, but also maintains the temporal consistency of the target subject. On the contrary, the author's method blurs the background and changes the target subject in the temporal sequence (the car window turns red)."},"rating":{"value":"5: marginally below the acceptance threshold"},"details_of_ethics_concerns":{"value":"No ethics concerns."},"code_of_conduct":{"value":"Yes"},"presentation":{"value":"3 good"},"contribution":{"value":"2 fair"},"strengths":{"value":"Long video editing, consistency maintenance, and structural fidelity are fundamental issues in text guided long video editing. The authors propose a simple and systematic solution for training free long video editing. The authors present a good number of experiments validating the effectiveness of their approach. The paper is overall well written and easy to follow."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The authors have pieced together too many other people's methods to achieve the goals, so their model capabilities are limited, such as being unable to perform shape transformations, object additions or deletions. So their products are limited to color or style changes.\n\n2. The innovation of the methods proposed by the authors is limited, as they have not effectively established the temporal information of the video. The so-called cross-window is naive (a minor change of ControlVideo **[1]**) and may be helpful for local video smoothing, but it still cannot truly establish the temporal dependence of long videos. This is reflected in the experimental results that although the generated video actions are locally coherent, the overall appearance is somewhat strange.\n\n**[1]** Zhang, Yabo, Yuxiang Wei, Dongsheng Jiang, Xiaopeng Zhang, Wangmeng Zuo and Qi Tian. “ControlVideo: Training-free Controllable Text-to-Video Generation.” ArXiv abs/2305.13077 (2023): n. pag."}},"nonreaders":[],"tmdate":1699636306041,"tcdate":1698525115946,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission3522/Reviewer_WhZ8"],"signatures":["ICLR.cc/2024/Conference/Submission3522/Reviewer_WhZ8"],"forum":"9ux2cgxw6O","number":1,"license":"CC BY 4.0","cdate":1698525115946,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission3522/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636306041,"domain":"ICLR.cc/2024/Conference","replyto":"9ux2cgxw6O","id":"SCluH7MCAn","forumContent":{"venue":{"value":"Submitted to ICLR 2024"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video editing","Diffusion models","Training-free"]},"supplementary_material":{"value":"/attachment/dc950b2a21404113366bc65150acc9194a6c097c.zip"},"primary_area":{"value":"generative models"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Leveraging pre-trained conditional diffusion models for video editing without further tuning has gained increasing attention due to its promise in film production, advertising, etc. Yet, seminal works in this line fall short in generation length, temporal coherence, or fidelity to the source video. This paper aims to bridge the gap, establishing a simple and effective baseline for training-free diffusion model-based long video editing. As suggested by prior arts, we build the pipeline upon ControlNet, which excels at various image editing tasks based on text prompts. To break down the length constraints caused by limited computational memory, we split the long video into consecutive windows and develop a novel cross-window attention mechanism to ensure the consistency of global style and maximize the smoothness among windows. To achieve more accurate control, we extract the information from the source video via DDIM inversion and integrate the outcomes into the latent feature maps of the generations. We also incorporate a video frame interpolation model to mitigate frame-level flickering issues further. Extensive empirical studies verify the superior efficacy of our method over competing baselines across scenarios, including replacing attributes of foreground objects, style transfer, and background replacement. In particular, our method manages to edit videos with up to 128 frames according to user requirements."},"_bibtex":{"value":"@misc{\nliao2024lovecon,\ntitle={{LOVEC}on: Text-driven Training-free Long Video Editing with ControlNet},\nauthor={Zhenyi Liao and Zhijie Deng},\nyear={2024},\nurl={https://openreview.net/forum?id=9ux2cgxw6O}\n}"},"title":{"value":"LOVECon: Text-driven Training-free Long Video Editing with ControlNet"},"pdf":{"value":"/pdf/12fce61dbe762b0442b6381cfd16a5bcd151aa49.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Rejected_Submission"},"paperhash":{"value":"liao|lovecon_textdriven_trainingfree_long_video_editing_with_controlnet"},"authorids":{"value":["~Zhenyi_Liao1","~Zhijie_Deng1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Zhenyi Liao","Zhijie Deng"]}},"version":2},{"content":{"summary":{"value":"This paper systematically studies the rated-distortion-perception tradeoff in neural video compression. Since videos are made of consecutive frames, this paper considers two different perceptual metrics, including the joint distribution of all video frames (PLF-JD), and per-frame perceptual loss (framewise marginal distribution, PLF-FDM). Some conclusions are proposed regarding the choice of PLF-JD or PLF-FDM. \n\nIn addition, motivated by the universal encoded representations in previous RDP papers from the field of image compression, the authors in this paper demonstrate the video-version universal encoded representations hold when the perceptual constraint is PLF-FDM. While a similar result does not hold for PLF-JD in general, it is satisfied for a special class of encoders which operate in the low-rate regime. The universal encoded representations in video compression have several advantages for video compression such as when it needs to extract motion information from an MSE-based reconstruction.  \n"},"presentation":{"value":"4 excellent"},"contribution":{"value":"3 good"},"soundness":{"value":"4 excellent"},"strengths":{"value":"Originality and quality of this paper are pretty good. The PLF-JD and PLF-FDM are two different but both reasonable perceptual metrics used for optimization of video compression models. This is the first paper that studies the RDP issue in video compression and the results and conclusions are insightful. Some important conclusions regarding RDP include:\n\n(1) There is a significant penalty in distortion when using PLF-JD in the low-rate regime.\n\n(2) While PLF-JD preserves better temporal consistency across video frames, it suffers from the permanence of error phenomenon in which the mistakes in reconstructions propagate to future frames. \n\nThe universal encoded representations for video compression are studied as well, which are demonstrates to be always held when using PLF-FDM (similar to compressing images sequentially). \n\nBesides, the effectiveness of RDP in video compression models is well verified by theoretical analysis in Gaussian-Markov sources. \n"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"(1)\t(Perhaps not as a weakness) I am wondering whether these conclusions of RDP in video compression can be extended to distortion-perception tradeoff in other video generation / restoration tasks. The joint-frame loss PLF-JD and the per-frame loss PLF-FDM should hold similar conclusions for video generation.\n\n(2)\tThe state-of-the-art neural video compression framework contains a motion encoder/decoder and a residual encoder/decoder, which are effective for high-resolution videos. It seems this paper studies the other neural video compression framework that does not explicitly transmit motion information. How to apply the proposed theory to such motion-residual-separated video compression framework?\n"},"confidence":{"value":"4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work."},"questions":{"value":"See the abovementioned “weaknesses”."},"rating":{"value":"7: Accept: Technically solid paper, with high impact on at least one sub-area, or moderate-to-high impact on more than one areas, with good-to-excellent evaluation, resources, reproducibility, and no unaddressed ethical considerations."},"code_of_conduct":{"value":"Yes"},"limitations":{"value":"Limitations are discussed in the final section of Appendix. This paper is in the field of video compression which does not involve strong negative societal impact."}},"nonreaders":[],"tmdate":1702410996649,"tcdate":1688571626330,"writers":["NeurIPS.cc/2023/Conference","NeurIPS.cc/2023/Conference/Submission5260/Reviewer_pcBB"],"signatures":["NeurIPS.cc/2023/Conference/Submission5260/Reviewer_pcBB"],"forum":"kLIieSS2P3","number":4,"license":"CC BY 4.0","cdate":1688571626330,"mdate":1702410996649,"readers":["everyone"],"invitations":["NeurIPS.cc/2023/Conference/Submission5260/-/Official_Review","NeurIPS.cc/2023/Conference/-/Edit"],"domain":"NeurIPS.cc/2023/Conference","replyto":"kLIieSS2P3","id":"ih1Q4qSUwy","forumContent":{"venue":{"value":"NeurIPS 2023 poster"},"keywords":{"value":["Video Compression","Information Theory","Neural Compression"]},"supplementary_material":{"value":"/attachment/7d704447f10ec10bb26885f927ad184fd091874d.pdf"},"_bibtex":{"value":"@inproceedings{\nsalehkalaibar2023on,\ntitle={On the choice of Perception Loss Function for Learned Video Compression},\nauthor={Sadaf Salehkalaibar and Truong Buu Phan and Jun Chen and Wei Yu and Ashish J Khisti},\nbooktitle={Thirty-seventh Conference on Neural Information Processing Systems},\nyear={2023},\nurl={https://openreview.net/forum?id=kLIieSS2P3}\n}"},"title":{"value":"On the choice of Perception Loss Function for Learned Video Compression"},"paperhash":{"value":"salehkalaibar|on_the_choice_of_perception_loss_function_for_learned_video_compression"},"TLDR":{"value":"Our work establishes the rate-distortion-perception (RDP) tradeoff and a principle of universality for causal low-delay lossy video compression."},"abstract":{"value":"We study causal, low-latency, sequential video compression when the output is subjected to both a mean squared-error (MSE) distortion loss as well as a perception loss to target realism. Motivated by prior approaches, we consider two different perception loss functions (PLFs). The first, PLF-JD,  considers the joint distribution (JD) of all the video frames up to the current one, while the second metric, PLF-FMD,  considers the framewise marginal distributions (FMD) between the source and reconstruction. Using information theoretic analysis and deep-learning based experiments, we demonstrate that the choice of PLF can have a significant effect on the reconstruction, especially at low-bit rates. In particular, while the reconstruction based on PLF-JD can better preserve the temporal correlation across frames, it also imposes a significant penalty in distortion  compared to PLF-FMD and further makes it more difficult to recover from errors  made in the earlier output frames. Although the choice of PLF decisively affects  reconstruction quality, we also demonstrate that it may not be essential to commit to a particular PLF during encoding and the choice of PLF can be delegated to the decoder. In particular, encoded representations generated by training a system to minimize the MSE (without requiring either PLF) can be  {\\em near universal}  and can generate close to optimal reconstructions for either choice of PLF at the decoder.  We validate our results using (one-shot) information-theoretic analysis, detailed study of the rate-distortion-perception tradeoff of the Gauss-Markov source model as well as deep-learning based experiments on moving MNIST and KTH datasets."},"pdf":{"value":"/pdf/90cb373591ab63151f6adff1a35deda6a5964bfd.pdf"},"venueid":{"value":"NeurIPS.cc/2023/Conference"},"authorids":{"value":["~Sadaf_Salehkalaibar1","~Truong_Buu_Phan1","~Jun_Chen8","~Wei_Yu20","~Ashish_J_Khisti1"]},"authors":{"value":["Sadaf Salehkalaibar","Truong Buu Phan","Jun Chen","Wei Yu","Ashish J Khisti"]}},"version":2},{"content":{"summary":{"value":"The paper presents Neus-E, which bases on Neus-V to identify segments in the video that does not comply with the prompt in terms of handling complex or sequential events. The method first converts the text prompt into propositions using LLMs, and builds a video automaton to find the weakest proposition that does not align with the video. Based on the proposition, the method further identifies the video segment to be re-generated and prompts the video generation model with revised prompt from LLM to replace the original segment with better alignment."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"1. Why would the step-by-step fall so short compared to the proposed method? Considering that both method rely on the point that smaller video segments generated from propositions are expected have better alignment, the best achievable alignment for both methods looks to be the same in theory. Could this imply that the \"strong\" segments, opposed to the edited segments from NeuS-E, to have benefitted from the overall context of the video and should be kept?\n\n2. How often would a re-generated segment be selected again in the next iteration? This relates to the previous question, to rule out the method from simply benefitting from having more trials within the iterations. \n\n3. Considering that the method combines multiple models, involving LLMs and VLMs, would it be also possible to identify misaligned video segments in a more simplistic manner with modern large VLMs? For example, it would be possible to simply input the video and the prompt a LVLM and request the model to identify the misaligned frames. Demonstrating that such understanding is still difficult even with state-of-the-art LVLMs would be interesting, and would further strengthen the motivation of the work for introducing neuro-symbolic feedback."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The paper introduces an interesting idea of using neuro-symbolic feedback for identifying video segments that misalign to the given prompt, which is then is re-generated for better alignment. The overall motivation, and the method for identifying such segments - using temporal logic and video automatons, are very sound and convincing.\n\n2. The method is designed to improve videos in a zero-shot manner, without fine-tuning the video for better alignment. Furthermore, the method is virtually model-agnostic asides from the part for accepting image inputs for stitching video segments."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The actual method for improving the problematic video segments however, relies on iterations of re-prompting the video generation model. While it would be reasonable for the model to better understand simple, short prompts from the atomic propositions, it would still depend on blind luck that the new segment would be better. Some study showing that the re-generated segments are much more aligned to the prompts would better support the method.\n\n2. The evaluation mostly focuses in NeuS-V, where the methodology for identifying misaligned video segments, and the evaluation metric, shares a lot in common. If the method is basically identifying regions that bottlenecks the NeuS-V score, re-generating such segments would trivially result in better NeuS-V scores. A more thorough examination, such as evaluating with T2VCompBench[1], would be helpful.\n\n3. The paper is lacking qualitative results, and no further results could be found in the appendix or the supplementary materials. This adds up with the concerns in the quantitative results, making it unclear how the much the proposed method can improve the original videos.\n\n[1]Sun, Kaiyue, et al. \"T2v-compbench: A comprehensive benchmark for compositional text-to-video generation.\" Proceedings of the Computer Vision and Pattern Recognition Conference. 2025."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762920051522,"tcdate":1761981921788,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission8061/Reviewer_f7yB"],"signatures":["ICLR.cc/2026/Conference/Submission8061/Reviewer_f7yB"],"forum":"ifJ91JSLhq","number":2,"license":"CC BY 4.0","cdate":1761981921788,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission8061/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762920051522,"domain":"ICLR.cc/2026/Conference","replyto":"ifJ91JSLhq","id":"bFHEeqM2ip","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"TLDR":{"value":"We propose a training free method to improve the temporal fidelity text to video generation."},"keywords":{"value":["Neuro-symbolic AI","Text-to-Video Synthesis and Generation","Explainable Computer Vision"]},"supplementary_material":{"value":"/attachment/b22f935dcfed4ca204abdfb8b0254521cd2af223.zip"},"primary_area":{"value":"neurosymbolic & hybrid AI systems (physics-informed, logic & formal reasoning, etc.)"},"abstract":{"value":"Current text-to-video (T2V) generation models are increasingly popular due to their ability to produce coherent videos from textual prompts. However, these models often struggle to generate semantically and temporally consistent videos when dealing with longer, more complex prompts involving multiple objects or sequential events. Additionally, the high computational costs associated with training or fine-tuning make direct improvements impractical. To overcome these limitations, we introduce NeuS-E, a novel zero-training video refinement pipeline that leverages neuro-symbolic feedback to automatically enhance video generation, achieving superior alignment with the prompts. Our approach first derives the neuro-symbolic feedback by analyzing a formal video representation and pinpoints semantically inconsistent events, objects, and their corresponding frames. This feedback then guides targeted edits to the original video. Extensive empirical evaluations on both open-source and proprietary T2V models demonstrate that NeuS-E significantly enhances temporal and logical alignment across diverse prompts by almost 40%."},"_bibtex":{"value":"@misc{\nchoi2026well,\ntitle={We'll Fix it in Post: Improving Text-to-Video Generation with Zero Training},\nauthor={Minkyu Choi and S P Sharan and Harsh Goel and Sahil Shah and Sandeep P. Chinchali},\nyear={2026},\nurl={https://openreview.net/forum?id=ifJ91JSLhq}\n}"},"title":{"value":"We'll Fix it in Post: Improving Text-to-Video Generation with Zero Training"},"pdf":{"value":"/pdf/70da27db733ed9eb78e57fb621cb21d3cad59875.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"choi|well_fix_it_in_post_improving_texttovideo_generation_with_zero_training"},"authorids":{"value":["~Minkyu_Choi2","~S_P_Sharan1","~Harsh_Goel1","~Sahil_Shah3","~Sandeep_P._Chinchali1"]},"authors":{"value":["Minkyu Choi","S P Sharan","Harsh Goel","Sahil Shah","Sandeep P. Chinchali"]}},"version":2},{"content":{"summary":{"value":"This paper studied different video encoders for imitation learning in modern video games. The motivation is that existing pre-trained models are usually trained on real-world images, while the impact of distributional shift on video-game images remains unknown. The paper conducted a systematic research that compared different pre-trained visual encoders and from-scratch trained visual encoders in three video games. The observations suggest that pre-trained self-supervised models are worth trying in video game agent development."},"presentation":{"value":"3 good"},"contribution":{"value":"2 fair"},"soundness":{"value":"2 fair"},"strengths":{"value":"- The writing is brilliant. The paper is very easy to follow.\n- The study is systematic and leads to some interesting observations. The paper also gives insightful analysis for these observations, which may shed some light on the research of video-game agent development.\n- The motivation is clear, and the identified problem (video-game image distribution is different from pre-training distribution) is meaningful for the community."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Although the paper offered many insights and potential analysis from the emprical observations, the paper lacks enough decisive conclusions. To be specific, I find that the following claims are not convincing:\n  - In the last sentence from section 5.1: \"while ViTs do not guarantee improvement over ResNets, they can provide significant improvement.\" This conclusion is drawn from the observation that ResNets are comparable with ViTs in Minecraft Dungeons, while are outperformed by ViTs for a large margin in Minecraft. However, the observation in Table 4 (the experiments in CS:GO) shows that ResNet outperforms ViT significantly. Therefore, it still remains unknown which of these two types of networks should be chosen as visual encoder.\n  - In section 5.3, \"This finding suggests that, if high-quality data is available for the specific task, it might be beneficial to consider training visual encoders end-to-end for BC agents, even in situations with less available data.\" As the finding shows the end-to-end encoders is comparable to pre-trained encoders, why do you say that end-to-end is beneficial? Moreover, as the pre-trained backbone is fixed during imitation learning in the paper, the trainable parameters for pre-trained settings are significantly less than end-to-end setting. I am curious if the performance will be better or worse if we don't fix the pre-trained visual encoders.\n  - In the last paragraph of section 5.5, it states that the pre-trained visual encoders fail to generalize when the input image size shifts. There are three related questions:\n    - The resize operation seems unreasonable. As the pre-trained encoder is fixed during BC training, the feature extracted from the image is fixed, which is distorted during resizing. What about padding the image to 280x280 and then resize it to 224x224?\n    - How about unfreezing the visual encoders during training? It may address the last point I raised as the visual encoder can adapt to new input size during fine-tuning.\n    - The conclusion is drawn from only one experiment, how about other cases that the pre-trained model fails when input image sizes are different?\n\n- The conclusions are drawn without controlling some critical variables. For example, the effect of network size is overlooked in the paper.\n- The experiments are constrained in a limited number of video game tasks. For example, there are thousands of tasks in Minecraft as shown by [1], and this paper only tested on the \"Treechop\" task. Also, the paper only studied the task-specific imitation learning, while large-scale pre-training adopted by VPT [2] or multi-task imitation learning [3] are not examined.\n- (minor) A key motivation of this paper is that, the images are often related to real-world scenes, which differs from video games. But there seems not enough suppotive evidence in the paper, and the readers usually don't know whether video game images are used during pre-training. Maybe a summary table that presents the pre-training sources of different models will be clear.\n\n \n\n[1] Minedojo: Building open-ended embodied agents with internet-scale knowledge. In NeurIPS, 2022.\n\n[2] Video pretraining (vpt): Learning to act by watching unlabeled online videos. In NeurIPS, 2022.\n\n[3] Open-World Multi-Task Control Through Goal-Aware Representation Learning and Adaptive Horizon Prediction. In CVPR 2022."},"confidence":{"value":"5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully."},"questions":{"value":"- How do you select the hyper-parameters for different experiments? \n- What are the original image sizes for Minecraft Dungeons and Minecraft?\n- Why does the paper use FocalNet as classification supervised pre-trained encoders? ImageNet pre-trained models such as ViT / DeiT are also popular, which share the same architectures as in the other categories (language contrastive / self-supervised pre-trained) and thus are easy to be compared."},"rating":{"value":"3: reject, not good enough"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1700463429066,"tcdate":1698405402719,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission5210/Reviewer_bAy4"],"signatures":["ICLR.cc/2024/Conference/Submission5210/Reviewer_bAy4"],"forum":"6CetUU9FSt","number":1,"license":"CC BY 4.0","cdate":1698405402719,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission5210/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1700463429066,"domain":"ICLR.cc/2024/Conference","replyto":"6CetUU9FSt","id":"4YzQG1huOA","forumContent":{"venue":{"value":"Submitted to ICLR 2024"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Imitation Learning","Visual Encoders"]},"primary_area":{"value":"reinforcement learning"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Video games have served as useful benchmarks for the decision making community, but going beyond Atari games towards training agents in modern games has been prohibitively expensive for the vast majority of the research community. Recent progress in the research, development and open release of large vision models has the potential to amortize some of these costs across the community. However, it is currently unclear which of these models have learnt representations that retain information critical for sequential decision making. Towards enabling wider participation in the research of gameplaying agents in modern games, we present a systematic study of imitation learning with publicly available visual encoders compared to the typical, task-specific, end-to-end training approach in Minecraft, Minecraft Dungeons and Counter-Strike: Global Offensive."},"_bibtex":{"value":"@misc{\nsch{\\\"a}fer2024visual,\ntitle={Visual Encoders for Data-Efficient Imitation Learning in Modern Video Games},\nauthor={Lukas Sch{\\\"a}fer and Logan Jones and Anssi Kanervisto and Yuhan Cao and Tabish Rashid and Raluca Georgescu and David Bignell and Siddhartha Sen and Andrea Trevi{\\~n}o Gavito and Sam Devlin},\nyear={2024},\nurl={https://openreview.net/forum?id=6CetUU9FSt}\n}"},"title":{"value":"Visual Encoders for Data-Efficient Imitation Learning in Modern Video Games"},"pdf":{"value":"/pdf/9f452e95a412eaac9a3e9f73be64e59d14bfb9cd.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Rejected_Submission"},"paperhash":{"value":"schäfer|visual_encoders_for_dataefficient_imitation_learning_in_modern_video_games"},"authorids":{"value":["~Lukas_Schäfer1","~Logan_Jones1","~Anssi_Kanervisto1","~Yuhan_Cao1","~Tabish_Rashid1","~Raluca_Georgescu1","~David_Bignell1","~Siddhartha_Sen1","~Andrea_Treviño_Gavito1","~Sam_Devlin2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Lukas Schäfer","Logan Jones","Anssi Kanervisto","Yuhan Cao","Tabish Rashid","Raluca Georgescu","David Bignell","Siddhartha Sen","Andrea Treviño Gavito","Sam Devlin"]}},"version":2},{"content":{"venue":{"value":"CoRR 2026"},"pdf":{"value":"https://arxiv.org/pdf/2603.23868v2"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"geng|mleuvad_minimal_latent_entropy_autoencoder_for_fully_unsupervised_video_anomaly_detection"},"html":{"value":"https://doi.org/10.48550/arXiv.2603.23868"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2603-23868,\n  publtype={informal},\n  author={Yuang Geng and Junkai Zhou and Kang Yang and Pan He and Zhuoyang Zhou and José C. Príncipe and Joel B. Harley and Ivan Ruchkin},\n  title={MLE-UVAD: Minimal Latent Entropy Autoencoder for Fully Unsupervised Video Anomaly Detection},\n  year={2026},\n  month={March},\n  cdate={1772323200000},\n  journal={CoRR},\n  volume={abs/2603.23868},\n  url={https://doi.org/10.48550/arXiv.2603.23868}\n}\n"},"abstract":{"value":"In this paper, we address the challenging problem of single-scene, fully unsupervised video anomaly detection (VAD), where raw videos containing both normal and abnormal events are used directly for training and testing without any labels. This differs sharply from prior work that either requires extensive labeling (fully or weakly supervised) or depends on normal-only videos (one-class classification), which are vulnerable to distribution shifts and contamination. We propose an entropy-guided autoencoder that detects anomalies through reconstruction error by reconstructing normal frames well while making anomalies reconstruct poorly. The key idea is to combine the standard reconstruction loss with a novel Minimal Latent Entropy (MLE) loss in the autoencoder. Reconstruction loss alone maps normal and abnormal inputs to distinct latent clusters due to their inherent differences, but also risks reconstructing anomalies too well to detect. Therefore, MLE loss addresses this by minimizing the entropy of latent embeddings, encouraging them to concentrate around high-density regions. Since normal frames dominate the raw video, sparse anomalous embeddings are pulled into the normal cluster, so the decoder emphasizes normal patterns and produces poor reconstructions for anomalies. This dual-loss design produces a clear reconstruction gap that enables effective anomaly detection. Extensive experiments on two widely used benchmarks and a challenging self-collected driving dataset demonstrate that our method achieves robust and superior performance over baselines."},"title":{"value":"MLE-UVAD: Minimal Latent Entropy Autoencoder for Fully Unsupervised Video Anomaly Detection"},"authors":{"value":[{"fullname":"Yuang Geng","username":""},{"fullname":"Junkai Zhou","username":""},{"fullname":"Kang Yang","username":""},{"fullname":"Pan He","username":"~Pan_He1"},{"fullname":"Zhuoyang Zhou","username":""},{"fullname":"José C. Príncipe","username":""},{"fullname":"Joel B. Harley","username":""},{"fullname":"Ivan Ruchkin","username":""}]}},"tmdate":1780154573105,"pdate":1798675200000,"externalIds":["dblp:journals/corr/abs-2603-23868"],"tcdate":1780154569603,"writers":["~","OpenReview.net/Public_Article/DBLP.org","OpenReview.net/Support"],"signatures":["~Pan_He1"],"forum":"piVj3sLD0N","license":"CC BY-SA 4.0","number":40540,"cdate":1772323200000,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/DBLP.org/-/Record","OpenReview.net/Public_Article/-/Edit"],"mdate":1780154573105,"domain":"OpenReview.net/Public_Article","id":"piVj3sLD0N","version":2},{"content":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Video Anomaly Detection","Fully Unsupervised Learning","Autoencoder","Entropy Regularization","Latent Distribution Collapse"]},"supplementary_material":{"value":"/attachment/9c5c3011a3698414a990dc826217df11919b7f2c.zip"},"primary_area":{"value":"unsupervised, self-supervised, semi-supervised, and supervised representation learning"},"abstract":{"value":"In this paper, we address the challenging problem of single-scene, fully unsupervised video anomaly detection (VAD), where raw videos containing both normal and abnormal events are used directly for training and testing without any labels. This differs sharply from prior work that either requires extensive labeling (fully or weakly supervised) or depends on normal-only videos (one-class classification), which are vulnerable to distribution shifts and contamination. We propose an entropy-guided autoencoder that detects anomalies through reconstruction error by reconstructing normal frames well while making anomalies reconstruct poorly. The key idea is to combine the standard reconstruction loss with a novel Minimal Latent Entropy (MLE) loss in the autoencoder. Reconstruction loss alone maps normal and abnormal inputs to distinct latent clusters due to their inherent differences, but also risks reconstructing anomalies too well to detect. Therefore, MLE addresses this by minimizing the entropy of latent embeddings, encouraging them to concentrate around high-density regions. Since normal frames dominate the raw video, sparse anomalous embeddings are pulled into the normal cluster, so the decoder emphasizes normal patterns and produces poor reconstructions for anomalies. This dual-loss design produces a clear reconstruction gap that enables effective anomaly detection. Extensive experiments on two widely used benchmarks and a challenging self-collected driving dataset demonstrate that our method achieves robust and superior performance over baselines."},"_bibtex":{"value":"@misc{\nzhou2026mleuvad,\ntitle={{MLE}-{UVAD}: Minimal Latent Entropy Autoencoder for Fully Unsupervised Video Anomaly Detection},\nauthor={Junkai Zhou and Yuang Geng and Kang Yang and Pan He and Zhuoyang Zhou and Jose C Principe and Joel Harley and Ivan Ruchkin},\nyear={2026},\nurl={https://openreview.net/forum?id=Hf0h5mTd4K}\n}"},"title":{"value":"MLE-UVAD: Minimal Latent Entropy Autoencoder for Fully Unsupervised Video Anomaly Detection"},"pdf":{"value":"/pdf/23ffcbabc18492d25a08d5223bfc378d0a66c01e.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"zhou|mleuvad_minimal_latent_entropy_autoencoder_for_fully_unsupervised_video_anomaly_detection"},"authorids":{"value":["~Junkai_Zhou5","~Yuang_Geng1","~Kang_Yang15","~Pan_He1","~Zhuoyang_Zhou3","~Jose_C_Principe1","~Joel_Harley1","~Ivan_Ruchkin1"]},"authors":{"value":["Junkai Zhou","Yuang Geng","Kang Yang","Pan He","Zhuoyang Zhou","Jose C Principe","Joel Harley","Ivan Ruchkin"]}},"tmdate":1770805004464,"tcdate":1758222707254,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission13792/Authors"],"signatures":["ICLR.cc/2026/Conference/Submission13792/Authors"],"forum":"Hf0h5mTd4K","license":"CC BY 4.0","number":13792,"cdate":1758222707254,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/-/Submission","ICLR.cc/2026/Conference/-/Post_Submission","ICLR.cc/2026/Conference/Submission13792/-/Full_Submission","ICLR.cc/2026/Conference/Submission13792/-/Rebuttal_Revision","ICLR.cc/2026/Conference/-/Edit"],"mdate":1770805004464,"odate":1759896705795,"domain":"ICLR.cc/2026/Conference","id":"Hf0h5mTd4K","version":2},{"content":{"venue":{"value":"IEEE Trans. Consumer Electron. 2004"},"pdf":{"value":"https://ieeexplore.ieee.org/iel5/30/29562/01341692.pdf"},"venueid":{"value":"dblp.org/journals/TCE/2004"},"paperhash":{"value":"he|networkaware_multicasting_for_videoondemand_services"},"authorids":{"value":["","","","~Guangtao_Xue1"]},"html":{"value":"https://doi.org/10.1109/TCE.2004.1341692"},"_bibtex":{"value":"@article{DBLP:journals/tce/HeTYX04,\n  author={Xiaojian He and Xinhuai Tang and Jinyuan You and Guangtao Xue},\n  title={Network-aware multicasting for video-on-demand services},\n  year={2004},\n  cdate={1072915200000},\n  journal={IEEE Trans. Consumer Electron.},\n  volume={50},\n  number={3},\n  pages={864-869},\n  url={https://doi.org/10.1109/TCE.2004.1341692}\n}\n"},"abstract":{"value":"In this paper, we propose a network-aware multicast scheme to supply instantaneous VOD (video-on-demand) services. Being different from traditional VOD systems, this system is implemented based on peer-to-peer grid. By taking advantage of the large storage and the powerful processing capability in client-side devices, the user host serves as both a client and a non-dedicated video server. The system capability is enhanced due to the contribution of user hosts. Peer groups are formed as a collection of peers that share similar interest, and workload may be fairly apportioned among autonomous groups. In this VOD system, the adaptive video delivery mainly employs a new dynamic buffering algorithm and an improved video multicast strategy to achieve an optimal utilization of system resources, and the network-aware adaptation makes the system more robust. In cooperating with each other, the relevant servers can supply instantaneous video services for local users."},"title":{"value":"Network-aware multicasting for video-on-demand services"},"authors":{"value":["Xiaojian He","Xinhuai Tang","Jinyuan You","Guangtao Xue"]}},"tmdate":1772602075061,"pdate":1104451200000,"externalIds":["dblp:journals/tce/HeTYX04"],"tcdate":1772602051378,"writers":["~"],"signatures":["~Guangtao_Xue1"],"forum":"lDTiu7iten","license":"CC BY-SA 4.0","number":847397,"cdate":1072915200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1772602075061,"domain":"DBLP.org","id":"lDTiu7iten","version":2},{"content":{"summary":{"value":"This paper focuses on multimodal prompt optimization for Multimodal Large Language Models (MLLMs), addressing two critical gaps in existing text-only Automatic Prompt Optimization (APO): underutilization of MLLMs’ multimodal capabilities, and inherent challenges (cross-modal inconsistency, sparse high-quality candidates) in multimodal prompt spaces.\n\nKey contributions:  \n1. Formalizes the multimodal prompt optimization problem, defining optimal prompts as text-nontext pairs \\((t,m)\\) that maximize MLLM performance on target tasks.  \n2. Proposes the **MPO framework**:  \n   - *Alignment-Preserving Exploration* (via cohesive backpropagation, joint multimodal update, and 3 complementary operators) ensures text-image semantic consistency;  \n   - *Prior-Inherited Bayesian UCB* leverages parent-child prompt performance correlation to solve cold-start, reducing evaluation budget by 42%-70%.  \n3. Validates on 10 datasets across 3 modalities (image/video/molecule), outperforming text-only APO (e.g., +8.6% accuracy on CUB-200-2011).  \n\nMPO fills the multimodal APO gap for MLLMs, with strong cross-model/multimodal generalization and practical efficiency."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"Same as the section of **Weaknesses**."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"- **Originality**: Breaks the text-only limitation of existing APO methods, first formalizing the multimodal prompt optimization problem (defining prompts as text-nontext pairs \\((t,m)\\)). It also creatively improves Bayesian UCB by leveraging parent-child prompt performance correlation to solve cold-start, extending bandit-based selection to multimodal scenarios innovatively.  \n- **Quality**: Conducts rigorous validation—covering 10 datasets across 3 modalities (image/video/molecule), cross-model tests (Qwen2.5-VL, Gemma3), and ablation studies (verifying alignment mechanisms/operators’ necessity). Results are reliable and generalize well.  \n- **Clarity**: Clearly presents problem formulation, MPO’s two core components (with formulas and flowcharts), and experimental design. The logical structure is straightforward, enabling easy understanding of the framework.  \n- **Significance**: Fills the gap of multimodal prompt optimization for MLLMs, reduces evaluation budget by 42%-70% for practicality, and provides a foundational framework for future multimodal prompt research."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. **Lack of Validation on Mainstream MLLM General Benchmarks, Limiting Evidence of Universal Adaptability**  \nThe current experiments rely solely on custom task-specific datasets (e.g., CUB-200-2011, PlantVillage) and fail to validate MPO on widely recognized MLLM multimodal benchmarks, leaving its ability to enhance MLLMs’ general capabilities unsubstantiated. Specifically:  \n- It omits **static multimodal foundational benchmarks** (e.g., MME, MMBench), which focus on core perceptual capabilities of MLLMs (e.g., image-text matching, attribute recognition)—there is no evidence that MPO can optimize prompt performance for these fundamental tasks.  \n- It excludes **long-video dynamic modality benchmarks** (e.g., VideoMME), which require handling temporal alignment between long-sequence videos and text (e.g., locating specific clips, understanding temporal logic). The paper’s existing alignment mechanism, designed for static images, remains untested for such long-video scenarios, casting doubt on its effectiveness.  \n- It neglects **cross-modal complex reasoning benchmarks** (e.g., ScienceQA, MathVista, Geometry3k)—benchmarks that demand MLLMs integrate multimodal information to solve logical reasoning or mathematical problems. Since the paper only validates classification/prediction tasks, there is no proof that MPO-optimized prompts can improve complex reasoning performance, leaving MPO’s adaptability to general MLLM scenarios unconfirmed.  \n\n2. **Unaddressed Implicit Deployment Costs, Missing Cost Comparison with Text-only APO**  \nWhile the paper emphasizes that the prior-inherited Bayesian UCB reduces evaluation budget by 42%–70%, it overlooks the significant computational overhead of generating multimodal candidates. Modal-specific generators (e.g., GPT-Image, video editing models) used for image/long-video clip generation incur much higher costs than text prompt generation: for instance, single-image generation takes 2–5 seconds (vs. 0.1 seconds for text generation), and diffusion-based image generators require 3–5 times more GPU memory than text models. Critically, the paper fails to compare the **total deployment cost of MPO** (including iterative multimodal generation overhead) with that of text-only APO. If the implicit costs of multimodal generation offset or even exceed the saved evaluation budget, MPO’s practical utility in real-world deployment would be severely undermined—this key cost trade-off is entirely unaddressed."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762942643488,"tcdate":1761912698288,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission23396/Reviewer_oPXv"],"signatures":["ICLR.cc/2026/Conference/Submission23396/Reviewer_oPXv"],"forum":"M5MfDi4gJO","number":4,"license":"CC BY 4.0","cdate":1761912698288,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission23396/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762942643488,"domain":"ICLR.cc/2026/Conference","replyto":"M5MfDi4gJO","id":"by3ZgcFE5g","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"TLDR":{"value":"We introduce the problem of multimodal prompt optimization and propose the multimodal prompt optimizer, to harness the full capacity of multimodal large language models beyond text."},"keywords":{"value":["Multimodal Large Language Models","Multimodal Prompt Optimization","Prompt Optimization"]},"supplementary_material":{"value":"/attachment/66ab2ee3be539964201f36e0e110e43d8a79b323.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Large Language Models (LLMs) have shown remarkable success, and their multimodal expansions (MLLMs) further unlock capabilities spanning images, videos, and other modalities beyond text. However, despite this shift, prompt optimization approaches, designed to reduce the burden of manual prompt crafting while maximizing performance, remain confined to text, ultimately limiting the full potential of MLLMs. Motivated by this gap, we introduce the new problem of multimodal prompt optimization, which expands the prior definition of prompt optimization to the multimodal space defined by the pairs of textual and non-textual prompts. To tackle this problem, we then propose the Multimodal Prompt Optimizer (MPO), a unified framework that not only performs the joint optimization of multimodal prompts through alignment-preserving updates but also guides the selection process of candidate prompts by leveraging earlier evaluations as priors in a Bayesian-based selection strategy. Through extensive experiments across diverse modalities that go beyond text, such as images, videos, and even molecules, we demonstrate that MPO outperforms leading text-only optimization methods, establishing multimodal prompt optimization as a crucial step to realizing the potential of MLLMs."},"_bibtex":{"value":"@inproceedings{\nchoi2026multimodal,\ntitle={Multimodal Prompt Optimization: Why Not Leverage Multiple Modalities for {MLLM}s},\nauthor={Yumin Choi and Dongki Kim and Jinheon Baek and Sung Ju Hwang},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=M5MfDi4gJO}\n}"},"title":{"value":"Multimodal Prompt Optimization: Why Not Leverage Multiple Modalities for MLLMs"},"pdf":{"value":"/pdf/ce29ceb6726b04a0dd0434ea2d04b437662520db.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"choi|multimodal_prompt_optimization_why_not_leverage_multiple_modalities_for_mllms"},"authorids":{"value":["~Yumin_Choi1","~Dongki_Kim1","~Jinheon_Baek1","~Sung_Ju_Hwang1"]},"authors":{"value":["Yumin Choi","Dongki Kim","Jinheon Baek","Sung Ju Hwang"]}},"version":2},{"content":{"venue":{"value":"ICASSP 2023"},"pdf":{"value":"https://ieeexplore.ieee.org/iel7/10094559/10094560/10096346.pdf"},"venueid":{"value":"dblp.org/conf/ICASSP/2023"},"paperhash":{"value":"ahmadian|enhancing_representation_learning_with_deep_classifiers_in_presence_of_shortcut"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Amirhossein_Ahmadian:","~Fredrik_Lindsten1"]},"html":{"value":"https://doi.org/10.1109/ICASSP49357.2023.10096346"},"_bibtex":{"value":"@inproceedings{DBLP:conf/icassp/AhmadianL23,\n  author={Amirhossein Ahmadian and Fredrik Lindsten},\n  title={Enhancing Representation Learning with Deep Classifiers in Presence of Shortcut},\n  year={2023},\n  cdate={1672531200000},\n  pages={1-5},\n  url={https://doi.org/10.1109/ICASSP49357.2023.10096346},\n  booktitle={ICASSP},\n  crossref={conf/icassp/2023}\n}\n"},"abstract":{"value":"A deep neural classifier trained on an upstream task can be leveraged to boost the performance of another classifier in a related downstream task through the representations learned in hidden layers. However, presence of shortcuts (easy-to-learn features) in the upstream task can considerably impair the versatility of intermediate representations and, in turn, the downstream performance. In this paper, we propose a method to improve the representations learned by deep neural image classifiers in spite of a shortcut in upstream data. In our method, the upstream classification objective is augmented with a type of adversarial training where an auxiliary network, so called lens, fools the classifier by exploiting the shortcut in reconstructing images. Empirical comparisons in self-supervised and transfer learning problems with three shortcut-biased datasets suggest the advantages of our method in terms of downstream performance and/or training time."},"title":{"value":"Enhancing Representation Learning with Deep Classifiers in Presence of Shortcut"},"authors":{"value":["Amirhossein Ahmadian","Fredrik Lindsten"]}},"tmdate":1756904856843,"pdate":1672531200000,"externalIds":["dblp:conf/icassp/AhmadianL23"],"tcdate":1756904839806,"writers":["~"],"signatures":["~Fredrik_Lindsten1"],"forum":"UVz5Q8HGyT","license":"CC BY-SA 4.0","number":620074,"cdate":1672531200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1756904856843,"domain":"DBLP.org","id":"UVz5Q8HGyT","version":2},{"content":{"TLDR":{"value":"We use RL to learn shortcuts in the abstract planning graph induced by predefined options."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Robot Planning","Reinforcement Learning"]},"supplementary_material":{"value":"/attachment/963c7ff7b5a8b143b0ebb6b42b4266d8ba9dbeab.zip"},"primary_area":{"value":"applications to robotics, autonomy, planning"},"abstract":{"value":"Long-horizon decision-making with sparse rewards and continuous states and actions remains a fundamental challenge in AI and robotics. Task and motion planning (TAMP) is a model-based framework that addresses this challenge by planning hierarchically with abstract actions (options). These options are manually defined, limiting the agent to behaviors that we as human engineers know how to program (pick, place, move). In this work, we propose Shortcut Learning for Abstract Planning (SLAP), a method that leverages existing TAMP options to automatically discover new ones. Our key idea is to use model-free reinforcement learning (RL) to learn *shortcuts* in the abstract planning graph induced by the existing options in TAMP. Without any additional assumptions or inputs, shortcut learning leads to shorter solutions than pure planning, and higher task success rates than flat and hierarchical RL. Qualitatively, SLAP discovers dynamic physical improvisations (e.g., slap, wiggle, wipe) that differ significantly from the manually-defined ones. In experiments in four simulated robotic environments, we show that SLAP solves and generalizes to a wide range of tasks, reducing overall plan lengths by over 50\\% and consistently outperforming planning and RL baselines."},"_bibtex":{"value":"@inproceedings{\nliu2026slap,\ntitle={{SLAP}: Shortcut Learning for Abstract Planning},\nauthor={Y. Isabel Liu and Bowen Li and Benjamin Eysenbach and Tom Silver},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=enprG5H9aD}\n}"},"title":{"value":"SLAP: Shortcut Learning for Abstract Planning"},"pdf":{"value":"/pdf/4f54aa8a88c1a1d21a3c9878134690e1432728d8.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"liu|slap_shortcut_learning_for_abstract_planning"},"authorids":{"value":["~Y._Isabel_Liu1","~Bowen_Li7","~Benjamin_Eysenbach1","~Tom_Silver1"]},"authors":{"value":["Y. Isabel Liu","Bowen Li","Benjamin Eysenbach","Tom Silver"]}},"tmdate":1778643402109,"pdate":1769436587790,"tcdate":1758326359801,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission22118/Authors"],"signatures":["ICLR.cc/2026/Conference/Submission22118/Authors"],"forum":"enprG5H9aD","license":"CC BY 4.0","number":22118,"cdate":1758326359801,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/-/Submission","ICLR.cc/2026/Conference/-/Post_Submission","ICLR.cc/2026/Conference/Submission22118/-/Full_Submission","ICLR.cc/2026/Conference/Submission22118/-/Rebuttal_Revision","ICLR.cc/2026/Conference/-/Edit","ICLR.cc/2026/Conference/Submission22118/-/Camera_Ready_Revision"],"mdate":1778643402109,"odate":1759896705795,"domain":"ICLR.cc/2026/Conference","id":"enprG5H9aD","version":2},{"content":{"venue":{"value":"CoRR 2025"},"pdf":{"value":"https://arxiv.org/pdf/2511.01107v1"},"venueid":{"value":"dblp.org/journals/CORR/2025"},"paperhash":{"value":"liu|slap_shortcut_learning_for_abstract_planning"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Y._Isabel_Liu:","https://dblp.org/search/pid/api?q=author:Bowen_Li:","~Benjamin_Eysenbach1","https://dblp.org/search/pid/api?q=author:Tom_Silver:"]},"html":{"value":"https://doi.org/10.48550/arXiv.2511.01107"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2511-01107,\n  publtype={informal},\n  author={Y. Isabel Liu and Bowen Li and Benjamin Eysenbach and Tom Silver},\n  title={SLAP: Shortcut Learning for Abstract Planning},\n  year={2025},\n  month={November},\n  cdate={1761955200000},\n  journal={CoRR},\n  volume={abs/2511.01107},\n  url={https://doi.org/10.48550/arXiv.2511.01107}\n}\n"},"abstract":{"value":"Long-horizon decision-making with sparse rewards and continuous states and actions remains a fundamental challenge in AI and robotics. Task and motion planning (TAMP) is a model-based framework that addresses this challenge by planning hierarchically with abstract actions (options). These options are manually defined, limiting the agent to behaviors that we as human engineers know how to program (pick, place, move). In this work, we propose Shortcut Learning for Abstract Planning (SLAP), a method that leverages existing TAMP options to automatically discover new ones. Our key idea is to use model-free reinforcement learning (RL) to learn shortcuts in the abstract planning graph induced by the existing options in TAMP. Without any additional assumptions or inputs, shortcut learning leads to shorter solutions than pure planning, and higher task success rates than flat and hierarchical RL. Qualitatively, SLAP discovers dynamic physical improvisations (e.g., slap, wiggle, wipe) that differ significantly from the manually-defined ones. In experiments in four simulated robotic environments, we show that SLAP solves and generalizes to a wide range of tasks, reducing overall plan lengths by over 50% and consistently outperforming planning and RL baselines."},"title":{"value":"SLAP: Shortcut Learning for Abstract Planning"},"authors":{"value":["Y. Isabel Liu","Bowen Li","Benjamin Eysenbach","Tom Silver"]}},"tmdate":1769378469195,"pdate":1767139200000,"externalIds":["dblp:journals/corr/abs-2511-01107"],"tcdate":1769378458460,"writers":["~"],"signatures":["~Benjamin_Eysenbach1"],"forum":"Np9RNCb73m","license":"CC BY-SA 4.0","number":801114,"cdate":1761955200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1769378469195,"domain":"DBLP.org","id":"Np9RNCb73m","version":2},{"content":{"venue":{"value":"PRCV (5) 2024"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-981-97-8620-6_22.pdf"},"venueid":{"value":"dblp.org/conf/PRCV/2024"},"paperhash":{"value":"song|uncertaintyaware_with_negative_samples_for_videotext_retrieval"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Weitao_Song:","~Weiran_Chen1","https://dblp.org/search/pid/api?q=author:Jialiang_Xu:","https://dblp.org/search/pid/api?q=author:Yi_Ji_0001:","https://dblp.org/search/pid/api?q=author:Ying_Li:","~Chunping_Liu1"]},"html":{"value":"https://doi.org/10.1007/978-981-97-8620-6_22"},"_bibtex":{"value":"@inproceedings{DBLP:conf/prcv/SongCXJLL24,\n  author={Weitao Song and Weiran Chen and Jialiang Xu and Yi Ji and Ying Li and Chunping Liu},\n  title={Uncertainty-Aware with Negative Samples for Video-Text Retrieval},\n  year={2024},\n  cdate={1704067200000},\n  pages={318-332},\n  url={https://doi.org/10.1007/978-981-97-8620-6_22},\n  booktitle={PRCV (5)},\n  crossref={conf/prcv/2024-5}\n}\n"},"abstract":{"value":"Video-text retrieval is a crucial and challenging task that involves searching for the most semantically relevant items in a given query. Existing works typically adopt symmetric contrastive learning loss, which encourage positive sample pairs to be pulled together and the negative to be pushed away in a shared embedding space. However, it solely treats one caption manually collected for a specific video as relevant. Indeed, numerous videos are semantically similar to the caption, yet they are disregarded as negatives. Therefore, the binary correlation-based learning strategy causes semantically similar samples to be irrationally separated, which results in confusing learning objectives. In this paper, we first explore the behavior of symmetric contrastive learning loss from the perspective of gradients, and reveal that the model may suffer from uncertain semantic biases, and the accumulation of the biases would arise to semantic collisions. Motivated by the discovery, an effective and efficient Uncertain SEmantic Consistency constraint (USEC) is proposed, which leverages the uncertain latent semantic correlation between all the two cross-matched negative sample pairs as a supervisory signal to minimize the distance between semantically similar negative sample pairs in a self-supervised manner, while ensuring semantic consistency of positive sample pairs. Based on the gradient analysis and visualization of USEC, we theoretically demonstrate the rationality of exploiting neglected negative sample semantics during training. Evaluation on a series of benchmarks shows that USEC improves retrieval performance and exhibits merits with state-of-the-art methods."},"title":{"value":"Uncertainty-Aware with Negative Samples for Video-Text Retrieval"},"authors":{"value":["Weitao Song","Weiran Chen","Jialiang Xu","Yi Ji","Ying Li","Chunping Liu"]}},"tmdate":1753059672519,"pdate":1704067200000,"tcdate":1731489873422,"writers":["~"],"signatures":["~Chunping_Liu1"],"forum":"06gvN8hwS4","license":"CC BY-SA 4.0","number":220236,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1753059672519,"domain":"DBLP.org","id":"06gvN8hwS4","version":2},{"content":{"summary":{"value":"This paper introduces two baselines and a novel approach called the Multimodal Video Understanding (MVU) framework for video understanding tasks. The baselines explore using either a single frame or no visual input at all. In contrast, MVU aggregates multimodal information relevant to the video to enhance understanding."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"As mentioned in the weakness."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"+ It is interesting that this work introduces three key attributes essential for video understanding: Global Object Information, Object Spatial Location, and Object Motion Trajectory. These attributes contribute significantly to a more comprehensive analysis of video content."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Although this paper focuses on long-video understanding, it lacks specific design elements to address long-video scenarios. Challenges such as context length limiting input frames and modeling long-range temporal information are not directly addressed. The attributes introduced—Global Object Information (GOI), Object Spatial Location (OSL), and Object Motion Trajectory (MOT)—are not tailored to tackle these issues in long-video understanding.\n\n- In Table 1, the comparisons may be weaker due to the differing training setups and base models used across experiments. As such, the superior performance of SF-LLM over the state-of-the-art does not necessarily imply that the number of frames is irrelevant for video understanding.\n\n- Since the model is based on LLaVA-v1.5-13B, an important baseline is missing: the use of multiple frames as input for LLaVA-v1.5-13B.\n\n- Certain details are lacking, such as the method for using LLaVA-v1.5-13B for frame sampling, as mentioned in Figure 3.\n\n- Likelihood Selection is widely used in the MCQ benchmark as an additional track. For a fairer comparison, other methods should also incorporate this strategy for comparison results.\n\n- The benchmark in this paper only includes mid-length videos, roughly under three minutes. To more competitively demonstrate MVU's capabilities in long-video understanding, it would be beneficial to evaluate on benchmarks like VideoMME and LongVideoBench."}},"nonreaders":[],"tmdate":1732679012934,"tcdate":1730784488203,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission811/Reviewer_V8Po"],"signatures":["ICLR.cc/2025/Conference/Submission811/Reviewer_V8Po"],"forum":"OxKi02I29I","number":3,"license":"CC BY 4.0","cdate":1730784488203,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission811/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732679012934,"domain":"ICLR.cc/2025/Conference","replyto":"OxKi02I29I","id":"xq2lMZptVm","forumContent":{"TLDR":{"value":"Investigates effects of LLM strengths on Long Video QnA tasks. Introduces Multimodal Video Understanding (MVU) framework that incorporates object-centric data from pre-trained models and sets a new state-of-the-art in long-video tasks."},"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["long-video","visual question answering","interpretability"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Large Language Models (LLMs) have allowed recent LLM-based approaches to achieve excellent performance on long-video understanding benchmarks. We investigate how extensive world knowledge and strong reasoning skills of underlying LLMs influence this strong performance. Surprisingly, we discover that LLM-based approaches can yield surprisingly good accuracy on long-video tasks with limited video information, sometimes even with no video-specific information. Building on this, we explore injecting video-specific information into an LLM-based framework. We utilize off-the-shelf vision tools to extract three object-centric information modalities from videos, and then leverage natural language as a medium for fusing this information. Our resulting Multimodal Video Understanding (MVU) framework demonstrates state-of-the-art performance across multiple video understanding benchmarks. Strong performance also on robotics domain tasks establishes its strong generality. Code: github.com/kahnchana/mvu"},"_bibtex":{"value":"@inproceedings{\nranasinghe2025understanding,\ntitle={Understanding Long Videos with Multimodal Language Models},\nauthor={Kanchana Ranasinghe and Xiang Li and Kumara Kahatapitiya and Michael S Ryoo},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=OxKi02I29I}\n}"},"title":{"value":"Understanding Long Videos with Multimodal Language Models"},"pdf":{"value":"/pdf/d58b32f8c632ee668136f84612189566ebfa6f0d.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"ranasinghe|understanding_long_videos_with_multimodal_language_models"},"authorids":{"value":["~Kanchana_Ranasinghe1","~Xiang_Li27","~Kumara_Kahatapitiya1","~Michael_S_Ryoo1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Kanchana Ranasinghe","Xiang Li","Kumara Kahatapitiya","Michael S Ryoo"]}},"version":2},{"content":{"summary":{"value":"The paper addresses the low sampling efficiency of diffusion/flow models by enabling one- or few-step generation. Building on “shortcut” samplers that learn large step sizes via a self-consistency loss against a many-step base flow-matching model, the authors reinterpret this loss through optimal control, treating few-step sampling as a controlled base process. From this view, they propose a cumulative self-consistency loss (CSL) that penalizes misalignment not only at the current step but also along future steps of the trajectory, encouraging large yet reliable steps that maintain downstream sample quality. They further draw connections to reinforcement learning. Experiments indicate improved one-/few-step generation quality under the same training budget, and the proposed shortcut-CSL variant outperforms baselines on certain tasks."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"My primary concern is the limited practical performance gains demonstrated (see Weaknesses)."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"• The paper is clearly written with precise notation; figures and tables are informative and easy to follow.\n\n• The proposed shortcut-CSL shows better performance than prior shortcut variants on some tasks, supporting the value of the cumulative alignment objective."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Metrics are too narrow. Evaluating primarily with FID is insufficient, and the observed FID gains are minor in some cases. Please broaden the evaluation metrics like F1, CLIP-score, and qualitative comparisons, and, if feasible, a small human preference study. This multifaceted assessment would make the empirical case more solid.\n\n2. The teaser (Fig. 1) visualizations look only slightly improved over baselines. \n\n3. The dataset scope is limited. Results are reported only on CelebA and CIFAR-10. While broader evaluation is costly, for fair comparisons with other methods, you should include additional datasets—e.g., ImageNet (64×64 and/or 256×256), FFHQ, or another modern benchmark—while keeping ablations on a single dataset to manage workload. This would better test generality and strengthen claims about scalability and robustness."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1764358986215,"tcdate":1761566586536,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission19781/Reviewer_sxLT"],"signatures":["ICLR.cc/2026/Conference/Submission19781/Reviewer_sxLT"],"forum":"cZqAk87Lu4","number":1,"license":"CC BY 4.0","cdate":1761566586536,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission19781/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1764358986215,"domain":"ICLR.cc/2026/Conference","replyto":"cZqAk87Lu4","id":"a3nvJCpj85","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["One-Step Diffusion","Optimal Control","Shortcut Diffusion Models"]},"primary_area":{"value":"generative models"},"abstract":{"value":"Although iterative denoising (i.e., diffusion/flow) methods offer strong generative performance, they suffer from low generation efficiency, requiring hundreds of steps of network forward passes to simulate a single sample. Mitigating this requires taking larger step-sizes during simulation, thereby allowing one- or few-step generation. Recently proposed shortcut model learns larger step-sizes by enforcing alignment between its direction and the path defined by a base many-step flow-matching model through a self-consistency loss. However, its generation quality is significantly lower than the base model. In this paper, we formulate few-step generation as a controlled base generative process, and show that self-consistency loss can be understood through the lens of optimal control. This perspective naturally motivates its generalization to the proposed cumulative self-consistency loss that cumulatively penalizes misalignment along the entire trajectory. This encourages larger step-sizes that not only align with the base model at the current time step but also support alignment in the subsequent steps, facilitating high-quality generation. Furthermore, we draw a connection between our approach and reinforcement learning, potentially opening the door to a new set of approaches for few-step generation. Experiments show that we significantly improve one- and few-step generation quality under the same training budget. Implementation is available at: [https://github.com/paribeshregmi/Shortcut-CSL](https://github.com/paribeshregmi/Shortcut-CSL)"},"_bibtex":{"value":"@inproceedings{\nregmi2026shortcut,\ntitle={Shortcut Diffusion Training with Cumulative Consistency Loss: An Optimal Control View},\nauthor={Paribesh Regmi and Sandesh Ghimire and Rui Li},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=cZqAk87Lu4}\n}"},"title":{"value":"Shortcut Diffusion Training with Cumulative Consistency Loss: An Optimal Control View"},"pdf":{"value":"/pdf/1b88b1373e822c88adbcbd6921499915ee5a276b.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"regmi|shortcut_diffusion_training_with_cumulative_consistency_loss_an_optimal_control_view"},"authorids":{"value":["~Paribesh_Regmi1","~Sandesh_Ghimire2","~Rui_Li3"]},"authors":{"value":["Paribesh Regmi","Sandesh Ghimire","Rui Li"]}},"version":2},{"content":{"summary":{"value":"The paper proposes Video-Infinity, a distributed inference pipeline designed for long-form video generation using diffusion models. The framework leverages two main mechanisms: Clip parallelism, which distributes video segments across multiple GPUs to improve processing efficiency, and Dual-scope attention, which balances local and global temporal contexts across devices. Together, these components enable Video-Infinity to generate lengthy, coherent videos with reduced memory overhead. On an 8 × Nvidia 6000 Ada GPU setup, the framework can produce videos up to 2,300 frames in approximately 5 minutes."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"Has “Video-Infinity” been tested on different types of video content, such as fast-moving scenes or varying lighting conditions, which might challenge the coherence of frame transitions? How robust is the model across these diverse scenarios?"},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. This work brings incremental novelty by adapting distributed parallelism specifically for long-form video generation. It introduces a dual-scope attention mechanism to balance local and global temporal interactions, ensuring coherence across extended sequences. The clip parallelism approach further enables efficient processing of video clips across GPUs, effectively handling the unique scalability and memory demands of video data. These adaptations, including optimizations for temporal continuity, showcase Video-Infinity’s tailored application of distributed inference to the distinct challenges of generating coherent long videos.\n\n2. Speed up performance is great. The proposed Clip parallelism and Dual-scope attention mechanisms optimize inter-device communication and memory requirements, leading to faster processing times and scalability for generating extended video sequences. It could reduce the inference time by up to 52%."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Performance. In the Table 2 under 64 frames settings, although the proposed work got the highest overall score, it did not showed dominating better results than other baselines.\n\n2. Results on longer context. This work claims capability to generate longer video clips, while it only shows results for a maximum of 192 frames in Table 2. Since it emphasis the long video generation ability, I would suggest putting more quantitive results on longer video. \n\n3. Results on memory usage comparison. This work lacks of comparison of reduced memory overhead to demonstrate the efficiency of the method."}},"nonreaders":[],"tmdate":1731428584700,"tcdate":1730707883201,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission6150/Reviewer_pci8"],"signatures":["ICLR.cc/2025/Conference/Submission6150/Reviewer_pci8"],"forum":"DxT3e2f1jc","number":4,"license":"CC BY 4.0","cdate":1730707883201,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission6150/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428584700,"domain":"ICLR.cc/2025/Conference","replyto":"DxT3e2f1jc","id":"ZaQU1OFbVk","forumContent":{"TLDR":{"value":"Generating extremely long video by multi-GPU parallelism and Dual-Scope attention"},"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["diffusion model","video generation"]},"supplementary_material":{"value":"/attachment/cd5540265b662f4abc497ea9d29308b63cde51b5.zip"},"primary_area":{"value":"generative models"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Diffusion models have recently achieved remarkable results for video generation. Despite the encouraging performances, the generated videos are typically constrained to a small number of frames, resulting in clips lasting merely a few seconds. The primary challenges in producing longer videos include the substantial memory requirements and the extended processing time required on a single GPU. A straightforward solution would be to split the workload across multiple GPUs, which, however, leads to two issues: (1) ensuring all GPUs communicate effectively to share timing and context information, and (2) modifying existing video diffusion models, which are usually trained on short sequences, to create longer videos without additional training. To tackle these, in this paper we introduce Video-Infinity, a distributed inference pipeline that enables parallel processing across multiple GPUs for long-form video generation. Specifically, we propose two coherent mechanisms: Clip parallelism and Dual-scope attention. Clip parallelism optimizes the gathering and sharing of context information across GPUs which minimizes communication overhead, while Dual-scope attention modulates the temporal self-attention to balance local and global contexts efficiently across the devices. Together, the two mechanisms join forces to distribute the workload and enable the fast generation of long videos. Under an 8 x Nvidia 6000 Ada GPU (48G) setup, our method generates videos up to 2,300 frames in approximately 5 minutes."},"_bibtex":{"value":"@misc{\ntan2024videoinfinity,\ntitle={Video-Infinity: Distributed Long Video Generation},\nauthor={Zhenxiong Tan and Xingyi Yang and Songhua Liu and Xinchao Wang},\nyear={2024},\nurl={https://openreview.net/forum?id=DxT3e2f1jc}\n}"},"title":{"value":"Video-Infinity: Distributed Long Video Generation"},"pdf":{"value":"/pdf/ab872845c013dd410396d58681b1fb4f11fbc56f.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"tan|videoinfinity_distributed_long_video_generation"},"authorids":{"value":["~Zhenxiong_Tan1","~Xingyi_Yang1","~Songhua_Liu2","~Xinchao_Wang1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Zhenxiong Tan","Xingyi Yang","Songhua Liu","Xinchao Wang"]}},"version":2},{"content":{"summary":{"value":"This submission summarizes and organizing algorithmic design space of one-step diffusion models including consistency training, shortcut diffusion models, and mean flows.\nAlthough differences in derivation and wordings of the methods, they share computational framework of training by aligning one-step predictions toward targets that are acquired by two-step computation, which can be viewed as learning \"shortcut\".\nThe manuscript includes theoretical analyses claim that 1) both two-step DTSC and CTSC have bounded errors w.r.t. Lipschitz constants of the velocity field, and 2) inference error of  shortcut models measured Wasserstein-2 distance are bounded using bias and variance defined using targets and losses.\nFollowing the insights,  improved training method Explicit&easier Shortcut Model (ESC) is proposed. It uses techniques from the summarized literatures and new plug-in velocity, which aggregates in-minibatch velocity.\nExperiments using CIFAR and ImageNet shows that ESC can make improvements over MeanFlow slightly (-0.1 -- -0.3 in FID)."},"soundness":{"value":4},"confidence":{"value":2},"questions":{"value":"- In algorithm 1, some lines are too much pythonic to interrupt non-programmer readers. for example, x[:,None,:] and logp_fn = Normal(0, 1).log_prob (function treated as an object) may be replaced if possible."},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":4},"strengths":{"value":"- One-step diffusion models are a hot topic in 2025 after Shortcut diffusion and Mean flow. This submission is a timely and nice follow-up to help understanding.\n- Theoretical analyses are deep and well motivated for understanding shortcut behavior of continuous and discrete models. (although I could not check all of the proofs in the near-30-page appendix.)\n- In the method part, plug-in velocity (Algorithm 1) is novel and shown to be useful in the experiments.\n- The overall paper are well organized."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Practical impact: surely the SOTA results in FID are achieved but the FID gains provided by ESC may be marginal. I think that the differences of FID ranging within 0.1 -- -0.3 hardly impacts on human opinions on the quality of images.\n\n- Experimental/theoretical supports of design choices are relatively weak: Table 2 shows the design choice picked from the design space, but I think these selections looks empirical and I could not grasp how the theoretical results are exploited. Please point if I missed some parts."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762915538082,"tcdate":1761966481677,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission524/Reviewer_rWuT"],"signatures":["ICLR.cc/2026/Conference/Submission524/Reviewer_rWuT"],"forum":"k6q8rRYVQR","number":2,"license":"CC BY 4.0","cdate":1761966481677,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission524/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762915538082,"domain":"ICLR.cc/2026/Conference","replyto":"k6q8rRYVQR","id":"v4WhLYylR5","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Few-step Diffusion","Shortcut Model","Diffusion Model","Flow Matching"]},"primary_area":{"value":"generative models"},"abstract":{"value":"Recent advances in few-step diffusion models have demonstrated their efficiency and effectiveness by shortcutting the probabilistic paths of diffusion models, especially in training one-step diffusion models from scratch (\\emph{a.k.a.} shortcut models). However, their theoretical derivation and practical implementation are often closely coupled, which obscures the design space.\nTo address this, we propose a common design framework for representative shortcut models. This framework provides theoretical justification for their validity and disentangles concrete component-level choices, thereby enabling systematic identification of improvements. With our proposed improvements, the resulting one-step model achieves a new state-of-the-art FID50k of 2.85 on ImageNet-256×256 under the classifier-free guidance setting with one step generation, and further reaches FID50k of 2.53 with 2× training steps. Remarkably, the model requires no pre-training, distillation, or curriculum learning.\nWe believe our work lowers the barrier to component-level innovation in shortcut models and facilitates principled exploration of their design space."},"_bibtex":{"value":"@inproceedings{\nlin2026on,\ntitle={On the Design of One-step Diffusion via Shortcutting Flow Paths},\nauthor={Haitao Lin and Peiyan Hu and Minsi Ren and Zhifeng Gao and Zhi-Ming Ma and Guolin Ke and Tailin Wu and Stan Z. Li},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=k6q8rRYVQR}\n}"},"title":{"value":"On the Design of One-step Diffusion via Shortcutting Flow Paths"},"pdf":{"value":"/pdf/307adb67b39e0428b67251381129f3e15f5bc083.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"lin|on_the_design_of_onestep_diffusion_via_shortcutting_flow_paths"},"authorids":{"value":["~Haitao_Lin2","~Peiyan_Hu1","~Minsi_Ren1","~Zhifeng_Gao1","~Zhi-Ming_Ma1","~Guolin_Ke3","~Tailin_Wu1","~Stan_Z._Li2"]},"authors":{"value":["Haitao Lin","Peiyan Hu","Minsi Ren","Zhifeng Gao","Zhi-Ming Ma","Guolin Ke","Tailin Wu","Stan Z. Li"]}},"version":2},{"content":{"venue":{"value":"CoRR 2024"},"pdf":{"value":"http://arxiv.org/pdf/2405.18148v1"},"venueid":{"value":"dblp.org/journals/CORR/2024"},"paperhash":{"value":"kwon|learning_to_detour_shortcut_mitigating_augmentation_for_weakly_supervised_semantic_segmentation"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Junehyoung_Kwon:","","https://dblp.org/search/pid/api?q=author:Yunsung_Cho:","https://dblp.org/search/pid/api?q=author:YoungBin_Kim:"]},"html":{"value":"https://doi.org/10.48550/arXiv.2405.18148"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2405-18148,\n  publtype={informal},\n  author={Junehyoung Kwon and Eunju Lee and Yunsung Cho and YoungBin Kim},\n  title={Learning to Detour: Shortcut Mitigating Augmentation for Weakly Supervised Semantic Segmentation},\n  year={2024},\n  cdate={1704067200000},\n  journal={CoRR},\n  volume={abs/2405.18148},\n  url={https://doi.org/10.48550/arXiv.2405.18148}\n}\n"},"abstract":{"value":"Weakly supervised semantic segmentation (WSSS) employing weak forms of labels has been actively studied to alleviate the annotation cost of acquiring pixel-level labels. However, classifiers trained on biased datasets tend to exploit shortcut features and make predictions based on spurious correlations between certain backgrounds and objects, leading to a poor generalization performance. In this paper, we propose shortcut mitigating augmentation (SMA) for WSSS, which generates synthetic representations of object-background combinations not seen in the training data to reduce the use of shortcut features. Our approach disentangles the object-relevant and background features. We then shuffle and combine the disentangled representations to create synthetic features of diverse object-background combinations. SMA-trained classifier depends less on contexts and focuses more on the target object when making predictions. In addition, we analyzed the behavior of the classifier on shortcut usage after applying our augmentation using an attribution method-based metric. The proposed method achieved the improved performance of semantic segmentation result on PASCAL VOC 2012 and MS COCO 2014 datasets."},"title":{"value":"Learning to Detour: Shortcut Mitigating Augmentation for Weakly Supervised Semantic Segmentation"},"authors":{"value":["Junehyoung Kwon","Eunju Lee","Yunsung Cho","YoungBin Kim"]}},"tmdate":1767353434092,"pdate":1704067200000,"tcdate":1739446099471,"writers":["~"],"signatures":["~Eunju_Lee1"],"forum":"PNzGqtGC78","license":"CC BY-SA 4.0","number":320996,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1767353434092,"domain":"DBLP.org","id":"PNzGqtGC78","version":2},{"content":{"venue":{"value":"WACV 2024"},"pdf":{"value":"https://ieeexplore.ieee.org/iel7/10483279/10483282/10483779.pdf"},"venueid":{"value":"dblp.org/conf/WACV/2024"},"paperhash":{"value":"kwon|learning_to_detour_shortcut_mitigating_augmentation_for_weakly_supervised_semantic_segmentation"},"authorids":{"value":["~Junehyoung_Kwon1","~Eunju_Lee1","https://dblp.org/search/pid/api?q=author:Yunsung_Cho:","https://dblp.org/search/pid/api?q=author:YoungBin_Kim:"]},"html":{"value":"https://doi.org/10.1109/WACV57701.2024.00087"},"_bibtex":{"value":"@inproceedings{DBLP:conf/wacv/KwonLCK24,\n  author={Junehyoung Kwon and Eunju Lee and Yunsung Cho and YoungBin Kim},\n  title={Learning to Detour: Shortcut Mitigating Augmentation for Weakly Supervised Semantic Segmentation},\n  year={2024},\n  cdate={1704067200000},\n  pages={808-817},\n  url={https://doi.org/10.1109/WACV57701.2024.00087},\n  booktitle={WACV},\n  crossref={conf/wacv/2024}\n}\n"},"abstract":{"value":"Weakly supervised semantic segmentation (WSSS) employing weak forms of labels has been actively studied to alleviate the annotation cost of acquiring pixel-level labels. However, classifiers trained on biased datasets tend to exploit shortcut features and make predictions based on spurious correlations between certain backgrounds and objects, leading to a poor generalization performance. In this paper, we propose shortcut mitigating augmentation (SMA) for WSSS, which generates synthetic representations of object-background combinations not seen in the training data to reduce the use of shortcut features. Our approach disentangles the object-relevant and background features. We then shuffle and combine the disentangled representations to create synthetic features of diverse object-background combinations. SMA-trained classifier depends less on contexts and focuses more on the target object when making predictions. In addition, we analyzed the behavior of the classifier on shortcut usage after applying our augmentation using an attribution method-based metric. The proposed method achieved the improved performance of semantic segmentation result on PASCAL VOC 2012 and MS COCO 2014 datasets."},"title":{"value":"Learning to Detour: Shortcut Mitigating Augmentation for Weakly Supervised Semantic Segmentation"},"authors":{"value":["Junehyoung Kwon","Eunju Lee","Yunsung Cho","YoungBin Kim"]}},"tmdate":1762666232375,"pdate":1704067200000,"tcdate":1739446099537,"writers":["~"],"signatures":["~Eunju_Lee1"],"forum":"NFDccPLfdz","license":"CC BY-SA 4.0","number":320998,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1762666232375,"domain":"DBLP.org","id":"NFDccPLfdz","version":2},{"content":{"summary":{"value":"This paper investigates whether general-purpose vision-language models can effectively learn biomedical knowledge from publicly available educational videos on platforms like YouTube. It introduces OpenBiomedVid, a large-scale, diverse instruction-tuning dataset comprising 1,031 hours of biomedical video content, curated through a multi-step human-in-the-loop pipeline that filters frames, refines captions, and generates question-answer pairs. For evaluation, two new benchmarks are released: Surgery VideoQA and MIMICEchoQA."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"please see the weakness"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"The effort to create these resources is commendable. The field lacks large-scale biomedical video datasets, and OpenBiomedVid, along with the two new benchmarks, represents a significant contribution."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The paper is undermined by several major flaws in its methodology and results, which make the central claim (that the models are \"learning medicine\") unconvincing.\n\n1.\tPotential Data Contamination and LLM Bias: A major concern is the pervasive use of GPT-4o throughout the pipeline (frame annotation, caption refinement, Q/A generation, and evaluation). Although the authors implement safeguards, there is a non-trivial risk of the model learning a \"stylistic bias\" or patterns specific to GPT-4o's output, rather than genuine biomedical reasoning. The evaluation with Gemini mitigates this but does not fully eliminate the concern, as the fine-tuned models' training data was itself shaped by GPT-4o. \n\n2.\tAbsolute Performance is Still Low: Despite impressive relative gains, the absolute performance on the video benchmarks remains low (e.g., 25.1% for the 7B model on SurgeryVideoQA). The paper correctly notes that these tasks are challenging, but the low scores highlight that the problem is far from solved and that the models are not yet reliable. This should be more prominently discussed in the context of the claim that models can \"learn medicine.\" \n\n3.\tThe paper claims the models are \"learning medicine,\" but their performance on the standard MedQA text benchmark significantly decreased after fine-tuning. For example, the Qwen-2-VL-7B-Biomed model’s accuracy dropped from a baseline of 52.6% to 47.4%. This suggests the informal, \"noisy\" knowledge from the videos may be conflicting with, or causing the model to \"forget,\" formal textbook knowledge. This directly contradicts the paper's primary narrative.\n\n4.\tLimited Analysis of \"Why\" and Failure Modes: The paper excellently demonstrates that the method works but provides less insight into why and how it fails. A deeper analysis of the types of questions or video segments where the model struggles (e.g., temporal reasoning in long surgical videos vs. static frame understanding in echocardiograms) would be highly valuable. The qualitative examples are good but are primarily success cases. \n\n5.\tClarity on Training-Test Split and Overlap: While the authors state there is no video ID overlap between OpenBiomedVid and SurgeryVideoQA, both are sourced from YouTube. There is a potential for concept or stylistic overlap. A more detailed discussion on how the \"cleanliness\" and focus of the evaluation set differ from the training data would clarify the generalization claim."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762942036875,"tcdate":1761933019834,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission22059/Reviewer_bp18"],"signatures":["ICLR.cc/2026/Conference/Submission22059/Reviewer_bp18"],"forum":"u4PmZOmtko","number":3,"license":"CC BY 4.0","cdate":1761933019834,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission22059/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762942036875,"domain":"ICLR.cc/2026/Conference","replyto":"u4PmZOmtko","id":"Lts6YS5utc","forumContent":{"TLDR":{"value":"Instruction-tuning Qwen-2-VL on 1,031 hours of pedagogical biomedical videos dramatically boosts video and image understanding and includes new expert-curated benchmarks, with all data and code released."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["vision-language models","biomedicine","datasets","evaluations"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Publicly available biomedical videos, such as those on YouTube, serve as valuable educational resources for medical students. Unlike standard machine learning datasets, these videos are designed for human learners, often mixing medical imagery with narration, explanatory diagrams, and contextual framing. In this work, we investigate whether such pedagogically rich, yet non-standardized and heterogeneous videos can effectively teach general-domain vision-language models biomedical knowledge. To this end, we introduce OpenBiomedVid, a biomedical video instruction tuning dataset comprising 1031 hours of video-caption and Q/A pairs, curated through a multi-step human-in-the-loop pipeline. Diverse biomedical video datasets are rare, and OpenBiomedVid fills an important gap by providing instruction-style supervision grounded in real-world educational content. Surprisingly, despite the informal and heterogeneous nature of these videos, the fine-tuned Qwen-2-VL models exhibit substantial performance improvements across most benchmarks. The 2B model achieves gains of 98.7% on video tasks, 71.2% on image tasks, and 0.2% on text tasks. The 7B model shows improvements of 37.09% on video and 11.2% on image tasks, with a slight degradation of 2.7% on text tasks compared to their respective base models. To address the lack of standardized biomedical video evaluation datasets, we also introduce two new expert curated benchmarks, MIMICEchoQA and SurgeryVideoQA. On these benchmarks, the 2B model achieves gains of 99.1% and 98.1%, while the 7B model shows gains of 22.5% and 52.1%, respectively, demonstrating the models' ability to generalize and perform biomedical video understanding on cleaner and more standardized datasets than those seen during training. These results suggest that educational videos created for human learning offer a surprisingly effective training signal for biomedical VLMs. We release OpenBiomedVid, MIMICEchoQA, SurgeryVideoQA, the fine-tuned models, and the complete codebase to support future research."},"_bibtex":{"value":"@misc{\nthapa2026how,\ntitle={How Well Can General Vision-Language Models Learn Medicine By Watching Public Educational Videos?},\nauthor={Rahul Thapa and Andrew Li and Qingyang Wu and Bryan He and Yuki Sahashi and Christina Binder and Angela Zhang and Ben Athiwaratkun and Shuaiwen Leon Song and David Ouyang and James Zou},\nyear={2026},\nurl={https://openreview.net/forum?id=u4PmZOmtko}\n}"},"title":{"value":"How Well Can General Vision-Language Models Learn Medicine By Watching Public Educational Videos?"},"pdf":{"value":"/pdf/fcf643dd9df9746824583d9bb992dfdc21d7b0d9.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"thapa|how_well_can_general_visionlanguage_models_learn_medicine_by_watching_public_educational_videos"},"authorids":{"value":["~Rahul_Thapa1","~Andrew_Li4","~Qingyang_Wu1","~Bryan_He1","~Yuki_Sahashi1","~Christina_Binder1","~Angela_Zhang1","~Ben_Athiwaratkun1","~Shuaiwen_Leon_Song1","~David_Ouyang1","~James_Zou1"]},"authors":{"value":["Rahul Thapa","Andrew Li","Qingyang Wu","Bryan He","Yuki Sahashi","Christina Binder","Angela Zhang","Ben Athiwaratkun","Shuaiwen Leon Song","David Ouyang","James Zou"]}},"version":2},{"content":{"summary":{"value":"This paper tackles the challenge of mitigating shortcut learning by introducing an ensemble framework named DiffDiv, which leverages Diffusion Probabilistic Models (DPMs). DiffDiv generates synthetic counterfactual data through DPMs to weaken shortcut dependencies, enhancing prediction diversity within the ensemble. Experimental results demonstrate that DiffDiv successfully guides the model to perform as intended."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"- The authors consider many objectives for diversificiation such as div, cross, and kl. Is there any reason for this design? If there is an analysis on the selection of diversification objective, it would be better.\n\n- I wonder if the authors have considered the options of using GANs or optimal transport distance to generate synthetic counterfactuals."},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":1},"contribution":{"value":2},"strengths":{"value":"- This paper presents a straightforward approach addressing shortcut learning by synthesizing counterfactuals, leveraging the strengths of modern DPMs.\n\n- An in-depth analysis of the impact of sample fidelity on the OODness is nice."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- While DPMs are central to the proposed method, additional details—such as the training scheme specific to this framework, the model's architecture, and the integration of synthetic data during training—would enhance readability. The current visualization results alone feel insufficient. At least, please provide clear and concise explanations of the overall process for integrating synthetic data into training, ensuring ease of understanding.\n\n- The presentation of the manuscript could be significantly improved. An overview figure of the proposed framework seems essential. Despite the complex components like DPMs in this framework and strategies such as early stopping sampling, the lack of additional explanations makes it challenging for readers to grasp the framework as a whole. Furthermore, the content in Table 1 is difficult to follow. Why are the fractions in Table 1 selected as they are? Also, the reviewer suggests splitting the main quantitative table to highlight the superiority of the proposed method and clearly presenting the ablation study on different diversification objectives for improved readability.\n\n- The paper presents a straightforward solution, but a lack of novelty remains a key concern. While insights like the influence of sample fidelity on OODness and the early-stopping sampling strategy are good, the manuscript would benefit from theoretical analysis if possible, or novel regularizers for DPMs to generate enhanced counterfactuals against spurious correlationsfrom a generative perspective.\n\n- Lastly, the proposed framework requires manual tuning for each specific domain or dataset, making it impractical for real-world application and reproducibility. The manuscript would benefit from introducing a universal cheat sheet or automated tuning strategies that perform effectively across diverse scenarios."}},"nonreaders":[],"tmdate":1731427375257,"tcdate":1730726619880,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission1114/Reviewer_EWrp"],"signatures":["ICLR.cc/2025/Conference/Submission1114/Reviewer_EWrp"],"forum":"SvydqVoHrp","number":3,"license":"CC BY 4.0","cdate":1730726619880,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission1114/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427375257,"domain":"ICLR.cc/2025/Conference","replyto":"SvydqVoHrp","id":"lnNDdmBYZe","forumContent":{"TLDR":{"value":"Bias Mitigation with Ensambles and Diffusion Counterfactuals"},"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Shortcut Learning","Simplicity Bias","Bias Mitigation","Debias","Ensembles","Diffusion Models","counterfactuals","Diversification"]},"primary_area":{"value":"alignment, fairness, safety, privacy, and societal considerations"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Spurious correlations in the data, where multiple cues are predictive of the target labels, often lead to a phenomenon known as shortcut learning, where a model relies on erroneous, easy-to-learn cues while ignoring reliable ones. In this work, we propose DiffDiv an ensemble diversification framework exploiting Diffusion Probabilistic Models (DPMs) to mitigate this form of bias. We show that at particular training intervals, DPMs can generate images with novel feature combinations, even when trained on samples displaying correlated input features. We leverage this crucial property to generate synthetic counterfactuals to increase model diversity via ensemble disagreement. We show that DPM-guided diversification is sufficient to remove dependence on shortcut cues, without a need for additional supervised signals. We further empirically quantify its efficacy on several diversification objectives, and finally show improved generalization and diversification on par with prior work that relies on auxiliary data collection."},"_bibtex":{"value":"@misc{\nscimeca2025mitigating,\ntitle={Mitigating Shortcut Learning with Diffusion Counterfactuals and Diverse Ensembles},\nauthor={Luca Scimeca and Alexander Rubinstein and Damien Teney and Seong Joon Oh and Armand Mihai Nicolicioiu and Yoshua Bengio},\nyear={2025},\nurl={https://openreview.net/forum?id=SvydqVoHrp}\n}"},"title":{"value":"Mitigating Shortcut Learning with Diffusion Counterfactuals and Diverse Ensembles"},"pdf":{"value":"/pdf/d6e22031bdf1e18f3d74d71011d9ea7f42ca84b5.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"scimeca|mitigating_shortcut_learning_with_diffusion_counterfactuals_and_diverse_ensembles"},"authorids":{"value":["~Luca_Scimeca1","~Alexander_Rubinstein1","~Damien_Teney1","~Seong_Joon_Oh1","~Armand_Mihai_Nicolicioiu1","~Yoshua_Bengio1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Luca Scimeca","Alexander Rubinstein","Damien Teney","Seong Joon Oh","Armand Mihai Nicolicioiu","Yoshua Bengio"]}},"version":2},{"content":{"summary":{"value":"This paper presents M3OOD, a meta-learning framework for zero-shot selection of Out-of-Distribution detectors in multimodal settings. The approach leverages historical performance data and combines multimodal embeddings with handcrafted meta-features to recommend suitable detectors for new datasets without requiring expensive evaluations. Experiments on video and optical flow data show M3OOD achieves competitive performance against multiple baselines with minimal overhead."},"soundness":{"value":2},"confidence":{"value":3},"questions":{"value":"1. If I understand correctly, Table 3 is central to understanding both the problem and your method's training data. Why is it not discussed in the main text?\n2. The performance of the Mahalanobis detector is catastrophic on some datasets (e.g., ~43% on EPIC). Does this indicate instability in the feature extraction or training process for certain detector-dataset pairs? Could such outliers negatively bias the meta-learner?\n3. How does M3OOD's performance degrade when the test dataset is highly dissimilar from all meta-training datasets? Please provide an analysis of its sensitivity to the meta-training distribution.\n4. Beyond concatenation, were more advanced fusion strategies for multimodal embeddings explored? Given the importance of cross-modal interactions, this seems a significant missed opportunity."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. The paper tackles the relevant challenge of automated OOD detector selection, extending it to multimodal contexts. The zero-shot selection paradigm is highly valuable for real-world deployment.\n2. The runtime analysis convincingly demonstrates that M3OOD's overhead is negligible compared to the cost of running the OOD detectors themselves, highlighting its practical utility."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The Role of Table 3 is Unexplained and Misleading: Table 3, which presents the full performance matrix of all detectors on all dataset pairs, is the foundation of the meta-learning process and the primary evidence for the problem's existence. Yet, it is never mentioned or explained in the main text. This creates a critical disconnect for the reader.\n2. The claim of a general \"multimodal\" framework is significantly undermined by the exclusive use of video and optical flow—two highly correlated modalities. This limited scope makes the contribution feel more like an incremental adaptation of established meta-learning model selection to a specific data pair, rather than a broader breakthrough for heterogeneous multimodal data (e.g., vision-language, audio-text).\n3. Critical design choices lack thorough investigation. The justification for using XGBoost over more sophisticated neural meta-predictors is weak, and the simple concatenation for modality fusion ignores potentially more effective cross-modal interaction mechanisms."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762942121782,"tcdate":1761479100046,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission22222/Reviewer_i7wL"],"signatures":["ICLR.cc/2026/Conference/Submission22222/Reviewer_i7wL"],"forum":"OpFf3ZjiFa","number":1,"license":"CC BY 4.0","cdate":1761479100046,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission22222/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762942121782,"domain":"ICLR.cc/2026/Conference","replyto":"OpFf3ZjiFa","id":"FHTlPb36x5","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Out-of-Distribution Detection","Multimodal Learning","Automatied Machine Learning","Model Embeddings"]},"primary_area":{"value":"transfer learning, meta learning, and lifelong learning"},"abstract":{"value":"Out-of-distribution (OOD) robustness is a critical challenge for modern machine learning systems, particularly as they increasingly operate in multimodal settings involving inputs like video, audio, and sensor data. Currently, many OOD detectors have been proposed, each with different designs targeting various distribution shifts. A single OOD detector may not prevail across all the scenarios; therefore, how can we automatically select an ideal OOD detection model for different distribution shifts? Due to the inherent unsupervised nature of the OOD detection task, it is difficult to predict model performance and find a universally best model. Also, systematically comparing models on the new unseen data is costly or even impractical. To address this challenge, we introduce M3OOD, a meta-learning-based framework for OOD detector selection in multimodal settings. Meta learning offers a solution by learning from historical model behaviors, enabling rapid adaptation to new data distribution shifts with minimal supervision. Our approach combines multimodal embeddings with handcrafted meta-features that capture distributional and cross-modal characteristics to represent datasets. By leveraging historical performance across diverse multimodal benchmarks, M3OOD can recommend suitable detectors for a new data distribution shift. Experimental evaluation demonstrates that M3OOD consistently outperforms 10 competitive baselines across 12 test scenarios with minimal computational overhead."},"_bibtex":{"value":"@misc{\nqin2025selecting,\ntitle={Selecting Out-of-Distribution Detector for Multiple Modalities},\nauthor={Yuehan Qin and Li Li and Defu Cao and Tiankai Yang and Yue Zhao},\nyear={2025},\nurl={https://openreview.net/forum?id=OpFf3ZjiFa}\n}"},"title":{"value":"Selecting Out-of-Distribution Detector for Multiple Modalities"},"pdf":{"value":"/pdf/edb4e7143aaa7c692e6c69e794e478cd62c3893a.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"qin|selecting_outofdistribution_detector_for_multiple_modalities"},"authorids":{"value":["~Yuehan_Qin1","~Li_Li18","~Defu_Cao1","~Tiankai_Yang1","~Yue_Zhao13"]},"authors":{"value":["Yuehan Qin","Li Li","Defu Cao","Tiankai Yang","Yue Zhao"]}},"version":2},{"content":{"summary":{"value":"This paper proposes AdaRD-Key, a keyframe sampling framework for query-driven long-form video understanding. The method introduces a joint Relevance–Diversity Max-Volume (RD-MV) objective that simultaneously encourages semantic relevance and visual diversity during keyframe selection. Two auxiliary modules enhance adaptability: Variability–Budget Scaling (VB-Scale) automatically adjusts the trade-off coefficient $\\lambda$ according to the variability of query–frame relevance scores and the video length; a lightweight relevance-aware gate detects weak query–video alignment and switches to a diversity-only objective to preserve coverage. AdaRD-Key is fully training-free, efficient on a single GPU, and applicable to existing VLMs such as LLaVA-Video and Qwen2-VL. Experiments on LongVideoBench, Video-MME, and VCapsBench show consistent performance gains over uniform sampling, AKS, and MaxInfo baselines."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"- Accuracy suffices for baseline comparison, but **additional metrics** reflecting frame diversity, redundancy, or query–frame alignment would better validate the proposed RD-MV objective.\n- Expand the **literature review** to cover recent 2024–2025 works in adaptive keyframe selection, temporal retrieval, and long-video reasoning. Discuss how AdaRD-Key differs from these temporal search and related retrieval-driven approaches.\n- A **sensitivity analysis** of $\\lambda$, $\\tau$, and the diversity budget $\\rho$—along with qualitative visualizations of both successful and failed selections—would strengthen the empirical evidence. How stable are the results across different parameter values?\n- An explicit analysis of **computational complexity** and **runtime scaling** would help justify the added design components of AdaRD-Key relative to its modest accuracy gains.\n- Could query complexity or confidence be used to **dynamically adjust** the number of selected frames rather than fixing $K$?\n- How does the **diversity-only fallback** behave in videos containing **multiple semantically relevant events**? Are there cases where this fallback leads to missing query-relevant evidence?\n- Are there **qualitative examples** where AdaRD-Key **fails** under noisy, ambiguous, or semantically diffuse queries? Such an analysis would clarify the model’s practical limitations.\n\n_If the above concerns are addressed, I would be inclined to raise my score._"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- **Practical design.** The framework is training-free and compatible with various pretrained multimodal LLMs, providing immediate utility for long-video reasoning pipelines. The method achieves near-linear computational cost and real-time inference on a single GPU, which is valuable for large-scale video datasets.\n- **Technical soundness.** The RD-MV objective is mathematically well-defined and optimized efficiently through greedy submodular maximization with Sherman–Morrison updates. \n- **Clear modular analysis.** The ablation study transparently evaluates the contributions of relevance, diversity, VB-Scale, and gating, confirming the incremental benefits of each."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- **Marginal improvement over baselines.** Although AdaRD-Key slightly outperforms prior training-free methods (e.g., AKS, MaxInfo), the gains are small and within variance. The paper lacks comparisons with adaptive temporal search methods (e.g., _LV-Haystack (T*)_, _Videotree_) and an analysis of complexity–efficiency trade-offs to justify the added design overhead.\n- **Heuristic parameterization.** Key parameters—such as the gating threshold $\\tau {=},0.4$ and the $\\lambda$ bounds in VB-Scale—are heuristic with no sensitivity analysis. The VB-Scale mechanism itself lacks theoretical grounding or quantitative evaluation, and reliance on BLIP-2 relevance scores may reduce adaptability.\n- **Lack of qualitative error analysis.** Results are reported only as averaged accuracies without qualitative or failure analyses. Including more diverse datasets would strengthen generalization evidence.\n- **Presentation clarity.** The related-work section reads as a citation list rather than a structured taxonomy, and the paper does not clearly justify why combining relevance with $\\log\\det$-based diversity provides new insights beyond existing methods like MaxInfo or AKS.\n- **Limited evaluation coverage.** The evaluation covers LongVideoBench and VideoMME, but omits some well-known long-video benchmarks such as MLVU, LVBench.\n- **Incomplete coverage of related work.** The discussion omits several recent and directly relevant studies on keyframe selection and long-video reasoning, such as _Logic-in-Frames_, _LV-Haystack (T*)_, _Videotree_ (CVPR 2025) ... Without contextualizing against these advances, the novelty appears incremental."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762916011494,"tcdate":1760481989832,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission2069/Reviewer_xNLc"],"signatures":["ICLR.cc/2026/Conference/Submission2069/Reviewer_xNLc"],"forum":"joo22YJkDN","number":1,"license":"CC BY 4.0","cdate":1760481989832,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission2069/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762916011494,"domain":"ICLR.cc/2026/Conference","replyto":"joo22YJkDN","id":"e2g7x8R5XN","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["video understanding","Long-video understanding"]},"primary_area":{"value":"foundation or frontier models, including LLMs"},"abstract":{"value":"Understanding long-form videos remains a significant challenge for vision-language models (VLMs) due to their extensive temporal length and high information density. Most current multimodal large language models (MLLMs) rely on uniform sampling, which often overlooks critical moments, leading to incorrect responses to queries. In parallel, many keyframe selection approaches impose rigid temporal spacing: once a frame is chosen, an exclusion window suppresses adjacent timestamps to reduce redundancy. While effective at limiting overlap, this strategy frequently misses short, fine-grained cues occurring near important events. Other methods instead emphasize visual diversity, but neglect to consider query relevance. We propose \\textbf{AdaRD-Key}, a training-free keyframe sampling module for query-driven long-form video understanding. AdaRD-Key maximizes a unified Relevance–Diversity Max-Volume (RD-MV) objective, which combines a query-conditioned relevance score with a log-determinant diversity component to yield informative yet non-redundant frames. To handle broad queries with weak alignment to the video, AdaRD-Key employs a lightweight relevance-aware gating mechanism. When the relevance distribution indicates weak alignment with the video, the method seamlessly shifts into a diversity-only mode, thereby enhancing coverage without requiring additional supervision. Our entire pipeline is training-free, computationally efficient (running in real time on a single GPU), and compatible with existing VLMs in a plug-and-play manner. Extensive experiments on LongVideoBench and Video-MME further demonstrate that AdaRD-Key achieves state-of-the-art \nperformance, particularly on long-form videos."},"_bibtex":{"value":"@misc{\nzhang2025adardkey,\ntitle={Ada{RD}-Key: Adaptive Relevance-Diversity Keyframe Sampling for Long-form Video Understanding},\nauthor={Xian Zhang and Zexi Wu and Zinuo Li and Hongming Xu and Luqi Gong and Farid Boussaid and Naoufel Werghi and Mohammed Bennamoun},\nyear={2025},\nurl={https://openreview.net/forum?id=joo22YJkDN}\n}"},"title":{"value":"AdaRD-Key: Adaptive Relevance-Diversity Keyframe Sampling for Long-form Video Understanding"},"pdf":{"value":"/pdf/048dfb2b8ac22b8f8e03815621c921d7f7e53b37.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"zhang|adardkey_adaptive_relevancediversity_keyframe_sampling_for_longform_video_understanding"},"authorids":{"value":["~Xian_Zhang5","~Zexi_Wu1","~Zinuo_Li1","~Hongming_Xu3","~Luqi_Gong2","~Farid_Boussaid1","~Naoufel_Werghi1","~Mohammed_Bennamoun1"]},"authors":{"value":["Xian Zhang","Zexi Wu","Zinuo Li","Hongming Xu","Luqi Gong","Farid Boussaid","Naoufel Werghi","Mohammed Bennamoun"]}},"version":2},{"content":{"research_area_keywords":{"value":"robustness, probing, hardness of samples, data shortcuts/artifacts"},"keywords":{"value":["grammatical acceptability","minimal pairs","language model evaluation","lexical frequency","long-tail generalization","syntax","morphology","argument structure"]},"B_use_or_create_scientific_artifacts":{"value":"Yes"},"languages_studied":{"value":"English"},"D4_ethics_review_board_approval":{"value":"N/A"},"C4_parameters_for_packages":{"value":"Yes"},"B6_elaboration":{"value":"Section 3.3 and Appendix C report the number of examples, frequency statistics, lexical diversity, sentence length, tokenizer length, and controlled content-word statistics."},"B1_cite_creators_of_artifacts":{"value":"Yes"},"B2_discuss_the_license_for_artifacts":{"value":"Yes"},"B4_elaboration":{"value":"Ethical considerations."},"B5_elaboration":{"value":"Sections 3.1–3.3 and Appendices A, C document the construction, lexical overlay, frequency regimes, covered BLiMP paradigms, and dataset statistics."},"D_human_subjects_including_annotators":{"value":"No"},"A2_elaboration":{"value":"Limitations and Ethical Considerations sections."},"C3_elaboration":{"value":"Sections 5.1–5.5 and Appendices D–G report mean accuracies, margins, head–extreme-tail changes, confidence intervals over paradigm means, and full model/paradigm breakdowns."},"B2_elaboration":{"value":"Appendix B"},"B3_elaboration":{"value":"Appendix B"},"C2_elaboration":{"value":"Section 4.2 and Appendix D describe the evaluation setup, scoring methods, prompts/templates, inference settings, and statistical aggregation. No model training, decoding, or hyperparameter search was performed."},"C4_elaboration":{"value":"Appendix D reports the main preprocessing and evaluation packages, versions, and implementation details."},"E_ai_assistants_in_research_or_writing":{"value":"Yes"},"C1_elaboration":{"value":"Section 4.1 and Appendix D.3 report model scales, infrastructure, and approximate H100 GPU-hour budget. The experiments are inference-only and do not train models."},"B1_elaboration":{"value":"Sections 3–5 and Appendix A cite the creators of the main artifacts used, including BLiMP, Open English WordNet, Wikidata, VerbNet, wordfreq, spaCy, COCA, and the evaluated open-weight models."},"C2_experimental_setup_and_hyperparameters":{"value":"Yes"},"E1_elaboration":{"value":"AI assistants (GPT5.5, Codex) were used for writing/editing suggestions, LaTeX polishing, and coding/debugging assistance. All paper content, experimental decisions, results, and claims were reviewed and verified by the authors."},"C1_model_size_and_budget":{"value":"Yes"},"venue":{"value":"ACL ARR 2026 May Submission"},"D3_data_consent":{"value":"N/A"},"_bibtex":{"value":"@inproceedings{\nanonymous2026freqblimp,\ntitle={Freq{BL}i{MP}: Frequency-Controlled Minimal Pairs Reveal Robustness and Fragility of {LLM}s Under Lexical Rarity},\nauthor={Anonymous},\nbooktitle={Submitted to ACL Rolling Review - May 2026},\nyear={2026},\nurl={https://openreview.net/forum?id=cmd4o2Wxzz},\nnote={under review}\n}"},"title":{"value":"FreqBLiMP: Frequency-Controlled Minimal Pairs Reveal Robustness and Fragility of LLMs Under Lexical Rarity"},"C3_descriptive_statistics":{"value":"Yes"},"contribution_types":{"value":["Model analysis & interpretability","Publicly available software and/or pre-trained models","Data resources","Data analysis"]},"A1_limitations_section":{"value":"This paper has a limitations section."},"B6_statistics_for_data":{"value":"Yes"},"B4_data_contains_personally_identifying_info_or_offensive_content":{"value":"Yes"},"B3_artifact_use_consistent_with_intended_use":{"value":"Yes"},"A2_potential_risks":{"value":"Yes"},"E1_information_about_use_of_ai_assistants":{"value":"Yes"},"abstract":{"value":"Minimal-pair benchmarks such as BLiMP evaluate linguistic knowledge by testing whether language models (LMs) prefer acceptable sentences over minimally different unacceptable ones. However, these benchmarks largely ignore lexical frequency variation, despite lexical frequency being a pervasive and highly skewed property of natural language use. Consequently, existing evaluations do not test whether grammatical preferences remain stable when contrasts involve rare lexical items. We introduce FreqBLiMP, a frequency-controlled extension of BLiMP that regenerates all $67$ paradigms under explicit Zipf-frequency regimes while preserving each minimal-pair's grammatical contrast. Evaluating multiple open-weight LLM families across scales, we find that decreasing lexical frequency produces a consistent, monotonic decrease in sentence likelihood, but only a modest reduction in overall contrastive acceptability accuracy.  However, this aggregate stability masks substantial variation across linguistic phenomena, with LLMs remaining robust on overt morphosyntactic generalization while degrading on phenomena that require lemma-specific licensing."},"D1_instructions_given_to_participants":{"value":"N/A"},"paper_type":{"value":"Long"},"C_computational_experiments":{"value":"Yes"},"pdf":{"value":"/pdf/f064b80c8856332847bf92a4e4dca4b143b2f781.pdf"},"research_area":{"value":"Interpretability and Analysis of Models for NLP"},"B5_documentation_of_artifacts":{"value":"Yes"},"EMNLP_2026_AI_Reviewing_Experiment":{"value":"yes"},"D2_recruitment_and_payment":{"value":"N/A"},"venueid":{"value":"aclweb.org/ACL/ARR/2026/May/Submission"}},"tmdate":1790697024749,"tcdate":1779788062637,"writers":["aclweb.org/ACL/ARR/2026/May","aclweb.org/ACL/ARR/2026/May/Submission14081/Authors"],"signatures":["aclweb.org/ACL/ARR/2026/May/Submission14081/Authors"],"forum":"cmd4o2Wxzz","license":"CC BY 4.0","number":14081,"cdate":1779788062637,"readers":["everyone"],"invitations":["aclweb.org/ACL/ARR/2026/May/-/Submission","aclweb.org/ACL/ARR/2026/May/-/Edit","aclweb.org/ACL/ARR/2026/May/-/Post_Submission","aclweb.org/ACL/ARR/2026/May/-/Preprint_Post_Submission"],"mdate":1790697024749,"odate":1780385598983,"domain":"aclweb.org/ACL/ARR/2026/May","id":"cmd4o2Wxzz","version":2},{"content":{"summary":{"value":"Captain Cinema proposes a two-stage framework that first performs top-down keyframe planning from a detailed movie storyline and then uses bottom-up video synthesis conditioned on those keyframes, training a Multimodal Diffusion Transformer (MM-DiT) with an interleaved strategy on a curated cinematic dataset to produce short multi-scene movies with long-range coherence. GoldenMem is the paper’s long-range memory mechanism that stores a small set of high-quality “golden” keyframe representations (visual and/or multimodal embeddings) and uses them as persistent anchors during interleaved training and inference to maintain character, scene, and stylistic consistency across many scenes and long context window."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. What exactly is the interleaving schedule and curriculum for MM-DiT (sequence lengths, mixing ratios, optimization schedule), and how sensitive are results to these hyperparameters?\n\n2. Does the system include mechanisms (e.g., latent sampling, clarification prompts, fallback heuristics) when the storyline lacks detail, and how does output diversity vs. fidelity trade off in those cases?\n\n3. Are there explicit modules or losses that enforce frame-to-frame continuity, or is it emergent from training?\n\n4. What are the common, reproducible failure modes (e.g., broken limbs, character swapping, sudden lighting jumps), and are there proposed mitigation strategies?\n\n5. How well does the model adapt to different cinematic genres, lighting conditions, actor types, or animation styles; is fine-tuning required per style?\n\n6. What are the sources and licensing terms of the curated cinematic dataset, how diverse is it, and how might dataset bias affect what stories/styles the model can generate?\n\n7. How does the system prevent or detect generation that imitates living actors or copyrighted content without consent, and are there guardrails against misuse?"},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. Clear decomposition into planning + synthesis, separating long-range narrative/keyframe planning from local spatio-temporal generation addresses coherence at different timescales and makes the overall task more tractable.\n\n2. Long-context modeling focus of adapting MM-DiT with an interleaved training strategy targets one of the core failure modes of generative video (losing global narrative or character consistency across scenes).\n\n3. Use of keyframes as conditioning signals where explicit keyframes act as strong structural constraints that likely improve framing, character consistency, and scene layout compared with unconstrained long-video diffusion.\n\n4. Building a dataset of interleaved data pairs tailored to multi-scene narrative generation helps align training data to the intended output domain and probably improves qualitative results.\n\n5. Practical pipeline for short-movie use case, the paper presents an end-to-end system that maps a storyline to a finished short movie, which is a meaningful step toward usable creative tools for filmmakers and storytellers\n\n6. GoldenMem keeps long-range visual anchors that directly enforce consistency for recurring characters and settings, which is a central failure mode in multi-scene movie generation. Compact and computationally efficient relative to storing entire past frames or attending over extremely long context lengths, improving scalability for multi-scene generation. It is easy to integrate into a two-stage pipeline (planner → keyframes → generator) because golden entries are naturally derived from planned keyframes and can be updated or replaced as the storyline evolves"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The method’s success hinges on the generated keyframes accurately reflecting the narrative and desired visual style; failure modes or hallucinated keyframes could produce hard-to-correct artifacts in the final video. Analysis on that could be useful. \n\n2. Conditioning on keyframes constrains endpoints, but plausible and temporally coherent transitions between them remain challenging; motion realism, continuity of lighting, and consistent character actions may still suffer. Any comments on that could be useful. \n\n3. Cinematic quality and narrative consistency are subjective; the paper relies heavily on qualitative examples and user studies rather than standardized, reproducible metrics, making comparisons and objective evaluation harder.\n\n4. MM-DiT long-context training with interleaved pairs and high-resolution cinematic data  requires substantial compute and curated data, raising barriers to replication and broader adoption."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762922182940,"tcdate":1761682597214,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission10998/Reviewer_2sQo"],"signatures":["ICLR.cc/2026/Conference/Submission10998/Reviewer_2sQo"],"forum":"zlNZBxQZIC","number":2,"license":"CC BY 4.0","cdate":1761682597214,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission10998/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762922182940,"domain":"ICLR.cc/2026/Conference","replyto":"zlNZBxQZIC","id":"YmGawiteGs","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"TLDR":{"value":"We present a method  for short movie generation covering data, system design and evaluation."},"keywords":{"value":["Video Generation","Diffusion Transformer"]},"supplementary_material":{"value":"/attachment/9a7a1f469f3e3a06edae240743c43fea875e60c0.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"We present **Captain Cinema**, a generation framework for short movie generation.\nGiven a detailed textual description of a movie storyline, our approach firstly generates a sequence of keyframes that outline the entire narrative, which ensures long-range coherence in both the storyline and visual appearance (e.g., scenes and characters). We refer to this step as top-down keyframe planning. These keyframes then serve as conditioning signals for a video synthesis model, which supports long context learning, to produce the spatio-temporal dynamics between them. This step is referred to as bottom-up video synthesis. To support stable and efficient generation of multi-scene long narrative cinematic works, we introduce an interleaved training strategy for Multimodal Diffusion Transformers (MM-DiT), specifically adapted for long-context video data. Our model is trained on a curated cinematic dataset consisting of interleaved samples for video generation. Our experiments demonstrate that Captain Cinema performs favorably in the automated creation of visually coherent and narratively consistent short films."},"_bibtex":{"value":"@inproceedings{\nxiao2026captain,\ntitle={Captain Cinema: Towards Short Movie Generation},\nauthor={Junfei Xiao and Ceyuan Yang and Lvmin Zhang and Shengqu Cai and Yang Zhao and Yuwei Guo and Gordon Wetzstein and Maneesh Agrawala and Alan Yuille and Lu Jiang},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=zlNZBxQZIC}\n}"},"title":{"value":"Captain Cinema: Towards Short Movie Generation"},"pdf":{"value":"/pdf/008032fa5c54abbf92e4a3bc7da69fb37304e00f.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"xiao|captain_cinema_towards_short_movie_generation"},"authorids":{"value":["~Junfei_Xiao1","~Ceyuan_Yang2","~Lvmin_Zhang2","~Shengqu_Cai1","~Yang_Zhao14","~Yuwei_Guo1","~Gordon_Wetzstein3","~Maneesh_Agrawala2","~Alan_Yuille1","~Lu_Jiang1"]},"authors":{"value":["Junfei Xiao","Ceyuan Yang","Lvmin Zhang","Shengqu Cai","Yang Zhao","Yuwei Guo","Gordon Wetzstein","Maneesh Agrawala","Alan Yuille","Lu Jiang"]}},"version":2},{"content":{"venue":{"value":"IEEE Wirel. Commun. 2013"},"pdf":{"value":"https://ieeexplore.ieee.org/iel7/7742/6549270/06549285.pdf"},"venueid":{"value":"dblp.org/journals/WC/2013"},"paperhash":{"value":"wang|cloudassisted_adaptive_video_streaming_and_socialaware_video_prefetching_for_mobile_users"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Xiaofei_Wang_0001:","https://dblp.org/search/pid/api?q=author:Ted_Taekyoung_Kwon:","https://dblp.org/search/pid/api?q=author:Yanghee_Choi:","https://dblp.org/search/pid/api?q=author:Haiyang_Wang:","~Jiangchuan_Liu1"]},"html":{"value":"https://doi.org/10.1109/MWC.2013.6549285"},"_bibtex":{"value":"@article{DBLP:journals/wc/WangKCWL13,\n  author={Xiaofei Wang and Ted Taekyoung Kwon and Yanghee Choi and Haiyang Wang and Jiangchuan Liu},\n  title={Cloud-assisted adaptive video streaming and social-aware video prefetching for mobile users},\n  year={2013},\n  cdate={1356998400000},\n  journal={IEEE Wirel. Commun.},\n  volume={20},\n  number={3},\n  pages={1-0},\n  url={https://doi.org/10.1109/MWC.2013.6549285}\n}\n"},"abstract":{"value":"While the demands of video streaming services over the mobile networks have been souring over these years, the wireless link capacity cannot practically keep up with the growing traffic load. The gap between the traffic demand and the link capacity, along with time-varying link conditions, results in poor service quality of video streaming services over the mobile networks, such as intermittent disruptions and long buffering delays. Leveraging the current cloud computing technology, we propose and discuss a framework to improve the quality of video services for mobile users, which includes two parts: cloud-assisted adaptive video streaming, and social-aware video prefetching. For each active mobile user, a private agent is constructed in the cloud center to adaptively adjust the video quality (bit rate) by the scalable video coding technique based on the feedback of link condition. Meanwhile, the online social network interactions among mobile users are monitored by the cloud-based agents, so that the videos that are shared among users will be effectively prefetched to mobile users in advance. The adaptability of the video streaming and the effectiveness of the social-aware prefetching supported by the cloud computing are evaluated based on a prototype implementation of the framework."},"title":{"value":"Cloud-assisted adaptive video streaming and social-aware video prefetching for mobile users"},"authors":{"value":["Xiaofei Wang","Ted Taekyoung Kwon","Yanghee Choi","Haiyang Wang","Jiangchuan Liu"]}},"tmdate":1768544849777,"pdate":1356998400000,"externalIds":["dblp:journals/wc/WangKCWL13"],"tcdate":1768544818455,"writers":["~"],"signatures":["~Jiangchuan_Liu1"],"forum":"IbP9WUDSYO","license":"CC BY-SA 4.0","number":767873,"cdate":1356998400000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1768544849777,"domain":"DBLP.org","id":"IbP9WUDSYO","version":2},{"content":{"summary":{"value":"This paper introduces an inverse reinforcement learning approach that enables the extraction of a dense reward signal from video demonstrations with minimal human intervention. The reward model is designed to consider both the observation and subtask information. This connection between the reward and subtask information intuitively breaks down the intricate long-term task into manageable subtasks. Empirical findings further validate the benefits of integrating subtask information into the reward learning process."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"What if we utilize the text embedding directly as input during the inference phase? Does this approach exhibit a substantial performance advantage over the method that employs video embedding?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"i) In this paper, subtask information is integrated into the reward learning process. During training, the subtask is included in the reward function input as a text embedding that provides instructions on completing a specific subtask. In the inference phase, the text embedding is substituted with a video embedding as an additional input.\n\nii) The approach employs the EPIC loss function to reduce the disparity between the predicted reward sequence and the ground truth reward. Experimental results demonstrate superiority of this approach."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"i) In the training phase, this paper decomposes the overall task into multiple subtasks based on domain knowledge. However, the reliance on predefined instructions from the environment for task decomposition raises concerns about practical applicability. Some environments may lack such predefined knowledge, necessitating human annotations or the need for a learned task decomposition model when extending this approach to new environments. This raises doubts about the novelty and scalability of this methodology. \n\nii) This paper introduces the function $\\psi$, which maps observations to subtasks. However, providing more details about this function is essential. Specifically, elucidating the form of the subtask output by $\\psi$ is crucial since it serves as the ground truth for reward learning, a key component for the success of this method. Therefore, the authors should elaborate on the detail of the subtasks produced by $\\psi$.\n\niii) During the reward learning, this method utilizes the EPIC loss function to quantify the disparity between the estimated reward and the proposed ground truth. While EPIC was initially introduced in a prior work, the rationale behind its selection and comparisons with previous loss functions for reward learning should be further discussed by the authors. This explanation should include both intuitive reasoning and empirical evidence.\n\niv) During the inference phase, the substitution of text embedding with video embedding for subtask information is addressed. Although an alignment mechanism is employed to bring video and text embeddings closer in the latent space, the loss function appears to overlook the maximization of distances between irrelevant embeddings. This oversight could be crucial when dealing with subtasks that share similar text descriptions or video content. It is advisable for the authors to investigate the effectiveness of learned embeddings within tasks involving similar subtasks, although this aspect is not a primary concern."}},"nonreaders":[],"tmdate":1732990879857,"tcdate":1730476029981,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission3969/Reviewer_Vi9Z"],"signatures":["ICLR.cc/2025/Conference/Submission3969/Reviewer_Vi9Z"],"forum":"mqKVe6F3Up","number":3,"license":"CC BY 4.0","cdate":1730476029981,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission3969/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1732990879857,"domain":"ICLR.cc/2025/Conference","replyto":"mqKVe6F3Up","id":"rOpOJkaiLs","forumContent":{"TLDR":{"value":"We propose a novel reward learning framework utilizing action-free videos with minimal guidance for long-horizon complex robotic tasks."},"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Reinforcement Learning","Reward Learning","Robotic Manipulation"]},"supplementary_material":{"value":"/attachment/2be10b8d059956b891d8f31da453fd417569f388.zip"},"primary_area":{"value":"applications to robotics, autonomy, planning"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Reinforcement Learning (RL) agents have demonstrated their potential across various robotic tasks. However, they still heavily rely on human-engineered reward functions, requiring extensive trial-and-error and access to target behavior information, often unavailable in real-world settings. This paper introduces REDS: REward learning from Demonstration with Segmentations, a novel reward learning framework that leverages action-free videos with minimal supervision. Specifically, REDS employs video demonstrations segmented into subtasks from diverse sources and treats these segments as ground-truth rewards. We train a dense reward function conditioned on video segments and their corresponding subtasks to ensure alignment with ground-truth reward signals by minimizing the Equivalent-Policy Invariant Comparison distance. Additionally, we employ contrastive learning objectives to align video representations with subtasks, ensuring precise subtask inference during online interactions. Our experiments show that REDS significantly outperforms baseline methods on complex robotic manipulation tasks in Meta-World and more challenging real-world tasks, such as furniture assembly in FurnitureBench, with minimal human intervention. Moreover, REDS facilitates generalization to unseen tasks and robot embodiments, highlighting its potential for scalable deployment in diverse environments."},"_bibtex":{"value":"@inproceedings{\nkim2025subtaskaware,\ntitle={Subtask-Aware Visual Reward Learning from Segmented Demonstrations},\nauthor={Changyeon Kim and Minho Heo and Doohyun Lee and Honglak Lee and Jinwoo Shin and Joseph J Lim and Kimin Lee},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=mqKVe6F3Up}\n}"},"title":{"value":"Subtask-Aware Visual Reward Learning from Segmented Demonstrations"},"pdf":{"value":"/pdf/18d7fe7853d07ace357a88ffd892f59889d69123.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"kim|subtaskaware_visual_reward_learning_from_segmented_demonstrations"},"authorids":{"value":["~Changyeon_Kim1","~Minho_Heo1","~Doohyun_Lee2","~Honglak_Lee2","~Jinwoo_Shin1","~Joseph_J_Lim1","~Kimin_Lee1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Changyeon Kim","Minho Heo","Doohyun Lee","Honglak Lee","Jinwoo Shin","Joseph J Lim","Kimin Lee"]}},"version":2},{"content":{"summary":{"value":"The paper introduces a novel, training-free framework to optimize video diffusion models. This framework, consisting of Feature Slicer, Operator Grouping, and Step Rehash, significantly reduces peak memory usage and computational overhead while maintaining video quality. Extensive experiments demonstrate that the approach can cut memory usage by up to 70% and improve inference speed by 1.6 times compared to baseline methods. The framework is compatible with existing models like AnimateDiff and SVD, enabling high-quality video generation on consumer-grade GPUs. The research paves the way for more efficient video diffusion models, making advanced video generation accessible on standard hardware."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"- Are there any specific scenarios where the overhead introduced by slicing and grouping operations could outweigh the benefits?\n- Have you considered conducting user studies to provide a more comprehensive assessment of the generated video quality?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"- The introduction of a training-free framework that optimizes video diffusion models seems a novel approach.\n- The method is compatible with existing video diffusion models like AnimateDiff and SVD, ensuring broad applicability.\n- Comprehensive experiments and detailed analysis demonstrate the framework's effectiveness and robustness."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Although the paper claims that video quality is maintained, the extent of quality degradation, if any, is not fully quantified. \n- Although some ablation studies are provided, a more detailed breakdown of the contributions of each component (Feature Slicer, Operator Grouping, and Step Rehash) would strengthen the understanding of their individual and combined impacts.\n- The evaluation relies heavily on FVD and CLIP-Scores, which, while useful, may not capture all dimensions of video quality and user satisfaction. Including additional metrics or user studies could provide a more holistic assessment of the generated video quality."},"limitations":{"value":"Authors have adequately addressed the limitations."}},"nonreaders":[],"tmdate":1730879007975,"tcdate":1720995039127,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission5521/Reviewer_mL5N"],"signatures":["NeurIPS.cc/2024/Conference/Submission5521/Reviewer_mL5N"],"forum":"iNvXYQrkpi","number":4,"license":"CC BY 4.0","cdate":1720995039127,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission5521/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730879007975,"domain":"NeurIPS.cc/2024/Conference","replyto":"iNvXYQrkpi","id":"u5FTv2a5Dp","forumContent":{"venue":{"value":"NeurIPS 2024 poster"},"keywords":{"value":["video diffusion models","training free","inference framework","Efficiency"]},"primary_area":{"value":"diffusion_based_models"},"abstract":{"value":"The rapid progress in artificial intelligence-generated content (AIGC), especially with diffusion models, has significantly advanced development of high-quality video generation. However, current video diffusion models exhibit demanding computational requirements and high peak memory usage, especially for generating longer and higher-resolution videos. These limitations greatly hinder the practical application of video diffusion models on standard hardware platforms. To tackle this issue, we present a novel, training-free framework named Streamlined Inference, which leverages the temporal and spatial properties of video diffusion models. Our approach integrates three core components: Feature Slicer, Operator Grouping, and Step Rehash. Specifically, Feature Slicer effectively partitions input features into sub-features and Operator Grouping processes each sub-feature with a group of consecutive operators, resulting in significant memory reduction without sacrificing the quality or speed. Step Rehash further exploits the similarity between adjacent steps in diffusion, and accelerates inference through skipping unnecessary steps. Extensive experiments demonstrate that our approach significantly reduces peak memory and computational overhead, making it feasible to generate high-quality videos on a single consumer GPU (e.g., reducing peak memory of Animatediff from 42GB to 11GB, featuring faster inference on 2080Ti)."},"_bibtex":{"value":"@inproceedings{\nzhan2024fast,\ntitle={Fast and Memory-Efficient Video Diffusion Using Streamlined Inference},\nauthor={Zheng Zhan and Yushu Wu and Yifan Gong and Zichong Meng and Zhenglun Kong and Changdi Yang and Geng Yuan and Pu Zhao and Wei Niu and Yanzhi Wang},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=iNvXYQrkpi}\n}"},"title":{"value":"Fast and Memory-Efficient Video Diffusion Using Streamlined Inference"},"pdf":{"value":"/pdf/746d0dded936a447f6abe89c86574a1936c3d8bd.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"zhan|fast_and_memoryefficient_video_diffusion_using_streamlined_inference"},"authorids":{"value":["~Zheng_Zhan3","~Yushu_Wu1","~Yifan_Gong2","~Zichong_Meng1","~Zhenglun_Kong1","~Changdi_Yang1","~Geng_Yuan1","~Pu_Zhao1","~Wei_Niu3","~Yanzhi_Wang3"]},"authors":{"value":["Zheng Zhan","Yushu Wu","Yifan Gong","Zichong Meng","Zhenglun Kong","Changdi Yang","Geng Yuan","Pu Zhao","Wei Niu","Yanzhi Wang"]}},"version":2},{"content":{"venue":{"value":"IEEE Trans. Multim. 2006"},"pdf":{"value":"https://ieeexplore.ieee.org/iel5/6046/33775/01608108.pdf"},"venueid":{"value":"dblp.org/journals/TMM/2006"},"paperhash":{"value":"yin|casm_a_contentaware_protocol_for_secure_video_multicast"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Hao_Yin:","https://dblp.org/search/pid/api?q=author:Chuang_Lin_0002:","https://dblp.org/search/pid/api?q=author:Feng_Qiu:","~Jiangchuan_Liu1","https://dblp.org/search/pid/api?q=author:Geyong_Min:","~Bo_Li33"]},"html":{"value":"https://doi.org/10.1109/TMM.2005.864316"},"_bibtex":{"value":"@article{DBLP:journals/tmm/YinLQLML06,\n  author={Hao Yin and Chuang Lin and Feng Qiu and Jiangchuan Liu and Geyong Min and Bo Li},\n  title={CASM: a content-aware protocol for secure video multicast},\n  year={2006},\n  cdate={1136073600000},\n  journal={IEEE Trans. Multim.},\n  volume={8},\n  number={2},\n  pages={270-277},\n  url={https://doi.org/10.1109/TMM.2005.864316}\n}\n"},"abstract":{"value":"Information security has been a critical issue in the design and development of reliable distributed communication systems and has attracted significant research efforts. A challenging task is how to maintain information security at a high level for multiple-destination video applications with the huge volume of data and dynamic property of clients. This paper proposes a novel Content-Aware Secure Multicast (CASM) protocol for video distribution that seamlessly integrates three important modules: 1)a scalable light-weight algorithm for group key management; 2) a content-aware key embedding algorithm that can make video quality distortion imperceptible and is reliable for clients to detect embedded keys; and 3) a smart two-level video encryption algorithm that can selectively encrypt a small set of video data only, and yet ensure the video as well as the embedded keys unrecognizable without a genuine key. The implementation of the CASM protocol is independent of the underlying multicast mechanism and is fully compatible with existing coding standards. Performance evaluation studies built upon a CASM prototype have demonstrated that CASM is highly robust and scalable in dynamic multicast environments. Moreover, it ensures secure distribution of key and video data with minimized communication and computation overheads. The proposed content-aware key embedding and encryption algorithms are fast enough to support real-time video multicasting."},"title":{"value":"CASM: a content-aware protocol for secure video multicast"},"authors":{"value":["Hao Yin","Chuang Lin","Feng Qiu","Jiangchuan Liu","Geyong Min","Bo Li"]}},"tmdate":1768544824057,"pdate":1136073600000,"tcdate":1723126475053,"writers":["~"],"signatures":["~Bo_Li33"],"forum":"aVfy25JLSP","license":"CC BY-SA 4.0","number":62609,"cdate":1136073600000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1768544824057,"domain":"DBLP.org","id":"aVfy25JLSP","version":2},{"content":{"summary":{"value":"The paper introduces MPS, a benchmark built from 8 text classification datasets to evaluate model robustness under five spurious correlation types, namely SCS, CCS, NBS, QBS, and WFS. It defines two analysis metrics, which are δ (change in worst-group accuracy when balancing a given shortcut type) and Δ (headroom from worst-group under shortcut to overall accuracy after balancing), to quantify each shortcut’s impact. Across models and mitigation methods, results show no single method is robust across all five types, and the newly defined QBS is particularly challenging."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"1. Table 7’s descriptions do not make the distinction between SCS and CCS sufficiently clear. Could you provide precise definitions and concrete, dataset-specific examples for each to clarify the difference?\n2. When δ < 0 (e.g., strong results on an imbalanced split), how do you conclude that the model exploits spurious correlations rather than demonstrating genuine understanding?\n3. How to explain the large negative δ of TCM on Ag News in Table 1?\n4. Could the low W-ACC simply because of small sample sizes in certain (𝑦, 𝑎) groups, rather than true vulnerability to the spurious attribute."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"The paper’s core strength lies in its scope and standardization. A unified benchmark across eight datasets and five shortcut types), with paired balanced/imbalanced splits, and clear robustness metrics (W-ACC, δ/Δ) that isolate and quantify shortcut effects. Empirically, it’s broad and careful, covering classic baselines, pretrained encoders, multiple mitigation families, and LLM backbones, plus a human reference, yielding actionable findings."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The distinction between sentence-level concepts (SCS) and core-word concepts (CCS) is not fully transparent from the main text.\n2. It is not clear how static attributes of SCS are selected and the relationship with their corresponding labels.\n3. Some ambiguous parts, like the usage of W-ACC and δ, are listed in the following Questions section."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762919556459,"tcdate":1761629510960,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission7442/Reviewer_3pFQ"],"signatures":["ICLR.cc/2026/Conference/Submission7442/Reviewer_3pFQ"],"forum":"s633mMGHYZ","number":2,"license":"CC BY 4.0","cdate":1761629510960,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission7442/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762919556459,"domain":"ICLR.cc/2026/Conference","replyto":"s633mMGHYZ","id":"hfUEDGxv90","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["NLP; Spurious Correlations; Benchmark"]},"supplementary_material":{"value":"/attachment/39d3a4f93df94a01d3045d9745a5488f7450905e.zip"},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Text classification is especially susceptible to diverse spurious correlations, such as those related to word-frequency and concept-level patterns. Nevertheless, there is a lack of a comprehensive and standardized benchmark for evaluating the robustness of models against these spurious correlations. To address this crucial issue, we present MPS (Multi - Perspective Benchmark For Assessing Spurious Correlations in Text Classification). To construct this benchmark, we collect eight widely used text classification datasets and introduce five categories of spurious correlations for each of them, producing 40 variants of datasets for comprehensively evaluating spurious correlations in diverse settings.We then extensively evaluate various text classification models and state-of-the-art anti-spurious correlation methods on this benchmark, which uncovers the vulnerabilities of these models and methods to diverse spurious correlations. A follow-up comparative analysis on this benchmark is performed to assess the performance of these anti-spurious correlation methods and humans in diverse settings."},"_bibtex":{"value":"@misc{\nzou2026mps,\ntitle={{MPS}: A Multi-Perspective Benchmark For Assessing Spurious Correlations in Text Classification},\nauthor={Liangjun Zou and Pinze Ren and Zhongliang Yang and Yinjun Wu and Jiehe Dong and Zhengyang Fang and Zhen Chen and Linna Zhou},\nyear={2026},\nurl={https://openreview.net/forum?id=s633mMGHYZ}\n}"},"title":{"value":"MPS: A Multi-Perspective Benchmark For Assessing Spurious Correlations in Text Classification"},"pdf":{"value":"/pdf/d62f7893563317ef6056527bed841d1fa83e6789.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"zou|mps_a_multiperspective_benchmark_for_assessing_spurious_correlations_in_text_classification"},"authorids":{"value":["~Liangjun_Zou1","~Pinze_Ren1","~Zhongliang_Yang2","~Yinjun_Wu1","~Jiehe_Dong1","~Zhengyang_Fang1","~Zhen_Chen25","~Linna_Zhou2"]},"authors":{"value":["Liangjun Zou","Pinze Ren","Zhongliang Yang","Yinjun Wu","Jiehe Dong","Zhengyang Fang","Zhen Chen","Linna Zhou"]}},"version":2},{"content":{"summary":{"value":"The paper introduces HiFi-Foley, an end-to-end text-video-to-audio framework that synthesizes high-fidelity audio precisely aligned with visual dynamics and semantic context.\n\nThe main paper contributions are:\n- a novel multimodal diffusion transformer that addresses semantic response imbalance between video and text modalities through dual-stream audio-video fusion via joint attention and balanced textual semantic injection via cross-attention.\n- a representation alignment training strategy that employs self-supervised audio features to guide latent diffusion training, thereby improving audio quality and semantic consistency.\n- a scalable data pipeline leveraging open-source tools for cleaning raw data and constructing training datasets.\n\nExtensive evaluations demonstrate that HiFi-Foley achieves state-of-the-art performance across audio fidelity, visual semantic alignment, temporal alignment, and distribution matching."},"soundness":{"value":1},"confidence":{"value":5},"questions":{"value":"What is the \"HunyuanVideo-Foley\" model mentioned in the Figure 2 caption?\nWhat is the ATST-Frame model mentioned throughout the paper?\nWhat is the parallel cross attention ablation in the Table 5? A figure would be welcome."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":1},"strengths":{"value":"The main strength of the paper is the data curation pipeline designed to build a high quality training dataset for the video to audio generation task."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"I believe that the proposed contributions lack novelty or significance.\n- Injecting text via cross attention in video to audio generation has been proposed before (MovieGen). Moreover, the motivation for this architecture design is unconvincing. The imbalance between conditioning signals can usually be addressed by employing different guidance weights during inference, which is not considered in this paper (at least as a baseline). The paper does not even mention the inference parameters used in the experiments.\n- The REPA loss has been introduced in prior works and this paper just applies it to a new task.\n- The proposed data curation pipeline is mostly descriptive and the resulting dataset is not published. The size of the resulting curated dataset is one order of magnitude than the biggest non curated open source dataset (AudioSet), which is probably the main explanation for the appealing results of the Tables 2 and 4. Moreover, the curation pipeline is undocumented and thus non reproducible. For example, no extensive description of the thresholds used at the different filtering steps are given.\n\nThe results presented in the tables, which do not report confidence intervals, yield minimal differences between the different methods (such as the Tables 5, 6 and 7). Thus they are unconvincing to the reader. For what it is worth, according to my experience, absolute variations of less than 0.2 in A4 scores are not significant. \n\nThe Figure 2 provides the same information as the Tables 2, 3, 4."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762928038463,"tcdate":1762066261782,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission18324/Reviewer_zgNX"],"signatures":["ICLR.cc/2026/Conference/Submission18324/Reviewer_zgNX"],"forum":"72xWFIzG15","number":4,"license":"CC BY 4.0","cdate":1762066261782,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission18324/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762928038463,"domain":"ICLR.cc/2026/Conference","replyto":"72xWFIzG15","id":"opGqpci01m","forumContent":{"TLDR":{"value":"HiFi-Foley is a novel Text-Video-to-Audio model that generates high-quality, perfectly synced sound for videos, outperforming SOTA methods."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Video-to-audio","Audio Generation","Foley Generation","Multi-modal","Representation Alignment"]},"supplementary_material":{"value":"/attachment/2fb59dffd89b64f5ae7d1cc1389979f01f451e7e.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Recent advances in video generation produce visually realistic content, yet the absence of synchronized audio severely compromises immersion. To address key challenges in video-to-audio generation, including multimodal data scarcity, modal semantic response imbalance, and limited audio quality in existing methods, we propose HiFi-Foley, an end-to-end text-video-to-audio framework that synthesizes high-fidelity audio precisely aligned with visual dynamics and semantic context. Our approach incorporates three core innovations: (1) a novel multimodal diffusion transformer that addresses semantic response imbalance between video and text modalities through dual-stream audio-video fusion via joint attention and balanced textual semantic injection via cross-attention; (2) a representation alignment training strategy that employs self-supervised audio features to guide latent diffusion training, thereby improving audio quality and semantic consistency; (3) a scalable data pipeline leveraging open-source tools for cleaning raw data and constructing training datasets. Extensive evaluations demonstrate that HiFi-Foley achieves state-of-the-art performance across audio fidelity, visual-semantic alignment, temporal alignment, and distribution matching."},"_bibtex":{"value":"@misc{\nshan2026hififoley,\ntitle={HiFi-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation},\nauthor={Sizhe Shan and Qiulin Li and Yutao Cui and Miles Yang and Zhao Zhong and Yuehai Wang and Jin Zhou},\nyear={2026},\nurl={https://openreview.net/forum?id=72xWFIzG15}\n}"},"title":{"value":"HiFi-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation"},"pdf":{"value":"/pdf/624378c966b88562072a14ac6e5a30083525d7d6.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"shan|hififoley_multimodal_diffusion_with_representation_alignment_for_highfidelity_foley_audio_generation"},"authorids":{"value":["~Sizhe_Shan1","~Qiulin_Li1","~Yutao_Cui1","~Miles_Yang2","~Zhao_Zhong1","~Yuehai_Wang1","~Jin_Zhou10"]},"authors":{"value":["Sizhe Shan","Qiulin Li","Yutao Cui","Miles Yang","Zhao Zhong","Yuehai Wang","Jin Zhou"]}},"version":2},{"content":{"venue":{"value":"SMC 2022"},"pdf":{"value":"https://ieeexplore.ieee.org/iel7/9945068/9945069/09945377.pdf"},"venueid":{"value":"dblp.org/conf/SMC/2022"},"paperhash":{"value":"luo|temporalaware_mechanism_with_bidirectional_complementarity_for_video_qa"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Yuanmao_Luo:","https://dblp.org/search/pid/api?q=author:Ruomei_Wang_0001:","https://dblp.org/search/pid/api?q=author:Fuwei_Zhang:","https://dblp.org/search/pid/api?q=author:Fan_Zhou_0001:","~Shujin_Lin1"]},"html":{"value":"https://doi.org/10.1109/SMC53654.2022.9945377"},"_bibtex":{"value":"@inproceedings{DBLP:conf/smc/LuoWZZL22,\n  author={Yuanmao Luo and Ruomei Wang and Fuwei Zhang and Fan Zhou and Shujin Lin},\n  title={Temporal-aware Mechanism with Bidirectional Complementarity for Video Q&A},\n  year={2022},\n  cdate={1640995200000},\n  pages={3273-3278},\n  url={https://doi.org/10.1109/SMC53654.2022.9945377},\n  booktitle={SMC},\n  crossref={conf/smc/2022}\n}\n"},"abstract":{"value":"Video question answering (Video Q&A) is a challenging task as it requires a sufficient understanding of the video and question information. Video is composed of frame sequence, which contains multi-scale temporal relationships and corresponding contextual information. A model competently tackle Video Q&A task that needs to be able to: 1) construct long-term and neighborhood dependencies in frame sequences to extract global and local contextual features that can reflect multi-scale temporal dependencies, and deduce the temporal-aware refined features, and 2) identify static and dynamic features from pertinent moments of a video, while filtering away question-irrelated dependencies of feature sequences, to yield the most precise and reasonable temporal-aware overall contextual features. In response to the above requirements, we propose a novel Video Q&A mechanism which consists of Bidirectional Complementary Attention(BCA) module and Adaptive Temporal-aware(ATA) module. Bidirectional complementary attention module stacks multi-head self-attention layer and convolutional layer in different orders to designed two kinds of attention units, which is able to make bidirectional multi-step reasoning based on complete global information and accurate local information to obtain temporal-aware refined features. Adaptive temporal-aware module is used to filter away question-irrelated dependencies in the feature sequence to yield the most precise and reasonable temporal-aware overall contextual features. Comprehensive comparative experiments are conducted on publicly available benchmark datasets. An extended ablation study is further conducted to show the usefulness of each module of the solution in acquiring its computational Q&A capabilities."},"title":{"value":"Temporal-aware Mechanism with Bidirectional Complementarity for Video Q&A"},"authors":{"value":["Yuanmao Luo","Ruomei Wang","Fuwei Zhang","Fan Zhou","Shujin Lin"]}},"tmdate":1740044018012,"pdate":1640995200000,"tcdate":1740044015654,"writers":["~"],"signatures":["~Shujin_Lin1"],"forum":"NcjGTani4s","license":"CC BY-SA 4.0","number":336443,"cdate":1640995200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1740044018012,"domain":"DBLP.org","id":"NcjGTani4s","version":2},{"content":{"summary":{"value":"The authors propose a fascinating approach to optimizing transformer-based large language models (LLMs). They delve into the intricacies of position encoding, particularly focusing on the widely used rotary position embedding (RoPE) technique. By examining how the angle between weight vector pairs impacts attention scores, they reveal that non-orthogonal pairs are crucial for processing basic syntax, while orthogonal pairs handle higher-level semantics. Their experiments show that fine-tuning LLMs predominantly alters these orthogonal pairs, allowing them to propose a new method (QK-IPM) to reduce fine-tuning overhead. This method effectively trims down the number of trainable parameters without sacrificing performance, as evidenced by their tests on Alpaca fine-tuned Llama-2. Overall, their work offers a fresh perspective on position encoding and presents a practical solution for more efficient LLM fine-tuning."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"See weaknesses."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The proposed Query-Key Internal Pair Masking (QK-IPM) method stands out for its efficiency. By identifying that non-orthogonal weight vector pairs don't need updating during fine-tuning, the method significantly reduces the number of trainable parameters, streamlining the fine-tuning process and saving computational resources.\n\n2. The approach is backed by solid theoretical insights. The paper explains the relationship between the angles of weight vector pairs and their roles in processing syntactic versus semantic information, providing a robust foundation for the proposed method. This depth of understanding adds credibility and makes the findings more convincing.\n\n3. The empirical evidence provided, particularly through tests on widely used models like Alpaca fine-tuned Llama-2, demonstrates the practical benefits of the method. This real-world validation shows that the technique is not just theoretically sound but also effective in improving model performance with reduced overhead."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The paper primarily tests the proposed method on specific models and datasets. While the results are promising, a broader evaluation across various LLM architectures and more diverse datasets would strengthen the generalizability of the findings and ensure the method's robustness in different contexts.\n\nAlthough the method reduces the number of trainable parameters, the process of calculating the angles between weight vector pairs and determining which pairs to update introduces additional computational steps. This could offset some of the efficiency gains, especially in large-scale applications, and might need further optimization to ensure overall net benefits."},"limitations":{"value":"No"}},"nonreaders":[],"tmdate":1730879322122,"tcdate":1720791278136,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission9554/Reviewer_CgW8"],"signatures":["NeurIPS.cc/2024/Conference/Submission9554/Reviewer_CgW8"],"forum":"e5Mv7iWfVW","number":4,"license":"CC BY 4.0","cdate":1720791278136,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission9554/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730879322122,"domain":"NeurIPS.cc/2024/Conference","replyto":"e5Mv7iWfVW","id":"nhRQ3SfNGv","forumContent":{"venue":{"value":"NeurIPS 2024 poster"},"keywords":{"value":["Large Language Model.+Rotary Position Embedding.+Self Attention"]},"primary_area":{"value":"natural_language_processing"},"abstract":{"value":"Transformer-based large language models (LLMs) have successfully handled various tasks. As one fundamental module in Transformers, position encoding encodes the positional information of tokens in a sequence. Specifically, rotary position embedding (RoPE), one of the most widely used techniques, encodes the positional information by dividing the query or key value with $d$ elements into $d/2$ pairs and rotating the 2d vectors corresponding to each pair of elements. Therefore, the direction of each pair and the position-related rotation jointly determine the attention score. In this paper, we show that the direction of the 2d pair is largely affected by the angle between the corresponding weight vector pair. We theoretically show that non-orthogonal weight vector pairs lead to great attention on tokens at a certain relative position and are less sensitive to the input which may correspond to basic syntactic information. Meanwhile, the orthogonal weight vector pairs are more flexible regarding the relative position, which may correspond to high-level syntactic information. Empirical evidence supports the hypothesis that shallow layers of LLMs focus more on local syntax and deep layers focus more on high-level semantics. Furthermore, we show that LLMs fine-tuning mainly changes the pairs of weight vectors that are nearly orthogonal, i.e., the weight corresponding to high-level semantics, which enables the reduction of the number of trainable parameters during fine-tuning without sacrificing performance. We propose a method namely Angle-based Weight Selection (AWS) to reduce the fine-tuning overhead and verify the effectiveness of the proposed method on widely used Alpaca fine-tuned Llama-2."},"_bibtex":{"value":"@inproceedings{\nchen2024what,\ntitle={What Rotary Position Embedding Can Tell Us: Identifying Query and Key Weights Corresponding to Basic Syntactic or High-level Semantic Information},\nauthor={Yiting Chen and Junchi Yan},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=e5Mv7iWfVW}\n}"},"title":{"value":"What Rotary Position Embedding Can Tell Us: Identifying Query and Key Weights Corresponding to Basic Syntactic or High-level Semantic Information"},"pdf":{"value":"/pdf/5d5a69ecf75e1413e8a4cbe55a98b16144a08473.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"chen|what_rotary_position_embedding_can_tell_us_identifying_query_and_key_weights_corresponding_to_basic_syntactic_or_highlevel_semantic_information"},"authorids":{"value":["~Yiting_Chen1","~Junchi_Yan2"]},"authors":{"value":["Yiting Chen","Junchi Yan"]}},"version":2},{"content":{"summary":{"value":"The paper proposes InterActHuman, a diffusion transformer based human video generation framework that aims to synthesize multi person or human object interaction videos in which each identity keeps its own appearance, motion style, and voice, given multiple reference images, a text description, and per speaker audio tracks. The core idea is to drop the usual single identity assumption and instead bind each conditioning signal to the correct spatial region over time. To do this, the model predicts for each reference concept a spatiotemporal mask via a lightweight cross attention head that matches the noisy latent video tokens to that concept. These masks are refined iteratively across denoising steps and then used to inject audio features only into the region of the speaker at the next step, which addresses the chicken and egg problem of needing masks to apply localized audio while the video is still being generated."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. How are individual audio streams associated with the correct identity during training. The paper explains that during inference, audio cross attention is only applied to tokens whose mask corresponds to that speaker, using cached masks from the previous denoising step.\n\n2. The data pipeline aligns audio segments to identities through lip synchronization and produces per-frame masks using Grounding SAM2 with a person query. Could you elaborate on how you ensure temporal identity consistency across frames in multi-person scenes, especially during occlusion or when two people have similar appearances? Do you rely on optical flow tracking or an ID re-identification module?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. The paper tackles a practically important and under explored setting: multi-person, audio-conditioned human animation where each identity must keep its own appearance and voice, instead of the common single identity assumption in prior audio-driven portrait or OmniHuman style models.\n\n2. The proposed iterative mask prediction and cached layout guided audio injection mechanism is elegant. By predicting per identity spatiotemporal masks using cross attention between reference appearance tokens and noisy video latents, then using the previous step mask to gate the next step audio cross attention, the model effectively solves the chicken and egg problem of aligning local audio before the final frames exist.\n\n3. Quantitative and qualitative results are compelling. The user study shows a strong preference for the proposed method in both lip sync realism and subject consistency."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Runtime cost and scalability claims are mostly deferred to the appendix. The main text asserts minimal overhead and compatibility with long video generation, but does not quantify inference speed or memory usage when conditioning on multiple identities.\n\n2. Multi speaker audio assignment appears to rely on injecting each audio stream only into the spatial region indicated by that speaker’s cached mask at inference time. It is unclear whether the model is ever explicitly trained on multi speaker scenes with multiple simultaneous audio streams, or whether this is essentially a zero shot composition of single speaker training. This could limit robustness when speakers interrupt each other or overlap."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762924203439,"tcdate":1761799040945,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission13626/Reviewer_xdLa"],"signatures":["ICLR.cc/2026/Conference/Submission13626/Reviewer_xdLa"],"forum":"rJilRU8D3c","number":2,"license":"CC BY 4.0","cdate":1761799040945,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission13626/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762924203439,"domain":"ICLR.cc/2026/Conference","replyto":"rJilRU8D3c","id":"PSbd19lBl1","forumContent":{"TLDR":{"value":"could generate 2-3 people dialogue videos, or single-person talking videos with human-object interactions"},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["talking person video generation","multi-concept video customization"]},"supplementary_material":{"value":"/attachment/100c21e0d2a409daf79e79f29329a4ae1ba96db3.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"End-to-end human animation with rich multi-modal conditions, e.g., text, image and audio has achieved remarkable advancements in recent years. However, most existing methods could only animate a single subject and inject conditions in a global manner, ignoring scenarios where multiple concepts could appear in the same video with rich human-human interactions and human-object interactions. Such a global assumption prevents precise and per-identity control of multiple concepts including humans and objects, therefore hinders applications. In this work, we discard the single-entity assumption and introduce a novel framework that enforces strong, region‑specific binding of conditions from modalities to each identity's spatiotemporal footprint. Given reference images of multiple concepts, our method could automatically infer layout information by leveraging a mask predictor to match appearance cues between the denoised video and each reference appearance. Furthermore, we inject local audio condition into its corresponding region to ensure layout-aligned modality matching in an iterative manner. This design enables the high-quality generation of human dialogue videos between two to three people or video customization from multiple reference images. Empirical results and ablation studies validate the effectiveness of our explicit layout control for multi-modal conditions compared to implicit counterparts and other existing methods."},"_bibtex":{"value":"@inproceedings{\nwang2026interacthuman,\ntitle={InterActHuman: Multi-Concept Human Animation with Layout-Aligned Audio Conditions},\nauthor={Zhenzhi Wang and Jiaqi Yang and Jianwen Jiang and Chao Liang and Gaojie Lin and Zerong Zheng and Ceyuan Yang and Yuan Zhang and Mingyuan Gao and Dahua Lin},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=rJilRU8D3c}\n}"},"title":{"value":"InterActHuman: Multi-Concept Human Animation with Layout-Aligned Audio Conditions"},"pdf":{"value":"/pdf/09cc87825e4cacbc45a1584c64713755b7a7fbc7.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"wang|interacthuman_multiconcept_human_animation_with_layoutaligned_audio_conditions"},"authorids":{"value":["~Zhenzhi_Wang1","~Jiaqi_Yang8","~Jianwen_Jiang2","~Chao_Liang4","~Gaojie_Lin1","~Zerong_Zheng1","~Ceyuan_Yang2","~Yuan_Zhang6","~Mingyuan_Gao2","~Dahua_Lin1"]},"authors":{"value":["Zhenzhi Wang","Jiaqi Yang","Jianwen Jiang","Chao Liang","Gaojie Lin","Zerong Zheng","Ceyuan Yang","Yuan Zhang","Mingyuan Gao","Dahua Lin"]}},"version":2},{"content":{"summary":{"value":"This paper introduces the OmniEval benchmark, designed for evaluating a model's understanding of synchronized audio-visual content. The videos are sourced from existing datasets (Finevideo and Youku-mplug). The authors' pipeline involves first using automated tools to generate visual captions and ASR transcripts for the videos. These textual annotations are then used as input for LLMs to generate question-answer pairs, which are subsequently refined by human annotators. The paper also introduces a \"Grounding\" task to test precise spatio-temporal understanding."},"soundness":{"value":2},"confidence":{"value":5},"questions":{"value":"- What is the performance of strong vision-only LLMs (provided with subtitles as an additional input) on the OmniEval benchmark?\n- Which LLM is used for annotation in Section 3?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"- Bilingual Benchmark: OmniEval is a bilingual video understanding benchmark that includes both English and Chinese videos and questions, which is valuable for evaluating multilingual models.\n- Audio-Visual Grounding Task: OmniEval introduces the \"Grounding\" task, which is an important capability for audio-visual understanding."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Missing Comparison with Existing Benchmarks: For the evaluation of audio-visual video understanding, there are already established benchmarks (e.g., AVUT, DailyOmni), which are not discussed or compared in the paper.\n- Limitations of the Data Generation Methodology: The method of using an LLM to generate questions based on video captions and audio subtitles, while cost-effective, has several critical limitations:\n  - Based on empirical evidence, Q&A pairs generated by LLMs tend to be of limited difficulty.\n  - Since the visual and audio information are provided as decoupled text streams (captions and subtitles), the LLM is likely to generate questions that are superficial, touching only on the surface-level content of each modality independently. This method fails to produce questions that probe a deeper, more challenging understanding derived from the interplay between audio and video.\n  - Furthermore, the synchronicity between the visual and auditory events is not considered during question generation, which undermines the benchmark's core purpose of evaluating omni-modal inputs. This oversight could even lead to questions with errors or ambiguous answers.\n- Narrow Focus on ASR for Audio: Regarding the audio modality, the benchmark appears to focus almost exclusively on ASR (speech content), neglecting the role of general audio events (e.g., environmental sounds, music, sound effects), which are crucial for comprehensive scene understanding.\n- Lack of Detail on Grounding Annotation: The annotation process for the Grounding task is not described. The paper provides few details on how these temporal annotations were created or how their accuracy and consistency were ensured.\n- Incomplete Experiments: The experimental section is insufficient. The paper only compares several omni-LLMs but lacks a crucial baseline of strong vision-only LLMs with subtitles as an additional text input. Moreover, the paper fails to compare against other powerful audio-visual LLMs such as Video-LLaMA 2 and Video-SALMONN 2."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762926646798,"tcdate":1761735307268,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission16567/Reviewer_sdL2"],"signatures":["ICLR.cc/2026/Conference/Submission16567/Reviewer_sdL2"],"forum":"B9iMn59jFE","number":1,"license":"CC BY 4.0","cdate":1761735307268,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission16567/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762926646798,"domain":"ICLR.cc/2026/Conference","replyto":"B9iMn59jFE","id":"hoXrbk1t98","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Omni models","Benchmark","Multimodality"]},"supplementary_material":{"value":"/attachment/47b7d0119eed424cb9eabba953ad3fb8d88b92fd.zip"},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"In this paper, we introduce OmniEval, a benchmark for evaluating multimodal Chinese and English video understanding, which encompasses visual, auditory, and textual inputs. Compared with existing benchmarks, our OmniEval has several distinctive features: (i) Full-modal collaboration: We design evaluation tasks that highlight the strong coupling between audio and video, requiring models to effectively leverage the collaborative perception of all modalities; (ii) Diversity of videos and tasks: OmniEval includes 1,000 audio-visual synchronized videos, with 307 Chinese videos and 558 English videos, systematically categorized into four major domains. (iii) Diversity and granularity of tasks: OmniEval contains 2783 question-answer pairs, comprising 1412 open-ended questions and 1371 multiple-choice questions. These questions are divided into four major task types and 12 subtask types to achieve comprehensive evaluation. Among them, we have introduced a more granular video localization task, which named as Grounding. Based on our OmniEval, we have extensively evaluated a variety of state-of-the-art models. The experimental results indicate that existing models face significant challenges in understanding the real world, with the best accuracy rate being only 10%. We hope that our OmniEval can provide a platform for evaluating the ability to construct and understand coherence from the context of all modalities."},"_bibtex":{"value":"@misc{\nzhang2025omnieval,\ntitle={OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs},\nauthor={Yiman Zhang and Ziheng Luo and Qiangyu YAN and Wei He and Borui Jiang and Xinghao Chen and Kai Han},\nyear={2025},\nurl={https://openreview.net/forum?id=B9iMn59jFE}\n}"},"title":{"value":"OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs"},"pdf":{"value":"/pdf/9debe9aea7038533619a28253d923ccec9e3b3fb.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"zhang|omnieval_a_benchmark_for_evaluating_omnimodal_models_with_visual_auditory_and_textual_inputs"},"authorids":{"value":["~Yiman_Zhang1","~Ziheng_Luo1","~Qiangyu_YAN1","~Wei_He10","~Borui_Jiang1","~Xinghao_Chen1","~Kai_Han2"]},"authors":{"value":["Yiman Zhang","Ziheng Luo","Qiangyu YAN","Wei He","Borui Jiang","Xinghao Chen","Kai Han"]}},"version":2},{"content":{"summary":{"value":"This paper presents a unified image and video model geared towards chat applications. Visual tokens from a frozen CLIP-backed vision \nencoder are dynamically merged using clustering at multiple-scales, and are then passed to a large language model after a simple projection layer. The projection layers are first fine-tuned on captioning datasets and later instruction-tuned along with the LLM on multimodal image and video datasets. Several experiments relating to image and video understanding, zero-shot question answering, object hallucination etc. along with some human evaluation results are presented and the proposed method outperforms similar sized video and image specific models."},"presentation":{"value":"3 good"},"contribution":{"value":"3 good"},"soundness":{"value":"3 good"},"strengths":{"value":"This paper is well-written and easy to follow. This paper demonstrates that a simple parameter-free clustering method can effectively sparsify the image and video features and such features can serve as strong inputs to LLM. Furthermore, the paper shows that they can jointly instruction-tune the LLM on both image and video datasets. Experimental results demonstrate superior performance against competing image and video based methods. The clustering visualizations show a clear semantic structure and the visualized chat applications in the appendix show a very good grounding."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The paper makes a strong claim that it is the first successful unified vision-language model that can outperform image and video models. However, there is broad work on such vision-language models [1, 2, 3, 4] that can do well on both image and video tasks. \n2. Only GPT based evaluations are considered making it difficult to situate with SOTA. \n- Image and video captioning results are not presented. See [1]\n- Standard image and video classification results such as on ImageNet and Kinetics are not presented. See [1] \n- GPT-based score for QA tasks such as MSRVTT, ActivityNet are presented making it difficult to compare with other methods.\n3. The approach of sparsifying transformer inputs is not novel [2, 5, 6, 7]. Note that [5] also uses clustering.\n\nRefs:\n[1]: Flamingo: a Visual Language Model for Few-Shot Learning\n[2]: MaMMUT: A Simple Architecture for Joint Learning for MultiModal Tasks\n[3]: Alternating Gradient Descent and Mixture-of-Experts for Integrated Multimodal Perception\n[4]: OmniVL: One Foundation Model for Image-Language and Video-Language Tasks\n[5]: Not All Tokens Are Equal: Human-centric Visual Analysis via Token Clustering Transformer\n[6]: Rethinking Video ViTs: Sparse Video Tubes for Joint Image and Video Learning\n[7]: DynamicViT: Efficient Vision Transformers with Dynamic Token Sparsification"},"confidence":{"value":"4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work."},"questions":{"value":"Is the token sparsification necessary for efficiency purpose or for performance?\nHow do the QA results compare with Flamingo?\nIs there a way to do captioning and compare ROUGE, CiDER scores?\nHow does the model do on ImageNet and Kinetics classification?"},"rating":{"value":"5: marginally below the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636154130,"tcdate":1698716330764,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission2201/Reviewer_ThZs"],"signatures":["ICLR.cc/2024/Conference/Submission2201/Reviewer_ThZs"],"forum":"fdfQYxMZEG","number":2,"license":"CC BY 4.0","cdate":1698716330764,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission2201/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636154130,"domain":"ICLR.cc/2024/Conference","replyto":"fdfQYxMZEG","id":"ZCsLVJj2rI","forumContent":{"venue":{"value":"ICLR 2024 Conference Withdrawn Submission"},"TLDR":{"value":"We introduce Chat-UniVi, a unified multimodal large language model that represents images and videos using a collection of dynamic visual tokens."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["vision and language","large language models","image and video understanding"]},"supplementary_material":{"value":"/attachment/3bcec8e6a32c2d07d5a1ec4850158851af7dfef0.zip"},"primary_area":{"value":"representation learning for computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Large language models have demonstrated impressive universal capabilities across a wide range of open-ended tasks and have extended their utility to encompass multimodal conversations. In this study, we introduce Chat-UniVi, a unified vision-language model capable of comprehending and engaging in conversations involving images and videos. Specifically, Chat-UniVi uniformly represents images and videos using a collection of dynamic visual tokens. This novel representation framework empowers the model to efficiently utilize a limited number of visual tokens to simultaneously capture the spatial details necessary for images and the comprehensive temporal relationship required for videos. Besides, we leverage a multi-scale representation that equips large language models to perceive both high-level semantic concepts and low-level visual details. More encouragingly, Chat-UniVi is trained on a mixed dataset containing both images and videos, making it directly applicable to tasks involving both mediums without the need for any modifications. Extensive experimental results demonstrate that Chat-UniVi, as a unified model, consistently surpasses even the existing methods exclusively designed for either images or videos. To the best of our knowledge, Chat-UniVi represents the first successful unified multimodal large language model that consistently outperforms both dedicated image and video models."},"_bibtex":{"value":"@misc{\njin2024chatunivi,\ntitle={Chat-UniVi: A Unified Vision-Language Model for Image and Video Understanding},\nauthor={Peng Jin and Ryuichi Takanobu and Cai Wan Zhang and Xiaochun Cao and Li Yuan},\nyear={2024},\nurl={https://openreview.net/forum?id=fdfQYxMZEG}\n}"},"title":{"value":"Chat-UniVi: A Unified Vision-Language Model for Image and Video Understanding"},"pdf":{"value":"/pdf/74f8e468ae179e68155db766bc1de7900003e85c.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Withdrawn_Submission"},"paperhash":{"value":"jin|chatunivi_a_unified_visionlanguage_model_for_image_and_video_understanding"},"authorids":{"value":["~Peng_Jin4","~Ryuichi_Takanobu1","~Cai_Wan_Zhang1","~Xiaochun_Cao3","~Li_Yuan2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Peng Jin","Ryuichi Takanobu","Cai Wan Zhang","Xiaochun Cao","Li Yuan"]}},"version":2},{"content":{"summary":{"value":"This paper presents DiffPhy, a framework aimed at improving the physical realism of text-to-video (T2V) diffusion models. While existing models excel at generating visually high-quality videos, they often ignore physical laws such as gravity, force, and motion consistency. DiffPhy addresses this gap by introducing a physics-aware fine-tuning paradigm that integrates reasoning from Large Language Models (LLMs) and Multimodal LLMs (MLLMs). The LLM first performs chain-of-thought reasoning on text prompts to infer relevant physical attributes, phenomena, and enhanced contextual descriptions. The MLLM then verifies whether the generated video aligns with these inferred rules, producing differentiable supervision signals that guide the diffusion model’s updates. The training combines three main objectives—physical phenomena loss, physical commonsense loss, and semantic consistency loss—and further employs an attention-injection mechanism to correct physically implausible generations.\n\nTo support this learning process, the authors construct a new dataset, HQ-Phy, consisting of roughly 8,000 real-world videos emphasizing physical interactions and realistic motion. Through extensive experiments on VideoPhy2 and PhyGenBench benchmarks, DiffPhy demonstrates measurable gains over several open and closed-source baselines, including Wan 2.1-14B, Kling, and CogVideoX. Both automated and human evaluations show that the proposed method produces videos with higher semantic alignment and stronger physical plausibility. Overall, the paper provides a technically coherent and empirically validated approach for enhancing diffusion-based video generation with explicit physical reasoning, offering a meaningful step toward bridging the gap between visual fidelity and physical correctness, though some improvements remain moderate in magnitude."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"N/A"},"rating":{"value":4},"details_of_ethics_concerns":{"value":"N/A"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":1},"strengths":{"value":"1. The paper is clearly written and logically organized. Each component of the proposed framework—LLM reasoning, MLLM verification, loss formulation, and failure-aware refinement—is explained in a step-by-step manner, supported by intuitive figures (e.g., Figure 2 and Figure 7). The motivation for combining symbolic reasoning with diffusion training is easy to follow, making the work accessible to both machine learning and vision audiences.\n2. By infusing physical rules into video diffusion, the paper addresses one of the most pressing limitations of current T2V systems—the lack of physical realism and commonsense consistency. This direction has high potential impact not only for video generation but also for downstream domains such as robotics simulation, digital content creation, and physics-based reasoning benchmarks. The work can be viewed as a meaningful step toward unifying generative AI with structured world modeling. Overall, the contribution is both timely and relevant to the evolving landscape of physics-informed generative models.\n3. The introduction of the HQ-Phy dataset represents an additional and meaningful contribution. By curating approximately 8 000 real-world videos covering diverse physical interactions—such as gravity-driven motion, collisions, and fluid dynamics—the authors address a key limitation of existing benchmarks, which are often synthetic or too small to support effective fine-tuning. HQ-Phy provides valuable training material for future research in physics-aware video generation and helps bridge the current data gap in real-world physical phenomena. This dataset substantially strengthens the paper’s practical impact and long-term significance to the community."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The discussion in Section 2 (Video Physics Reasoning) is relatively narrow. It mainly contrasts simulator-based and representation-learning approaches but overlooks a growing line of research that leverages post-training or preference optimization techniques to enhance physical reasoning in generative models. Recent works such as [1-3] are highly relevant and should be discussed. These studies demonstrate that physics awareness can also be introduced during post-training, offering an important comparative context for DiffPhy.\n2. The method is only evaluated on a single backbone (Wan 2.1-14B). Although this model is a strong open-source baseline, improvements demonstrated on one architecture do not fully establish the generality or robustness of the proposed framework. To make the empirical evidence more convincing, DiffPhy should be applied to additional diffusion backbones to verify that the approach generalizes across different architectures and data. \n3. In lines 221–223, the authors state that they decode the predicted latent clip $x_t\\in R^{m\\times c\\times h \\times w}$ at a sampled timestep $t$ into pixel space for MLLM evaluation. However, for larger $t$, these intermediate latents are typically highly noisy and lack meaningful semantic structure. It remains unclear how an MLLM can reliably evaluate alignment or physical correctness from such noisy decoded clips. Without additional clarification or ablation showing the sensitivity of the evaluation to timestep noise, this procedure may introduce unstable or unreliable supervision signals.\n4. This paper strongly depends on physical reasoning ability of LLM and MLLM. Therefore, the entire framework implicitly assumes that the underlying LLM and MLLM possess strong physical reasoning and evaluation capabilities. However, the paper does not include any diagnostic experiments or analysis to validate this assumption. If these models fail to accurately reason about or judge physical phenomena, the provided feedback could be misleading, thereby undermining the claimed improvements. A more systematic assessment of how the quality of the LLM/MLLM affects the overall performance would strengthen the paper’s empirical credibility.\n\n[1] Ipo: Iterative preference optimization for text-to-video generation\n\n[2] Pisa experiments: Exploring physics post-training for video diffusion models by watching stuff drop\n\n[3] RDPO: Real Data Preference Optimization for Physics Consistency Video Generation"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762918559428,"tcdate":1761192649320,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission6229/Reviewer_YtuE"],"signatures":["ICLR.cc/2026/Conference/Submission6229/Reviewer_YtuE"],"forum":"lPKsPBstHg","number":1,"license":"CC BY 4.0","cdate":1761192649320,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission6229/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762918559428,"domain":"ICLR.cc/2026/Conference","replyto":"lPKsPBstHg","id":"nXRE8iCc4Y","forumContent":{"TLDR":{"value":"We propose DiffPhy, a generic framework that enables physically-correct and semantically coherent video generation by fine-tuning a pre-trained video diffusion model."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Video Diffusion","Video Generation","Physical Commensense"]},"supplementary_material":{"value":"/attachment/dd73e8deab01feb18bdae06f4f6ecc4dca24d57c.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Recent video diffusion models have demonstrated their great capability in generating visually-pleasing results, while synthesizing the correct physical effects in generated videos remains challenging. The complexity of real-world motions, interactions, and dynamics introduce great difficulties when learning physics from data. In this work, we propose DiffPhy, a generic framework that enables physically-correct and photo-realistic video generation by fine-tuning a pre-trained video diffusion model. Our method leverages large language models (LLMs) to infer rich physical context from the text prompt. To incorporate this context into the video diffusion model, we use a multimodal large language model (MLLM) to verify intermediate latent variables against the inferred physical rules, guiding the model’s gradient updates accordingly. MLLM’s textual output is transformed into continuous signals. We then formulate a set of training objectives that jointly ensure physical accuracy and semantic alignment with the input text.  Additionally, failure facts of physical phenomena are corrected via attention injection. We also establish a high-quality physical video dataset containing diverse phyiscal actions and events to facilitate effective finetuning. Extensive experiments on public benchmarks demonstrate that DiffPhy is able to produce state-of-the-art results across diverse physics-related scenarios. Code and data will be made available post-review."},"_bibtex":{"value":"@misc{\nzhang2026think,\ntitle={Think Before You Diffuse: Infusing Physical Rules into Video Diffusion},\nauthor={Ke Zhang and Cihan Xiao and Jiacong Xu and Yiqun Mei and Vishal M. Patel},\nyear={2026},\nurl={https://openreview.net/forum?id=lPKsPBstHg}\n}"},"title":{"value":"Think Before You Diffuse: Infusing Physical Rules into Video Diffusion"},"pdf":{"value":"/pdf/abed66725f92f4403a0e6585ebdac7432376d20d.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"zhang|think_before_you_diffuse_infusing_physical_rules_into_video_diffusion"},"authorids":{"value":["~Ke_Zhang17","~Cihan_Xiao1","~Jiacong_Xu1","~Yiqun_Mei1","~Vishal_M._Patel1"]},"authors":{"value":["Ke Zhang","Cihan Xiao","Jiacong Xu","Yiqun Mei","Vishal M. Patel"]}},"version":2},{"content":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Spurious Correlations","Stochastic Gradient Descent (SGD)","Implicit Regularization"]},"primary_area":{"value":"learning theory"},"abstract":{"value":"Training with stochastic gradient descent (SGD) at moderately large learning rates has been observed to improve robustness against spurious correlations, strong correlation between non-predictive features and target labels. Yet, the mechanism underlying this effect remains unclear. In this work, we identify batch size as an additional critical factor and show that robustness gains arise from the implicit regularization of SGD, which intensifies with larger learning rates and smaller batch sizes. This implicit regularization reduces reliance on spurious or shortcut features, thereby enhancing robustness while preserving accuracy. Importantly, this effect appears unique to SGD: gradient descent (GD) does not confer the same benefit and may even exacerbate shortcut reliance. Theoretically, we establish this phenomenon in linear models by leveraging statistical formulations of spurious correlations, proving that SGD systematically suppresses spurious feature dependence. Empirically, we demonstrate that the effect extends to deep neural networks across multiple benchmarks. Our code is available at\n\\href{https://github.com/mirzanahal/sgd-implicit-regularization-shortcuts}{https://github.com/mirzanahal/sgd-implicit-regularization-shortcuts}."},"_bibtex":{"value":"@inproceedings{\nmirzaie2026implicit,\ntitle={Implicit Regularization of {SGD} Reduces Shortcut Learning},\nauthor={Nahal Mirzaie and Alireza Alipanah and Ali Abbasi and Amirmahdi Farzane and Hossein Jafarinia and Erfan Sobhaei and Mahdi Ghaznavi and Amir Najafi and Mahdieh Soleymani Baghshah and Mohammad Hossein Rohban},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=CPdAB7H8mU}\n}"},"title":{"value":"Implicit Regularization of SGD Reduces Shortcut Learning"},"pdf":{"value":"/pdf/3289eea3b6c492d78183b7689285f2dba88a122f.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"mirzaie|implicit_regularization_of_sgd_reduces_shortcut_learning"},"authorids":{"value":["~Nahal_Mirzaie1","~Alireza_Alipanah1","~Ali_Abbasi2","~Amirmahdi_Farzane2","~Hossein_Jafarinia1","~Erfan_Sobhaei1","~Mahdi_Ghaznavi1","~Amir_Najafi1","~Mahdieh_Soleymani_Baghshah1","~Mohammad_Hossein_Rohban1"]},"authors":{"value":["Nahal Mirzaie","Alireza Alipanah","Ali Abbasi","Amirmahdi Farzane","Hossein Jafarinia","Erfan Sobhaei","Mahdi Ghaznavi","Amir Najafi","Mahdieh Soleymani Baghshah","Mohammad Hossein Rohban"]}},"tmdate":1779006880368,"pdate":1769436657086,"tcdate":1758349948137,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission23895/Authors"],"signatures":["ICLR.cc/2026/Conference/Submission23895/Authors"],"forum":"CPdAB7H8mU","license":"CC BY 4.0","number":23895,"cdate":1758349948137,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/-/Submission","ICLR.cc/2026/Conference/-/Post_Submission","ICLR.cc/2026/Conference/Submission23895/-/Full_Submission","ICLR.cc/2026/Conference/Submission23895/-/Rebuttal_Revision","ICLR.cc/2026/Conference/-/Edit","ICLR.cc/2026/Conference/Submission23895/-/Camera_Ready_Revision"],"mdate":1779006880368,"odate":1759896705795,"domain":"ICLR.cc/2026/Conference","id":"CPdAB7H8mU","version":2},{"content":{"summary":{"value":"The paper proposes GLaVE-Cap, a novel global-local aligned video captioning framework integrating vision experts to improve fine-grained and contextually consistent video captions. It introduces two key modules:\n- TrackFusion integrates Grounding DINO and SAM 2 for cross-frame object tracking and uses a dual-stream architecture (differential+detailed captions) to capture both static and dynamic details.\n- CaptionBridge injects a global overview caption to guide local captioning and performs adaptive scene-level summarization for coherent video-level captions.\n\nAdditionally, the authors introduce:\n- GLaVE-Bench, a new benchmark with ~5x more queries per video than existing datasets.\n- GLaVE-1.2M, a large-scale dataset (16K videos, 1.2M QA pairs).\n\nExperiments on GLaVE-Bench, Video-MME, VidCapBench, MVBench show state-of-the-art (SOTA) results. A student model (GLaVE-7B) fine-tuned on GLaVE-1.2M also outperforms baselines, validating dataset quality."},"soundness":{"value":4},"confidence":{"value":4},"questions":{"value":"1. How is \"adaptive scene segmentation\" quantitatively validated beyond accuracy gains---are the segmentation boundaries human-verified?\n\n2. Were hallucination rates or factual alignment metrics (e.g., VideoHallucer subset) computed per-stage to confirm reduced propagation?\n\n3. Can the dual-stream design be generalized to text-only summarization tasks?\n\n4. How much of the performance gain remains if using smaller backbones ($\\leq$7B), without GPT-4o supervision?\n\n5. How does GLaVE-Cap handle videos with dense camera motion or overlapping actions (e.g., sports broadcasts)?\n\n6. Finally, I suggest resolving my concerns in Weaknesses."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":4},"strengths":{"value":"1. The paper has conceptual novelty by moving beyond the local-to-global paradigm by explicitly modeling bidirectional local-global interaction.\n\n2. Integration of vision experts:\nClever combination of Grounding DINO+SAM2 for robust object tracking, plus dual-stream design effectively captures both dynamic and static details.\n\n3. The paper proposes new benchmark and dataset.\nGLaVE-Bench and GLaVE-1.2M fill an evaluation and data gap in fine-grained video captioning. \nThe QA-per-video (118 vs. 16) and multi-scene structure make it far more comprehensive.\n\n4. Evaluations on 4+ major benchmarks with both closed (GPT-4o) and open (Qwen2.5-VL-72B) models show consistent SOTA performance.\n\n5. Detailed ablations quantify the contribution of each module. Appendix provides thorough prompt designs and reproducibility materials.\n\n6. Both quantitative and qualitative results (including user study in video generation) convincingly demonstrate improved caption granularity and contextual consistency."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. **Limited novelty in components:**\nThough integration is well-engineered, both TrackFusion (based on Set-of-Mark+vision experts) and CaptionBridge (context injection+summarization) build on known elements rather than introducing fundamentally new algorithms.\n\n2. **Dependence on powerful LLMs:**\nMost evaluations rely on GPT-4o and Qwen2.5-VL-72B. It's unclear how much gain comes from the model scale versus method design. Smaller models (e.g., Qwen2.5-VL-7B) perform notably worse.\n\n3. Since GLaVE-Bench and GLaVE-1.2M are built using GLaVE-Cap captions, potential data leakage or alignment bias could inflate results, even if mitigated by re-captioning and human verification.\n\n4. The framework requires multiple calls to large VLMs and expert models (Grounding DINO, SAM2), making it computationally heavy and less practical for large-scale or real-time applications.\n\n5. While video generation results are interesting, the paper could better relate fine-grained caption quality to downstream generation metrics quantitatively."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762918035752,"tcdate":1761934730872,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission5394/Reviewer_626Y"],"signatures":["ICLR.cc/2026/Conference/Submission5394/Reviewer_626Y"],"forum":"EFOMez9cBW","number":4,"license":"CC BY 4.0","cdate":1761934730872,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission5394/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762918035752,"domain":"ICLR.cc/2026/Conference","replyto":"EFOMez9cBW","id":"84EeRO3Psu","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Video Captioning","Fine-Grained Video Description","Global-Local Alignment","Vision Experts","Video Benchmark","Video QA Dataset","Video-LLM"]},"supplementary_material":{"value":"/attachment/52afa223b0d290f940715ab012a1de0251cb3ba2.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Video detailed captioning aims to generate comprehensive video descriptions to facilitate video understanding. Recently, most efforts in the video detailed captioning community have been made towards a local-to-global paradigm, which first generates local captions from video clips and then summarizes them into a global caption. However, we find this paradigm leads to less detailed and contextual-inconsistent captions, which can be attributed to (1) no mechanism to ensure fine-grained captions, and (2) weak interaction between local and global caption generation.\nTo remedy the above two issues, we propose **GLaVE-Cap**, a **G**lobal-**L**ocal **a**ligned framework with **V**ision **E**xpert integration for **Cap**tioning, which consists of two core modules: TrackFusion enables comprehensive local caption generation, by leveraging vision experts to acquire cross-frame visual prompts, coupled with a dual-stream structure; while CaptionBridge establishes a local-global interaction, by using global context to guide local captioning, and adaptively summarizing local captions into a coherent global caption. Besides, we construct **GLaVE-Bench**, a comprehensive video captioning benchmark featuring 5$\\times$ more queries per video than existing benchmarks, covering diverse visual dimensions to facilitate reliable evaluation.\nWe further provide a training dataset **GLaVE-1.2M** containing 16K high-quality fine-grained video captions and 1.2M related question-answer pairs. Extensive experiments on four benchmarks show that our GLaVE-Cap achieves state-of-the-art performance. Besides, the ablation studies and student model analyses further validate the effectiveness of the proposed modules and the contribution of GLaVE-1.2M to the video understanding community. The source code, model weights, benchmark, and dataset will be open-sourced."},"_bibtex":{"value":"@misc{\nxu2026glavecap,\ntitle={{GL}a{VE}-Cap: Global-Local Aligned Video Captioning with Vision Expert Integration},\nauthor={Wan Xu and Feng Zhu and Yihan Zeng and Yuanfan Guo and Ming Liu and Hang Xu and Wangmeng Zuo},\nyear={2026},\nurl={https://openreview.net/forum?id=EFOMez9cBW}\n}"},"title":{"value":"GLaVE-Cap: Global-Local Aligned Video Captioning with Vision Expert Integration"},"pdf":{"value":"/pdf/fb38182042babe0bc8c48f7c9759cb3d36e68798.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"xu|glavecap_globallocal_aligned_video_captioning_with_vision_expert_integration"},"authorids":{"value":["~Wan_Xu1","~Feng_Zhu17","~Yihan_Zeng1","~Yuanfan_Guo1","~Ming_Liu10","~Hang_Xu1","~Wangmeng_Zuo3"]},"authors":{"value":["Wan Xu","Feng Zhu","Yihan Zeng","Yuanfan Guo","Ming Liu","Hang Xu","Wangmeng Zuo"]}},"version":2},{"content":{"summary":{"value":"The paper proposed a safety aware video understanding benchmark, including 2M human verified videos. The unsafe scenarios are separated into 6 classes. The authors also design a pipeline for automatically data generation. For the video understanding model, the authors propose the Parallel Equivalent Policy Encoding and Policy-Aware Adaptive Pruning to encode the Safety Policy Guidelines and reduce the redundancy. The result of the trained model is good comparing to both close- and open- source models."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. The author mentioned that Human Verification is used for curating the 2M video benchmark. Can the authors make a detailed description of the Human Verification procedure?\n\n2. More procedure should be enclosed in the appendix."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The proposed dataset is novel and the data curation procedure is well-organized.\n2. The proposed Parallel Equivalent Policy Encoding and Policy-Aware Adaptive Pruning can effectively encode the Safety Policy Guidelines and reduce the redundancy of video tokens.\n3. The results are good."},"flag_for_ethics_review":{"value":["Yes, Privacy, security and safety","Yes, Potentially harmful insights, methodologies and applications","Yes, Responsible research practice (e.g., human subjects, data release)"]},"weaknesses":{"value":"1. Dataset Construction: \n       (1) The release and separation of the dataset is a concern. The author can only provide links to publicly available sources and annotations. But it is common that the link may fail.I understand that this is something unavoidable, but it will undoubtedly reduce the frequency of use and impact of this dataset.\n       (2) 2M videos for human verification is a huge effort, the authors don't provide any details of the procedure.\n\n2. Model training：\n      (1) Some of the training procedure is ambiguous, there should be more details about Preference Post-tuning procedure."}},"nonreaders":[],"tmdate":1731897265384,"tcdate":1730137369995,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission9949/Reviewer_PM58"],"signatures":["ICLR.cc/2025/Conference/Submission9949/Reviewer_PM58"],"forum":"xjKz6IxgCX","number":1,"license":"CC BY 4.0","cdate":1730137369995,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission9949/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731897265384,"domain":"ICLR.cc/2025/Conference","replyto":"xjKz6IxgCX","id":"oKsNtQMQXj","forumContent":{"TLDR":{"value":"We propose an efficient MLLM-based video guardrail model and a large-scale video safety benchmark dataset to enforce safety policy compliance and transparent explanations."},"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video Guardrail Model","Safe Foundation Models","Efficient LLMs Inference","LLM Safety","Multimodal Foundation Models"]},"primary_area":{"value":"alignment, fairness, safety, privacy, and societal considerations"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"With the rise of generative AI and rapid growth of high-quality video generation, video guardrails have become more crucial than ever to ensure safety and security across platforms. Current video guardrails, however, are either overly simplistic, relying on pure classification models trained on simple policies with limited unsafe categories, which lack detailed explanations, or prompting multimodal large language models (MLLMs) with long safety guidelines, which are inefficient and impractical for guardrailing real-world content. To bridge this gap, we propose SafeWatch, an efficient MLLM-based video guardrail model designed to follow customized safety policies and provide multi-label video guardrail outputs with content-specific explanations in a zero-shot manner. In particular, unlike traditional MLLM-based guardrails that encode all safety policies autoregressively, causing inefficiency and bias, SafeWatch uniquely encodes each policy chunk in parallel and eliminates their position bias such that all policies are attended simultaneously with equal importance. In addition, to improve efficiency and accuracy, SafeWatch incorporates a policy-aware visual token pruning algorithm that adaptively selects the most relevant video tokens for each policy, discarding noisy or irrelevant information. This allows for more focused, policy-compliant guardrail with significantly reduced computational overhead. Considering the limitations of existing video guardrail benchmarks, we propose SafeWatch-Bench, a large-scale video guardrail benchmark comprising over 2M videos spanning six safety categories which covers over 30 tasks to ensure a comprehensive coverage of all potential safety scenarios. We have conducted extensive experiments, showing that SafeWatch outperforms all SOTA video guardrails on SafeWatch-Bench by 28.2%, and achieves a 13.6% improvement on existing benchmarks, all while reducing inference costs by an average of 10%. SafeWatch also demonstrates strong policy-following abilities and outperforms previous SOTAs by 5.6% and 15.6% in zero-shot generalizability to new policies and new prompting tasks. Additionally, both LLM-as-a-judge and human evaluators confirm the high quality of the explanations provided by SafeWatch. Our project is open-sourced at https://safewatch-aiguard.github.io."},"_bibtex":{"value":"@inproceedings{\nchen2025safewatch,\ntitle={SafeWatch: An Efficient Safety-Policy Following Video Guardrail Model with Transparent Explanations},\nauthor={Zhaorun Chen and Francesco Pinto and Minzhou Pan and Bo Li},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=xjKz6IxgCX}\n}"},"title":{"value":"SafeWatch: An Efficient Safety-Policy Following Video Guardrail Model with Transparent Explanations"},"pdf":{"value":"/pdf/bebab0da9d02c81b7ccbdfc1c7db6ab2f3a21c53.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"chen|safewatch_an_efficient_safetypolicy_following_video_guardrail_model_with_transparent_explanations"},"authorids":{"value":["~Zhaorun_Chen1","~Francesco_Pinto1","~Minzhou_Pan1","~Bo_Li19"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Zhaorun Chen","Francesco Pinto","Minzhou Pan","Bo Li"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a new image to video adaptation method. First, it generates more realistic videos to mitigate the modality gap between source images and target videos. Then, to mitigate the gap between the flows extracted from the generated videos and target videos, it then proposes the category-aware flow memory bank, by replacing the optical flow in a generated source video with the optical flow selected from target videos, and then generate new composed videos for training. They also leverage the video pace prediction task to enhance the model's perception of speed. The proposed method achieves state-of-the-art performance in the field of image to video adaptation."},"suitability":{"value":3},"strengths":{"value":"1. The proposed several approaches, such as the diverse camera movements simulation and the category-aware flow memory bank, are effective for unsupervised image to video adaption with good illustrations.\n2. The proposed method achieves superior performance in the field of unsupervised image to video adaption. \n3. The overall structure of this paper is well organized."},"confidence":{"value":1},"rating":{"value":4},"limitations":{"value":"1. Is it possible for the authors to report the performance of \"SO [51] + flow\" and \"SO (RGB) + CFB\" to better demonstrate the effectiveness of the proposed category-aware flow memory bank? \n2. It seems that the generated videos with 3D scenes are only used during the warm-up epochs if the category-aware flow memory bank will be used. If we could approach other pretrained image to video adaption models, are there any ways to directly use a category-aware flow memory bank to generate videos for training?"}},"nonreaders":[],"tmdate":1721657247305,"tcdate":1716220454332,"writers":["acmmm.org/ACMMM/2024/Conference","acmmm.org/ACMMM/2024/Conference/Submission2362/Reviewer_ydNo"],"signatures":["acmmm.org/ACMMM/2024/Conference/Submission2362/Reviewer_ydNo"],"forum":"7uRTPwyUum","number":2,"license":"CC BY 4.0","cdate":1716220454332,"readers":["everyone"],"invitations":["acmmm.org/ACMMM/2024/Conference/Submission2362/-/Official_Review","acmmm.org/ACMMM/2024/Conference/-/Edit","acmmm.org/ACMMM/2024/Conference/Submission2362/Official_Review2/-/Final_Rating"],"mdate":1721657247305,"domain":"acmmm.org/ACMMM/2024/Conference","replyto":"7uRTPwyUum","id":"AHL0aFEw2i","forumContent":{"venue":{"value":"MM2024 Poster"},"supplementary_material":{"value":"/attachment/08aa6673c17dc168c4e6e03cf30b88cb6a02f17f.zip"},"abstract":{"value":"Image-to-Video adaptation is proposed to train a model using labeled images and unlabeled videos to facilitate the classification of unlabeled videos.\nThe latest work synthesizes videos using still images to mitigate the modality gap between images and videos. However, the synthesized videos are not realistic due to the camera movements are only simulated in 2D space. Therefore, we generate realistic videos by simulating arbitrary camera movements in 3D scenes, and then the model can be trained using the generated source videos.\nUnfortunately, the optical flows from the generated videos have unexpected negative impacts, resulting in suboptimal performance. To address this issue, we propose the Category-aware Flow Memory Bank, which replaces optical flows in source videos with real target flows, and the new composed videos are beneficial for training.\nIn addition, we leverage the video pace prediction task to enhance the speed awareness of the model in order to solve the problem that the model performs poorly in handling some categories with similar appearances but significant speed differences. Our method achieves state-of-the-art performance and comparable performance on three Image-to-Video benchmarks."},"relevance_to_conference":{"value":"Training a video recognition model from scratch is costly and time comsuming. However,there are  numerous image datasets accessible and images are easier to be collected and annotated than videos. Therefore, Image-to-Video adaptation becomes an active research direction in the field of multimedia. We propose a novel method which can be used to train a video recognition model using labeled images and unlabeled videos. And the trained model can be used to classify unlabeled videos effectively. Specifically, we simulate the movements of the camera in 3D space and save new perspective images as video frames. In this way, the generated video are more realistic and beneficial for training a discriminative spatio-temporal model. In order to mitigate the significant discrepancies between the flow data in source and target videos, we construct a Category-aware Flow Memory Bank and replace the optical flows in source videos with real target flows under the guidance of pseudo labels. The new composed videos greatly improve the performance of the model. Finally, we leverage the video pace prediction task to enhance the speed awareness of the model in order to solve the problem that the model performs poorly in handling some categories with similar appearances but significant speed differences. We have validated the effectiveness of our method through extensive experiments.In particular, we achieved state-of-the-art performance on the challenging task E→H."},"_bibtex":{"value":"@inproceedings{\nhuang2024unsupervised,\ntitle={Unsupervised Image-to-Video Adaptation via Category-aware Flow Memory Bank and Realistic Video Generation},\nauthor={Kenan Huang and Junbao Zhuo and Shuhui Wang and Chi Su and Qingming Huang and Huimin Ma},\nbooktitle={ACM Multimedia 2024},\nyear={2024},\nurl={https://openreview.net/forum?id=7uRTPwyUum}\n}"},"title":{"value":"Unsupervised Image-to-Video Adaptation via Category-aware Flow Memory Bank and Realistic Video Generation"},"pdf":{"value":"/pdf/496dd3f41c18bce24571849b4fb69ff6342dd45f.pdf"},"venueid":{"value":"acmmm.org/ACMMM/2024/Conference"},"paperhash":{"value":"huang|unsupervised_imagetovideo_adaptation_via_categoryaware_flow_memory_bank_and_realistic_video_generation"},"primary_subject_area":{"value":"[Content] Media Interpretation"},"authorids":{"value":["~Kenan_Huang1","~Junbao_Zhuo1","~Shuhui_Wang1","~Chi_Su2","~Qingming_Huang2","~Huimin_Ma1"]},"authors":{"value":["Kenan Huang","Junbao Zhuo","Shuhui Wang","Chi Su","Qingming Huang","Huimin Ma"]}},"version":2},{"content":{"summary":{"value":"The paper presents the SEINE model, for generating \"story-level\" long videos from short clips. It introduces a unique problem in generative transition and prediction. Using a random-mask video diffusion based on textual descriptions, the model shows smooth transitions between scenes. To evaluate its efficacy, the authors provide three new criteria: temporal consistency, semantic similarity, and video-text alignment. Results show its potential for generating coherent long videos."},"presentation":{"value":"3 good"},"contribution":{"value":"3 good"},"soundness":{"value":"3 good"},"strengths":{"value":"- The method of using masks was proposed in [1] and [2], but as far as I know, this is the first time it has been used in video transition. It could be novel.\n\n- The proposed method shows better performance on the metric compared to the baseline.\n\n- The proposed method can be applied in various areas such as long video generation and image-to-video animation.\n\n\nReferences\n\n[1] Voleti, Vikram, Alexia Jolicoeur-Martineau, and Chris Pal. \"MCVD-masked conditional video diffusion for prediction, generation, and interpolation.\" Advances in Neural Information Processing Systems 35 (2022): 23371-23385.\n\n[2] Fu, Tsu-Jui, et al. \"Tell me what happened: Unifying text-guided video completion via multimodal masked video generation.\" Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2023."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- [Major] My main concern is that there is not enough quantitative evaluation of video transitions. This paper conducted quantitative experiments by randomly selecting one caption from MSRVTT and determining CLIP-TEXT. However, since video transitions occur when scenes change, it does not seem appropriate to evaluate video semantic correlation. Also, no video quality evaluation metrics (such as FVD etc.) have been considered. This makes it difficult to quantify the exact quality of generation.\n\n- [Major] Several details related to the human evaluation are missing. (such as number of frames in the generated video, the dataset used, and the questions posed in the user study.)\nWas the user study appropriately reflective of temporal coherence, text-video alignment, and semantic similarity?\n\n- [Minor] For transitions to be applied in the real-world, it would require generating more than 16 frames. Would the quality be maintained if more frames are generated?\n\n- [Minor] In Figure 5 related to video transition, the frame numbers and details are omitted."},"confidence":{"value":"3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked."},"questions":{"value":"How long does the inference take? Is it capable of handling transitions with multiple objects across more than two scenes?"},"rating":{"value":"5: marginally below the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636813292,"tcdate":1698848627748,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission6964/Reviewer_GChA"],"signatures":["ICLR.cc/2024/Conference/Submission6964/Reviewer_GChA"],"forum":"FNq3nIvP4F","number":4,"license":"CC BY 4.0","cdate":1698848627748,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission6964/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636813292,"domain":"ICLR.cc/2024/Conference","replyto":"FNq3nIvP4F","id":"Uik1TO8KOa","forumContent":{"venue":{"value":"ICLR 2024 poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["generative model; video generation; diffusion model"]},"supplementary_material":{"value":"/attachment/0aca679cba688b36d8b9e374ec4f09022aa9a1fa.zip"},"primary_area":{"value":"generative models"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Recently video generation has achieved substantial progress with realistic results. Nevertheless, existing AI-generated videos are usually very short clips (\"shot-level'') depicting a single scene. To deliver a coherent long video (\"story-level''), it is desirable to have creative transition and prediction effects across different clips. This paper presents a short-to-long video diffusion model, SEINE, that focuses on generative transition and prediction. The goal is to generate high-quality long videos with smooth and creative transitions between scenes and varying lengths of shot-level videos. Specifically, we propose a random-mask video diffusion model to automatically generate transitions based on textual descriptions. By providing the images of different scenes as inputs, combined with text-based control, our model generates transition videos that ensure coherence and visual quality. Furthermore, the model can be readily extended to various tasks such as image-to-video animation and autoregressive video prediction. To conduct a comprehensive evaluation of this new generative task, we propose three assessing criteria for smooth and creative transition: temporal consistency, semantic similarity, and video-text semantic alignment. Extensive experiments validate the effectiveness of our approach over existing methods for generative transition and prediction, enabling the creation of story-level long videos."},"_bibtex":{"value":"@inproceedings{\nchen2024seine,\ntitle={{SEINE}: Short-to-Long Video Diffusion Model for Generative Transition and Prediction},\nauthor={Xinyuan Chen and Yaohui Wang and Lingjun Zhang and Shaobin Zhuang and Xin Ma and Jiashuo Yu and Yali Wang and Dahua Lin and Yu Qiao and Ziwei Liu},\nbooktitle={The Twelfth International Conference on Learning Representations},\nyear={2024},\nurl={https://openreview.net/forum?id=FNq3nIvP4F}\n}"},"title":{"value":"SEINE: Short-to-Long Video Diffusion Model for Generative Transition and Prediction"},"pdf":{"value":"/pdf/3f1069641eb7f0b47e27967faa59a2b59b6097bb.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference"},"paperhash":{"value":"chen|seine_shorttolong_video_diffusion_model_for_generative_transition_and_prediction"},"authorids":{"value":["~Xinyuan_Chen1","~Yaohui_Wang1","~Lingjun_Zhang1","~Shaobin_Zhuang1","~Xin_Ma3","~Jiashuo_Yu1","~Yali_Wang1","~Dahua_Lin1","~Yu_Qiao1","~Ziwei_Liu1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Xinyuan Chen","Yaohui Wang","Lingjun Zhang","Shaobin Zhuang","Xin Ma","Jiashuo Yu","Yali Wang","Dahua Lin","Yu Qiao","Ziwei Liu"]}},"version":2},{"content":{"summary":{"value":"In this paper, the authors propose a framework called VideoDiT for extending a text-to-image diffusion model into text-to-video model without much new learnable parameters. In particular, this work has two main contributions. One is called DP-VAE, which extends a 2D image VAE into a 3D video VAE. The main difference from previous video VAE is that it tries to align the latent distribution of video latent with image latent, so that we can seamlessly leverage the image pretrain on video training. The second contribution is on how to extend 2D DiT-based image diffusion model into 3D video modeling. Experiments are conducted on both VAE reconstruction evaluation and video generation quality comparison."},"soundness":{"value":2},"confidence":{"value":5},"questions":{"value":"1. The authors mentioned that the video encoding is decomposed into key frame and residuals. What's the meaning of this? If I have a video of 10 frames and regard the first frame as key frame, how do we derive residual for the 2-10 frames? By substracting each frame with the first frame? If the authors just input x_1 to image VAE and x_2, ..., x_10 to the video VAE, then why do the authors call it as residual? Is it simply because the encoded latent is added to the image latent?\n2. How is the key frame decided at the training and inference stage? Is it always the first frame? If not, does this mean we have to additionally choose the keyframe before auto-encoding.\n3. I'm still quite confused by the compression ratio of the proposed DP-VAE. The authors claim to have temporal compression ratio of 4, but this is without the considering of key frame. For example, if I have 13-frame video, and the first frame is regarded as keyframe. The actual compression ratio is 1+ (13 -1 ) // 4 = 4, which is less than 4, and heavily depends on the encoded video sequence length."},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. This paper focuses on an important research problem, which has potential benefit to the downstream applications. Since the image/video generation will all need a good video tokenizer.\n2. The intuition of this work makes sense to me, where it's better to train a video generation model based on a pretrained image model, so that we can maximize the utilization of all existing research practices in building strong image generation models. \n3. The experiments are conducted on both auto-encoding reconstruction and generation, which can reflect both the reconstruction quality and whether it can indeed boost the generation process, which is quite good."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The method illustration is quite unclear to me, where I'm confused by many parts of the approach. For example, how do we choose the key frame? Is the key frame always the first frame? What if there're several scene changes, so that the encoded image latent is largely different from the video latent. Why is the compression ratio 4? More questions are elaborated in the next section. Please provide a more detailed explanation or diagram of the key frame selection process and how it handles different scenarios like scene changes. Additionally, please clarify on how the compression ratio is determined and applied consistently across different video lengths.\n2. I would suggest the authors to also include the inference speed (wall-clock time) and memory cost on encoding-decoding a video. Since the proposed framework actually has more FLOPs - one image VAE + one video VAE - I think the inference speed might be slow and memory cost would be higher compared to baselines. Please provide a table comparing inference speed and memory usage for the proposed method versus baselines on standardized video lengths and resolutions.\n3. I would suggest the authors to highlight the compression ratio for baselines and the proposed video VAE. Since it's unclear to me whether all methods have the same compression ratio, hence whether this is a fair comparison. Please add a column in the comparison tables that explicitly states the compression ratio for each method, or to clarify in the text how you ensured comparable compression ratios across methods.\n4. Novelty. The DP-VAE part is generally OK to me, conditioned on the authors have clarified my unclear points. However, the extension from image DiT to video DiT has no clear novelty to me. I think it's already a quite common practice for all the recent video DiT generators. They would even incorporate better position embedding extension like 2D RoPE -> 3D RoPE, so the section 3.2 is actually of no novelty to me. When changing from image to video under the DiT architecture, except for the position embedding change, it's always the full attention and I see nothing new.\n5. The generation results still have motion blur and not that high-quality to me.\n6. I'm wondering whether this method performs well on higher compression ratio. Since 4x8x8 compression ratio is still too low compared to current SOTA video generators (e.g., Meta MovieGen uses 8x8x8), the sequence length would be unaffordable long. Please include an experiment or discussion on how your method performs with higher compression ratios, specifically comparing to the 8x8x8 ratio."}},"nonreaders":[],"tmdate":1731427236607,"tcdate":1729734536983,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission325/Reviewer_xV9d"],"signatures":["ICLR.cc/2025/Conference/Submission325/Reviewer_xV9d"],"forum":"lvgsPjRtLM","number":1,"license":"CC BY 4.0","cdate":1729734536983,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission325/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427236607,"domain":"ICLR.cc/2025/Conference","replyto":"lvgsPjRtLM","id":"6kO5kSfUtH","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"TLDR":{"value":"We introduce VideoDiT, a framework that integrates a Distribution-Preserving VAE and 3D Diffusion Transformers into pre-trained T2I models, enabling efficient joint image-video training and high-quality synthesis with minimal additional parameters."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Text-to-Video Generation","Diffusion Models","Image Diffusion Transformer"]},"supplementary_material":{"value":"/attachment/86b735f0f961c6a2b5ddd7953f5e6ecbc841435f.zip"},"primary_area":{"value":"generative models"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"We present VideoDiT, a streamlined video generation framework adapted from pre-trained image generation models. Unlike previous methods that simply add temporal layers to image diffusion models, we enhance both the tokenizer, implemented with the variational autoencoder (VAE), and the diffusion model. We emphasize the importance of combining 3D VAE compression with knowledge from pre-trained image diffusion models to achieve efficient video generation, though the tight coupling between image diffusion models and 2D VAEs poses significant challenges. To address this, we introduce the Distribution-Preserving VAE (DP-VAE), which encodes key frames in a video clip using the original 2D VAE while compressing non-key frames with a 3D VAE for spatiotemporal modeling. A regularization term ensures alignment between the 3D video latent space and the 2D image latent space, facilitating seamless transfer of pre-trained diffusion models. Leveraging the Diffusion Image Transformers (DiT) architecture and incorporating 3D positional embeddings, we extend 2D attention into 3D with negligible increased parameters. Furthermore, leveraging our proposed DP-VAE, VideoDiT supports joint image-video training, preserving the spatial modeling capabilities of the base model while excelling in both image and video generation. Extensive experiments validate the effectiveness of our approach."},"_bibtex":{"value":"@misc{\nfeng2025videodit,\ntitle={VideoDiT: Bridging Image Diffusion Transformers for Streamlined Video Generation},\nauthor={Ruoyu Feng and Tiankai Hang and Tianyu He and Kai Qiu and Qi Dai and Jianmin Bao and Zhibo Chen and Chong Luo},\nyear={2025},\nurl={https://openreview.net/forum?id=lvgsPjRtLM}\n}"},"title":{"value":"VideoDiT: Bridging Image Diffusion Transformers for Streamlined Video Generation"},"pdf":{"value":"/pdf/6049f25acd7425df8cc44ac9ac4ba2146233d536.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"feng|videodit_bridging_image_diffusion_transformers_for_streamlined_video_generation"},"authorids":{"value":["~Ruoyu_Feng1","~Tiankai_Hang1","~Tianyu_He1","~Kai_Qiu1","~Qi_Dai4","~Jianmin_Bao1","~Zhibo_Chen1","~Chong_Luo1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Ruoyu Feng","Tiankai Hang","Tianyu He","Kai Qiu","Qi Dai","Jianmin Bao","Zhibo Chen","Chong Luo"]}},"version":2},{"content":{"summary":{"value":"This paper proposes Video-MTR, a reinforced multi-turn reasoning framework designed to enable iterative key video segment selection and question comprehension. Video-MTR performs reasoning in multiple turns, selecting video segments progressively based on the evolving understanding of previously processed segments and the current question. Extensive experiments on benchmarks like VideoMME, MLVU, LVBench, and EgoSchema demonstrate that VideoMTR outperforms existing methods in both accuracy and efficiency, advancing the state-of-the-art in long video understanding."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"Please reply to Weaknesses."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"1. As claimed by the authors, this could be the first attempt to incorporate multi-turn reasoning in the context of long video understanding.\n2. The proposed method can adaptively select important frames for the question.\n3. Experiments on multiple long-video benchmarks show that Video-MTR outperforms its backbone Qwen2.5-VL with the same number of frames."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Qwen2.5-VL can support up to 768 frames and outperforms the proposed Video-MTR with the input of 64 frames. More experiments should be conducted to investigate whether Video-MTR can outperform Qwen2.5-VL with more frames.\n2. It's weired that the QA accuracy in Figure 4 exceed 1. Some explanations should be provided.\n3. The information of compared baseline models is not given, especially their backbone models.\n4. While introducing multi-turn reasoning, the efficiency compared to baselines should be analyzed.\n5. The training framework is complicated with various tricks, which may be unstable in other scenarios."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762920289105,"tcdate":1761931773311,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission8384/Reviewer_R4Fn"],"signatures":["ICLR.cc/2026/Conference/Submission8384/Reviewer_R4Fn"],"forum":"7AXP2RYw2N","number":2,"license":"CC BY 4.0","cdate":1761931773311,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission8384/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762920289105,"domain":"ICLR.cc/2026/Conference","replyto":"7AXP2RYw2N","id":"EamHWpfgHe","forumContent":{"TLDR":{"value":"leveraging end-to-end RL to enable MLLMs to perform multi-turn reasoning."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Long-form video understanding;MLLM; multi-turn reasoning"]},"supplementary_material":{"value":"/attachment/e2d1f3a2047d795ea554b32aac830167bd6226b1.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Long-form video understanding, characterized by long-range temporal dependencies and multiple events, remains a challenge. Existing methods often rely on static reasoning or external visual-language models (VLMs), which face issues like complexity and sub-optimal performance due to the lack of end-to-end training. In this paper, we propose Video-MTR, a reinforced multi-turn reasoning framework designed to enable iterative key video segment selection and question comprehension. Unlike traditional video reasoning pipeline, which generate predictions in a single turn, Video-MTR performs reasoning in multiple turns, selecting video segments progressively based on the evolving understanding of previously processed segments and the current question. This iterative process allows for a more refined and contextually aware analysis of the video. To ensure intermediate reasoning process, we introduce a novel gated bi-level reward system, combining trajectory-level rewards based on answer correctness and turn-level rewards emphasizing frame-query relevance. This system optimizes both video segment selection and question comprehension, eliminating the need for external VLMs and allowing end-to-end training. Extensive experiments on benchmarks like VideoMME, MLVU, and EgoSchema demonstrate that Video-MTR outperforms existing methods in both accuracy and efficiency, advancing the state-of-the-art in long-form video understanding."},"_bibtex":{"value":"@misc{\nxie2026videomtr,\ntitle={Video-{MTR}: Reinforced Multi-Turn Reasoning for Long Video Understanding},\nauthor={Yuan Xie and Tianshui Chen and Zheng Ge and Lionel Ni},\nyear={2026},\nurl={https://openreview.net/forum?id=7AXP2RYw2N}\n}"},"title":{"value":"Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding"},"pdf":{"value":"/pdf/19355b43c7bf6fb2e7aa7c2b0821d95b89f53837.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"xie|videomtr_reinforced_multiturn_reasoning_for_long_video_understanding"},"authorids":{"value":["~Yuan_Xie7","~Tianshui_Chen1","~Zheng_Ge1","~Lionel_Ni1"]},"authors":{"value":["Yuan Xie","Tianshui Chen","Zheng Ge","Lionel Ni"]}},"version":2},{"content":{"venue":{"value":"AAAI 2026"},"pdf":{"value":"https://ojs.aaai.org/index.php/AAAI/article/download/37607/41569"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"li|dttnet_improving_video_shadow_detection_via_darkaware_guidance_and_tokenized_temporal_modeling"},"html":{"value":"https://doi.org/10.1609/aaai.v40i8.37607"},"_bibtex":{"value":"@inproceedings{DBLP:conf/aaai/LiSYZHZSZ26,\n  author={Zhicheng Li and Kunyang Sun and Rui Yao and Hancheng Zhu and Fuyuan Hu and Jiaqi Zhao and Zhiwen Shao and Yong Zhou},\n  title={DTTNet: Improving Video Shadow Detection via Dark-Aware Guidance and Tokenized Temporal Modeling},\n  year={2026},\n  cdate={1767225600000},\n  pages={6753-6761},\n  url={https://doi.org/10.1609/aaai.v40i8.37607},\n  booktitle={AAAI},\n  crossref={conf/aaai/2026}\n}\n"},"abstract":{"value":"Video shadow detection confronts two entwined difficulties: distinguishing shadows from complex backgrounds and modeling dynamic shadow deformations under varying illumination. To address shadow-background ambiguity, we leverage linguistic priors through the proposed Vision-language Match Module (VMM) and a Dark-aware Semantic Block (DSB), extracting text-guided features to explicitly differentiate shadows from dark objects. Furthermore, we introduce adaptive mask reweighting to downweight penumbra regions during training and apply edge masks at the final decoder stage for better supervision. For temporal modeling of variable shadow shapes, we propose a Tokenized Temporal Block (TTB) that decouples spatiotemporal learning. TTB summarizes cross-frame shadow semantics into learnable temporal tokens, enabling efficient sequence encoding with minimal computation overhead. Comprehensive Experiments on multiple benchmark datasets demonstrate state-of-the-art accuracy and real-time inference efficiency."},"title":{"value":"DTTNet: Improving Video Shadow Detection via Dark-Aware Guidance and Tokenized Temporal Modeling"},"authors":{"value":[{"fullname":"Zhicheng Li","username":""},{"fullname":"Kunyang Sun","username":""},{"fullname":"Rui Yao","username":"~Rui_Yao1"},{"fullname":"Hancheng Zhu","username":""},{"fullname":"Fuyuan Hu","username":""},{"fullname":"Jiaqi Zhao","username":""},{"fullname":"Zhiwen Shao","username":""},{"fullname":"Yong Zhou","username":""}]}},"tmdate":1784618258192,"pdate":1798675200000,"externalIds":["dblp:conf/aaai/LiSYZHZSZ26"],"tcdate":1784618251741,"writers":["~","OpenReview.net/Public_Article/DBLP.org","OpenReview.net/Support"],"signatures":["~Rui_Yao1"],"forum":"f9vEgvlNdi","license":"CC BY-SA 4.0","number":80208,"cdate":1767225600000,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/DBLP.org/-/Record","OpenReview.net/Public_Article/-/Edit"],"mdate":1784618258192,"domain":"OpenReview.net/Public_Article","id":"f9vEgvlNdi","version":2},{"content":{"summary":{"value":"The paper aims to build a more reliable benchmark for composed image retrieval (CIR) that ensures (1) visually and semantically consistent reference–target image pairs and (2) a truly zero-shot evaluation setting. To this end, the authors propose ZeroSight, a novel benchmark derived from video-sourced datasets that provides consistent reference–target pairs, multiple positive samples, and hard negatives per query. To address the challenges posed by this benchmark, the paper introduces SC4CIR (Symmetric Consistency for CIR); a training-free, MLLM-driven method designed to effectively identify hard negative targets through three symmetric consistency checks. The results show that existing CLIP-based models may have overestimated performance under previous benchmarks."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"Wrote above"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. The motivation of the paper is solid. As the authors point out, current CIR evaluation benchmarks are relatively coarse. Incorporating partial edits under consistent reference and target pairs is needed for genuinely evaluating CIR performance. This research direction is both needed for this community.\n\n2. The proposed pipeline that leverages video datasets to obtain consistent reference–target image pairs is well-motivated, and the overall generation process is clearly structured. Moreover, maintaining a true zero-shot setting adds further credibility and robustness to the proposed benchmark.\n\n3. The proposed training-free method is simple yet effective, demonstrating strong performance under the newly established benchmark."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. To claim that “existing CLIP-based models may have overestimated performance under previous benchmarks,” more detailed analyses comparing the proposed and prior benchmarks are needed. Currently, the paper only reports results on the newly proposed benchmark. Furthermore, even within the proposed benchmark, CLIP-based methods (SEARLE, ...) still perform better compared to the proposed approach. Additional comparative results and analyses on both benchmarks would strengthen this claim.  \n\n2. Although the proposed method is training-free, the additional generation processes involved may introduce substantial computational overhead. A more detailed analysis of the computational cost would be beneficial. Additionally, it would be valuable to evaluate the proposed method on existing benchmarks (CIRR, CIRCO, FashionIQ) to assess its generalizability beyond the proposed dataset.\n\n3. The definition and motivation for the Positive–Negative Ranking mAP metric are unclear. Specifically, it is not evident how hard negatives are ranked within this formulation. \n\n4. The paper’s writing can be significantly improved for clarity and readability. The notations, while formal, are overly complex ( St_i, Tt_i, Ts_i, St, ... ) and sometimes redundant, particularly in Section 3.2. In my personal opinion, complex mathematical notations would not be necessary for explaining the generation procedure.  Simplifying the mathematical notation and refining explanations would make the paper easier to follow.\n\nI believe the benchmark itself is meaningful, but the explanation and metrics, and analyses might not be sufficient in the current stage, but I think it would be improved in the rebuttal phase."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762919699627,"tcdate":1762016615213,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission7620/Reviewer_dwp7"],"signatures":["ICLR.cc/2026/Conference/Submission7620/Reviewer_dwp7"],"forum":"MRfPHIJi94","number":3,"license":"CC BY 4.0","cdate":1762016615213,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission7620/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762919699627,"domain":"ICLR.cc/2026/Conference","replyto":"MRfPHIJi94","id":"nRJTg4Oc2Z","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Composed image retrieval","zero-shot learning","multi-modal retrieval"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Zero-Shot Composed Image Retrieval (ZS-CIR) aims to retrieve a target image based on a query composed of a reference image and a relative caption without training samples. Existing ZS-CIR datasets often suffer from inconsistencies between reference and target images due to noisy image sources, and do not achieve a true zero-shot scenario as they use public image datasets that models like CLIP have been trained on. To tackle these challenges, we introduce ZeroSight, a novel benchmark for ZS-CIR. It includes a dataset with consistent reference-target pairs sourced from videos, a data construction pipeline, and evaluation methods that consider the ranking of multiple positive and negative target images. We ensure visually and semantically consistent reference-target pairs by extracting frames from a single video and generating relative captions using LLM-assisted methods. To ensure a true zero-shot scenario, we use video data published after March 31, 2022, ensuring it was not included in CLIP's pre-training data. Additionally, we propose a training-free MLLM-driven method, SC4CIR (Symmetric Consistency for CIR), which can effectively identify hard negative targets through 3 symmetric consistency checks. This method is plug-and-play, seamlessly integrating with various CIR methods and significantly improving performance. Our experimental results from 27 methods reveal that current ZS-CIR datasets and evaluation metrics result in inflated retrieval performance, exaggerating the capabilities of CIR methods. Our benchmark and models can be accessed at https://anonymous.4open.science/r/ZeroSight-CFE1."},"_bibtex":{"value":"@misc{\nyang2026never,\ntitle={Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets},\nauthor={Zhenyu Yang and Zemin Du and Baoyuan Qi and Shengsheng Qian and Weiming Dong and Changsheng Xu},\nyear={2026},\nurl={https://openreview.net/forum?id=MRfPHIJi94}\n}"},"title":{"value":"Never Seen Before: Benchmarking Genuine Zero-Shot Composed Image Retrieval with Consistent Video-Sourced Datasets"},"pdf":{"value":"/pdf/4421a0990abb194a9036e9c5b7d667a092d8a964.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"yang|never_seen_before_benchmarking_genuine_zeroshot_composed_image_retrieval_with_consistent_videosourced_datasets"},"authorids":{"value":["~Zhenyu_Yang6","~Zemin_Du2","~Baoyuan_Qi1","~Shengsheng_Qian1","~Weiming_Dong1","~Changsheng_Xu1"]},"authors":{"value":["Zhenyu Yang","Zemin Du","Baoyuan Qi","Shengsheng Qian","Weiming Dong","Changsheng Xu"]}},"version":2},{"content":{"summary":{"value":"This paper introduces ConFiner, a method that uses image diffusion models to improve video generation. ConFiner first uses a text-to-video model to generate a rough structure of the video. Then, noise is added to this video, and it's passed through image generation models for spatial refinement and the video model for temporal refinement. They also propose ConFiner-Long, which is designed for making long videos by using three strategies to keep segments consistent and transitions smooth. Experimental results show that ConFiner improves video quality, and ConFiner-Long successfully generates longer videos."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"1. The writing of the paper needs to correspond to the training-free long video generation in the title, including motivation, related work, method, and experiments.\n2. In the related work, the omission of these most related studies is puzzling. \n3. The authors should explain how their method is different from current methods and what makes it stand out, including ConFiner and ConFiner-Long"},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":3},"strengths":{"value":"a. Clarity and Simplicity: The approach presented is straightforward, and the method section is generally clear and easy to follow.\nb. Comprehensive and Convincing Experiments: The experiments are thorough, with results showing that ConFiner effectively improves video quality and coherence compared to existing models.\nc. Long Video Generation Capability: ConFiner-Long introduces three strategies that help maintain consistency and smooth transitions in longer videos, allowing for the generation of high-quality videos with extended frame lengths."},"flag_for_ethics_review":{"value":["No ethics review needed.","Yes, Other reasons (please specify below)"]},"weaknesses":{"value":"1. The title suggests a focus on “training-free long video generation”, but the main content is mainly about introducing ConFiner’s ability to enhance video quality. And most experiments also focus on ConFiner’s improvements, creating a bit of a disconnect between the title and the paper's main content.\n2. Limited Novelty: ConFiner’s approach of using T2I models to enhance T2V quality isn’t new. Similar ideas have already been explored in works like VideoElevator[1]. This reduces the novelty of the proposed method.\n3. Missing Related Work: The paper is notably lacking in its discussion of long video generation and related work using T2I as a video generation refiner, such as [1,2,3] This aspect is vital as it forms the basis of the research's motivation. The omission of these most related studies is puzzling. \n4. The experiments mainly focus on ConFiner's comparison and analysis and lacks comparison with existing long video generation methods, like StoryDiffusion, StreamingT2V, SEINE. \n5. The ablation study on three strategies of ConFiner-Long is missing quantitative results. Fig.4 cannot fully prove the effectiveness of three strategies.\n\n\n[1] VideoElevator: Elevating Video Generation Quality with Versatile Text-to-Image Diffusion Models\n[2] StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video Generation\n[3] SEINE: Short-to-Long Video Diffusion Model for Generative Transition and Prediction"}},"nonreaders":[],"tmdate":1731427195195,"tcdate":1730012921536,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission91/Reviewer_Py7V"],"signatures":["ICLR.cc/2025/Conference/Submission91/Reviewer_Py7V"],"forum":"3hc2ESNU6n","number":1,"license":"CC BY 4.0","cdate":1730012921536,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission91/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427195195,"domain":"ICLR.cc/2025/Conference","replyto":"3hc2ESNU6n","id":"Ro9jyX1AIR","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["generative models","diffusion models","video generation"]},"supplementary_material":{"value":"/attachment/6e2226f8b5f4ca9b4a9788025382fe717b612dbe.zip"},"primary_area":{"value":"generative models"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Video generation models hold substantial potential in areas such as filmmaking. However, current video diffusion models need high computational costs and produce suboptimal results due to high complexity of video generation task. In this paper, we propose \\textbf{ConFiner}, an efficient high-quality video generation framework that decouples video generation into easier subtasks: structure \\textbf{con}trol and spatial-temporal re\\textbf{fine}ment. It can generate high-quality videos with chain of off-the-shelf diffusion model experts, each expert responsible for a decoupled subtask. During the refinement, we introduce coordinated denoising, which can merge multiple diffusion experts' capabilities into a single sampling. Furthermore, we design ConFiner-Long framework, which can generate long coherent video with three constraint strategies on ConFiner. Experimental results indicate that with only 10\\% of the inference cost, our ConFiner surpasses representative models like Lavie and Modelscope across all objective and subjective metrics. And ConFiner-Long can generate high-quality and coherent videos with up to 600 frames."},"_bibtex":{"value":"@misc{\nli2024trainingfree,\ntitle={Training-free Long Video Generation with Chain of Diffusion Model Experts},\nauthor={Wenhao Li and Yichao Cao and Xiu Su and Xi Lin and Shan You and Mingkai Zheng and Yi Chen and Chang Xu},\nyear={2024},\nurl={https://openreview.net/forum?id=3hc2ESNU6n}\n}"},"title":{"value":"Training-free Long Video Generation with Chain of Diffusion Model Experts"},"pdf":{"value":"/pdf/62fe9cdb0c4e4a640f28ae22bf0867f590ed2f9c.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"li|trainingfree_long_video_generation_with_chain_of_diffusion_model_experts"},"authorids":{"value":["~Wenhao_Li14","~Yichao_Cao1","~Xiu_Su1","~Xi_Lin4","~Shan_You3","~Mingkai_Zheng1","~Yi_Chen18","~Chang_Xu4"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Wenhao Li","Yichao Cao","Xiu Su","Xi Lin","Shan You","Mingkai Zheng","Yi Chen","Chang Xu"]}},"version":2},{"content":{"summary":{"value":"This study introduces a video-language foundation model that harnesses the potential of multi-modal fusion and intricate modality alignment, aiming to refine the video-to-text generation endeavor. It finds its foundational roots in the image-centric BLIP-2 model, further enriched by the integration of the cascaded Q-Former, a mechanism adept at consolidating information from video frames and ASR transcripts. Empirical evaluations across tasks such as video captioning, summarization, and retrieval provide insightful results, suggesting the relative efficacy of the model when benchmarked against contemporaneous contributions, notably VideoChat and Video-LLaMA."},"presentation":{"value":"2 fair"},"contribution":{"value":"2 fair"},"soundness":{"value":"3 good"},"strengths":{"value":"1. The initiative to incorporate LLM in the creation of an extensive video-language foundation is discernibly a forward-thinking approach with significant potential.\n\n2. The model's rigorous empirical assessment across diverse tasks, such as video captioning, summarization, and retrieval, offers a well-rounded perspective. This extensive evaluation underscores the model's versatility, ensuring its capabilities are thoroughly understood and aptly validated for a variety of applications.\n\n3. The integration of the cascaded Q-Former marks a noteworthy advancement. Crafted to fluidly amalgamate information from both video frames and ASR transcripts, this mechanism suggests promising prospects for enhancing the quality of information retrieval and representation within the video-language domain."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. A primary point of contention pertains to the element of novelty. At its core, this study appears to be a straightforward adaptation of BLIP2 tailored for the video realm. The inclusion of the Q-former, while intriguing, is not entirely groundbreaking, as it's an integral component of BLIP2. Furthermore, the training objectives, which encompass intermediate contrastive learning and auto-regressive text generation, have been comprehensively explored in both BLIP2 and CoCa.\n\n2. In terms of practical outcomes, there's room for enhancement. Specifically, when observing video summarization, the zero-shot outcomes do not stand out, especially when benchmarked against solutions like VideoChat. Additionally, in the context of video captioning on MSRVTT, the results lag behind the performance metrics set by GIT2.\n\n3. Why not incorporate a comparative analysis with task-specific methodologies? This would provide a clearer perspective on the relative efficacy of the video foundation model, shedding light on whether it indeed offers superior results in specific contexts.\n\n4.  The review did not mention any in-depth error analysis, which is crucial in understanding where the model falls short and how it can be improved."},"confidence":{"value":"4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work."},"questions":{"value":"1. Figure 1 appears to echo the content of Figure 2, leading to redundancy in visual representation. A more impactful approach might be to leverage Figure 1 to delve deeper into task-specific motivations or insights. This would not only differentiate the content of the two figures but also offer readers a more granular understanding, rather than just presenting an overarching framework overview. Such an enhancement could streamline the content and enrich the narrative of the research.\n\n2. Missing reference like ChatVideo [1].\n\n[1]https://arxiv.org/pdf/2304.14407.pdf"},"rating":{"value":"5: marginally below the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636251434,"tcdate":1699189008687,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission3062/Reviewer_5do8"],"signatures":["ICLR.cc/2024/Conference/Submission3062/Reviewer_5do8"],"forum":"ia5wG0Fp2U","number":4,"license":"CC BY 4.0","cdate":1699189008687,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission3062/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636251434,"domain":"ICLR.cc/2024/Conference","replyto":"ia5wG0Fp2U","id":"Oh3aGjUaIC","forumContent":{"venue":{"value":"ICLR 2024 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video Understanding;Modality alignment;Representation decoupling;Fine-grained"]},"primary_area":{"value":"representation learning for computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"This paper proposes Video-Teller, a Video-Text foundational model that leverages multi modal fusion and fine-grained modality alignment to significantly enhance the cross-modal generation task. Video-Teller boosts the training efficiency by utilizing frozen pre-trained vision and language modules. Furthermore, it capitalizes on the robust linguistic capabilities of large language model, enabling the generation of more nuanced descriptions for videos (video summaries). To effectively integrate visual and auditory information and improve the model's understanding of videos, Video-Teller employs cascaded Q-Former to fuse information from different frames and modalities. In addition to conventional loss functions, we introduce an additional Text Auto-Encoder to decouple the target text for fine-grained modality alignment, further optimizing the model. Experimental results demonstrate the efficacy of our proposed video foundational model in accurately comprehending videos and generating coherent and precise language descriptions. It is worth noting that the fine-grained alignment enhances the model's capabilities with a relatively minimal increase in cost."},"_bibtex":{"value":"@misc{\nliu2024videoteller,\ntitle={Video-Teller: Enhancing Cross-Modal Generation with Fusion and Decoupling},\nauthor={Haogeng Liu and Qihang Fan and Tingkai Liu and Linjie Yang and Yunzhe Tao and Huaibo Huang and Ran He and Hongxia Yang},\nyear={2024},\nurl={https://openreview.net/forum?id=ia5wG0Fp2U}\n}"},"title":{"value":"Video-Teller: Enhancing Cross-Modal Generation with Fusion and Decoupling"},"pdf":{"value":"/pdf/7f22d5615efd2648bdc809ea533f15f67a9f6668.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Withdrawn_Submission"},"paperhash":{"value":"liu|videoteller_enhancing_crossmodal_generation_with_fusion_and_decoupling"},"authorids":{"value":["~Haogeng_Liu1","~Qihang_Fan1","~Tingkai_Liu1","~Linjie_Yang4","~Yunzhe_Tao2","~Huaibo_Huang1","~Ran_He1","~Hongxia_Yang2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Haogeng Liu","Qihang Fan","Tingkai Liu","Linjie Yang","Yunzhe Tao","Huaibo Huang","Ran He","Hongxia Yang"]}},"version":2},{"content":{"summary":{"value":"This paper presents BLrVR, a training-free framework for Bengali long-range video question answering. \nThe approach adapts the EgoSchema dataset into Bengali through machine translation and human validation, and evaluates two reasoning frameworks: 1. CeAS, a structured close-ended prompting strategy with role specification and task cues\n2. RAG, a retrieval-augmented generation variant leveraging external textual evidence.\n\nExperiments benchmark various multi-modal LLMs (Gemini-2.0/1.5, Gemma2-9B) as captioners and reasoning modules, reporting moderate gains in precision and recall for CeAS over RAG on the translated Bengali EgoSchema subset."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Please see the weakness section, overall, I have doubts about the area of contribution is aligned with ICLR requirement or not."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. This paper adapts long-range video reasoning to Bengali is appreciated and valuable for inclusive AI research, given the paucity of multimodal datasets in South Asian languages.\n\n2. The paper provides some empirical comparisons across captioners, LLMs, and prompting variants (CeAS, CoT, Plan-and-Solve), offering a practical reference for low-resource language setups.\n\n3. The framework is training-free and modular, making it easy to reproduce and extend to other low-resource settings."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Limited technical novelty: I think the core contributions (CeAS prompt design, RAG baseline) are largely incremental and primarily involve reusing existing prompting and retrieval frameworks in a Bengali setting. The novelty lies more in application and data translation rather than in algorithmic innovation.\n\n2. Dataset contribution is minimal: The “Bengali EgoSchema” is simply a translated version of an existing benchmark with limited linguistic validation; there is no evidence of new video content, annotation schema, or task definition.\n\nalso, results hover around 69% accuracy, with small differences between methods. The evaluation lacks ablations on reasoning depth, temporal grounding fidelity, or cross-language generalization, which would be necessary for an ICLR-style contribution + missing important comparison/discussion with the line of related work [1-4] (and more).\n\n\nOverall, I feel the focus on prompt structure, translation quality, and linguistic adaptation aligns more naturally with ACL/EMNLP tracks (e.g., multilingual or low-resource multimodal reasoning), rather than ICLR’s emphasis on technical advances in learning representations or modeling architectures.\n\n\n[1] A simple llm framework for long-range video question-answering.  \n[2] Videotree: Adaptive tree-based video representation for llm reasoning on long videos.  \n[3] Lifelongmemory: Leveraging llms for answering queries in long-form egocentric videos.  \n[4] VideoLucy: Deep Memory Backtracking for Long Video Understanding"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762920309973,"tcdate":1762025830380,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission8409/Reviewer_GP3y"],"signatures":["ICLR.cc/2026/Conference/Submission8409/Reviewer_GP3y"],"forum":"8VthdYWjHt","number":3,"license":"CC BY 4.0","cdate":1762025830380,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission8409/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762920309973,"domain":"ICLR.cc/2026/Conference","replyto":"8VthdYWjHt","id":"3eUH89Obvz","forumContent":{"TLDR":{"value":"CeAS and RAG for Bengali Long-Range Video Reasoning"},"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Long-range Video Reasoning","Bengali Language Processing","Visual Question Answering","Large Language Models","Close-ended Answer Selection","Retrieval-Augmented Generation","Multimodal Captioning"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Long-range video question answering (VQA) remains a challenging task, especially\n in low-resource languages like Bengali, due to limited linguistic tools and\nthe need for multi-step temporal reasoning. To address these challenges, we propose\n a training-free framework for Bengali Long-range Video Reasoning (BLrVR).\nOur approach adapts the EgoSchema benchmark to Bengali through high-quality\ntranslation and contextual validation. We introduce a novel prompting strategy,\nCeAS (Close-ended Answer Selection), which integrates structured roles, task\ncues, and strict constraints to guide LLM reasoning. Additionally, we explore a\nRetrieval-Augmented Generation (RAG) variant that fuses relevant caption context \nwith external evidence for enriched inference. Empirical results show that\nCeAS achieves state-of-the-art performance, surpassing RAG in precision, recall,\nand runtime efficiency, despite matching in accuracy and F1-score. We further\nbenchmark different captioners, LLMs, retrievers, and prompting schemes, providing\n a comprehensive evaluation of components crucial to BLrVR success. Our\nfindings demonstrate that structured prompting can outperform retrieval-heavy \nmethods in both effectiveness and efficiency for low-resource multimodal reasoning.\n The code is publicly released at: https://github.com/anaxy-code/Bengali-Long-Range-Video-Reasoning"},"_bibtex":{"value":"@misc{\nrony2025visual,\ntitle={Visual Grounding Meets Language: Ce{AS} and {RAG} for Bengali Long-Range Video Reasoning},\nauthor={Mohammad Abu Tareq Rony and Farjana Kabir and Mohammad Shariful Islam},\nyear={2025},\nurl={https://openreview.net/forum?id=8VthdYWjHt}\n}"},"title":{"value":"Visual Grounding Meets Language: CeAS and RAG for Bengali Long-Range Video Reasoning"},"pdf":{"value":"/pdf/9a2703ba907ddf948d764859dbd81d077af70ee9.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"rony|visual_grounding_meets_language_ceas_and_rag_for_bengali_longrange_video_reasoning"},"authorids":{"value":["~Mohammad_Abu_Tareq_Rony1","~Farjana_Kabir2","~Mohammad_Shariful_Islam1"]},"authors":{"value":["Mohammad Abu Tareq Rony","Farjana Kabir","Mohammad Shariful Islam"]}},"version":2},{"content":{"summary":{"value":"The paper proposes to improve the trajectory guided video generation by using entity representations, rather than point/pixel/bounding box representations. Specifically, the paper utilizes stable video diffusion and CoTrack to segment the frames to identify some entities and track their motion trajectories, which allows users to drag them to control the video generation. Meanwhile, the proposed method also employs a position-aware relation module to model and constrain the relative positions or motion among multiple object parts or entities. The proposed method has been evaluated on 3 public datasets and compare favorably with the SOTA."},"suitability":{"value":3},"strengths":{"value":"The paper is well motivated since how to control the video generation by the motion of multiple objects is a critical problem for users to adopt the video generation technology.\n\nThe proposed entity representation and position-aware relation module are new and sound approaches to advance the trajectory guided video generation.\n\nThe experiments are convincing to demonstrate the clear advantage of the proposed method.\n\nOverall, this is a descent paper which is a clear step to advance  trajectory guided video generation."},"confidence":{"value":3},"rating":{"value":5},"limitations":{"value":"I wonder how sensitive of the RBF kernel and the learnable means and covariances, as well as the semantic relation in Eq.4 to the video generation performance. Even for the same kind of objects, the motion or relative positions among the parts could be dramatically different due to the camera view angles. Please elaborate on the generalizability of the position-aware relation, since the training videos may not cover all possible motion and view angles of objects.\n\nTypo: \"on on\" line 641."}},"nonreaders":[],"tmdate":1721657382613,"tcdate":1715896206413,"writers":["acmmm.org/ACMMM/2024/Conference","acmmm.org/ACMMM/2024/Conference/Submission4742/Reviewer_AxYa"],"signatures":["acmmm.org/ACMMM/2024/Conference/Submission4742/Reviewer_AxYa"],"forum":"H4E32wjshc","number":1,"license":"CC BY 4.0","cdate":1715896206413,"readers":["everyone"],"invitations":["acmmm.org/ACMMM/2024/Conference/Submission4742/-/Official_Review","acmmm.org/ACMMM/2024/Conference/-/Edit","acmmm.org/ACMMM/2024/Conference/Submission4742/Official_Review1/-/Final_Rating"],"mdate":1721657382613,"domain":"acmmm.org/ACMMM/2024/Conference","replyto":"H4E32wjshc","id":"ambhBVByBN","forumContent":{"venue":{"value":"MM2024 Oral"},"supplementary_material":{"value":"/attachment/3f130f38ab0ced7f4092eae9e8860520227e4a4f.zip"},"abstract":{"value":"In recent years, diffusion models have achieved tremendous success in the field of video generation, with controllable video generation receiving significant attention. However, existing control methods still face two limitations: Firstly, control conditions (such as depth maps, 3D Mesh) are difficult for ordinary users to obtain directly. Secondly, it’s challenging to drive multiple objects through complex motions with multiple trajectories simultaneously. In this paper, we introduce DragEntity, a video generation model that utilizes entity representation for controlling the motion of multiple objects. In comparison to previous methods, MotionCtrl offers two main advantages: 1) Trajectory-based methods are more user-friendly for interaction. Users only need to draw trajectories during the interaction to generate videos. 2) We use entity representation to represent any object in the image, and multiple objects can maintain relative spatial relationships. Therefore, we allow multiple trajectories to control multiple objects in the image with different levels of complexity simultaneously. Our experiments validate the effectiveness of DragEntity, demonstrating its superior performance in fine-grained control in video generation."},"relevance_to_conference":{"value":"The primary task of this work is video generation, achieving fine-grained control through the integration of images and trajectories. The input of our method is an image and trajectory, and the output is a video, involving multiple modalities such as images, trajectories, and videos. This aligns well with the themes of multimedia/multimodal. Furthermore, addressing the object distortion issue present in existing works, we employ interactive segmentation to select the control regions in the initial frame. We also use the masks of each entity in the first frame to extract the central coordinates of that entity and predict the entity's motion trajectory through CoTrack. In comparison with existing works, our method shows improvements in both FID and FVD metrics. Our work makes a significant contribution to the task of fine-grained control in video generation."},"_bibtex":{"value":"@inproceedings{\nzhang2024dragentitytrajectory,\ntitle={DragEntity:Trajectory Guided Video Generation using Entity and Positional Relationships},\nauthor={Wan Zhang and Sheng Tang and Jiawei Wei and Ruize Zhang and Juan Cao},\nbooktitle={ACM Multimedia 2024},\nyear={2024},\nurl={https://openreview.net/forum?id=H4E32wjshc}\n}"},"title":{"value":"DragEntity:Trajectory Guided Video Generation using Entity and Positional Relationships"},"secondary_subject_area":{"value":["[Generation] Social Aspects of Generative AI","[Generation] Generative Multimedia","[Experience] Multimedia Applications"]},"pdf":{"value":"/pdf/fad119a39324d517f96aba0eca41a2c1c9ebf90c.pdf"},"venueid":{"value":"acmmm.org/ACMMM/2024/Conference"},"paperhash":{"value":"zhang|dragentitytrajectory_guided_video_generation_using_entity_and_positional_relationships"},"primary_subject_area":{"value":"[Generation] Generative Multimedia"},"authorids":{"value":["~Wan_Zhang2","~Sheng_Tang1","~Jiawei_Wei1","~Ruize_Zhang2","~Juan_Cao1"]},"authors":{"value":["Wan Zhang","Sheng Tang","Jiawei Wei","Ruize Zhang","Juan Cao"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a novel way to implement long-context Video MLLM, which only requires long context training of plain text and image-text alignment training.\n- The trained long-context LLM is used for image-text training. Through the designed UniRes visual information encoding method, the long-context LLM can be extended to the long-context video MLLM without video-text training.\n- A Visual Needle-In-A-Haystavk (V-NIAH) benchmark is proposed to test the capabilities of long-context video MLLM.\n- By inputting more frames for reasoning, the proposed LongVA achieves SOTA performance on Video-MME and MLVU."},"soundness":{"value":2},"confidence":{"value":5},"questions":{"value":"(1) Can training only be performed on a single image? How is the long text obtained in image-text training? Because the trained LLM backbone has a context length of 224K. The training details are not provided.\n\n(2) According to the display of Figure 2, the author treats each frame of the video as a grid of the image during testing, so 576 tokens are formed for inference. In order to input more frames during inference, my understanding is that during image-text training, the image here needs to have a very high resolution? In order to segment more grids, so as to adapt to more frames during inference. How are high-resolution images obtained? For the collection process of the training set, this seems to be a method of exchanging spatial information annotation for temporal information annotation.\n\n(3) Unless the details of the training are given, including the amount of data, image resolution, and text length, it is difficult to understand why simple image-text training can be transferred to the reasoning of videos with very long frames?\n\n(4) Why can image training achieve long-context video MLLM? The author did not discuss and analyze the principle.\n\n(5) The experimental results in Table 4 show that inputting more frames into LongVA does not continuously enhance the performance of Video-MME? Especially in long duration videos, why is this? And as the number of frames continues to increase, it gets worse in short, medium and long. Although other experiments show that LongVA performs well in VNIAH, demonstrating the ability of information retrieval, it does not show the advantage of this model in long video understanding.\n\n(6) In the experimental results of Table 7, in 6 benchmarks, UniRes is better than AnyRes in 4 tasks, which does not prove that UniRes is a better visual encoding method. Why use UniRes instead of AnyRes?"},"rating":{"value":5},"details_of_ethics_concerns":{"value":"NA"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"(1) Explored a method to implement long-context video MLLM from the perspective of long-context text training. The idea is very interesting.  \n \n(2) The evaluation of long-context text retrieval (NIAH) and video retrieval (V-NIAH) is very complete and well presented.   \n\n(3) The experimental results are rich. LongVA was evaluated on different multi-modal benchmarks and achieved SOTA performance on Video-MME and MLVU."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"(1) The results of LongVA on some benchmarks are not ideal, such as ActivityNetQA and VideoChatGPT, which is reflected in the fact that more frames do not show better video understanding ability. This makes people suspect that LongVA only has the ability to retrieve long-context video information, but lacks the ability to understand long-context videos.\n\n(2) The paper does not give detailed training details, such as the training details of the long-context LLM of pure text, the training details of image-text, and the datasets used in these two training processes.\n\n(3) Some experimental results in the article do not show that a long-context MLLM with good video understanding ability has been achieved. And the proposed UniRes has not been proven to be better than the previous AnyRes."}},"nonreaders":[],"tmdate":1731427612175,"tcdate":1730721369722,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission2523/Reviewer_92h6"],"signatures":["ICLR.cc/2025/Conference/Submission2523/Reviewer_92h6"],"forum":"QETk0lBdVf","number":4,"license":"CC BY 4.0","cdate":1730721369722,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission2523/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427612175,"domain":"ICLR.cc/2025/Conference","replyto":"QETk0lBdVf","id":"vp1YTGRZUF","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"TLDR":{"value":"Long context capability can zero-shot transfer from language to vision. LongVA can process 2000 frames or over 200K visual tokens. It achieves state-of-the-art performance on Video-MME among 7B models."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Vision Language Model","Long Context Model"]},"primary_area":{"value":"foundation or frontier models, including LLMs"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Video sequences offer valuable temporal information, but existing large multimodal models (LMMs) fall short in understanding extremely long videos. Many works address this by reducing the number of visual tokens using visual resamplers. Alternatively, in this paper, we approach this problem from the perspective of the language model. By simply extrapolating the context length of the language backbone, we enable LMMs to comprehend orders of magnitude more visual tokens without any video training. We call this phenomenon long context transfer and carefully ablate its properties. To effectively measure LMMs' ability to generalize to long contexts in the vision modality, we develop V-NIAH (Visual Needle-In-A-Haystack), a purely synthetic long vision benchmark inspired by the language model's NIAH test. Our proposed Long Video Assistant (LongVA) can process 2000 frames or over 200K visual tokens without additional complexities. With its extended context length, LongVA achieves state-of-the-art performance on Video-MME among 7B-scale models by densely sampling more input frames."},"_bibtex":{"value":"@misc{\nzhang2025long,\ntitle={Long Context Transfer from Language to Vision},\nauthor={Peiyuan Zhang and Kaichen Zhang and Bo Li and Guangtao Zeng and Jingkang Yang and Yuanhan Zhang and Ziyue Wang and Haoran Tan and Chunyuan Li and Ziwei Liu},\nyear={2025},\nurl={https://openreview.net/forum?id=QETk0lBdVf}\n}"},"title":{"value":"Long Context Transfer from Language to Vision"},"pdf":{"value":"/pdf/9b7ba8c5ace9d49f442f099a7e2deff48ddff1c6.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"zhang|long_context_transfer_from_language_to_vision"},"authorids":{"value":["~Peiyuan_Zhang2","~Kaichen_Zhang2","~Bo_Li23","~Guangtao_Zeng1","~Jingkang_Yang1","~Yuanhan_Zhang1","~Ziyue_Wang5","~Haoran_Tan1","~Chunyuan_Li1","~Ziwei_Liu1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Peiyuan Zhang","Kaichen Zhang","Bo Li","Guangtao Zeng","Jingkang Yang","Yuanhan Zhang","Ziyue Wang","Haoran Tan","Chunyuan Li","Ziwei Liu"]}},"version":2},{"content":{"summary":{"value":"This paper introduces StreamV2V, a real-time video-to-video translation model that uses a feature bank and backward-looking mechanism to ensure temporal consistency. StreamV2V achieves significantly faster speeds than existing methods while maintaining consistent outputs."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"See Weaknesses."},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The paper proposes a novel approach to streaming video-to-video translation, introducing a feature bank mechanism to address temporal consistency, enabling each frame in the video stream to reference features from past frames in an innovative manner.\n\n2. This method achieves efficient real-time processing without model fine-tuning, outperforming many comparable methods in experimental results.\n\n3. StreamV2V requires no training or fine-tuning and seamlessly integrates with existing image diffusion models, demonstrating the potential for a wide range of video editing tasks.\n\n4. The paper includes extensive quantitative experiments and user studies, validating its performance.\n\n5. This paper is well written and easy to follow. The supplementary materials also include comprehensive experimental details, content, and discussions, demonstrating the effectiveness of the proposed method."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Does DyMe Bank still have advantages in video quality compared to queue-based banks when dealing with longer videos, more complex video content, or video motion?"}},"nonreaders":[],"tmdate":1731428628356,"tcdate":1730567640633,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission6311/Reviewer_wGbG"],"signatures":["ICLR.cc/2025/Conference/Submission6311/Reviewer_wGbG"],"forum":"AMkf7h7HER","number":2,"license":"CC BY 4.0","cdate":1730567640633,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission6311/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428628356,"domain":"ICLR.cc/2025/Conference","replyto":"AMkf7h7HER","id":"ocJWGJQs4R","forumContent":{"TLDR":{"value":"We present StreamV2V to support real-time video-to-video translation for streaming input."},"venue":{"value":"ICLR 2025 Poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Streaming video translation","diffusion models","feature banks"]},"supplementary_material":{"value":"/attachment/6636cc26116c9f1106fdd3c3b05da0232d939b8c.zip"},"primary_area":{"value":"generative models"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"This paper introduces StreamV2V, a diffusion model that achieves real-time streaming video-to-video (V2V) translation with user prompts. \n  Unlike prior V2V methods using batches to process limited frames, we opt to process frames in a streaming fashion, to support unlimited frames.\n  At the heart of StreamV2V lies a backward-looking principle that relates the present to the past. \n  This is realized by maintaining a feature bank, which archives information from past frames.\n  For incoming frames, StreamV2V extends self-attention to include banked keys and values, and directly fuses similar past features into the output.\n  The feature bank is continually updated by merging stored and new features, making it compact yet informative.\n  StreamV2V stands out for its adaptability and efficiency, seamlessly integrating with image diffusion models without fine-tuning.\n  It can run 20 FPS on one A100 GPU, being 15$\\times$, 46$\\times$, 108$\\times$, and 158$\\times$ faster than FlowVid, CoDeF, Rerender, and TokenFlow, respectively. \n  Quantitative metrics and user studies confirm StreamV2V's exceptional ability to maintain temporal consistency."},"_bibtex":{"value":"@inproceedings{\nliang2025looking,\ntitle={Looking Backward: Streaming Video-to-Video Translation with Feature Banks},\nauthor={Feng Liang and Akio Kodaira and Chenfeng Xu and Masayoshi Tomizuka and Kurt Keutzer and Diana Marculescu},\nbooktitle={The Thirteenth International Conference on Learning Representations},\nyear={2025},\nurl={https://openreview.net/forum?id=AMkf7h7HER}\n}"},"title":{"value":"Looking Backward: Streaming Video-to-Video Translation with Feature Banks"},"pdf":{"value":"/pdf/65906303ad329767664a6e8fca6ed65a144ae4eb.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference"},"paperhash":{"value":"liang|looking_backward_streaming_videotovideo_translation_with_feature_banks"},"authorids":{"value":["~Feng_Liang3","~Akio_Kodaira1","~Chenfeng_Xu1","~Masayoshi_Tomizuka2","~Kurt_Keutzer1","~Diana_Marculescu4"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Feng Liang","Akio Kodaira","Chenfeng Xu","Masayoshi Tomizuka","Kurt Keutzer","Diana Marculescu"]}},"version":2},{"content":{"summary":{"value":"This paper focuses on existing Figure-to-Caption models that fall short of metrics like helpfulness, explainability, and visual-descriptiveness. To this end, they introduce an RLHF framework for figure-to-caption generation with a small amount of actual human feedback for generating high-quality captions. They also propose a new benchmark of figure-caption pairs with caption quality scores for a better understanding of the read-aligned figure-caption pairs."},"presentation":{"value":"2 fair"},"contribution":{"value":"2 fair"},"soundness":{"value":"2 fair"},"strengths":{"value":"1. **The proposed method is effective**. As shown in Fig.2, the RLHF has improved the BLIP and Vit+GPT2 models with a clear improvement.\n\n2. **The benchmark of figure-caption pairs is helpful**. This paper provides a new benchmark of figure-caption pairs is helpful for the research community, and they have done a great release."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. **Time Complexity Analysis.** This paper claims that using offline reward-conditioned behavioral cloning for model optimization is computationally efficient. It is not convinced, that you should compare the time complexity analysis between the offline RLHF and online RLHF methods.\n\n2. **The proposed method is not novel enough.** The framework of the RLHF for figure-to-text can not provide more insight for the understanding of the read-aligned figure-caption pairs. It may not be enough for the technical contribution.\n\n3. **The writing needs to be improved**. There are many typos in the main text.\ne.g. figure-ti-caption -> figure-to-caption in Section 6."},"confidence":{"value":"3: You are fairly confident in your assessment. It is possible that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work. Math/other details were not carefully checked."},"questions":{"value":"As shown in weaknesses."},"rating":{"value":"5: marginally below the acceptance threshold"},"details_of_ethics_concerns":{"value":"N/A"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699636630050,"tcdate":1698784594062,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission5924/Reviewer_LPWT"],"signatures":["ICLR.cc/2024/Conference/Submission5924/Reviewer_LPWT"],"forum":"7pVIFJW2Hp","number":3,"license":"CC BY 4.0","cdate":1698784594062,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission5924/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699636630050,"domain":"ICLR.cc/2024/Conference","replyto":"7pVIFJW2Hp","id":"yPT4rpkuEW","forumContent":{"TLDR":{"value":"To enable the generation of high-quality figure captions, we introduce FigCaps-HF a new framework for figure-caption generation that can incorporate domain expert feedback in generating captions optimized for reader preferences."},"venue":{"value":"Submitted to ICLR 2024"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Figure Caption Generation","Image-to-Text Generation","Reinforcement Learning using Human Feedback","Figure-Caption Benchmark","Human Feedback"]},"primary_area":{"value":"datasets and benchmarks"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"Captions are crucial for understanding scientific visualizations and documents. Existing captioning methods for scientific figures rely on figure-caption pairs extracted from documents for training, many of which fall short with respect to metrics like helpfulness, explainability, and visual-descriptiveness leading to generated captions being misaligned with reader preferences. To enable the generation of high-quality figure captions, we introduce FigCaps-HF a new framework for figure-caption generation that can incorporate domain expert feedback in generating captions optimized for reader preferences. Our framework comprises of 1) an automatic method for evaluating quality of figure-caption pairs, 2) a novel reinforcement learning with human feedback (RLHF) method to optimize a generative figure-to-caption model for reader preferences. We demonstrate the effectiveness of our simple learning framework by improving performance over standard fine-tuning across different types of models. In particular, when using BLIP as the base model, our RLHF framework achieves a mean gain of 35.7%, 16.9%, and 9% in ROUGE, BLEU, and Meteor, respectively. Finally, we release a large-scale benchmark dataset with human feedback on figure-caption pairs to enable further evaluation and development of RLHF techniques for this problem"},"_bibtex":{"value":"@misc{\nsingh2024figcapshf,\ntitle={FigCaps-{HF}: A Figure-to-Caption Generative Framework and Benchmark with Human Feedback},\nauthor={Ashish Singh and Prateek Agarwal and Zixuan Huang and Arpita Singh and Tong Yu and Sungchul Kim and Victor Bursztyn and Nikos Vlassis and Ryan A. Rossi},\nyear={2024},\nurl={https://openreview.net/forum?id=7pVIFJW2Hp}\n}"},"title":{"value":"FigCaps-HF: A Figure-to-Caption Generative Framework and Benchmark with Human Feedback"},"pdf":{"value":"/pdf/c43660a76412a1223da4504adc5fd86847766864.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Rejected_Submission"},"paperhash":{"value":"singh|figcapshf_a_figuretocaption_generative_framework_and_benchmark_with_human_feedback"},"authorids":{"value":["~Ashish_Singh2","~Prateek_Agarwal1","~Zixuan_Huang5","~Arpita_Singh1","~Tong_Yu3","~Sungchul_Kim1","~Victor_Bursztyn1","~Nikos_Vlassis1","~Ryan_A._Rossi2"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Ashish Singh","Prateek Agarwal","Zixuan Huang","Arpita Singh","Tong Yu","Sungchul Kim","Victor Bursztyn","Nikos Vlassis","Ryan A. Rossi"]}},"version":2},{"content":{"summary":{"value":"The paper presents an improved formulation for Direct Preference Optimization in diffusion modles. Speciifcally, the paper shows how images from altered prompts can be used to construct improved preference pairs (i.e. pairing the right caption for the image vs mismatched captions). In addition ot this, there's also a mask applied at the pixel/latent level to ensure that important regions in the sample are highlighted in the optimization. Results on prompt following benchmarks such as T2I-Compbench, GenEval, and DPG-Bench indicate that the method is able to bring notable improvements to the base SDXL model."},"soundness":{"value":4},"confidence":{"value":4},"questions":{"value":"I think broadly the paper introduces an interesting concept for improving text-to-image models, and I'm inclined towards accepting it. I'd mostly want clarity from the authors on a) the effect of the data source and b) contextualizing the paper better in the contex to existing works on diffusion model alignment (e.g. there are 25 papers here on DPO https://github.com/xie-lab-ml/awesome-alignment-of-diffusion-models). While it's impossible to exhaustively cover all the works, there must atleast be some effort to summarize the existing work in the field."},"rating":{"value":4},"details_of_ethics_concerns":{"value":"N/A"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"The paper proposes a very interesting and compelling idea of a) preparing preference pairs in a targeted manner with precise modifications between the positive negatives and b) an improved loss fomrulation where mismatched captions can be used for additional supervision.\n\nThe paper also achieves promising results on SDXL and is reasonably well-presented."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"[Major]\n\nFollowing Diffusion-DPO, there's been a bunch of work incorportaitng additional structure both in terms of the data construction and the loss into the formulation [a,b,c]. DenseDPO for instance applies the loss at finer temporal granularity for video models, which is quite close to the idea of having a spatial mask in images. Similarly, CaPO and RankDPO (among many other works in this direction) also demonstrate good improvements over the original Diffusion-DPO formulation, including on these same benchmarks. It would be a good idea for the paper to acknowledge these. \n\nAdditionally, an important caveat is regarding the source of data/prompts. A lot of previous DPO works (Diffusion-DPO/KTO, CaPO) rely on either Pick-a-Pic or DiffusionDB which have images that are quite different from T2I-Compbench/GenEval. Here, the data directly includes prompts from T2I-Compbench training set among other sources. While this is clearly not overfitting, it also might explain why the performance improvements on T2I-Compbench are so outsized compared to the gains on other benchmarks (which are still good, but not as dramatic). \n\n\n\n[Minor]\n\nI'd also be curious to see how the BiDPO formulation performs beyond the SDXL LoRA fine-tuning setting with newer models (e.g. with SD3/Flux etc.)\n\nThe effect of the region-guidance seems quite minimal, I would be curious if this is really necessary, or is there some scenario where it can be particularly useful? \n\n\n[a] Lee et al. \"Calibrated Multi-Preference Optimization for Aligning Diffusion Models\", CVPR 2025\n\n[b] Karthik et al. \"Scalable Ranked Preference Optimization for Text-to-Image Generation\", ICCV 2025\n\n[c] Wu et al. \"DenseDPO: Fine-Grained Temporal Preference Optimization for Video Diffusion Models\", NeurIPS 2025"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762915867415,"tcdate":1761977439869,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission1719/Reviewer_qSVV"],"signatures":["ICLR.cc/2026/Conference/Submission1719/Reviewer_qSVV"],"forum":"EkzVuZj60c","number":4,"license":"CC BY 4.0","cdate":1761977439869,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission1719/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762915867415,"domain":"ICLR.cc/2026/Conference","replyto":"EkzVuZj60c","id":"4DtrZHZQv4","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Compositional Text-To-Image Generation","Direct Preference Optimization","Diffusion Models","Image Generation"]},"primary_area":{"value":"generative models"},"abstract":{"value":"Despite the rapid progress of text-to-image (T2I) models, generating images that accurately reflect complex compositional prompts (covering attribute bindings, object relationships, counting) still remains challenging. To address this, we propose \\ourmethod, a framework to enhance T2I model's capability of compositional text-to-image generation. We begin by introducing an carefully designed pipeline to construct a large-scale preference dataset, \\ourdataset, with strictly quality control. Then, we extend Diffusion DPO to jointly optimize image and text preferences, which is shown to greatly effective in improving the models to follow complex text prompt in generation. To further enhance the models for fine-grained alignment, we employ a region-level guidance method to focus on regions relevant to compositional concepts. Experimental results demonstrate that our \\ourmethod substantially improves compositional fidelity, consistently outperforming prior methods across multiple benchmarks. Our approach highlights the potential of preference-based fine-tuning for complex text-to-image tasks, offering a flexible and scalable alternative to existing techniques."},"_bibtex":{"value":"@misc{\nliu2025compositional,\ntitle={Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization},\nauthor={Zhuohan Liu and Wujian Peng and Yitong Chen and Zuxuan Wu},\nyear={2025},\nurl={https://openreview.net/forum?id=EkzVuZj60c}\n}"},"title":{"value":"Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization"},"pdf":{"value":"/pdf/6dd09afb0bb81cb6a83478560ee5a8be7cb32560.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"liu|compositional_texttoimage_generation_via_regionaware_bimodal_direct_preference_optimization"},"authorids":{"value":["~Zhuohan_Liu2","~Wujian_Peng1","~Yitong_Chen1","~Zuxuan_Wu1"]},"authors":{"value":["Zhuohan Liu","Wujian Peng","Yitong Chen","Zuxuan Wu"]}},"version":2},{"content":{"venue":{"value":"CoRR 2025"},"pdf":{"value":"http://arxiv.org/pdf/2502.01277v1"},"venueid":{"value":"dblp.org/journals/CORR/2025"},"paperhash":{"value":"nguyen|octopinf_workloadaware_inference_serving_for_edge_video_analytics"},"authorids":{"value":["~Thanh-Tung_Nguyen2","https://dblp.org/search/pid/api?q=author:Lucas_Liebe:","","~Yuheng_Wu1","https://dblp.org/search/pid/api?q=author:Jinghan_Cheng:","~Dongman_Lee1"]},"html":{"value":"https://doi.org/10.48550/arXiv.2502.01277"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2502-01277,\n  publtype={informal},\n  author={Thanh-Tung Nguyen and Lucas Liebe and Nhat-Quang Tau and Yuheng Wu and Jinghan Cheng and Dongman Lee},\n  title={OCTOPINF: Workload-Aware Inference Serving for Edge Video Analytics},\n  year={2025},\n  month={February},\n  cdate={1738368000000},\n  journal={CoRR},\n  volume={abs/2502.01277},\n  url={https://doi.org/10.48550/arXiv.2502.01277}\n}\n"},"abstract":{"value":"Edge Video Analytics (EVA) has gained significant attention as a major application of pervasive computing, enabling real-time visual processing. EVA pipelines, composed of deep neural networks (DNNs), typically demand efficient inference serving under stringent latency requirements, which is challenging due to the dynamic Edge environments (e.g., workload variability and network instability). Moreover, EVA pipelines also face significant resource contention caused by resource (e.g., GPU) constraints at the Edge. In this paper, we introduce OCTOPINF, a novel resource-efficient and workload-aware inference serving system designed for real-time EVA. OCTOPINF tackles the unique challenges of dynamic edge environments through fine-grained resource allocation, adaptive batching, and workload balancing between edge devices and servers. Furthermore, we propose a spatiotemporal scheduling algorithm that optimizes the co-location of inference tasks on GPUs, improving performance and ensuring service-level objectives (SLOs) compliance. Extensive evaluations on a real-world testbed demonstrate the effectiveness of our approach. It achieves an effective throughput increase of up to 10x compared to the baselines and shows better robustness in challenging scenarios. OCTOPINF can be used for any DNN-based EVA inference task with minimal adaptation and is available at https://github.com/tungngreen/PipelineScheduler."},"title":{"value":"OCTOPINF: Workload-Aware Inference Serving for Edge Video Analytics"},"authors":{"value":["Thanh-Tung Nguyen","Lucas Liebe","Nhat-Quang Tau","Yuheng Wu","Jinghan Cheng","Dongman Lee"]}},"tmdate":1768385167719,"pdate":1735689600000,"tcdate":1744089233268,"writers":["~"],"signatures":["~Thanh-Tung_Nguyen2"],"forum":"ZI2WnveimU","license":"CC BY-SA 4.0","number":380899,"cdate":1738368000000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1768385167719,"domain":"DBLP.org","id":"ZI2WnveimU","version":2},{"content":{"venue":{"value":"PerCom 2025"},"pdf":{"value":"https://ieeexplore.ieee.org/iel8/11018666/11018684/11018743.pdf"},"venueid":{"value":"dblp.org/conf/PERCOM/2025"},"paperhash":{"value":"nguyen|octopinf_workloadaware_inference_serving_for_edge_video_analytics"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Thanh-Tung_Nguyen:","https://dblp.org/search/pid/api?q=author:Lucas_Liebe:","https://dblp.org/search/pid/api?q=author:Nhat-Quang_Tau:","~Yuheng_Wu1","https://dblp.org/search/pid/api?q=author:Jinghan_Cheng:","~Dongman_Lee1"]},"html":{"value":"https://doi.org/10.1109/PerCom64205.2025.00032"},"_bibtex":{"value":"@inproceedings{DBLP:conf/percom/NguyenLTWCL24,\n  author={Thanh-Tung Nguyen and Lucas Liebe and Nhat-Quang Tau and Yuheng Wu and Jinghan Cheng and Dongman Lee},\n  title={OctopInf: Workload-Aware Inference Serving for Edge Video Analytics},\n  year={2025},\n  cdate={1735689600000},\n  pages={128-137},\n  url={https://doi.org/10.1109/PerCom64205.2025.00032},\n  booktitle={PerCom},\n  crossref={conf/percom/2025}\n}\n"},"abstract":{"value":"Edge Video Analytics (EVA) has become a major application of pervasive computing, enabling real-time visual processing. EVA pipelines, composed of deep neural networks (DNNs), typically demand efficient inference serving under stringent latency requirements, which is challenging due to the dynamic Edge environments (e.g., workload variability and network instability). Moreover, EVA pipelines face significant resource contention due to resource (e.g., GPU) constraints at the Edge. In this paper, we introduce OctopInf, a novel resource-efficient and workload-aware inference serving system designed for real-time EVA. OctopInf tackles the unique challenges of dynamic edge environments through fine-grained resource allocation, adaptive batching, and workload balancing between edge devices and servers. Furthermore, we propose a spatiotemporal scheduling algorithm that optimizes the co-location of inference tasks on GPUs, improving performance and ensuring service-level objectives (SLOs) compliance. Extensive evaluations on a real-world testbed demonstrate the effectiveness of our approach. It achieves an effective throughput increase of up to 10× compared to the baselines and shows better robustness in challenging scenarios. OctopInf can be used for any DNN-based EVA inference task with minimal adaptation and is available at https://github.com/tungngreen/PipelineScheduler."},"title":{"value":"OctopInf: Workload-Aware Inference Serving for Edge Video Analytics"},"authors":{"value":["Thanh-Tung Nguyen","Lucas Liebe","Nhat-Quang Tau","Yuheng Wu","Jinghan Cheng","Dongman Lee"]}},"tmdate":1768385166411,"pdate":1735689600000,"externalIds":["dblp:conf/percom/NguyenLTWCL24"],"tcdate":1762940341654,"writers":["~"],"signatures":["~Yuheng_Wu1"],"forum":"hjLlWYwMsQ","license":"CC BY-SA 4.0","number":693205,"cdate":1735689600000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1768385166411,"domain":"DBLP.org","id":"hjLlWYwMsQ","version":2},{"content":{"venue":{"value":"CoRR 2025"},"pdf":{"value":"http://arxiv.org/pdf/2502.01277v1"},"venueid":{"value":"dblp.org/journals/CORR/2025"},"paperhash":{"value":"nguyen|octopinf_workloadaware_inference_serving_for_edge_video_analytics"},"authorids":{"value":["~Thanh-Tung_Nguyen2","https://dblp.org/search/pid/api?q=author:Lucas_Liebe:","~Nhat-Quang_Tau1","https://dblp.org/search/pid/api?q=author:Yuheng_Wu:","https://dblp.org/search/pid/api?q=author:Jinghan_Cheng:","https://dblp.org/search/pid/api?q=author:Dongman_Lee:"]},"html":{"value":"https://doi.org/10.48550/arXiv.2502.01277"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2502-01277,\n  publtype={informal},\n  author={Thanh-Tung Nguyen and Lucas Liebe and Nhat-Quang Tau and Yuheng Wu and Jinghan Cheng and Dongman Lee},\n  title={OCTOPINF: Workload-Aware Inference Serving for Edge Video Analytics},\n  year={2025},\n  month={February},\n  cdate={1738368000000},\n  journal={CoRR},\n  volume={abs/2502.01277},\n  url={https://doi.org/10.48550/arXiv.2502.01277}\n}\n"},"abstract":{"value":"Edge Video Analytics (EVA) has gained significant attention as a major application of pervasive computing, enabling real-time visual processing. EVA pipelines, composed of deep neural networks (DNNs), typically demand efficient inference serving under stringent latency requirements, which is challenging due to the dynamic Edge environments (e.g., workload variability and network instability). Moreover, EVA pipelines also face significant resource contention caused by resource (e.g., GPU) constraints at the Edge. In this paper, we introduce OCTOPINF, a novel resource-efficient and workload-aware inference serving system designed for real-time EVA. OCTOPINF tackles the unique challenges of dynamic edge environments through fine-grained resource allocation, adaptive batching, and workload balancing between edge devices and servers. Furthermore, we propose a spatiotemporal scheduling algorithm that optimizes the co-location of inference tasks on GPUs, improving performance and ensuring service-level objectives (SLOs) compliance. Extensive evaluations on a real-world testbed demonstrate the effectiveness of our approach. It achieves an effective throughput increase of up to 10x compared to the baselines and shows better robustness in challenging scenarios. OCTOPINF can be used for any DNN-based EVA inference task with minimal adaptation and is available at https://github.com/tungngreen/PipelineScheduler."},"title":{"value":"OCTOPINF: Workload-Aware Inference Serving for Edge Video Analytics"},"authors":{"value":["Thanh-Tung Nguyen","Lucas Liebe","Nhat-Quang Tau","Yuheng Wu","Jinghan Cheng","Dongman Lee"]}},"tmdate":1744095584392,"pdate":1735689600000,"tcdate":1744089242609,"writers":["~"],"signatures":["~Thanh-Tung_Nguyen2"],"forum":"gx0nHiKpBn","license":"CC BY-SA 4.0","number":380900,"cdate":1738368000000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1744095584392,"domain":"DBLP.org","id":"gx0nHiKpBn","version":2},{"content":{"venue":{"value":"CoRR 2025"},"pdf":{"value":"https://arxiv.org/pdf/2504.10551v1"},"venueid":{"value":"dblp.org/journals/CORR/2025"},"paperhash":{"value":"zhao|mimu_mitigating_multiple_shortcut_learning_behavior_of_transformers"},"authorids":{"value":["~Lili_Zhao3","~Qi_Liu3","~Wei_Chen54","https://dblp.org/search/pid/api?q=author:Liyi_Chen_0001:","https://dblp.org/search/pid/api?q=author:Ruijun_Sun:","~Min_Hou1","https://dblp.org/search/pid/api?q=author:Yang_Wang_0001:","https://dblp.org/search/pid/api?q=author:Shijin_Wang_0001:"]},"html":{"value":"https://doi.org/10.48550/arXiv.2504.10551"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2504-10551,\n  publtype={informal},\n  author={Lili Zhao and Qi Liu and Wei Chen and Liyi Chen and Ruijun Sun and Min Hou and Yang Wang and Shijin Wang},\n  title={MiMu: Mitigating Multiple Shortcut Learning Behavior of Transformers},\n  year={2025},\n  month={April},\n  cdate={1743465600000},\n  journal={CoRR},\n  volume={abs/2504.10551},\n  url={https://doi.org/10.48550/arXiv.2504.10551}\n}\n"},"abstract":{"value":"Empirical Risk Minimization (ERM) models often rely on spurious correlations between features and labels during the learning process, leading to shortcut learning behavior that undermines robustness generalization performance. Current research mainly targets identifying or mitigating a single shortcut; however, in real-world scenarios, cues within the data are diverse and unknown. In empirical studies, we reveal that the models rely to varying extents on different shortcuts. Compared to weak shortcuts, models depend more heavily on strong shortcuts, resulting in their poor generalization ability. To address these challenges, we propose MiMu, a novel method integrated with Transformer-based ERMs designed to Mitigate Multiple shortcut learning behavior, which incorporates self-calibration strategy and self-improvement strategy. In the source model, we preliminarily propose the self-calibration strategy to prevent the model from relying on shortcuts and make overconfident predictions. Then, we further design self-improvement strategy in target model to reduce the reliance on multiple shortcuts. The random mask strategy involves randomly masking partial attention positions to diversify the focus of target model other than concentrating on a fixed region. Meanwhile, the adaptive attention alignment module facilitates the alignment of attention weights to the calibrated source model, without the need for post-hoc attention maps or supervision. Finally, extensive experiments conducted on Natural Language Processing (NLP) and Computer Vision (CV) demonstrate the effectiveness of MiMu in improving robustness generalization abilities."},"title":{"value":"MiMu: Mitigating Multiple Shortcut Learning Behavior of Transformers"},"authors":{"value":["Lili Zhao","Qi Liu","Wei Chen","Liyi Chen","Ruijun Sun","Min Hou","Yang Wang","Shijin Wang"]}},"tmdate":1769314574508,"pdate":1735689600000,"externalIds":["dblp:journals/corr/abs-2504-10551"],"tcdate":1768790904868,"writers":["~"],"signatures":["~Qi_Liu3"],"forum":"zP4Dy4SDag","license":"CC BY-SA 4.0","number":770641,"cdate":1743465600000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1769314574508,"domain":"DBLP.org","id":"zP4Dy4SDag","version":2},{"content":{"TLDR":{"value":"Potential shortcuts can be found by monitoring the easy features learned by the initial layers of a DNN early during the training using suitable instance difficulty metrics."},"venue":{"value":"SCIS 2023 Poster"},"pdf":{"value":"/pdf/2cb12364882b3cd8bf2a5f13428c163979ab70bd.pdf"},"keywords":{"value":["shortcut learning","spurious correlations","convolutional neural networks","deep learning","machine learning","computer vision","training dynamics"]},"venueid":{"value":"ICML.cc/2023/Workshop/SCIS"},"abstract":{"value":"Deep Neural Networks (DNNs) are prone to learning *shortcut* patterns that damage the generalization of the DNN during deployment. This paper aims to better understand shortcut learning through the lens of the learning dynamics of the internal neurons during the training process. We make the following observations: (1) While previous works treat shortcuts as synonymous with spurious correlations, we emphasize that not all spurious correlations are shortcuts. We show that shortcuts are only those spurious features that are \"easier\" than the core features. (2) We build upon this premise and use *instance difficulty* methods (like Prediction Depth) to quantify \"easy\" and to identify this behavior during the training phase. (3) We empirically show that shortcut learning can be detected by observing the learning dynamics of the DNN's *early layers*. In other words, easy features learned by the initial layers of a DNN early during the training are potential shortcuts. We verify our claims on medical and vision datasets, both simulated and real, and justify the empirical success of our hypothesis by showing the theoretical connections between Prediction Depth and information-theoretic concepts like $\\mathcal{V}$-usable information. Lastly, our experiments show the insufficiency of monitoring only accuracy plots during training (as is common in machine learning pipelines). We highlight the need for monitoring early training dynamics using example difficulty metrics."},"title":{"value":"Shortcut Learning Through the Lens of Training Dynamics"}},"tmdate":1690513107398,"pdate":1687228755900,"tcdate":1684512151823,"writers":["ICML.cc/2023/Workshop/SCIS","ICML.cc/2023/Workshop/SCIS/Submission4/Authors"],"signatures":["ICML.cc/2023/Workshop/SCIS/Submission4/Authors"],"forum":"35AJGbhnDe","number":4,"cdate":1684512151823,"mdate":1690513107398,"readers":["everyone"],"invitations":["ICML.cc/2023/Workshop/SCIS/-/Submission","ICML.cc/2023/Workshop/SCIS/-/Post_Submission","ICML.cc/2023/Workshop/SCIS/-/Edit","ICML.cc/2023/Workshop/SCIS/Submission4/-/Post-Decision_Revision"],"domain":"ICML.cc/2023/Workshop/SCIS","id":"35AJGbhnDe","version":2},{"content":{"summary":{"value":"This paper presents an efficient Video LLM, coined as BLIP-3-Video, by incorporating extra modules to transform dense video tokens into sparse tokens (e.g., # can be 32). Specifically, the authors use a sequential transformer (so called token turing machine) on top of frame-wise image tokens and perceiver-resampler to produce limited 32 tokens. Diverse video QA benchmarks show the competitive performance of BLIP-3-Video on video question answering tasks. The authors also test the captioning ability of BLIP-3-Video."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. Can one BLIP-3-Video model produce 32 and 128 tokens on demand as a hyper parameter? Or they are different models trained individually?\n\n2. What does text encoder in Fig 2 mean? The text tokenizer?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":3},"strengths":{"value":"This paper investigates an important topic, how to efficiently & effectively understand videos by LLMs, which is underexplored so far. This paper proposes the BLIP-3-Video to understand videos with just 32 tokens (in LLMs) based on Phi-3, and it outperforms both parameter-heavy or visual-token-heavy models (as shown in Fig 1) on QA and captioning benchmarks. Additionally, this paper presents a compelling finding (somewhat): a video can be effectively represented by just 32 tokens in LLMs for QA and captioning tasks. I think this research line is promising and can benefit several downstreaming tasks, e.g., captioning for text-to-video generation."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Although this is overall a good paper, several concerns are here:\n1. The presentation of the main method (Sec 2.2) somwhat presents confusion: Does BLIP-3-Video both use spatio-temporal attentional pooling and TTM? Is there a perceive resampler before temporal encoder in BLIP-3-Video (cannot be infered from Figure 2)?\n\n2. Compressing a video into 32 tokens is a compelling and exciting idea. However, I am worried that spatial-temporal details will be missing through compression, which is crucial for some detailed reasoning in LLMs. More evaluation of BLIP-3-Video on diverse tasks beyond captioning and MCQ are encouraged.\n(also, as the compression is not text query guided, the compression is solely dominated by the visual information itself. That is to say, 32 tokens per a video are fixed under different text query, which might not be appropriate in general)\n\n3. Relating with the weakness 1, extra modules introduced besides the visual encoder (SigLIP) and LLM sound too complicated. If I understand correctly, there are a perceiver-resampler and a temporal encoder (attention pooling or TTM). My idea is naive and simple, can we just finetune a perceiver-resampler in BLIP-3 into a temporal encoder, rather than just compressing tokens per frame? Given the strong performance of cross attention layers in the perceiver resampler, this seems to be a missing but promising ablation study in this paper."}},"nonreaders":[],"tmdate":1731428344367,"tcdate":1730555422495,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission5214/Reviewer_gZaZ"],"signatures":["ICLR.cc/2025/Conference/Submission5214/Reviewer_gZaZ"],"forum":"CKYsXi0dOV","number":2,"license":"CC BY 4.0","cdate":1730555422495,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission5214/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731428344367,"domain":"ICLR.cc/2025/Conference","replyto":"CKYsXi0dOV","id":"FVGFVCy5Ft","forumContent":{"TLDR":{"value":"We present BLIP-3-Video, which is a compact video-based VLM using 20x less visual tokens compared to the standard models."},"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["video representation","video foundation model","vlm","multimodal language model"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"We present BLIP-3-Video, a multimodal language model for videos, particularly designed to efficiently capture temporal information over multiple frames. BLIP-3-Video takes advantage of the `temporal encoder' in addition to the conventional visual tokenizer, which maps a sequence of tokens over multiple frames into a compact set of visual tokens. This enables BLIP-3-Video to use much fewer visual tokens than its competing models (e.g., 32 vs. 4608 tokens). We explore different types of temporal encoders, including learnable spatio-temporal pooling as well as sequential models like Token Turing Machines. We experimentally confirm that BLIP-3-Video obtains video question-answering accuracies comparable to much larger state-of-the-art models (e.g., 34B), while being much smaller (i.e., 4B) and more efficient by using fewer visual tokens."},"_bibtex":{"value":"@misc{\nryoo2025blipvideo,\ntitle={{BLIP}-3-Video: You Only Need 32 Tokens to Represent a Video Even in {VLM}s},\nauthor={Michael S Ryoo and Honglu Zhou and Shrikant Kendre and Can Qin and Le Xue and Manli Shu and Silvio Savarese and Ran Xu and Caiming Xiong and Juan Carlos Niebles},\nyear={2025},\nurl={https://openreview.net/forum?id=CKYsXi0dOV}\n}"},"title":{"value":"BLIP-3-Video: You Only Need 32 Tokens to Represent a Video Even in VLMs"},"pdf":{"value":"/pdf/51b0026137447689d180516b521f8c18a76edb70.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"ryoo|blip3video_you_only_need_32_tokens_to_represent_a_video_even_in_vlms"},"authorids":{"value":["~Michael_S_Ryoo1","~Honglu_Zhou1","~Shrikant_Kendre1","~Can_Qin1","~Le_Xue1","~Manli_Shu1","~Silvio_Savarese1","~Ran_Xu1","~Caiming_Xiong1","~Juan_Carlos_Niebles1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Michael S Ryoo","Honglu Zhou","Shrikant Kendre","Can Qin","Le Xue","Manli Shu","Silvio Savarese","Ran Xu","Caiming Xiong","Juan Carlos Niebles"]}},"version":2},{"content":{"venue":{"value":"CoRR 2024"},"pdf":{"value":"http://arxiv.org/pdf/2407.07478v1"},"venueid":{"value":"dblp.org/journals/CORR/2024"},"paperhash":{"value":"ma|eavtr_eventaware_videotext_retrieval"},"authorids":{"value":["","https://dblp.org/search/pid/api?q=author:Ziqi_Zhang:","https://dblp.org/search/pid/api?q=author:Yuxin_Chen:","https://dblp.org/search/pid/api?q=author:Zhongang_Qi:","https://dblp.org/search/pid/api?q=author:Chunfeng_Yuan:","~Bing_Li1","https://dblp.org/search/pid/api?q=author:Yingmin_Luo:","https://dblp.org/search/pid/api?q=author:Xu_Li_0015:","https://dblp.org/search/pid/api?q=author:Xiaojuan_Qi_0001:","https://dblp.org/search/pid/api?q=author:Ying_Shan:","~Weiming_Hu1"]},"html":{"value":"https://doi.org/10.48550/arXiv.2407.07478"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2407-07478,\n  publtype={informal},\n  author={Zongyang Ma and Ziqi Zhang and Yuxin Chen and Zhongang Qi and Chunfeng Yuan and Bing Li and Yingmin Luo and Xu Li and Xiaojuan Qi and Ying Shan and Weiming Hu},\n  title={EA-VTR: Event-Aware Video-Text Retrieval},\n  year={2024},\n  cdate={1704067200000},\n  journal={CoRR},\n  volume={abs/2407.07478},\n  url={https://doi.org/10.48550/arXiv.2407.07478}\n}\n"},"abstract":{"value":"Understanding the content of events occurring in the video and their inherent temporal logic is crucial for video-text retrieval. However, web-crawled pre-training datasets often lack sufficient event information, and the widely adopted video-level cross-modal contrastive learning also struggles to capture detailed and complex video-text event alignment. To address these challenges, we make improvements from both data and model perspectives. In terms of pre-training data, we focus on supplementing the missing specific event content and event temporal transitions with the proposed event augmentation strategies. Based on the event-augmented data, we construct a novel Event-Aware Video-Text Retrieval model, ie, EA-VTR, which achieves powerful video-text retrieval ability through superior video event awareness. EA-VTR can efficiently encode frame-level and video-level visual representations simultaneously, enabling detailed event content and complex event temporal cross-modal alignment, ultimately enhancing the comprehensive understanding of video events. Our method not only significantly outperforms existing approaches on multiple datasets for Text-to-Video Retrieval and Video Action Recognition tasks, but also demonstrates superior event content perceive ability on Multi-event Video-Text Retrieval and Video Moment Retrieval tasks, as well as outstanding event temporal logic understanding ability on Test of Time task."},"title":{"value":"EA-VTR: Event-Aware Video-Text Retrieval"},"authors":{"value":["Zongyang Ma","Ziqi Zhang","Yuxin Chen","Zhongang Qi","Chunfeng Yuan","Bing Li","Yingmin Luo","Xu Li","Xiaojuan Qi","Ying Shan","Weiming Hu"]}},"tmdate":1737949823592,"pdate":1704067200000,"tcdate":1730301041940,"writers":["~"],"signatures":["~Zongyang_Ma3"],"forum":"O0pUTxDe3Y","license":"CC BY-SA 4.0","number":164692,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1737949823592,"domain":"DBLP.org","id":"O0pUTxDe3Y","version":2},{"content":{"venue":{"value":"ECCV (52) 2024"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-031-72943-0_5.pdf"},"venueid":{"value":"dblp.org/conf/ECCV/2024"},"paperhash":{"value":"ma|eavtr_eventaware_videotext_retrieval"},"authorids":{"value":["~Zongyang_Ma3","https://dblp.org/search/pid/api?q=author:Ziqi_Zhang:","https://dblp.org/search/pid/api?q=author:Yuxin_Chen:","https://dblp.org/search/pid/api?q=author:Zhongang_Qi:","https://dblp.org/search/pid/api?q=author:Chunfeng_Yuan:","https://dblp.org/search/pid/api?q=author:Bing_Li_0001:","https://dblp.org/search/pid/api?q=author:Yingmin_Luo:","https://dblp.org/search/pid/api?q=author:Xu_Li_0015:","https://dblp.org/search/pid/api?q=author:Xiaojuan_Qi_0001:","https://dblp.org/search/pid/api?q=author:Ying_Shan:","https://dblp.org/search/pid/api?q=author:Weiming_Hu:"]},"html":{"value":"https://doi.org/10.1007/978-3-031-72943-0_5"},"_bibtex":{"value":"@inproceedings{DBLP:conf/eccv/MaZCQYLLLQSH24,\n  author={Zongyang Ma and Ziqi Zhang and Yuxin Chen and Zhongang Qi and Chunfeng Yuan and Bing Li and Yingmin Luo and Xu Li and Xiaojuan Qi and Ying Shan and Weiming Hu},\n  title={EA-VTR: Event-Aware Video-Text Retrieval},\n  year={2024},\n  cdate={1704067200000},\n  pages={76-94},\n  url={https://doi.org/10.1007/978-3-031-72943-0_5},\n  booktitle={ECCV (52)},\n  crossref={conf/eccv/2024-52}\n}\n"},"abstract":{"value":"Understanding the content of events occurring in the video and their inherent temporal logic is crucial for video-text retrieval. However, web-crawled pre-training datasets often lack sufficient event information, and the widely adopted video-level cross-modal contrastive learning also struggles to capture detailed and complex video-text event alignment. To address these challenges, we make improvements from both data and model perspectives. In terms of pre-training data, we focus on supplementing the missing specific event content and event temporal transitions with the proposed event augmentation strategies. Based on the event-augmented data, we construct a novel Event-Aware Video-Text Retrieval model, i.e., EA-VTR, which achieves powerful video-text retrieval ability through superior video event awareness. EA-VTR can efficiently encode frame-level and video-level visual representations simultaneously, enabling detailed event content and complex event temporal cross-modal alignment, ultimately enhancing the comprehensive understanding of video events. Our method not only significantly outperforms existing approaches on multiple datasets for Text-to-Video Retrieval and Video Action Recognition tasks, but also demonstrates superior event content perceive ability on Multi-event Video-Text Retrieval and Video Moment Retrieval tasks, as well as outstanding event temporal logic understanding ability on Test of Time task."},"title":{"value":"EA-VTR: Event-Aware Video-Text Retrieval"},"authors":{"value":["Zongyang Ma","Ziqi Zhang","Yuxin Chen","Zhongang Qi","Chunfeng Yuan","Bing Li","Yingmin Luo","Xu Li","Xiaojuan Qi","Ying Shan","Weiming Hu"]}},"tmdate":1737949775819,"pdate":1704067200000,"tcdate":1737949772531,"writers":["~"],"signatures":["~Zongyang_Ma3"],"forum":"V5yPE9iX7A","license":"CC BY-SA 4.0","number":282562,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1737949775819,"domain":"DBLP.org","id":"V5yPE9iX7A","version":2},{"content":{"summary":{"value":"The paper proposes using MCTS for planning-based long video generation, which expands an important direction in the TTT field. Through this approach, the paper even achieves long video generation results that surpass closed-source SOTA models, demonstrating the potential of TTT in long video generation."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"See the Weaknesses section."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":4},"strengths":{"value":"- The work has a certain degree of novelty and community value. The paper is the first to apply MCTS-based TTT to long video generation, showcasing the value of classical methods in the video domain.\n- The experimental results are impressive. The proposed method enables Cosmos-Predict2 to surpass or tie with closed-source SOTA models (Sora/Kling), which demonstrates the strong potential of TTT."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Tab. 5 should include a comparison of the computational cost.\n- Regarding the long-video baselines, the paper would be more sound if a more comprehensive set could be included [1,2]\n\n[1] FIFO-Diffusion: Generating Infinite Videos from Text without Training\n\n[2] Skyreels-v2: Infinite-length film generative model\n\n- The paper lacks discussion and comparison with several recently accepted works on long-video generation.\n\n[1] Zhao et al., Riflex: A Free Lunch for Length Extrapolation in Video Diffusion Transformers (ICML 2025).\n\n[2] Tan et al., FreePCA: Integrating Consistency Information Across Long-Short Frames in Training-Free Long Video Generation via Principal Component Analysis (CVPR 2025).\n\n[3] Lu et al., FreeLong: Training-Free Long Video Generation with SpectralBlend Temporal Attention (NeurIPS 2024).\n\n[4] Cai et al., DitCtrl: Exploring Attention Control in Multi-Modal Diffusion Transformer for Tuning-Free Multi-Prompt Longer Video Generation (CVPR 2025)."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762941983632,"tcdate":1761977464734,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission21926/Reviewer_UF3o"],"signatures":["ICLR.cc/2026/Conference/Submission21926/Reviewer_UF3o"],"forum":"ilir6A52vh","number":2,"license":"CC BY 4.0","cdate":1761977464734,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission21926/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762941983632,"domain":"ICLR.cc/2026/Conference","replyto":"ilir6A52vh","id":"agahc7Vbmo","forumContent":{"TLDR":{"value":"We leverage Monte Carlo Tree Search–based test-time scaling to select better continuations, enabling the generation of coherent long videos."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Video Generation","Long Video Generation","Test Time Scaling","Monte Carlo Tree Search"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Generating long videos with consistent content and visual quality remains a ma-\njor challenge, as existing one-shot and chunked methods often suffer from se-\nmantic drift and compounding artifacts. We explore Test-Time Scaling (TTS)\nas a framework for long video generation, formulating the task as a sequential\ndecision-making problem. Our approach uses Monte Carlo Tree Search (MCTS)\nto evaluate multiple continuations with look-ahead rollouts and backpropagated\nrewards, and we introduce a Multi-Tree MCTS variant that improves exploration\nin continuous generation spaces. The method is modular and can be applied to ex-\nisting backbones without retraining. Experiments on Cosmos-Predict2 and other\nmodels show consistent improvements in object permanence, temporal coherence,\nand text-video alignment over Best-of-N, Greedy, and Beam search. Furthermore,\nour method produces high-quality videos exceeding 20 seconds, surpassing the\noutput of leading models like Sora and Kling by 18% and 47% respectively, all\nwhile maintaining comparable visual fidelity. Although the results are limited\nby the quality of current generators and verifiers, our study highlights both the\npromise of search-based TTS and the limitations of today’s video generation and\nevaluation models."},"_bibtex":{"value":"@misc{\nbale2026planning,\ntitle={Planning at Inference: {MCTS} Test-Time Scaling for Long Video Generation},\nauthor={Ritvik Bale and Ethan He and Ashwath Aithal and Linnan Wang},\nyear={2026},\nurl={https://openreview.net/forum?id=ilir6A52vh}\n}"},"title":{"value":"Planning at Inference: MCTS Test-Time Scaling for Long Video Generation"},"pdf":{"value":"/pdf/f1fa4e58f6fc3e1b843f6d1596141449216660b9.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"bale|planning_at_inference_mcts_testtime_scaling_for_long_video_generation"},"authorids":{"value":["~Ritvik_Bale1","~Ethan_He1","~Ashwath_Aithal1","~Linnan_Wang2"]},"authors":{"value":["Ritvik Bale","Ethan He","Ashwath Aithal","Linnan Wang"]}},"version":2},{"content":{"venue":{"value":"MMM (2) 2025"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-981-96-2061-6_19.pdf"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"kwon|lightweight_motionaware_video_superresolution_for_compressed_videos"},"html":{"value":"https://doi.org/10.1007/978-981-96-2061-6_19"},"_bibtex":{"value":"@inproceedings{DBLP:conf/mmm/KwonLSP25,\n  author={Ilhwan Kwon and Jun Li and Rajiv Ratn Shah and Mukesh Prasad},\n  title={Lightweight Motion-Aware Video Super-Resolution for Compressed Videos},\n  year={2025},\n  cdate={1735689600000},\n  pages={254-267},\n  url={https://doi.org/10.1007/978-981-96-2061-6_19},\n  booktitle={MMM (2)},\n  crossref={conf/mmm/2025-2}\n}\n"},"abstract":{"value":"This study introduces a lightweight and efficient Video Super-Resolution (VSR) model that leverages codec prior information, focusing on motion. Unlike existing VSR models that aim for superior performance through complex structures, the proposed model is designed to be suitable for videos with significant object motion by conducting extensive training with a large dataset. Input of the model, in line with the real-time nature of video services, is compressed video data. By incorporating compressed information and affine transform considerations into the model structure, we have broadened the utilization of spatial-temporal information. Despite its lightweight parameters and minimal computational resource requirements, the proposed model demonstrates meaningful results when applied to videos with substantial motion. This marks a significant departure from other VSR models that demand extensive parameters and computation time, thereby highlighting the potential for implementing a small yet powerful VSR model with enhanced real-time applicability."},"title":{"value":"Lightweight Motion-Aware Video Super-Resolution for Compressed Videos"},"authors":{"value":[{"fullname":"Ilhwan Kwon","username":""},{"fullname":"Jun Li","username":""},{"fullname":"Rajiv Ratn Shah","username":"~Rajiv_Ratn_Shah1"},{"fullname":"Mukesh Prasad","username":""}]}},"tmdate":1785924403000,"pdate":1767139200000,"externalIds":["dblp:conf/mmm/KwonLSP25"],"tcdate":1785924389041,"writers":["~","OpenReview.net/Public_Article/DBLP.org","OpenReview.net/Support"],"signatures":["~Rajiv_Ratn_Shah1"],"forum":"z5lHCX6TcV","license":"CC BY-SA 4.0","number":122339,"cdate":1735689600000,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/DBLP.org/-/Record","OpenReview.net/Public_Article/-/Edit"],"mdate":1785924403000,"domain":"OpenReview.net/Public_Article","id":"z5lHCX6TcV","version":2},{"content":{"venue":{"value":"Crossref"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-981-96-2061-6_19.pdf"},"venueid":{"value":"OpenReview.net/Public_Article"},"paperhash":{"value":"kwon|lightweight_motionaware_video_superresolution_for_compressed_videos"},"html":{"value":"https://doi.org/10.1007/978-981-96-2061-6_19"},"abstract":{"value":"This study introduces a lightweight and efficient Video Super-Resolution (VSR) model that leverages codec prior information, focusing on motion. Unlike existing VSR models that aim for superior performance through complex structures, the proposed model is designed to be suitable for videos with significant object motion by conducting extensive training with a large dataset. Input of the model, in line with the real-time nature of video services, is compressed video data. By incorporating compressed information and affine transform considerations into the model structure, we have broadened the utilization of spatial-temporal information. Despite its lightweight parameters and minimal computational resource requirements, the proposed model demonstrates meaningful results when applied to videos with substantial motion. This marks a significant departure from other VSR models that demand extensive parameters and computation time, thereby highlighting the potential for implementing a small yet powerful VSR model with enhanced real-time applicability."},"title":{"value":"Lightweight Motion-Aware Video Super-Resolution for Compressed Videos"},"authors":{"value":[{"fullname":"Ilhwan Kwon","username":"https://orcid.org/orcid-search/search?searchQuery=Ilhwan%20Kwon"},{"fullname":"Jun Li","username":"~Jun_Li18"},{"fullname":"Rajiv Ratn Shah","username":"https://orcid.org/orcid-search/search?searchQuery=Rajiv%20Ratn%20Shah"},{"fullname":"Mukesh Prasad","username":"https://orcid.org/orcid-search/search?searchQuery=Mukesh%20Prasad"}]}},"tmdate":1778157178794,"pdate":1735689600000,"externalIds":["doi:10.1007/978-981-96-2061-6_19"],"tcdate":1778157173718,"writers":["~","OpenReview.net/Public_Article/ORCID.org","OpenReview.net/Support"],"signatures":["~Jun_Li18"],"forum":"JMl6MZT6Uv","license":"CC BY-SA 4.0","number":68624,"cdate":1738474697263,"readers":["everyone"],"invitations":["OpenReview.net/Public_Article/ORCID.org/-/Record","OpenReview.net/Public_Article/-/Edit"],"mdate":1778157178794,"domain":"OpenReview.net/Public_Article","id":"JMl6MZT6Uv","version":2},{"content":{"summary":{"value":"The authors train their own video diffusion models on synthetic video models and run thorough analysis on how the physics based video generation."},"soundness":{"value":1},"confidence":{"value":4},"questions":{"value":"* In what real world cases would the data be significantly outside of the distribution of the training data. Should the solution not be scaling the number of parameters or the number of data, but rather the diversity of the data then?\n\n* What aspects of the behavior you observation are purely artifacts of the latent encoder / decoder?\n\n* Why did you not finetune any existing video models on these tasks? Surely this task is entirely OOD from the training task since it's a synthetic environment. Furthermore, if as a researcher your goal is to train the best video world model possible, why would you not start with a pretrained model. If you are confident that these problems cannot be solved purely by scaling, than pretrained video models on real world images should not generalize to this simple benchmark and should observe similar biases as in this paper. If you are making a claim that this behavior can be observed on all video models, than you should be able to evaluate this lack of generalizations on existing pre-trained models right?\n\n\n* We design datasets which delibrately leave out some latent values, i.e. velocity. After training, we test model’s prediction on both seen and unseen scenarios. We mainly focus on uniform motion and collision processes\nFrom an optimization point of view, what sparsity level is needed for the model to be able to linearly extrapolate then? After all, a diffusion model must also know what NOT to generate. There is surely some level sparsity for which the video model will generalize? What level of sparsity is it?\n\n* \"color > size > velocity > shape\" how much of this is simply affected by this explicit instantiation of the video VAE? What about other ones like Stable Video Diffusions? Or more recently released video model like COGVideo's?"},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"* The paper train many video diffusion models from scratch on small physical datasets across model parameter sizes. \n* They have thorough evaluation of extrapolation behavior\n* The paper attempts to tackle a very complex problem, by simplifying it to a synthetic data setting.\n* The analysis between in distribution and OOD seems useful, and this distinction could inform the creation of future video models, particularly the experiments in 5.4"},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"* Instead of training a \"small\" video model from scratch, why not try finetuning SOTA models on these video datasets? One issue with this analysis it supposes that there is not a minimum threshold for the number of parameters needed for a useful diffusion video model or for generalization to hold. I would not be surprised if these models generalized when simply having more parameters and train on more data, even out of domain data. Finetuning a model like SVD should be doable on a similar level of compute.\n\n* The rendering of the synthetic examples are overly simplistic and do not have imaging artifacts that video models could exploit in real world use cases to generalize, like motion blur.\n\n* Many of these reasoning weakness of generative models have been brought before in other domains. Such as Arc-AGE Challenge - \"On the Measure of Intelligence\" by François Chollet.\n\n*  \"For example, in Figure 10, it is difficult to determine if a ball can pass through a gap based on vision alone when the size difference\nis at the pixel level, leading to visually plausible but incorrect results. Similarly, visual ambiguity in a ball’s horizontal position relative to a block can result in different outcomes. These findings suggest that relying solely on visual representations, may be inadequate for accurate physics modeling.\" Pixel level differences will be obliterated by the VAE encoder. It is a compressive architecture, information will be lost to train the model more cheaply. Pixel level criteria is not a motivating example as a result. If you want pixel level accuracy, you need pixel level diffusion models, not one trained on latent. I would disregard these experiments or rewrite this entire section 5.5 as this is a fundamental issue with the model architecture, and reveals no new information. Can you at least verify that the VAE is able to reconstruct different images with pixel level differences? I think you will find that it will not, as Stable Diffusion's VAE architecture cannot. The number of channels is way too limiting which is why it's been massively increased in more recent models like BlackForrest's FLUX.\n\n* \" This ranking could explain why current video generation models often struggle with maintaining object consistency\" Please support with evidence? Or at least specify with which video models?\n\n\n* Our in-depth analysis suggests that video model generalization relies more on referencing similar training examples rather than learning universal rules.\nAll generative models exhibit this behavior, and it is well known. Similar observations can be made on small language models, but they usually generalize way better when scaled up in terms of compute. What is the new observation here with respect to physics?\n\n\nRather a better way to frame this paper might be what role do VAE reconstruction issues prevent us from using existing video world models for physics in critical settings? How do the video VAE's and the diffusion model prioritize shape color and velocity? The main issue here is that some of the experiments are clearly demonstrating architectural failures of the VAEs and the authors are attempting to generalize it to all video models, which is an massive overclaim. \n\nThis paper has useful experimental and scientific data, but it needs to be rewritten to support the claims in the paper. Furthermore, it could massively benefit by examining which of these issues are coming from the VAE and which are coming from the latent video diffusion.\n\nClaims are made that scaling cannot solve these physics problem. The evidence this paper shows that scaling parameters and perhaps date in the video diffusion model is ineffective if the VAE removes important information for the video model."}},"nonreaders":[],"tmdate":1731427313191,"tcdate":1730663711328,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission759/Reviewer_ym7u"],"signatures":["ICLR.cc/2025/Conference/Submission759/Reviewer_ym7u"],"forum":"ZyLkNVHBZF","number":3,"license":"CC BY 4.0","cdate":1730663711328,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission759/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427313191,"domain":"ICLR.cc/2025/Conference","replyto":"ZyLkNVHBZF","id":"4DOJ51QmAi","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"TLDR":{"value":"we conduct systematic experiments to investigate \"how far is video generation model from world model\" from the physical law perspetive."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["video generation","diffusion model","world model"]},"primary_area":{"value":"generative models"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"OpenAI's Sora highlights the potential of video generation for developing world models that adhere to fundamental physical laws. \nHowever, the ability of video generation models to discover such laws purely from visual data without human priors can be questioned.\nA world model learning the true law should give predictions robust to nuances and correctly extrapolate on unseen scenarios.\nIn this work, we evaluate across three key scenarios: in-distribution, out-of-distribution, and combinatorial generalization.\nWe developed a 2D simulation testbed for object movement and collisions to generate videos deterministically governed by one or more classical mechanics laws.\nThis provides unlimited supply of data for large-scale experimentation, and enables quantitative evaluation for the law in generated videos. \nWe trained diffusion-based video generation models to predict object movements based on initial frames.\nOur scaling experiments show perfect generalization within the distribution, measurable scaling behavior for combinatorial generalization, but failure in out-of-distribution scenarios.\nFurther experiments reveal two key insights about the generalization mechanisms of these models: (1) the models fail to abstract general physical rules and instead exhibit ``case-based'' generalization behavior, \\textit{i.e.}, mimicking the closest training example; (2) when generalizing to new cases, models are observed to prioritize different factors when referencing training data: color $>$ size $>$ velocity $>$ shape.\nOur study suggests that scaling alone is insufficient for video generation models to uncover fundamental physical laws, despite its role in Sora's broader success."},"_bibtex":{"value":"@misc{\nkang2025how,\ntitle={How Far Is Video Generation from World Model: A Physical Law Perspective},\nauthor={Bingyi Kang and Yang Yue and Rui Lu and Zhijie Lin and Yang Zhao and Kaixin Wang and Gao Huang and Jiashi Feng},\nyear={2025},\nurl={https://openreview.net/forum?id=ZyLkNVHBZF}\n}"},"title":{"value":"How Far Is Video Generation from World Model: A Physical Law Perspective"},"pdf":{"value":"/pdf/61786c392aeb1743704ab9483f7f67d63ce837d2.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"kang|how_far_is_video_generation_from_world_model_a_physical_law_perspective"},"authorids":{"value":["~Bingyi_Kang1","~Yang_Yue1","~Rui_Lu2","~Zhijie_Lin1","~Yang_Zhao14","~Kaixin_Wang1","~Gao_Huang1","~Jiashi_Feng1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Bingyi Kang","Yang Yue","Rui Lu","Zhijie Lin","Yang Zhao","Kaixin Wang","Gao Huang","Jiashi Feng"]}},"version":2},{"content":{"venue":{"value":"CVPR 2026"},"abstract":{"value":"Existing video editing methods face a critical trade-off: expert models offer precision but rely on task-specific priors like masks, hindering unification; conversely, unified temporal in-context learning models are mask-free but lack explicit spatial cues, leading to weak instruction-to-region mapping and imprecise localization. To resolve this conflict, we propose VideoCoF, a novel Chain-of-Frames approach inspired by Chain-of-Thought reasoning. VideoCoF enforces a \"seeing, reasoning, then editing\" procedure by compelling the video diffusion model to first predict reasoning tokens (edit-region latents) before generating the target video tokens. This explicit reasoning step removes the need for user-provided masks while achieving precise instruction-to-region alignment and fine-grained video editing. Furthermore, we introduce a RoPE alignment strategy that leverages these reasoning tokens to ensure motion alignment and enable length extrapolation beyond the training duration. We demonstrate that with a minimal data cost of only 50k video pairs, VideoCoF achieves state-of-the-art performance on VideoCoF-Bench, validating the efficiency and effectiveness of our approach. Our code, weight, data are available at https://github.com/knightyxp/VideoCoF."},"_bibtex":{"value":"@inproceedings{\nyang2026videocof,\ntitle={VideoCoF: Unified Video Editing with Temporal Reasoner},\nauthor={Xiangpeng Yang and Ji Xie and Yiyuan Yang and Yue Ma and Yan Huang and Min Xu and Qiang Wu},\nbooktitle={Conference on Computer Vision and Pattern Recognition 2026},\nyear={2026},\nurl={https://openreview.net/forum?id=AluZ1htNKz}\n}"},"title":{"value":"VideoCoF: Unified Video Editing with Temporal Reasoner"},"pdf":{"value":"https://openaccess.thecvf.com/content/CVPR2026/papers/Yang_VideoCoF_Unified_Video_Editing_with_Temporal_Reasoner_CVPR_2026_paper.pdf"},"venueid":{"value":"thecvf.com/CVPR/2026/Conference"},"paperhash":{"value":"yang|videocof_unified_video_editing_with_temporal_reasoner"},"authorids":{"value":["~Xiangpeng_Yang1","~Ji_Xie2","~yiyuan_yang2","~Yue_Ma2","~Yan_Huang14","~Min_Xu5","~Qiang_Wu2"]},"authors":{"value":["Xiangpeng Yang","Ji Xie","Yiyuan Yang","Yue Ma","Yan Huang","Min Xu","Qiang Wu"]}},"tmdate":1789656706162,"pdate":1789656438790,"tcdate":1765221321167,"writers":["thecvf.com/CVPR/2026/Conference","thecvf.com/CVPR/2026/Conference/Submission35506/Authors"],"signatures":["thecvf.com/CVPR/2026/Conference/Submission35506/Authors"],"forum":"AluZ1htNKz","license":"CC BY 4.0","number":35506,"cdate":1765221321167,"readers":["everyone"],"invitations":["thecvf.com/CVPR/2026/Conference/-/Submission","thecvf.com/CVPR/2026/Conference/Submission35506/-/Full_Submission","thecvf.com/CVPR/2026/Conference/-/Post_Submission","thecvf.com/CVPR/2026/Conference/-/Edit","thecvf.com/CVPR/2026/Conference/Submission35506/-/Supplementary_Material","thecvf.com/CVPR/2026/Conference/-/Compute_Flag"],"mdate":1789656706162,"odate":1789656438790,"domain":"thecvf.com/CVPR/2026/Conference","id":"AluZ1htNKz","version":2},{"content":{"summary":{"value":"This paper introduces a novel and highly efficient method for zero-shot, subject-driven video generation. The core contribution is a training strategy termed \"proxy experience replay,\" which is founded on the hypothesis that learning a subject's identity from images and learning temporal dynamics from videos are orthogonal objectives. The authors experimentally validate this hypothesis, demonstrating that the gradient conflict between these tasks naturally converges to near-zero.\nBased on this theoretical analysis, the author conducted an experiment and reduced the training cost of subject to video by using subject to image data mixed with a small amount of video data for training."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"See Weakness"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":1},"contribution":{"value":2},"strengths":{"value":"1.  The paper's idea of using image data and a small amount of video data to accomplish Subject-to-Video generation is good.\n2.  The method of analyzing the feasibility of using image data and a small amount of video data for the Subject-to-Video generation task from the perspective of gradients is very interesting."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Although the author's method of analyzing from the perspective of gradients is very interesting, I still have some questions: \n\n    a. The paper only provides the cosine similarity of the gradients of the two tasks during training when using proxy replay, lacking corresponding control experiments, i.e., training models with the S2I task and the T2V task separately, and counting the gradients individually to see if there is a conflict.\n\n    b. The 'gradient conflict' metric, defined as the cosine similarity $\\phi(t)$ in Eq. 2, is theoretically invariant to the scale of the gradients. However, this property may not hold in practice during the later stages of training. As the model converges, the magnitudes (L2 norms) of the gradients, $\\|\\nabla L_{\\text{img}}\\|$ and $\\|\\nabla L_{\\text{vid}}\\|$, tend to decay towards zero. When these magnitudes become sufficiently small, the computation of $\\phi(t)$ can become numerically unstable. The observed trend of $\\phi(t) \\to 0$ might therefore be an artifact of vanishing gradients, where near-zero vectors can appear spuriously orthogonal due to floating-point inaccuracies, rather than a genuine indicator of directional orthogonality. Specifically, the author should add the monitoring results of the gradient magnitudes. If $\\phi(t)$ decreases while $\\|\\text{g}_1\\|$ and $\\|\\text{g}_2\\|$ remain stable, it indicates that the directional change is real.\n2. The technical contribution is weak. Fundamentally, the method proposed in this paper is to train a video generation model using a mix of image and video data. However, this operation has been widely used in the training of various Text-to-Video foundation models, such as VideoCrafter2[1], Open-Sora[2], etc. Applying it to the S2V task is just a simple change of the training task.\n3. Many of the compared methods are based on UNet, such as Videobooth, Stilling Moving, Customcrafter. The base models of these methods have a large gap with the CogVideoX-5b used by the author. There is a lack of comparison with some new methods based on DiT architecture with comparable base model performance, such as ConsisID[3], VACE[4], and Phantom[5]. The author could consider using the existing OpenS2V-Eval BenchMark[6] for comparison.\n4. The generated results do not look very good visually. Are there results based on other foundation models, for example, Wan2.2?\n\n5. Minor Suggestions\n* Equations (2) and (3) are missing closing punctuation.\n* Please use the `\\citep` command from the ICLR template correctly for citations. The citation format in the paper is severely disorganized.\n\nIf the author addresses my concerns and revises the article, I will increase my score.\n\n[1] Videocrafter2: Overcoming data limitations for high-quality video diffusion models\n\n[2] Open-Sora Plan: Open-Source Large Video Generation Model\n\n[3] Identity-Preserving Text-to-Video Generation by Frequency Decomposition\n\n[4] VACE: All-in-One Video Creation and Editing\n\n[5] Phantom: Subject-Consistent Video Generation via Cross-Modal Alignment\n\n[6] OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762920542453,"tcdate":1760597566540,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission8758/Reviewer_sJdV"],"signatures":["ICLR.cc/2026/Conference/Submission8758/Reviewer_sJdV"],"forum":"xAaW436epC","number":1,"license":"CC BY 4.0","cdate":1760597566540,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission8758/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762920542453,"domain":"ICLR.cc/2026/Conference","replyto":"xAaW436epC","id":"NmmHQcaelo","forumContent":{"TLDR":{"value":"We employ continual learning with video replay, and adjust replay ratio dynamically, achieving on-par subject fidelity and motion with lower compute compared to state-of-the-art models."},"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["video generation","customization","personalization","diffusion models","continual learning"]},"supplementary_material":{"value":"/attachment/a075028e26247ef5388ca11118150c5daa1eb0b3.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"We aim to enable efficient subject-to-video (S2V) learning, which otherwise requires expensive video-subject-pair datasets that require tens of thousands of GPU hours for training. While utilizing image-paired datasets to train video models could address this challenge, naively training with image pairs results in catastrophic loss of temporal ability due to gradient conflicts. We hypothesize that S2V generation decomposes into two orthogonal objectives of identity learning from images and temporal dynamics from videos. Based on this orthogonality assumption, we design a stochastic task-switching strategy that predominantly samples from image datasets while maintaining minimal video replay for temporal coherence. Our experiments validate this hypothesis by demonstrating that the gradient inner product between tasks converges exponentially to near-zero, confirming emergent orthogonalization without requiring explicit orthogonal projection. This validated orthogonality enables efficient image-dominant training while preventing catastrophic forgetting through proxy experience replay. We employ regularization techniques including random frame selection and token dropping during video replay to ensure efficient temporal learning. Extensive experiments demonstrate our approach achieves superior performance with  comparable compute to per-subject tuned methods for single subjects, while providing zero-shot capability and outperforming both per-subject tuned methods and some existing zero-shot approaches."},"_bibtex":{"value":"@misc{\nkim2025subjectdriven,\ntitle={Subject-driven Video Generation Emerges from Experience Replays},\nauthor={Daneul Kim and Jingxu Zhang and Wonjoon Jin and Sunghyun Cho and Qi Dai and Jaesik Park and Chong Luo},\nyear={2025},\nurl={https://openreview.net/forum?id=xAaW436epC}\n}"},"title":{"value":"Subject-driven Video Generation Emerges from Experience Replays"},"pdf":{"value":"/pdf/8049c6a47793f0babda63c028ea8ef3e70ae36cf.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"kim|subjectdriven_video_generation_emerges_from_experience_replays"},"authorids":{"value":["~Daneul_Kim1","~Jingxu_Zhang1","~Wonjoon_Jin1","~Sunghyun_Cho3","~Qi_Dai4","~Jaesik_Park3","~Chong_Luo1"]},"authors":{"value":["Daneul Kim","Jingxu Zhang","Wonjoon Jin","Sunghyun Cho","Qi Dai","Jaesik Park","Chong Luo"]}},"version":2},{"content":{"summary":{"value":"This paper proposes a method to unlock the long video generation ability of a pretrained text-to-video generation model. The major technical components of this method include: 1. The analysis of artifacts and causes when generating long videos. 2. A noise schedule for long video generation. 3. Windowed attention fusion to keep attention perception field while avoiding content jump between windows, 4. Motion injection for varied textual prompts. The experiments show FreeNoise is a competitive long video generator."},"presentation":{"value":"3 good"},"contribution":{"value":"2 fair"},"soundness":{"value":"3 good"},"strengths":{"value":"This is a technically solid paper. Long video generation is a tricky long-standing problem. The authors propose a series of insights and techniques that have sufficient novelty to address the difficulties:\n1.\tThe analysis of long video artifacts and causes is valuable for developing better video generators. \n2.\tThe noise scheduling and window-based attention fusion address the long video difficulties mentioned in the analysis. They are simple yet effective. Window-based attention fusion addresses the notorious content jump problem, which will likely help develop future video generation foundation models.\n3.\tFreeNoise does not require additional UNet forward propagation. Therefore, the inference cost overshoot is low. \n4.\tThe qualitative results are marvelous. In human preference evaluation, FreeNoise still achieves the best. The authors provide an anonymous website to show more visual results. The image definition and motion consistency of FreeNoise are both good.\n5.\tThe motion injection technique successfully preserves video contents and drives the video to follow a new text prompt.\n6.\tThe qualitative ablations show each technical component of FreeNoise is effective and important."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Major concerns: \n1.\tI’m interested in detailed experiment settings. Please include diffusion sampler configurations, sample resolutions, frame stride, etc. in your future revision.\n2.\tPlease add a pipeline figure. It is not very easy to fully understand how and where FreeNoise is working on generating long videos.\n3.\tIs direct inference, sliding, GenL, and FreeNoise sharing the same pretrained text-to-video model? If it is, then the evaluation is very convincing since they can all generate the same short video using the prompts but only FreeNoise can achieve good long video results.\n\nMinor concerns:\n1.\tPage 7. In the second line. A full stop is missing before ‘Obviously’.\n2.\tIn which case FreeNoise may fail? A discussion over the limitations is welcomed.\n3.\t100 evaluation prompts have limited diversity. If it is feasible, please add more evaluation prompts to make the comparison more convincing."},"confidence":{"value":"4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work."},"questions":{"value":"1.\tThe authors claim FreeNoise achieves the maximum long video of 512 frames due to GPU memory limit. Is it possible to unlock even long video generation ability by using the CPU offload technique? \n2.\tWith the help of ControlNet, is it possible to generate more diverse motion with FreeNoise?"},"rating":{"value":"6: marginally above the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1699635988667,"tcdate":1698630899771,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission611/Reviewer_pTfs"],"signatures":["ICLR.cc/2024/Conference/Submission611/Reviewer_pTfs"],"forum":"ijoqFqSC7p","number":3,"license":"CC BY 4.0","cdate":1698630899771,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission611/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1699635988667,"domain":"ICLR.cc/2024/Conference","replyto":"ijoqFqSC7p","id":"louOef1EG3","forumContent":{"TLDR":{"value":"A tuning-free and time-efficient paradigm for longer video generation based on pretrained video diffusion models"},"venue":{"value":"ICLR 2024 poster"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["diffusion","video diffusion","video generation","tuning-free"]},"supplementary_material":{"value":"/attachment/fba947c0e74d735c4102e8f42a1c6272d23f94f1.zip"},"primary_area":{"value":"generative models"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"With the availability of large-scale video datasets and the advances of diffusion models, text-driven video generation has achieved substantial progress. However, existing video generation models are typically trained on a limited number of frames, resulting in the inability to generate high-fidelity long videos during inference. Furthermore, these models only support single-text conditions, whereas real-life scenarios often require multi-text conditions as the video content changes over time. To tackle these challenges, this study explores the potential of extending the text-driven capability to generate longer videos conditioned on multiple texts. 1) We first analyze the impact of initial noise in video diffusion models. Then building upon the observation of noise, we propose FreeNoise, a tuning-free and time-efficient paradigm to enhance the generative capabilities of pretrained video diffusion models while preserving content consistency. Specifically, instead of initializing noises for all frames, we reschedule a sequence of noises for long-range correlation and perform temporal attention over them by window-based fusion. 2) Additionally, we design a novel motion injection method to support the generation of videos conditioned on multiple text prompts. Extensive experiments validate the superiority of our paradigm in extending the generative capabilities of video diffusion models. It is noteworthy that compared with the previous best-performing method which brought about 255% extra time cost, our method incurs only negligible time cost of approximately 17%. Generated video samples are available at our website: http://haonanqiu.com/projects/FreeNoise.html."},"_bibtex":{"value":"@inproceedings{\nqiu2024freenoise,\ntitle={FreeNoise: Tuning-Free Longer Video Diffusion via Noise Rescheduling},\nauthor={Haonan Qiu and Menghan Xia and Yong Zhang and Yingqing He and Xintao Wang and Ying Shan and Ziwei Liu},\nbooktitle={The Twelfth International Conference on Learning Representations},\nyear={2024},\nurl={https://openreview.net/forum?id=ijoqFqSC7p}\n}"},"title":{"value":"FreeNoise: Tuning-Free Longer Video Diffusion via Noise Rescheduling"},"pdf":{"value":"/pdf/bd47f35c18df619e675c737ccc56c1d802537b73.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference"},"paperhash":{"value":"qiu|freenoise_tuningfree_longer_video_diffusion_via_noise_rescheduling"},"authorids":{"value":["~Haonan_Qiu1","~Menghan_Xia1","~Yong_Zhang6","~Yingqing_He1","~Xintao_Wang1","~Ying_Shan2","~Ziwei_Liu1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Haonan Qiu","Menghan Xia","Yong Zhang","Yingqing He","Xintao Wang","Ying Shan","Ziwei Liu"]}},"version":2},{"content":{"TLDR":{"value":"A camera-aware reference encoder jointly controls persona, action, and camera motion in human video generation while keeping the video diffusion backbone frozen."},"venue":{"value":"ICLR 2027 Conference Submission"},"keywords":{"value":["human video generation","identity preservation","camera control","video diffusion","reference encoding","multimodal conditioning"]},"supplementary_material":{"value":"/attachment/024cee5a9ed2cbf7bb3f7a52dfcb2ab5bd24aa94.pdf"},"primary_area":{"value":"generative models"},"abstract":{"value":"Cinematic human video generation must jointly control who appears in frame, what the subject does, and how the camera moves, yet the literature has split into single-axis specialists whose serial composition compounds error on every axis. We argue that the right conditioning object is a chronologically ordered space-time lattice of micro-clips cut from a reference video, and present CinePersona, a camera-aware reference encoder for a frozen 14B video diffusion transformer. Its Chrono-Mosaic Lattice packs nine time-ordered micro-clips into one trunk-compatible tensor, and we show that this ordering is what keeps camera direction decodable: time-reversed trajectories are provably indistinguishable under any shuffled interface. A dual-pathway encoder pairs a diffusion-synchronized implicit stream with explicit persona, action and camera branches, Cine-RoPE writes the target trajectory into the trunk's rotary phases, and a tri-axis sampler exposes one guidance dial per control, all while training 1.5% of parameters. On CineHuman-Bench and OpenS2V-Eval, CinePersona natively controls all three axes and achieves state-of-the-art persona similarity, action alignment and camera fidelity, the latter verified with an estimator decoupled from training, on rendered scenes with ground truth, and by 41 human raters."},"_bibtex":{"value":"@inproceedings{\nanonymous2026cinepersona,\ntitle={CinePersona: Camera-Aware Dynamic Persona Encoding for Identity-Preserving Human Video Generation},\nauthor={Anonymous},\nbooktitle={Submitted to The Fifteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=IwOl43jtYT},\nnote={under review}\n}"},"title":{"value":"CinePersona: Camera-Aware Dynamic Persona Encoding for Identity-Preserving Human Video Generation"},"pdf":{"value":"/pdf/09248a4407780b3c311bc24cdf835ceff1a65ddd.pdf"},"venueid":{"value":"ICLR.cc/2027/Conference/Submission"}},"tmdate":1791227821842,"tcdate":1789356922460,"writers":["ICLR.cc/2027/Conference","ICLR.cc/2027/Conference/Submission16768/Authors"],"signatures":["ICLR.cc/2027/Conference/Submission16768/Authors"],"forum":"IwOl43jtYT","license":"CC BY 4.0","number":16768,"cdate":1789356922460,"readers":["everyone"],"invitations":["ICLR.cc/2027/Conference/-/Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Bidding","ICLR.cc/2027/Conference/Submission16768/-/Full_Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Reviewing","ICLR.cc/2027/Conference/-/Edit"],"mdate":1791227821842,"odate":1791031328286,"domain":"ICLR.cc/2027/Conference","id":"IwOl43jtYT","version":2},{"content":{"venue":{"value":"ICLR 2027 Conference Submission"},"keywords":{"value":["audio-video generation; visual shortcut; counterfactual invariance; structural causal model; multimodal generation"]},"supplementary_material":{"value":"/attachment/6839fbb1ab330a2e789c79606a7f281eb642abc4.zip"},"primary_area":{"value":"causal reasoning"},"abstract":{"value":"Joint audio--video (AV) generators are trained on data in which \\emph{what an event\nlooks like} and \\emph{what it sounds like} are spuriously correlated. We present a\n\\emph{controlled causal study} of the resulting failure mode. In an AV structural causal\nmodel where the audio is, by construction, independent of the video's nuisance appearance,\nmodels that let audio read video directly---through cross-attention or a shared\nlatent---learn a \\emph{visual shortcut}: they predict sound from appearance rather than the\ncausal event and, when the appearance--event correlation is broken at test time, synthesize\nthe wrong event's sound. Crucially, the popular remedy of routing both modalities through a\n\\emph{shared common-cause latent} does \\emph{not} fix this---a bottleneck, an unsupervised\nshared/private factorization, and a faithful shared-prior model all grab the appearance proxy\nand fail like the direct model. Blocking the shortcut instead requires an \\emph{intervention\non the nuisance}: under the stated assumptions we prove that counterfactual invariance is\nnecessary and sufficient to identify the causal predictor, and we verify the mechanism from\nfeature-vector SCMs to procedural pixel video, real images with spectrogram audio, moving\nreal digits, and a conditional generator. On a \\emph{real, pretrained} V2A generator\n(MMAudio), an input-intervention test shows the model is far from invariant to\nsound-irrelevant edits, though a generic-noise control reveals it is broadly input-brittle\nrather than specifically colour-shortcutting---clean isolation of the shortcut needs the\ncontrolled confounds our synthetic studies provide. We characterize \\emph{when} the shortcut\noccurs, compare the objective against supervised counterfactual augmentation, and isolate the\n\\emph{unknown-nuisance} regime---where the intervention cannot be applied---as the central\nopen problem."},"_bibtex":{"value":"@inproceedings{\nanonymous2026intervention,\ntitle={Intervention, Not Shared Latents:  Blocking Visual Shortcuts in Audio-Video Generation},\nauthor={Anonymous},\nbooktitle={Submitted to The Fifteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=YNfbxnx8yM},\nnote={under review}\n}"},"title":{"value":"Intervention, Not Shared Latents:  Blocking Visual Shortcuts in Audio-Video Generation"},"pdf":{"value":"/pdf/a859229fee6240d6dee32e278cd29e7c581066f6.pdf"},"venueid":{"value":"ICLR.cc/2027/Conference/Submission"}},"tmdate":1791221764924,"tcdate":1788252198082,"writers":["ICLR.cc/2027/Conference","ICLR.cc/2027/Conference/Submission7294/Authors"],"signatures":["ICLR.cc/2027/Conference/Submission7294/Authors"],"forum":"YNfbxnx8yM","license":"CC BY 4.0","number":7294,"cdate":1788252198082,"readers":["everyone"],"invitations":["ICLR.cc/2027/Conference/-/Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Bidding","ICLR.cc/2027/Conference/-/Submission_Change_Before_Reviewing","ICLR.cc/2027/Conference/-/Edit"],"mdate":1791221764924,"odate":1791031328286,"domain":"ICLR.cc/2027/Conference","id":"YNfbxnx8yM","version":2},{"content":{"summary":{"value":"This paper introduces Quicksviewer, an innovative large multimodal model (LMM) designed for efficient video understanding by addressing the issue of nonuniform temporal information density in videos. Using a novel cubing approach based on Gumbel Softmax, the model dynamically partitions video frames and resamples them, enabling significant spatiotemporal compression (45×) while maintaining a large receptive field. Trained with only 0.8M video-text samples, Quicksviewer outperforms baseline models by a notable margin (up to 8.72 in accuracy) and achieves state-of-the-art performance on Video-MME with reduced token usage. Additionally, the model demonstrates scalable improvements with increased input frames and provides insights into analyzing continuous events in videos."},"soundness":{"value":1},"confidence":{"value":3},"questions":{"value":"Please see the weakness part."},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. The paper is well-written, with the methodology and experimental results clearly and systematically presented.\n2. The exploration of video token compression in this paper is highly meaningful, as it contributes to both improving long-video understanding and accelerating video processing.\n3. The proposed method demonstrates notable improvements over the baseline, achieving up to 8.72 gains in accuracy."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. While the paper introduces a visual token compression approach for videos based on LMMs, it lacks comparisons with other compression methods (e.g., SlowFast, token merging) under equivalent conditions, such as using the same training dataset, to better validate the advantages of the proposed strategy.\n2. Although the paper focuses on token compression, it only discusses token counts without analyzing key metrics like latency and memory usage. In practice, even with the same token count, a more complex method could lead to different performance in terms of latency and memory efficiency.\n3. The paper claims that Quicksviewer is a video understanding model, but most of the evaluations are based on video QA tasks, which typically have lower token requirements. To provide a more comprehensive assessment of its video understanding capabilities, the authors should include comparisons on other tasks like video captioning.\n4. The paper claims that Quicksviewer achieves state-of-the-art performance on several benchmarks, but the baselines used for comparison are relatively weak. For instance, on Video-MME, many existing models (e.g., Qwen VL, Intern VL) have already achieved scores of 60 or even 65+, whereas Quicksviewer only achieves 56.9. While differences in conditions such as training data and token counts exist, the claim itself is not rigorous and does not convincingly demonstrate that the proposed approach achieves high accuracy."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762919002794,"tcdate":1760692726343,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission6713/Reviewer_6NpJ"],"signatures":["ICLR.cc/2026/Conference/Submission6713/Reviewer_6NpJ"],"forum":"AcnCQR2ElW","number":1,"license":"CC BY 4.0","cdate":1760692726343,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission6713/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762919002794,"domain":"ICLR.cc/2026/Conference","replyto":"AcnCQR2ElW","id":"zsTVSr6iF3","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Video Understanding","Large Multimodal Models"]},"primary_area":{"value":"foundation or frontier models, including LLMs"},"abstract":{"value":"Large Multimodal Models (LMMs) uniformly perceive video frames, creating computational inefficiency for videos with inherently varying temporal information density. This paper present Quicksviewer, an LMM with new perceiving paradigm that partitions a video of nonuniform density into varying cubes using Gumbel Softmax, followed by a unified resampling for each cube to achieve efficient video understanding. This simple and intuitive approach dynamically compress video online based on its temporal density, significantly reducing spatiotemporal redundancy (overall 45$\\times$ compression rate), while enabling efficient training with large receptive field. We train the model from a language backbone through three progressive stages, each incorporating lengthy videos on average of 420s/1fps thanks to the perceiving efficiency. With only 0.8M total video-text samples for training, our model outperforms the direct baseline employing a fixed partitioning strategy by a maximum of 8.72 in accuracy, demonstrating the effectiveness in performance. On Video-MME, Quicksviewer achieves competitive performance compared to models of similar size while utilizing just up to 5\\% of tokens per frame required by baselines. With this paradigm, scaling up the number of input frames reveals a clear power law of the model capabilities. It is also empirically verified that the segments generated by the cubing network can help for analyzing continuous events in videos."},"_bibtex":{"value":"@misc{\nqi2026quicksviewer,\ntitle={Quicksviewer: An {LMM} for Efficient Video Understanding via Reinforced Compression of Video Cubes},\nauthor={Ji Qi and Yuan Yao and Yushi Bai and Bin Xu and Juanzi Li and Zhiyuan Liu and Tat-Seng Chua},\nyear={2026},\nurl={https://openreview.net/forum?id=AcnCQR2ElW}\n}"},"title":{"value":"Quicksviewer: An LMM for Efficient Video Understanding via Reinforced Compression of Video Cubes"},"pdf":{"value":"/pdf/d0853ad0b2487eafe20922eb597ad8e827e08f36.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"qi|quicksviewer_an_lmm_for_efficient_video_understanding_via_reinforced_compression_of_video_cubes"},"authorids":{"value":["~Ji_Qi3","~Yuan_Yao12","~Yushi_Bai1","~Bin_Xu1","~Juanzi_Li1","~Zhiyuan_Liu1","~Tat-Seng_Chua2"]},"authors":{"value":["Ji Qi","Yuan Yao","Yushi Bai","Bin Xu","Juanzi Li","Zhiyuan Liu","Tat-Seng Chua"]}},"version":2},{"content":{"summary":{"value":"This paper presents Video Storyboarding, a training-free method that enhances pre-trained text-to-video models to generate multiple shots with consistent characters while maintaining high video quality and responsiveness to text prompts. By leveraging self-attention query features that capture motion and identity, the method addresses the trade-off between character consistency and video dynamics through a novel query injection strategy. Experimental results show significant improvements in character consistency and motion quality, offering insights into the video generation process and the interplay of structure and motion in diffusion models."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"1. The paper only demonstrates the performance of a single character across different videos, and the reviewer is curious about how the proposed  method performs with multiple characters.\n2. The prompt in the paper provides overly detailed descriptions of the character. Would a more concise description impact character consistency? For example, replace the \"Cinematic, middle-aged female athlete\" in Fig.8 with \"A woman\"."},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"1. This paper introduces a training-free method to ensure charactoer consistency and motion adherence in producing multi-shot video sequences.\n2. This paper presents a two-phase query injection strategy to balance encoding motion and identity.\n3. A benchmark and evalution protocol are proposed to evaluate consistency of video generation."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The conducted experiments are not comprehensive, including two aspects: (a) The paper only provides several comparison samples. (b) The paper misses some important baseline methods, e.g., VSTAR[1]\n2. Although the purpose of the paper is to maintain consistency of characters across different video clips, the results are not particularly good. For example, in Fig.3, the color and style of clothes change across different video shots.\n\n[1] Li, Yumeng, William Beluch, Margret Keuper, Dan Zhang, and Anna Khoreva. \"VSTAR: Generative Temporal Nursing for Longer Dynamic Video Synthesis.\" arXiv preprint arXiv:2403.13501 (2024)."}},"nonreaders":[],"tmdate":1731427408387,"tcdate":1730626291734,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission1316/Reviewer_kLfg"],"signatures":["ICLR.cc/2025/Conference/Submission1316/Reviewer_kLfg"],"forum":"0zRuk3QdiH","number":3,"license":"CC BY 4.0","cdate":1730626291734,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission1316/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427408387,"domain":"ICLR.cc/2025/Conference","replyto":"0zRuk3QdiH","id":"4uqYdnPexW","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"TLDR":{"value":"A training-free approach for generating video shots of the same characters, preserving identity and motion to prompt agreement."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["text to video","subject consistency","video personalization","motion alignment","feature injection","extended attention"]},"supplementary_material":{"value":"/attachment/621d1cc3ea225a3977b932fdc47bb1a0c469c29e.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Text-to-video models have made significant strides in generating short video clips from textual descriptions. Yet, a significant challenge remains: generating several video shots of the same characters, preserving their identity without hurting video quality, dynamics, and responsiveness to text prompts. We present Video Storyboarding, a training-free method to enable pretrained text-to-video models to generate multiple shots with consistent characters, by sharing features between them. Our key insight is that self-attention query features (Q) encode both motion and identity. This creates a hard-to-avoid trade-off between preserving character identity and making videos dynamic, when features are shared. To address this issue, we introduce a novel query injection strategy that balances identity preservation and natural motion retention. This approach improves upon naive consistency techniques applied to videos, which often struggle to maintain this delicate equilibrium. Our experiments demonstrate significant improvements in character consistency across scenes while maintaining high-quality motion and text alignment. These results offer insights into critical stages of video generation and the interplay of structure and motion in video diffusion models."},"_bibtex":{"value":"@misc{\natzmon2025multishot,\ntitle={Multi-Shot Character Consistency for Text-to-Video Generation},\nauthor={Yuval Atzmon and Rinon Gal and Yoad Tewel and Yoni Kasten and Gal Chechik},\nyear={2025},\nurl={https://openreview.net/forum?id=0zRuk3QdiH}\n}"},"title":{"value":"Multi-Shot Character Consistency for Text-to-Video Generation"},"pdf":{"value":"/pdf/309972798761b560e7adc61ec56cf8714136d46a.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"atzmon|multishot_character_consistency_for_texttovideo_generation"},"authorids":{"value":["~Yuval_Atzmon1","~Rinon_Gal1","~Yoad_Tewel1","~Yoni_Kasten1","~Gal_Chechik1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Yuval Atzmon","Rinon Gal","Yoad Tewel","Yoni Kasten","Gal Chechik"]}},"version":2},{"content":{"venue":{"value":"IEEE Access 2025"},"pdf":{"value":"https://ieeexplore.ieee.org/iel8/6287639/10820123/10902083.pdf"},"venueid":{"value":"dblp.org/journals/ACCESS/2025"},"paperhash":{"value":"choi|momentaware_video_retrieval_for_video_corpus_moment_retrieval"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Yura_Choi:","https://dblp.org/search/pid/api?q=author:Daechul_Ahn:","~Jonghyun_Choi1"]},"html":{"value":"https://doi.org/10.1109/ACCESS.2025.3542720"},"_bibtex":{"value":"@article{DBLP:journals/access/ChoiAC25,\n  author={Yura Choi and Daechul Ahn and Jonghyun Choi},\n  title={Moment-Aware Video Retrieval for Video Corpus Moment Retrieval},\n  year={2025},\n  cdate={1735689600000},\n  journal={IEEE Access},\n  volume={13},\n  pages={38574-38583},\n  url={https://doi.org/10.1109/ACCESS.2025.3542720}\n}\n"},"abstract":{"value":"Video corpus moment retrieval (VCMR) task aims to retrieve a specific moment from a large corpus of untrimmed videos. This task has been addressed by decomposing it into video retrieval and moment retrieval subtasks each with specialized heads (i.e., democratic decomposition), due to the computational complexity. However, this approach overlooks the interdependency between the subtasks, which is crucial for temporally fine-grained query-video alignment. To address this suboptimality, we propose an integrated learning framework that explicitly establishes a connection between the two subtasks, namely moment-aware video retrieval, allowing them to benefit from each other’s learning process. Furthermore, we employ a curriculum-based negative sampling strategy that gradually provides harder negative samples to enhance the discriminative ability between semantically similar negative videos. We empirically show that the proposed approach outperforms state-of-the-art methods on three benchmarks – TVR, ActivityNet, and DiDeMo. Notably, on TVR, our method achieves 10.04% in VCMR R@1 with tIoU=0.7, representing a 1.68% absolute improvement over prior work, and similarly demonstrates consistent gains on ActivityNet (4.98% in VCMR R@1 with tIoU=0.5) and DiDeMo, validating the effectiveness of our integrated approach."},"title":{"value":"Moment-Aware Video Retrieval for Video Corpus Moment Retrieval"},"authors":{"value":["Yura Choi","Daechul Ahn","Jonghyun Choi"]}},"tmdate":1759842142712,"pdate":1735689600000,"externalIds":["dblp:journals/access/ChoiAC25"],"tcdate":1759842134869,"writers":["~"],"signatures":["~Jonghyun_Choi1"],"forum":"pmDTMPcQPZ","license":"CC BY-SA 4.0","number":639204,"cdate":1735689600000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1759842142712,"domain":"DBLP.org","id":"pmDTMPcQPZ","version":2},{"content":{"summary":{"value":"This paper proposes TornadoAttention, a training-free sparse attention mechanism for accelerating video Diffusion Transformers (DiTs). The key innovation is applying fine-grained spatio-temporal permutations to reorder token sequences, concentrating attention scores around the main diagonal. Through offline search over permutation orders and tile granularities, the method identifies optimal static sparse masks for each attention head. Evaluated on HunyuanVideo and Wan2.1 models, TornadoAttention achieves 1.4× speedup with minimal quality degradation."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"See weakness."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"The permutation-based approach is novel - reordering tokens to align memory layout with spatio-temporal locality differs from existing fixed-pattern or dynamic sparse attention methods. The 4D search space (permutation order plus T/H/W tile sizes) provides a principled framework for exploring video data structure. The paper clearly motivates the problem with examples of diverse attention patterns that defy simple classification. The visualization of misalignment between logical locality and memory layout effectively illustrates why standard token ordering is suboptimal. Evaluation on large-scale state-of-the-art models (HunyuanVideo-14B, Wan2.1-14B) demonstrates practical applicability. The method maintains good fidelity metrics (PSNR, SSIM, VBench) while achieving measurable speedup, and the ablation validates the necessity of the full search space."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1.The paper only compares against one baseline (SVG). Missing comparisons with other recent sparse attention methods such as MInference, DiTFastAttn, PAB, DraftAttention, Radial Attention, CompactAttention, and AdaSpa that are discussed in related work.\n\n2.Ablation studies are insufficient. Only one ablation on the 4D search space is provided. Critical design choices lack analysis: the energy threshold τ=0.9, block size b=128, number of profiling samples, and the decision to skip first 25% denoising steps.\n\n3.Model coverage is narrow. Evaluation limited to two models (HunyuanVideo and Wan2.1). Missing validation on other video DiT architectures like CogVideoX which is widely adopted in the community.\n\n4.Quality metrics are incomplete. Reports only PSNR, SSIM, LPIPS and two VBench metrics. Missing comprehensive video quality evaluation including temporal consistency, motion smoothness, aesthetic quality, and other VBench dimensions that are standard for video generation assessment.\n\n5.No end-to-end latency breakdown is provided. The paper reports overall 1.4× speedup but lacks detailed analysis of where improvements come from - attention computation, memory access, kernel overhead, or other components. Missing kernel-level performance analysis showing actual performance of the permutation strategy."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762925956801,"tcdate":1761799779684,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission15710/Reviewer_XQYT"],"signatures":["ICLR.cc/2026/Conference/Submission15710/Reviewer_XQYT"],"forum":"fx3Ht2gYTR","number":3,"license":"CC BY 4.0","cdate":1761799779684,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission15710/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762925956801,"domain":"ICLR.cc/2026/Conference","replyto":"fx3Ht2gYTR","id":"7NEsXjibYE","forumContent":{"TLDR":{"value":"A training-free hardware-efficient sparse attention via fine-grained spatio-temporal permutation for video generation."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Sparse Attention","Video Generation"]},"primary_area":{"value":"other topics in machine learning (i.e., none of the above)"},"abstract":{"value":"Diffusion Transformers (DiTs) have demonstrated remarkable success in video generation. However, their core component, the self-attention mechanism, suffers from quadratic complexity. To alleviate this issue, sparse attention mechanisms have been proposed. Existing methods, however, often impose strong, handcrafted priors, restricting attention to a few fixed patterns, fail to capture the diverse, data-dependent attention patterns unique to each layer and head. Motivated by the spatio-temporal locality of self-attention in DiTs, we propose TornadoAttention, a \\textbf{training-free} sparse attention mechanism. Our key idea is to apply a fine-grained permutation of the query and key sequences that better matches the underlying attention structure. We applied TornadoAttention to advanced open-source video generation models. It reveal that the attention masks obtained through offline searching exhibit excellent generalization capabilities across a diverse range of prompts, which provides a crucial foundation for aggressive, hardware-specific kernel-level optimizations. On the HunyuanVedio model, our method achieves a 1.4$\\times$ speedup with negligible loss in fidelity."},"_bibtex":{"value":"@misc{\nzhang2026tornadoattention,\ntitle={TornadoAttention: Hardware-Efficient Sparse Attention via Fine-Grained Spatio-Temporal Permutation for Video Generation},\nauthor={Wentai Zhang and Hongyu Jia and Zichen Tang and Ronghui Xi and Haoran Luo and Haihong E},\nyear={2026},\nurl={https://openreview.net/forum?id=fx3Ht2gYTR}\n}"},"title":{"value":"TornadoAttention: Hardware-Efficient Sparse Attention via Fine-Grained Spatio-Temporal Permutation for Video Generation"},"pdf":{"value":"/pdf/1a48dc947d9d020fa140d3c0005f77cbde20c41d.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"zhang|tornadoattention_hardwareefficient_sparse_attention_via_finegrained_spatiotemporal_permutation_for_video_generation"},"authorids":{"value":["~Wentai_Zhang2","~Hongyu_Jia1","~Zichen_Tang1","~Ronghui_Xi2","~Haoran_Luo1","~Haihong_E1"]},"authors":{"value":["Wentai Zhang","Hongyu Jia","Zichen Tang","Ronghui Xi","Haoran Luo","Haihong E"]}},"version":2},{"content":{"summary":{"value":"The paper proposes ST-GridPool, a practical training-free method to enhance visual tokens for Video LLMs. The approach is technically sound, combining a hierarchical temporal gridding strategy (PTG) and a saliency-aware spatial pooling mechanism (NSP). The method demonstrates consistent performance improvements on LLaVA-family models across six benchmarks and compelling gains in computational efficiency. The empirical evaluation is extensive, including strong ablations and comparisons to other token reduction methods."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1.Is it possible to explore an adaptive mechanism that automatically determine the optimal values of non-trivial hyperparameters in your method?\\\n2.How does the proposed PTG interact with or differ from the built-in temporal reasoning mechanisms already present in modern Video LLMs such as NVILA?\\\n3.Could other unsupervised indicators serve as complementary or alternative signals to the token norm?\\\n4.How to interpret the observed trend of a larger improvement for Long Video Understanding tasks compared to General Video Understanding tasks?"},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1.The proposed ST-GridPool method is training-free and plug-and-play, offering an efficient solution to enhance existing Video LLMs without the need for costly retraining. It demonstrates consistent performance improvements across various backbones on all six evaluated video understanding benchmarks, with notable gains in Long Video Understanding.\\\n2.The motivation behind Norm-based Spatial Pooling is well-founded, supported by a clear analysis that establishes a positive correlation between token norms and semantic saliency, contributing to the method's effectiveness.\\\n3.The empirical evaluation is comprehensive, featuring extensive ablation studies on individual components, key hyperparameters, and design choices such as pooling shape and gridding strategy."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1.The performance depends on several non-trivial hyperparameters, including the number of temporal levels, norm order, and temperature. Balancing these parameters effectively is crucial for their successful application in real-world scenarios.\\\n2.The paper does not clearly show whether PTG provides complementary benefits or duplicates functionality already present in recent Video LLMs.\\\n3.The NSP module relies heavily on an empirical correlation between token norm and semantic importance, yet offers limited empirical comparison or grounding.\\\n4.The experimental section primarily focuses on quantitative results but lacks deeper analytical discussion. The paper does not explore why the proposed ST-GridPool achieves varying degrees of improvement across different benchmarks or tasks.\\\n5.There are some typos in the paper. E.g., the percentage gains reported for \"LLaVA-Video-7B + Ours\" are identical to those reported for \"LLaVA-OneVision-7B + Ours\"."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762943165738,"tcdate":1761840007221,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission24698/Reviewer_BLb5"],"signatures":["ICLR.cc/2026/Conference/Submission24698/Reviewer_BLb5"],"forum":"MZi9SYPVz5","number":3,"license":"CC BY 4.0","cdate":1761840007221,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission24698/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762943165738,"domain":"ICLR.cc/2026/Conference","replyto":"MZi9SYPVz5","id":"W2EI4Wt3bm","forumContent":{"TLDR":{"value":"Our training-free method, ST-GridPool, boosts Video LLM performance and efficiency by intelligently compressing visual tokens based on their spatiotemporal importance."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Visual Token Representation","Video Understanding","Multimodal Large Language Models"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Recent advances in Multimodal Large Language Models (MLLMs) have significantly advanced video understanding tasks, yet challenges remain in efficiently compressing visual tokens while preserving spatiotemporal interactions. Existing methods, such as LLaVA family, utilize simplistic pooling or interpolation techniques that overlook the intricate dynamics of visual tokens. To bridge this gap, we propose ST-GridPool, a novel training-free visual token enhancement method designed specifically for Video LLMs. Our approach integrates Pyramid Temporal Gridding (PTG), which captures multi-grained spatiotemporal interactions through hierarchical temporal gridding, and Norm-based Spatial Pooling (NSP), which preserves high-information visual regions by leveraging the correlation between token norms and semantic richness. Extensive experiments on various benchmarks demonstrate that ST-GridPool consistently enhances performance of Video LLMs without requiring costly retraining. \nOur method offers an efficient and plug-and-play solution for improving visual token representations.\nOur code is available in [https://github.com/bingjunluo/ST-GridPool](https://github.com/bingjunluo/ST-GridPool)."},"_bibtex":{"value":"@inproceedings{\nluo2026enhancing,\ntitle={Enhancing Visual Token Representations for Video Large Language Models via Training-free Spatial-Temporal Pooling and Gridding},\nauthor={Bingjun Luo and Tony Wang and Hanqi Chen and Xinpeng Ding},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=MZi9SYPVz5}\n}"},"title":{"value":"Enhancing Visual Token Representations for Video Large Language Models via Training-free Spatial-Temporal Pooling and Gridding"},"pdf":{"value":"/pdf/5b313af3ea8d0b93020227322bbf5578b292b992.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"luo|enhancing_visual_token_representations_for_video_large_language_models_via_trainingfree_spatialtemporal_pooling_and_gridding"},"authorids":{"value":["~Bingjun_Luo2","~Tony_Wang4","~Hanqi_Chen1","~Xinpeng_Ding1"]},"authors":{"value":["Bingjun Luo","Tony Wang","Hanqi Chen","Xinpeng Ding"]}},"version":2},{"content":{"summary":{"value":"This paper aims at text-video retrieval in the domain of vision-and-language pretraining and understanding. Similar to general information retrieval or multimedia retrieval system, the authors proposed a two-stage or coarse-to-fine retrieval framework. In the first stage, the consine similarity network is trained with contrastive learning loss. Thus, a general text and video matching score could be obtained by such a model. The main contribution is the proposed re-ranker for fine-grained ranking in the second stage. Specifically, the authors proposed a novel cross attention module called multi-grained text-video cross attention in order to compute text-frame level attention and text-video level attention respectively. Finally, this prototype has been verified on several popular text-video retrieval benchmarks."},"presentation":{"value":"3 good"},"contribution":{"value":"2 fair"},"soundness":{"value":"3 good"},"strengths":{"value":"1. The motivation of this work is clear and is easy to follow. \n\n2. I think the overall technology roadmap is in the right way which includes: (a) a course-to-fine retrieval framework for the sake of efficiacy and effiency trade-off; (b) freezing the image encoders to make the training process practical and affordable;\n\n3. The proposed  multi-grained text-video cross attention mechanism is brilliant and makes sense. Specifically, spatial text attention module is proposed at the frame level in order to discover small objects or entities. And temporal text attention is designed to capture subtle movement."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. It seems that the experimental results do not include retrieval performance (e.g. stage 1 / stage 2/ pre- and post-processing), which is important in terms of a cross-modal retrieval topic. \n\n2. Again, in terms of a retrieval task, it is important to verify the effectiveness under a huge database setting. However, the most large dataset only contains 118,081 videos which is much less than industiral scales such as YouTube. At least, more distractors should be included if scaling positive text-video pairs is a concern. \n\n3. I think the visualization is not enough. For example, the retrieval results of different work or the effect of proposed re-ranking could be included as supplementary materials if space limitation is a concern.  Besides, the visualization should verify the effectiveness of proposed saptial text attention and temporal text attention."},"confidence":{"value":"5: You are absolutely certain about your assessment. You are very familiar with the related work and checked the math/other details carefully."},"questions":{"value":"Besides the problems in the above weaknesses part, there are a few other questions listed as follows.\n\n[Q1] The token selector is employed as a trade-off between missing detailed information and bringing redundant information issues. The implementation of such a module is a MLP followed by a Softmax Layer, which predicts each token's importance score and selects the M most informative tokens as output. The question is what is the ground truth (GT) for such predictions? It seems that it is even impossible for human beings to label. Note that this is not a single label classfication problem. \n\n[Q2] From the ablation study Table 6, the hard negative mining module makes little contribution to the final results, which is below my expectation. Is there any possible explanations?\n\n[Q3] It is quite often to incorporate query expansion for the sake of increasing recall. Is it still effective after the 2nd stage re-ranking? It is not a necessary ablation study but could be considered as a option."},"rating":{"value":"6: marginally above the acceptance threshold"},"code_of_conduct":{"value":"Yes"}},"nonreaders":[],"tmdate":1701413304507,"tcdate":1697452175919,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission4495/Reviewer_qf6J"],"signatures":["ICLR.cc/2024/Conference/Submission4495/Reviewer_qf6J"],"forum":"Q4FmJPQwuJ","number":1,"license":"CC BY 4.0","cdate":1697452175919,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission4495/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1701413304507,"domain":"ICLR.cc/2024/Conference","replyto":"Q4FmJPQwuJ","id":"97pYyVZkMz","forumContent":{"venue":{"value":"Submitted to ICLR 2024"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Text-Video Retrieval","CLIP","Frozen","Multimodal"]},"supplementary_material":{"value":"/attachment/1a203262dc44b942656bfa2864dde71260b1d49f.pdf"},"primary_area":{"value":"representation learning for computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"State-of-the-art text-video retrieval (TVR) methods commonly use CLIP and cosine similarity for efficient retrieval. \nMeanwhile, cross attention methods, which employ a transformer decoder to compute attention between text queries and video frames, offer a more comprehensive interaction between multimodal information.\nComplementary to the existed one-stage text video retrieval approaches above, we propose a re-ranker called CrossTVR that further explores the fine-grained and comprehensive interaction between text and all the vision tokens of a given video at the frame level and the video (clips or segments) level. \nFurthermore, we employ the frozen CLIP model strategy for fine-grained retrieval, enabling scalability to larger pre-trained vision models like ViT-G and resulting in further improved retrieval performance.\nSubsequently,  a two-stage text-video retrieval architecture can be proposed.\nIn the first stage, we leverage existed TVR methods with cosine similarity network to efficently obtain text video candidate pairs.\nIn the second stage, the proposed re-ranker is applied for fine-grained retrieval. \nExperimental results on text-video retrieval datasets demonstrate the effectiveness and scalability of the proposed re-ranker when combined with existing mainstream one-stage text-video retrieval approaches."},"_bibtex":{"value":"@misc{\ndai2024crosstvr,\ntitle={Cross{TVR}: Multi-Grained Re-Ranker for Text Video Retrieval with Frozen Image Encoders},\nauthor={Zuozhuo Dai and Fangtao Shao and Qingkun Su and Zilong Dong and Yao Yao and Long Qin and Siyu Zhu},\nyear={2024},\nurl={https://openreview.net/forum?id=Q4FmJPQwuJ}\n}"},"title":{"value":"CrossTVR: Multi-Grained Re-Ranker for Text Video Retrieval with Frozen Image Encoders"},"pdf":{"value":"/pdf/f817b84177a75a83bb890f9dc207126c1d613cc3.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference/Rejected_Submission"},"paperhash":{"value":"dai|crosstvr_multigrained_reranker_for_text_video_retrieval_with_frozen_image_encoders"},"authorids":{"value":["~Zuozhuo_Dai2","~Fangtao_Shao1","~Qingkun_Su1","~Zilong_Dong2","~Yao_Yao1","~Long_Qin2","~Siyu_Zhu1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Zuozhuo Dai","Fangtao Shao","Qingkun Su","Zilong Dong","Yao Yao","Long Qin","Siyu Zhu"]}},"version":2},{"content":{"summary":{"value":"This paper introduces VisR-Bench, a synthetic benchmark dataset comprising 35K QA pairs across 1.2K documents in 16 languages for evaluating multilingual visual retrieval. Each PDF document is parsed using the Adobe Document Extract API, which extracts text, figures, and tables; GPT-4o then generates corresponding QA pairs. The dataset is designed to test retrieval and visual understanding across languages and modalities (text, tables, figures). The authors benchmark a wide range of retrieval methods — text-based, multimodal encoders, and multimodal large language models (MLLMs). Results show that retrieval performance degrades for low-resource languages and non-text modalities (figures and tables) relative to English text documents."},"soundness":{"value":2},"confidence":{"value":2},"questions":{"value":"1. On the multilingual gap:\nIf performance degradation in low-resource languages mirrors existing text-based multilingual retrieval gaps, what additional insights does VisR-Bench provide? Do results show cases where models are strong in text retrieval for a language but significantly weaker in multimodal (VQA) retrieval for the same language?\n\n2. On ColQwen2 variants:\nHow do you explain ColQwen2-M’s lower performance compared to the base ColQwen2-v0.1?\n\n3. Text vs. MLLM retrieval:\nIn multilingual settings, why do text-based retrievers sometimes outperform MLLMs?\n\n4. Dataset balance:\nDo all languages contain comparable proportions of text-, table-, and figure-based QA pairs?\n\n5. GPT-4o QA Accuracy: Since GPT-4o generated the QA, why can’t it achieve perfect accuracy?\n\n6. Table QA Filtering: Did the table-based QA undergo the same heuristic filtering as figure-based QA to ensure the question requires table understanding and can’t be answered from the surrounding text?\n\n7. Human QA Validation: Beyond filtering for harmful content and PII, was there any human validation to confirm that the synthetic questions are answerable and the answers are correct?\n\nL289:  English English split -->  English split\nL481: VisR-Benchpaves --> VisR-Bench paves"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. Comprehensive model coverage: The authors evaluate a diverse set of retrieval models, including text-only, multimodal encoders, and MLLM-based approaches.\n\n2. Strong clarity and structure: The paper is well written, logically organized, and easy to follow, making its experimental design and contributions accessible.\n\n3. Data curation: The paper notes that all documents underwent human validation to ensure exclusion of harmful content and PII, enhancing dataset safety and reliability.\n\n4. Dataset Diversity: The benchmark includes many QA pairs and documents, spanning 16 languages and incorporating figures, tables, and multilingual text, making it comprehensive."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Limited evaluation of reasoning-capable MLLMs.\nIt would strengthen the analysis to include recent reasoning-optimized MLLMs (e.g., OpenAI o3), even on a small, hard subset. These models could provide deeper insights into reasoning gaps and multimodal generalization.\n\n2. Over-reliance on Top-1 accuracy.\nThe use of Top-1 retrieval accuracy as the primary metric may overstate model weaknesses. While the paper reports a Top-1 accuracy of ~75.2%, the Top-5 accuracy reaches 94.1%, suggesting retrieval systems can already surface the correct document in most cases. Given that RAG pipelines typically include rerankers or pass multiple candidates to the generator, the claimed “large room for improvement” may be somewhat overstated.\n\n3. Benchmark Novelty: While multilinguality and visual context are both valuable, it’s unclear how the dataset offers fundamentally new insights over existing multilingual or visual-only benchmarks.\n\n4. Dataset balance and comparability.\nIt is unclear whether the proportions of document types (text / table / figure) are consistent across languages, which could affect the fairness of multilingual performance comparisons.\n\n5. Lack of Human QA Verification: I did not see the mention of human validation of the QA pairs beyond filtering for safety. This raises concerns about factual correctness and answerability.\n\n6. Limited Dataset Analysis and Insights:\nThe paper introduces a large dataset, but offers minimal exploratory analysis that could help the community understand its properties and challenges. Additional insights about Document-Type Composition in multilingual subsets, Question Type Taxonomy, Visual Density, Domain Diversity, and more would strengthen the paper."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762920003819,"tcdate":1761695798882,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission8002/Reviewer_1xa4"],"signatures":["ICLR.cc/2026/Conference/Submission8002/Reviewer_1xa4"],"forum":"7iFZ6uzILL","number":3,"license":"CC BY 4.0","cdate":1761695798882,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission8002/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762920003819,"domain":"ICLR.cc/2026/Conference","replyto":"7iFZ6uzILL","id":"Rsby4ErPph","forumContent":{"TLDR":{"value":"VisR-Bench is a comprehensive benchmark dataset for question-driven, multilingual, and multimodal document retrieval in long documents."},"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Document Retrieval","Vision Question Answering","Multimodal Large Language Models"]},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Most organizational data in this world are stored as documents, and visual retrieval plays a crucial role in unlocking the collective intelligence from all these documents. However, existing benchmarks focus on English-only document retrieval or only consider multilingual question-answering on a single-page image. To bridge this gap, we introduce VisR-Bench, a multilingual benchmark designed for question-driven multimodal retrieval in long documents. Our benchmark comprises over 35K high-quality QA pairs across 1.2K documents, enabling fine-grained evaluation of multimodal retrieval. VisR-Bench spans sixteen languages with three question types (figures, text, and tables), offering diverse linguistic and question coverage. Unlike prior datasets, we include queries without explicit answers, preventing models from relying on superficial keyword matching. We evaluate various retrieval models, including text-based methods, multimodal encoders, and MLLMs, providing insights into their strengths and limitations. Our results show that while MLLMs significantly outperform text-based and multimodal encoder models, they still struggle with structured tables and low-resource languages, highlighting key challenges in multilingual visual retrieval."},"_bibtex":{"value":"@misc{\nchen2026visrbench,\ntitle={VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding},\nauthor={Jian Chen and Ming Li and Chenguang Wang and Jihyung Kil and Tong Yu and Tianyi Zhou and Changyou Chen and Ruiyi Zhang},\nyear={2026},\nurl={https://openreview.net/forum?id=7iFZ6uzILL}\n}"},"title":{"value":"VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding"},"pdf":{"value":"/pdf/580faed25f2f6ed0e8f06c62e6714d6d4eb08fa5.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"chen|visrbench_an_empirical_study_on_visual_retrievalaugmented_generation_for_multilingual_long_document_understanding"},"authorids":{"value":["~Jian_Chen9","~Ming_Li18","~Chenguang_Wang4","~Jihyung_Kil1","~Tong_Yu3","~Tianyi_Zhou2","~Changyou_Chen1","~Ruiyi_Zhang3"]},"authors":{"value":["Jian Chen","Ming Li","Chenguang Wang","Jihyung Kil","Tong Yu","Tianyi Zhou","Changyou Chen","Ruiyi Zhang"]}},"version":2},{"content":{"summary":{"value":"This paper introduces a new framework for generating captions for sounding videos and proposes a new model architecture for sounding video generation. The captioning framework provides modality-specific captions for both audio and video. It first creates a detailed video caption from a given video clip using VLLM. Then, another LLM extracts sounding events from the video caption and creates a corresponding audio caption. These captions are used to train a new model called BridgeDiT, which is constructed by combining two pretrained DiT models for audio and video through dual cross-attention fusion. The experimental results show that the proposed model outperforms several baselines, including independent generation methods, sequential generation methods, and joint generation methods."},"soundness":{"value":1},"confidence":{"value":4},"questions":{"value":"- Did the authors examine the proposed architecture with other backbones?\n- How can we ensure that the generated audio caption accurately describes the content in the soundtrack of the source sounding video?\n- How is spatial information handled in the proposed captioning framework?"},"rating":{"value":2},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- This work addresses a challenging and important problem: the efficient construction of sounding video generation models. While numerous video and audio generation models have been released, models capable of jointly generating audio and video remain limited. Efficiently constructing such models is a hot topic in this field.\n- The design of the proposed model is simple and applicable to any DiT-based models, which are a standard choice for recent audio and video generation tasks.\n- The idea of creating dedicated text prompts for audio and video seems generally reasonable, though there are concerns about how to create them, which I will mention in the Weaknesses section."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The empirical findings obtained through this work are somewhat limited, and it is not clear if they are sufficiently general.\n  - Since the proposed architecture is widely applicable to DiT-based models, it is highly recommended to examine it with various backbones other than Wan + Stable Audio Open.\n  - Several empirical findings are interesting; for example, a few BridgeDiT blocks are enough to construct sounding video generation models, and dual cross-attention is more effective than full-attention or addition. However, it is not clear if these results depend on the choice of backbone models.\n- It would be helpful if the authors could show how to ensure that the generated audio caption accurately describes the content in the soundtrack of the source sounding video.\n  - Since the proposed captioning framework solely relies on visual content in a sounding video, the generated audio caption is essentially limited to on-screen audible events. This means that any off-screen sounds could degrade the quality of the audio caption, making the training data noisier.\n  - The datasets used in the experiments are all carefully filtered, so I think this may not be a significant issue in the experiments. However, it would be highly desirable in practice to have a validation stage in the framework to ensure that the generated audio caption aligns with the actual audio of the sounding video.\n- Regarding the proposed captioning framework, it is not clearly described how spatial information is handled.\n  - Since this study uses Stable Audio Open as a backbone for audio generation, the proposed model can generate stereo audio.\n  - Therefore, spatial information about audible objects in a scene is important for the corresponding caption to properly describe the audio content.\n  - However, this aspect is not mentioned in the instruction prompts shown in the appendix.\n- The demo page should include results from baselines for qualitative comparison."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762915541296,"tcdate":1761608585972,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission533/Reviewer_K86T"],"signatures":["ICLR.cc/2026/Conference/Submission533/Reviewer_K86T"],"forum":"jcVOMVkljY","number":2,"license":"CC BY 4.0","cdate":1761608585972,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission533/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762915541296,"domain":"ICLR.cc/2026/Conference","replyto":"jcVOMVkljY","id":"hbMP46YIbN","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"TLDR":{"value":"This work presents a systematic approach to text-to-sounding-video generation task, tackling modal conditioning and interaction challenges through disentangled prompting and symmetric feature fusion to achieve superior synchronization."},"keywords":{"value":["Multi-modal learning","sounding video generation"]},"primary_area":{"value":"generative models"},"abstract":{"value":"This study focuses on a challenging yet promising task, Text-to-Sounding-Video (T2SV) generation, which aims to generate a video with synchronized audio from text conditions, meanwhile ensuring both modalities are aligned with text. \nDespite progress in joint audio-video training, two critical challenges still remain unaddressed: (1) a single, shared text caption where the text for video is equal to the text for audio $(T_V = T_A)$ often creates modal interference, confusing the pretrained backbones, and (2) the optimal mechanism for cross-modal feature interaction remains unclear.\nTo address these challenges, we first propose the Hierarchical Visual-Grounded Captioning (HVGC) framework that generates pairs of disentangled captions, a video caption ($T_V$), and an audio caption ($T_A$), eliminating interference at the conditioning stage.\nBased on HVGC, we further introduce BridgeDiT, a novel dual-tower diffusion transformer, which employs a Dual CrossAttention (DCA) mechanism that acts as a robust ``bridge\" to enable a symmetric, bidirectional exchange of information, achieving both semantic and temporal synchronization.\nExtensive experiments on three benchmark datasets, supported by human evaluations, demonstrate that our method achieves state-of-the-art results on most metrics. Comprehensive ablation studies further validate the effectiveness of our contributions, offering key insights for the future T2SV task. All the codes and checkpoints will be publicly released."},"_bibtex":{"value":"@misc{\nguan2026taming,\ntitle={Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction},\nauthor={Kaisi Guan and Xihua Wang and Zhengfeng Lai and Xin Cheng and Peng Zhang and Xiaojiang Liu and Ruihua Song and Meng Cao},\nyear={2026},\nurl={https://openreview.net/forum?id=jcVOMVkljY}\n}"},"title":{"value":"Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction"},"pdf":{"value":"/pdf/44c683f892c30063f615cc3f57a5aaf3664fbcc4.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"guan|taming_texttosounding_video_generation_via_advanced_modality_condition_and_interaction"},"authorids":{"value":["~Kaisi_Guan1","~Xihua_Wang2","~Zhengfeng_Lai1","~Xin_Cheng6","~Peng_Zhang3","~Xiaojiang_Liu2","~Ruihua_Song1","~Meng_Cao2"]},"authors":{"value":["Kaisi Guan","Xihua Wang","Zhengfeng Lai","Xin Cheng","Peng Zhang","Xiaojiang Liu","Ruihua Song","Meng Cao"]}},"version":2},{"content":{"venue":{"value":"CVPR 2026"},"abstract":{"value":"Generative video models achieve high visual fidelity but often violate basic physical principles, limiting reliability in real‑world settings. Prior attempts to inject physics rely on conditioning: frame‑level signals are domain‑specific and short‑horizon, while global text prompts are coarse and noisy, missing fine‑grained dynamics. We present PhysVid, a physics‑aware local conditioning scheme that operates over temporally contiguous chunks of frames. Each chunk is annotated with physics‑grounded descriptions of states, interactions, and constraints, which are fused with the global prompt via chunk‑aware cross‑attention during training. At inference, we introduce negative physics prompts (descriptions of locally relevant law violations) to steer generation away from implausible trajectories. On VideoPhy, PhysVid improves physical commonsense scores by $\\approx 33$% over baseline video generators, and by up to $\\approx 8$% on VideoPhy2. These results show that local, physics‑aware guidance substantially increases physical plausibility in generative video and marks a step toward physics‑grounded video models."},"_bibtex":{"value":"@inproceedings{\npathak2026physvid,\ntitle={PhysVid: Physics Aware Local Conditioning for Generative Video Models},\nauthor={Saurabh Pathak and Elahe Arani and Mykola Pechenizkiy and Bahram Zonooz},\nbooktitle={Conference on Computer Vision and Pattern Recognition 2026},\nyear={2026},\nurl={https://openreview.net/forum?id=DWk15AhY2I}\n}"},"title":{"value":"PhysVid: Physics Aware Local Conditioning for Generative Video Models"},"pdf":{"value":"https://openaccess.thecvf.com/content/CVPR2026/papers/Pathak_PhysVid_Physics_Aware_Local_Conditioning_for_Generative_Video_Models_CVPR_2026_paper.pdf"},"venueid":{"value":"thecvf.com/CVPR/2026/Conference"},"paperhash":{"value":"pathak|physvid_physics_aware_local_conditioning_for_generative_video_models"},"authorids":{"value":["~Saurabh_Pathak1","~Elahe_Arani1","~Mykola_Pechenizkiy1","~Bahram_Zonooz1"]},"authors":{"value":["Saurabh Pathak","Elahe Arani","Mykola Pechenizkiy","Bahram Zonooz"]}},"tmdate":1789656732724,"pdate":1789656438951,"tcdate":1765223567205,"writers":["thecvf.com/CVPR/2026/Conference","thecvf.com/CVPR/2026/Conference/Submission44998/Authors"],"signatures":["thecvf.com/CVPR/2026/Conference/Submission44998/Authors"],"forum":"DWk15AhY2I","license":"CC BY 4.0","number":44998,"cdate":1765223567205,"readers":["everyone"],"invitations":["thecvf.com/CVPR/2026/Conference/-/Submission","thecvf.com/CVPR/2026/Conference/Submission44998/-/Full_Submission","thecvf.com/CVPR/2026/Conference/-/Post_Submission","thecvf.com/CVPR/2026/Conference/Submission44998/-/Supplementary_Material","thecvf.com/CVPR/2026/Conference/-/Edit","thecvf.com/CVPR/2026/Conference/-/Compute_Flag"],"mdate":1789656732724,"odate":1789656438951,"domain":"thecvf.com/CVPR/2026/Conference","id":"DWk15AhY2I","version":2},{"content":{"venue":{"value":"IEEE Access 2018"},"pdf":{"value":"https://ieeexplore.ieee.org/iel7/6287639/8274985/08315010.pdf"},"venueid":{"value":"dblp.org/journals/ACCESS/2018"},"paperhash":{"value":"li|an_improved_resnet_based_on_the_adjustable_shortcut_connections"},"authorids":{"value":["~Baoqi_Li1","https://dblp.org/search/pid/api?q=author:Yuyao_He:"]},"html":{"value":"https://doi.org/10.1109/ACCESS.2018.2814605"},"_bibtex":{"value":"@article{DBLP:journals/access/LiH18,\n  author={Baoqi Li and Yuyao He},\n  title={An Improved ResNet Based on the Adjustable Shortcut Connections},\n  year={2018},\n  cdate={1514764800000},\n  journal={IEEE Access},\n  volume={6},\n  pages={18967-18974},\n  url={https://doi.org/10.1109/ACCESS.2018.2814605}\n}\n"},"abstract":{"value":"ResNet can achieve deeper network and higher performance, but there is no good explanation for how identity shortcut connections solve the gradient fading problems. Moreover, it is not reasonable to adopt identity mapping for all layer parameters. In this paper, we first establish a simplified ResNet that is similar to the ResNet in principle, and deduce the back propagation of the networks. Second, according to the back propagation of the simplified ResNet, we indirectly explain how the identity shortcut connections solve the problems of gradient fading in convolutional neural networks. Third, we propose an improved ResNet via adjustable shortcut connections, and design a convex k strategy for the improved ResNet according to the different region parameters changing rules. Experimental results on the CIFAR-10 data set show that the test accuracy of the improved ResNet is 78.63%, which is 2.85% higher than that of ResNet. On the CIFAR-100 data set, the test accuracy of the improved ResNet is 42.53%, which is 3.66% higher than that of ResNet. More importantly, the improved ResNet does not increase the amount of computation compared with the classical ResNet."},"title":{"value":"An Improved ResNet Based on the Adjustable Shortcut Connections"},"authors":{"value":["Baoqi Li","Yuyao He"]}},"tmdate":1738815851872,"pdate":1514764800000,"tcdate":1738815849584,"writers":["~"],"signatures":["~Baoqi_Li1"],"forum":"sYZ0qBv89n","license":"CC BY-SA 4.0","number":300647,"cdate":1514764800000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1738815851872,"domain":"DBLP.org","id":"sYZ0qBv89n","version":2},{"content":{"venue":{"value":"CoRR 2025"},"pdf":{"value":"http://arxiv.org/pdf/2504.06022v2"},"venueid":{"value":"dblp.org/journals/CORR/2025"},"paperhash":{"value":"denninger|camcontexti2v_contextaware_controllable_video_generation"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Luis_Denninger:","~Sina_Mokhtarzadeh_Azar1","https://dblp.org/search/pid/api?q=author:Juergen_Gall:"]},"html":{"value":"https://doi.org/10.48550/arXiv.2504.06022"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2504-06022,\n  publtype={informal},\n  author={Luis Denninger and Sina Mokhtarzadeh Azar and Juergen Gall},\n  title={CamContextI2V: Context-aware Controllable Video Generation},\n  year={2025},\n  month={April},\n  cdate={1743465600000},\n  journal={CoRR},\n  volume={abs/2504.06022},\n  url={https://doi.org/10.48550/arXiv.2504.06022}\n}\n"},"abstract":{"value":"Recently, image-to-video (I2V) diffusion models have demonstrated impressive scene understanding and generative quality, incorporating image conditions to guide generation. However, these models primarily animate static images without extending beyond their provided context. Introducing additional constraints, such as camera trajectories, can enhance diversity but often degrade visual quality, limiting their applicability for tasks requiring faithful scene representation. We propose CamC2V, a context-to-video (C2V) model that integrates multiple image conditions as context with 3D constraints alongside camera control to enrich both global semantics and fine-grained visual details. This enables more coherent and context-aware video generation. Moreover, we motivate the necessity of temporal awareness for an effective context representation. Our comprehensive study on the RealEstate10K dataset demonstrates improvements in visual quality and camera controllability. We will publish our code upon acceptance."},"title":{"value":"CamContextI2V: Context-aware Controllable Video Generation"},"authors":{"value":["Luis Denninger","Sina Mokhtarzadeh Azar","Juergen Gall"]}},"tmdate":1762245919006,"pdate":1735689600000,"externalIds":["dblp:journals/corr/abs-2504-06022"],"tcdate":1762245870881,"writers":["~"],"signatures":["~Sina_Mokhtarzadeh_Azar1"],"forum":"q3u5zWv0Iq","license":"CC BY-SA 4.0","number":651646,"cdate":1743465600000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1762245919006,"domain":"DBLP.org","id":"q3u5zWv0Iq","version":2},{"content":{"venue":{"value":"CoRR 2024"},"pdf":{"value":"http://arxiv.org/pdf/2402.13566v1"},"venueid":{"value":"dblp.org/journals/CORR/2024"},"paperhash":{"value":"hou|eventaware_video_corpus_moment_retrieval"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Danyang_Hou:","~Liang_Pang1","~Huawei_Shen1","~Xueqi_Cheng1"]},"html":{"value":"https://doi.org/10.48550/arXiv.2402.13566"},"_bibtex":{"value":"@article{DBLP:journals/corr/abs-2402-13566,\n  publtype={informal},\n  author={Danyang Hou and Liang Pang and Huawei Shen and Xueqi Cheng},\n  title={Event-aware Video Corpus Moment Retrieval},\n  year={2024},\n  cdate={1704067200000},\n  journal={CoRR},\n  volume={abs/2402.13566},\n  url={https://doi.org/10.48550/arXiv.2402.13566}\n}\n"},"abstract":{"value":"Video Corpus Moment Retrieval (VCMR) is a practical video retrieval task focused on identifying a specific moment within a vast corpus of untrimmed videos using the natural language query. Existing methods for VCMR typically rely on frame-aware video retrieval, calculating similarities between the query and video frames to rank videos based on maximum frame similarity.However, this approach overlooks the semantic structure embedded within the information between frames, namely, the event, a crucial element for human comprehension of videos. Motivated by this, we propose EventFormer, a model that explicitly utilizes events within videos as fundamental units for video retrieval. The model extracts event representations through event reasoning and hierarchical event encoding. The event reasoning module groups consecutive and visually similar frame representations into events, while the hierarchical event encoding encodes information at both the frame and event levels. We also introduce anchor multi-head self-attenion to encourage Transformer to capture the relevance of adjacent content in the video. The training of EventFormer is conducted by two-branch contrastive learning and dual optimization for two sub-tasks of VCMR. Extensive experiments on TVR, ANetCaps, and DiDeMo benchmarks show the effectiveness and efficiency of EventFormer in VCMR, achieving new state-of-the-art results. Additionally, the effectiveness of EventFormer is also validated on partially relevant video retrieval task."},"title":{"value":"Event-aware Video Corpus Moment Retrieval"},"authors":{"value":["Danyang Hou","Liang Pang","Huawei Shen","Xueqi Cheng"]}},"tmdate":1769320794759,"pdate":1704067200000,"tcdate":1747748546808,"writers":["~"],"signatures":["~Huawei_Shen1"],"forum":"nzh5AX6QO6","license":"CC BY-SA 4.0","number":538126,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1769320794759,"domain":"DBLP.org","id":"nzh5AX6QO6","version":2},{"content":{"venue":{"value":"ICRA 2024"},"pdf":{"value":"https://ieeexplore.ieee.org/iel8/10609961/10609862/10610149.pdf"},"venueid":{"value":"dblp.org/conf/ICRA/2024"},"paperhash":{"value":"xu|tivode_a_neural_odebased_approach_for_controllable_video_generation_from_textimage_pairs"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Yucheng_Xu:","~Nanbo_Li1","https://dblp.org/search/pid/api?q=author:Arushi_Goel:","https://dblp.org/search/pid/api?q=author:Zonghai_Yao:","~Zijian_Guo1","https://dblp.org/search/pid/api?q=author:Hamidreza_Kasaei_0001:","https://dblp.org/search/pid/api?q=author:Mohammadreza_Kasaei_0001:","https://dblp.org/search/pid/api?q=author:Zhibin_Li_0001:"]},"html":{"value":"https://doi.org/10.1109/ICRA57147.2024.10610149"},"_bibtex":{"value":"@inproceedings{DBLP:conf/icra/XuLGYG00024,\n  author={Yucheng Xu and Nanbo Li and Arushi Goel and Zonghai Yao and Zijian Guo and Hamidreza Kasaei and Mohammadreza Kasaei and Zhibin Li},\n  title={TiV-ODE: A Neural ODE-based Approach for Controllable Video Generation From Text-Image Pairs},\n  year={2024},\n  cdate={1704067200000},\n  pages={14645-14652},\n  url={https://doi.org/10.1109/ICRA57147.2024.10610149},\n  booktitle={ICRA},\n  crossref={conf/icra/2024}\n}\n"},"abstract":{"value":"Videos capture the evolution of continuous dynamical systems over time in the form of discrete image sequences. Recently, video generation models have been widely used in robotic research. However, generating controllable videos from image-text pairs is an important yet underexplored research topic in both robotic and computer vision communities. This paper introduces an innovative and elegant framework named TiV-ODE, formulating this task as modeling the dynamical system in a continuous space. Specifically, our framework leverages the ability of Neural Ordinary Differential Equations (Neural ODEs) to model the complex dynamical system depicted by videos as a nonlinear ordinary differential equation. The resulting framework offers control over the generated videos’ dynamics, content, and frame rate, a feature not provided by previous methods. Experiments demonstrate the ability of the proposed method to generate highly controllable and visually consistent videos and its capability of modeling dynamical systems. Overall, this work is a significant step towards developing advanced controllable video generation models that can handle complex and dynamic scenes."},"title":{"value":"TiV-ODE: A Neural ODE-based Approach for Controllable Video Generation From Text-Image Pairs"},"authors":{"value":["Yucheng Xu","Nanbo Li","Arushi Goel","Zonghai Yao","Zijian Guo","Hamidreza Kasaei","Mohammadreza Kasaei","Zhibin Li"]}},"tmdate":1780168005621,"pdate":1704067200000,"tcdate":1747145224774,"writers":["~"],"signatures":["~Zijian_Guo1"],"forum":"lHTt1sEIgx","license":"CC BY-SA 4.0","number":448006,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1780168005621,"domain":"DBLP.org","id":"lHTt1sEIgx","version":2},{"content":{"summary":{"value":"This paper presents WonderFree, which is developed to enhance the visual quality and consistency of 3D worlds generated by WonderWorld. In particular, it includes a diffusion-based video restoration model and to improve quality of rendered views, and it also integrates a ConsistView mechanism to enhance consistency across multiple views. The author also built a large-scale WorldScopeDataset to train the restoration network."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"In additional to the weakness part, the authors are suggested to discuss analyze the computational complexity and efficiency. \n\nIs WorldRestorer designed only for refining WonderWorld-generated scenes, or can it generalize to other 3D reconstruction systems?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"The paper is well written, and the proposed pipeline is logically reasonable as an extension of WonderWorld. The implementation is solid, and the experimental results have shown improvements and been well presented. \n\nThe contribution of the WorldScopeDataset (23 M images, 6 K+ scenes) that combines real and synthetic multi-view video pairs is also valuable."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The overall contribution seems incremental. This paper is mainly to add an iterative restoration and fine-tuning stage on top of existing 3D reconstruction framework, i.e., WonderWorld, without fundamentally improving the core 3D generation mechanism. \n\nThe proposed distillation scheme that leverages diffusion or image priors for restoration has been explored in many prior works, and its novelty is limited.\n\nThe framework depends on both the quality of the initial 3D reconstruction and the performance of the diffusion model in correcting errors. This kind of dual dependency probably makes the restoration quality unstable and potentially introduces artifacts or degradation into the 3D model itself.\n\nThe complexity and limited scalability. The iterative multi-view video restoration and 3D refinement loop might be computationally expensive, so it's hard to generalize or deploy at scale."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762920217360,"tcdate":1761566018937,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission8287/Reviewer_LFTP"],"signatures":["ICLR.cc/2026/Conference/Submission8287/Reviewer_LFTP"],"forum":"IwjXdoum4W","number":3,"license":"CC BY 4.0","cdate":1761566018937,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission8287/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762920217360,"domain":"ICLR.cc/2026/Conference","replyto":"IwjXdoum4W","id":"Zh0j6WEfLY","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["3D scene generation; Coss-view consistency; Video restoration"]},"supplementary_material":{"value":"/attachment/04e260817005fefe6b06aa16630a328f18da07d5.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"3D scene generation from a single image has gained significant attention due to its potential to create immersive virtual worlds. However, a key challenge in current 3D generation methods is the limited explorability, which cannot render high-quality images during larger maneuvers beyond the original viewpoint, particularly when attempting to move forward into unseen areas. To address this challenge, we propose WonderFree, a model that enables users to generate 3D worlds with enhanced freedom to explore from diverse angles and directions. Specifically, we decouple this challenge into two key subproblems: novel view quality, which addresses visual artifacts and floating issues in novel views, and cross-view consistency, which ensures spatial consistency across different viewpoints. To enhance rendering quality in novel views, we introduce WorldRestorer, a data-driven video restoration model designed to eliminate floaters and artifacts. In addition, a data collection pipeline is presented to automatically gather training data for WorldRestorer, ensuring it can handle scenes with varying styles needed for 3D scene generation. Furthermore, to improve cross-view consistency, we propose ConsistView, a multi-view joint restoration mechanism that simultaneously restores multiple perspectives while maintaining spatiotemporal coherence. Qualitative visualization results demonstrate that WonderFree not only enhances rendering quality across diverse viewpoints but also improves global coherence and consistency. These improvements are further confirmed by CLIP-based metrics and a user study showing a 77.20\\% preference for WonderFree over WonderWorld."},"_bibtex":{"value":"@misc{\nni2025wonderfree,\ntitle={WonderFree: Enhancing 3D World Generation via Video Diffusion Prior with Multi-view Consistency},\nauthor={Chaojun Ni and Jie Li and Haoyun Li and Hengyu Liu and Xiaofeng Wang and Zheng Zhu and Guosheng Zhao and Boyuan Wang and Chenxin Li and Guan Huang and Wenjun Mei},\nyear={2025},\nurl={https://openreview.net/forum?id=IwjXdoum4W}\n}"},"title":{"value":"WonderFree: Enhancing 3D World Generation via Video Diffusion Prior with Multi-view Consistency"},"pdf":{"value":"/pdf/e80032b0855395cdba0cfa150c69e25ca8e3d187.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"ni|wonderfree_enhancing_3d_world_generation_via_video_diffusion_prior_with_multiview_consistency"},"authorids":{"value":["~Chaojun_Ni1","~Jie_Li41","~Haoyun_Li1","~Hengyu_Liu2","~Xiaofeng_Wang5","~Zheng_Zhu1","~Guosheng_Zhao1","~Boyuan_Wang7","~Chenxin_Li1","~Guan_Huang1","~Wenjun_Mei1"]},"authors":{"value":["Chaojun Ni","Jie Li","Haoyun Li","Hengyu Liu","Xiaofeng Wang","Zheng Zhu","Guosheng Zhao","Boyuan Wang","Chenxin Li","Guan Huang","Wenjun Mei"]}},"version":2},{"content":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["video generative model"]},"supplementary_material":{"value":"/attachment/eeae571b9112665efb8b991b44435036abdfa0ea.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Video DiTs have advanced video generation, yet they still struggle to model multi-instance or subject-object interactions. This raises a key question: How do these models internally represent interactions? To answer this, we curate MATRIX-11K, \na video dataset with interaction-aware captions and multi-instance mask tracks. Using this dataset, we conduct a systematic analysis that formalizes two perspectives of video DiTs: semantic grounding, via video-to-text attention, which evaluates whether noun and verb tokens capture instances and their relations; and semantic propagation, via video-to-video attention, which assesses whether instance bindings persist across frames. We find both effects concentrate in a small subset of interaction-dominant layers. Motivated by this, we introduce MATRIX, a simple and effective regularization that aligns attention in specific layers of video DiTs with multi-instance mask tracks from the MATRIX-11K dataset, enhancing both grounding and propagation. We further propose InterGenEval, an evaluation protocol for interaction-aware video generation. In experiments, MATRIX improves both interaction fidelity and semantic alignment while reducing drift and hallucination. Extensive ablations validate our design choices. Codes and weights will be released."},"_bibtex":{"value":"@inproceedings{\njin2026matrix,\ntitle={{MATRIX}: Mask Track Alignment for Interaction-aware Video Generation},\nauthor={Siyoon Jin and Seongchan Kim and Jae Ho Lee and Dahyun Chung and Hyunwook Choi and Jisu Nam and Jiyoung Kim and Seungryong Kim},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=lhVrFEssk5}\n}"},"title":{"value":"MATRIX: Mask Track Alignment for Interaction-aware Video Generation"},"pdf":{"value":"/pdf/b63f42e37db4e11444f6346e32d76682e26a112a.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"jin|matrix_mask_track_alignment_for_interactionaware_video_generation"},"authorids":{"value":["~Siyoon_Jin1","~Seongchan_Kim2","~Jae_Ho_Lee1","~Dahyun_Chung1","~Hyunwook_Choi2","~Jisu_Nam1","~Jiyoung_Kim2","~Seungryong_Kim1"]},"authors":{"value":["Siyoon Jin","Seongchan Kim","Jae Ho Lee","Dahyun Chung","Hyunwook Choi","Jisu Nam","Jiyoung Kim","Seungryong Kim"]}},"tmdate":1775876956616,"pdate":1769435779822,"tcdate":1757737896148,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission4661/Authors"],"signatures":["ICLR.cc/2026/Conference/Submission4661/Authors"],"forum":"lhVrFEssk5","license":"CC BY 4.0","number":4661,"cdate":1757737896148,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/-/Submission","ICLR.cc/2026/Conference/-/Post_Submission","ICLR.cc/2026/Conference/Submission4661/-/Full_Submission","ICLR.cc/2026/Conference/Submission4661/-/Rebuttal_Revision","ICLR.cc/2026/Conference/-/Edit","ICLR.cc/2026/Conference/Submission4661/-/Camera_Ready_Revision"],"mdate":1775876956616,"odate":1759896705795,"domain":"ICLR.cc/2026/Conference","id":"lhVrFEssk5","version":2},{"content":{"summary":{"value":"The paper introduces Video Active Perception (VAP), a method aimed at efficient inference-time understanding of long-form videos using vision-language models. VAP is designed to actively select keyframes during inference to reduce computational costs while maintaining performance."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"I have addressed this in the weakness section"},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"VAP enhances long-form video QA efficiency by selecting key frames that diverge from a lightweight video generation model, achieving up to 5.6× efficiency improvement and outperforming other models on multiple datasets, demonstrating effective use of prior knowledge for better performance."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"LVNet [1] introduced the hierarchical keyframe selection strategy for the very long-form video and achieves 61.1% accuracy using 12 frames with GPT-4o on the EgoSchema dataset. In contrast, Table 2 shows that VAP attains 58.6% accuracy with 16 frames using GPT-4o mini—a lower accuracy despite utilizing more frames. This discrepancy raises concerns about why VAP underperforms compared to LVNet on EgoSchema, even when processing more frames. I recommend that the authors provide an explanation for VAP's lower performance relative to LVNet. Additionally, comparing the performance on the NExT-QA dataset using approximately 12 frames would offer valuable insights. If VAP performs worse than LVNet when utilizing around 12 frames, it might suggest that the keyframe selection quality of VAP is less effective than that of LVNet. It would be beneficial for the authors to experiment with VAP using 12 frames and GPT-4o to directly compare its performance with LVNet. It would be the best if the author can show the performance of VAP on IntentQA and compare it to LVNet to prove its generalizability on diverse very long-form datasets.\n\nMoreover, to ensure a fair comparison, it would be appropriate to adopt GPT-4 instead of GPT-4o when evaluating VAP against VideoTree. Otherwise, it remains unclear whether any performance boost is due to the VAP architecture or the use of a more powerful language model.\n\nAnother concern is the lack of analysis regarding computational cost and inference speed. Since VAP aims to be efficient at inference time, addressing this concern can be beneficial. Including metrics such as the number of inserted captions and inference speed in Table 1 would allow for a direct comparison of both accuracy and complexity among models.\n\n[1] Park, Jongwoo, et al. \"Too Many Frames, Not All Useful: Efficient Strategies for Long-Form Video QA.\" arXiv preprint arXiv:2406.09396 (2024).\n[2] Li, Jiapeng, et al. \"IntentQA: Context-Aware Video Intent Reasoning.\" Proceedings of the IEEE/CVF International Conference on Computer Vision. 2023."}},"nonreaders":[],"tmdate":1731427375855,"tcdate":1730676040191,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission1117/Reviewer_bXhP"],"signatures":["ICLR.cc/2025/Conference/Submission1117/Reviewer_bXhP"],"forum":"KtqZrNjvjd","number":2,"license":"CC BY 4.0","cdate":1730676040191,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission1117/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427375855,"domain":"ICLR.cc/2025/Conference","replyto":"KtqZrNjvjd","id":"o3nrDUPXuQ","forumContent":{"TLDR":{"value":"The paper proposes an \"active perception\" method that selects key frames using generative models to improve efficiency and performance of vision-language models in long-form video question answering."},"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Video question answering","Vision Language Model"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Large vision-language models (VLMs) have advanced multimodal tasks such as video question answering (QA), yet they struggle with long-form videos due to the computational burden of processing excessive tokens. Inspired by active perception theory, which posits that models gain information by acquiring data that differ from their expectations, we introduce Video Active Perception (VAP), a training-free method to enhance long-form video QA using VLMs. Our approach treats key frame selection as data acquisition in active perception and leverages a lightweight text-conditioned video generation model to represent prior world knowledge. Empirically, VAP achieves state-of-the-art zero-shot results on long-form video QA datasets such as EgoSchema, NExT-QA, ActivityNet-QA and CLEVRER, achieving an increase of up to 5.6 X efficiency by frames per question over standard GPT-4o, Gemini 1.5 Pro, and LLaVA-OV. Moreover, VAP shows stronger reasoning abilities than previous methods and effectively selects key frames relevant to questions. These findings highlight the potential of leveraging active perception to improve efficiency and effectiveness of long-form video QA."},"_bibtex":{"value":"@misc{\nma2025video,\ntitle={Video Active Perception: Efficient Inference-Time Long-Form Video Understanding with Vision-Language Models},\nauthor={Martin Q. Ma and Willis Guo and Aditya Agrawal and Ankit Gupta and Paul Pu Liang and Russ Salakhutdinov and Louis-Philippe Morency},\nyear={2025},\nurl={https://openreview.net/forum?id=KtqZrNjvjd}\n}"},"title":{"value":"Video Active Perception: Efficient Inference-Time Long-Form Video Understanding with Vision-Language Models"},"pdf":{"value":"/pdf/0c35130c0a6315516545bebd5ff63f0df45726d6.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"ma|video_active_perception_efficient_inferencetime_longform_video_understanding_with_visionlanguage_models"},"authorids":{"value":["~Martin_Q._Ma1","~Willis_Guo1","~Aditya_Agrawal3","~Ankit_Gupta7","~Paul_Pu_Liang1","~Russ_Salakhutdinov1","~Louis-Philippe_Morency1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Martin Q. Ma","Willis Guo","Aditya Agrawal","Ankit Gupta","Paul Pu Liang","Russ Salakhutdinov","Louis-Philippe Morency"]}},"version":2},{"content":{"venue":{"value":"CoRR 2015"},"pdf":{"value":"http://arxiv.org/pdf/1503.01070v1"},"venueid":{"value":"dblp.org/journals/CORR/2015"},"paperhash":{"value":"torabi|using_descriptive_video_services_to_create_a_large_data_source_for_video_annotation_research"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Atousa_Torabi:","https://dblp.org/search/pid/api?q=author:Christopher_J._Pal:","~Hugo_Larochelle1","https://dblp.org/search/pid/api?q=author:Aaron_C._Courville:"]},"html":{"value":"http://arxiv.org/abs/1503.01070"},"_bibtex":{"value":"@article{DBLP:journals/corr/TorabiPLC15,\n  publtype={informal},\n  author={Atousa Torabi and Christopher J. Pal and Hugo Larochelle and Aaron C. Courville},\n  title={Using Descriptive Video Services to Create a Large Data Source for Video Annotation Research},\n  year={2015},\n  cdate={1420070400000},\n  journal={CoRR},\n  volume={abs/1503.01070},\n  url={http://arxiv.org/abs/1503.01070}\n}\n"},"abstract":{"value":"In this work, we introduce a dataset of video annotated with high quality natural language phrases describing the visual content in a given segment of time. Our dataset is based on the Descriptive Video Service (DVS) that is now encoded on many digital media products such as DVDs. DVS is an audio narration describing the visual elements and actions in a movie for the visually impaired. It is temporally aligned with the movie and mixed with the original movie soundtrack. We describe an automatic DVS segmentation and alignment method for movies, that enables us to scale up the collection of a DVS-derived dataset with minimal human intervention. Using this method, we have collected the largest DVS-derived dataset for video description of which we are aware. Our dataset currently includes over 84.6 hours of paired video/sentences from 92 DVDs and is growing."},"title":{"value":"Using Descriptive Video Services to Create a Large Data Source for Video Annotation Research"},"authors":{"value":["Atousa Torabi","Christopher J. Pal","Hugo Larochelle","Aaron C. Courville"]}},"tmdate":1745671562531,"pdate":1420070400000,"tcdate":1745671513159,"writers":["~"],"signatures":["~Hugo_Larochelle1"],"forum":"ipJ5pMyz1g","license":"CC BY-SA 4.0","number":406456,"cdate":1420070400000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1745671562531,"domain":"DBLP.org","id":"ipJ5pMyz1g","version":2},{"content":{"summary":{"value":"This paper proposes the Physics-RW benchmark to measure the understanding of physical laws by video understanding models and video generation models. It conducts experiments on some open-source and closed-source models and analyzes the results."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"Please refer to details in weaknesses."},"rating":{"value":5},"code_of_conduct":{"value":"Yes"},"presentation":{"value":2},"contribution":{"value":2},"strengths":{"value":"The authors established a relatively comprehensive dataset to measure the understanding of physical laws by video understanding models and video generation models, analyzed the results in detail, and designed certain experimental explorations to improve the model."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- line 099 \"Encompassing the four major categories xxx\" , why you choose this four categories?\n  - line 103 “xxx generate subsequant videos\", Why did the author choose to model it as a video2video task? Can't text2video measure the physics understanding ability of models like sora? The author's previous motivation started with sora, but this is a T2V model, and now a V2V benchmark is designed. I don't understand the logic here.\n  - line 138: some papers like VideoPhy, VideoScore, PhyGenBench already evaluate capabilities of video generation models to simulate the world。\n\n  - line 193: In order to check whether the generated video conforms to the laws of the physical world, it is not enough to just use FVD. This is explained in VideoPhy/PhyGenBench, so it is not reasonable for the author to only use FVD to judge the video quality as a criterion for measuring the ability of video generation models to simulate the world. In other words, if the author designs for understanding the laws of physics, this is not to only evaluate video quality.\n  - line 259: “we manually review...\" , In some cases, the author manually evaluates the response of the model, which makes the results difficult to reproduce and makes it difficult for others to use this benchmark. Why not use tools like VLMEvalKit?\n  - The author said that the results in Table 4 show that the current model lacks understanding of physical laws, which is an overclaim. Because most of the training tasks of video generation models are T2V, and V2V itself is even more difficult, this will affect the author's evaluation of the correctness of physical laws.\n  - line 329，As shown in Figure 2, many models tend to answer yes directly (Video-LLaVA), which is similar to cheating and will lead to errors in the evaluation results. The author should design certain robust evaluation methods to avoid misjudgment caused by the model always answering yes/no. For example, for a question, ask its affirmative description (the answer is yes) and negative description (the answer is no) at the same time, and both must be answered correctly to be considered correct.\n  - I wonder if the author has explored the experiment of using Video-VLM to extract the caption of the video and then give it to LLM for Yes/No QA judgment. I think using the common sense of LLM should be able to achieve a significant improvement on this benchmark."}},"nonreaders":[],"tmdate":1731427812208,"tcdate":1730947889309,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission8650/Reviewer_YX57"],"signatures":["ICLR.cc/2025/Conference/Submission8650/Reviewer_YX57"],"forum":"vsYt8UHGzI","number":4,"license":"CC BY 4.0","cdate":1730947889309,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission8650/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427812208,"domain":"ICLR.cc/2025/Conference","replyto":"vsYt8UHGzI","id":"toruaAwNNG","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Physical Reasoning","General World Models","Zero-shot Inference"]},"supplementary_material":{"value":"/attachment/8a9d30d7834d6ec1ec200284ecd4b782f179245e.zip"},"primary_area":{"value":"datasets and benchmarks"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"While general world models have demonstrated excellent capability in modeling and simulating the world through video understanding and generation, their ability to reason about physical phenomena beyond mechanics remains underexplored. This includes crucial aspects like thermodynamics, electromagnetism, and optics, all of which are fundamental for simulating and predicting real-world dynamics. Existing benchmarks for evaluating physical reasoning in models often rely on datasets consisting solely of simulator-generated, virtual videos, limiting their generalizability to real-world scenarios.  This limitation hinders the comprehensive evaluation of general world models' physical reasoning in real-world scenarios. To bridge this gap, we introduce the Physics-RW benchmark, a physical reasoning dataset constructed from real-world videos. Encompassing a broad spectrum of real-world phenomena—mechanics, thermodynamics, electromagnetism, and optics—Physics-RW offers a comprehensive evaluation platform. We conducted extensive experiments on the Physics-RW benchmark, and the results indicate that there is still significant room for improvement in the physical reasoning abilities of general world models. We further analyzed the experimental results and explored several avenues for improvement. Virtual environment finetuning and physical knowledge injection via prompts demonstrate the potential for enhancing zero-shot physical reasoning ability."},"_bibtex":{"value":"@misc{\nzhao2024bridging,\ntitle={Bridging the Reality Gap: A Benchmark for Physical Reasoning in General World Models with Various Physical Phenomena beyond Mechanics},\nauthor={Pengyu Zhao and Ning Cheng and Huiqi Hu and Xue Zhang and Xiuwen Xu and Zijian Jin and Fandong Meng and Jie Zhou and Jinan Xu and Wenjuan Han},\nyear={2024},\nurl={https://openreview.net/forum?id=vsYt8UHGzI}\n}"},"title":{"value":"Bridging the Reality Gap: A Benchmark for Physical Reasoning in General World Models with Various Physical Phenomena beyond Mechanics"},"pdf":{"value":"/pdf/116f8bbe3a9e86bbe3daa01fc8bba295ed595389.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"zhao|bridging_the_reality_gap_a_benchmark_for_physical_reasoning_in_general_world_models_with_various_physical_phenomena_beyond_mechanics"},"authorids":{"value":["~Pengyu_Zhao3","~Ning_Cheng1","~Huiqi_Hu1","~Xue_Zhang3","~Xiuwen_Xu1","~Zijian_Jin1","~Fandong_Meng3","~Jie_Zhou8","~Jinan_Xu1","~Wenjuan_Han1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Pengyu Zhao","Ning Cheng","Huiqi Hu","Xue Zhang","Xiuwen Xu","Zijian Jin","Fandong Meng","Jie Zhou","Jinan Xu","Wenjuan Han"]}},"version":2},{"content":{"summary":{"value":"This paper introduces the Multi-human Interactive Talking (MIT) dataset, designed for multi-person talking video generation, and presents CovOG, a baseline model that integrates both pose and audio cues to produce natural and realistic multi-human talking videos. The resulting dataset comprises 12 hours of high-resolution footage, each featuring two to four speakers, with fine-grained annotations of body poses and speech interactions."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"See weaknesses."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1.This work presents the first dataset specifically designed for multi-human talking video generation, addressing the scarcity of multi-human interactive data.\n2. It develops an automatic data collection pipeline and construct the first dataset for multi-human talking video generation, featuring annotations of pose and speech interaction.\n3. A baseline model is proposed for this task, which supports a flexible number of human speakers and captures the dynamics of speech interactions. We further conduct extensive studies to benchmark our baseline against existing methods and analyze its performance."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Although this paper is the first to propose a dataset for multi-human talking video generation, several recent works have also addressed this task, including: \n[1] HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters\n[2] Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation\n[3] Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router\n\n2. Current video generation models typically require a large amount of training data. However, the proposed dataset contains only about 2 hours of video, which is relatively small and may limit its applicability. Moreover, the videos are sourced from only two channels, resulting in limited diversity.\n3. Most talking head generation methods employ evaluation metrics such as Sync-C and Sync-D (as introduced in[4]) to measure the synchronization between audio and lip movements. However, this paper does not include any evaluation of audio–lip synchronization, which makes it difficult to quantitatively assess the alignment quality of the generated videos. \n[4]Out of time: automated lip sync in the wild.\n4. It is unclear whether the proposed method supports single-person talking head generation. If it does, a comparison with other single-person talking head methods should be provided.\n5. As far as I am aware, TalkNet’s scores can be highly unreliable when the speaker’s speech intervals are short. Additionally, the presence of background noise in the video may lead to evaluation errors. I would appreciate it if the authors could elaborate on how their data annotation pipeline addresses these issues."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762933739587,"tcdate":1761719968564,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission20255/Reviewer_iKXf"],"signatures":["ICLR.cc/2026/Conference/Submission20255/Reviewer_iKXf"],"forum":"MC5KhOUT7r","number":1,"license":"CC BY 4.0","cdate":1761719968564,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission20255/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762933739587,"domain":"ICLR.cc/2026/Conference","replyto":"MC5KhOUT7r","id":"jnEz4oq6ML","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Audio-driven Video Generation","Multi-human Interactive Talk Video Generation"]},"supplementary_material":{"value":"/attachment/a37a6bdba8f5b0c600b718aba5c05b4473fad9cf.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Existing studies on talking video generation have predominantly focused on single-person monologues or isolated facial animations, limiting their applicability to realistic multi-human interactions. To bridge this gap, we introduce MIT, a large-scale dataset specifically designed for multi-human talking video generation. To this end, we develop an automatic pipeline that collects and annotates multi-person conversational videos. The resulting dataset comprises 12 hours of high-resolution footage, each featuring two to four speakers, with fine-grained annotations of body poses and speech interactions. It captures natural conversational dynamics in multi-speaker scenario, offering a rich resource for studying interactive visual behaviors. To demonstrate the potential of MIT, we furthur propose CovOG, a baseline model for this novel task. It integrates a Multi-Human Pose Encoder (MPE) to handle varying numbers of speakers by aggregating individual pose embeddings, and an Interactive Audio Driver (IAD) to modulate head dynamics based on speaker-specific audio features. Together, these components showcase the feasibility and challenges of generating realistic multi-human talking videos, establishing MIT as a valuable benchmark for future research. The dataset and code will be public available."},"_bibtex":{"value":"@misc{\nzhu2026multihuman,\ntitle={Multi-Human Interactive Talking Dataset},\nauthor={Zeyu Zhu and Weijia Wu and Mike Zheng Shou},\nyear={2026},\nurl={https://openreview.net/forum?id=MC5KhOUT7r}\n}"},"title":{"value":"Multi-Human Interactive Talking Dataset"},"pdf":{"value":"/pdf/e4492e9a6877d1e7bffde01b16a3fab4e6cf5459.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"zhu|multihuman_interactive_talking_dataset"},"authorids":{"value":["~Zeyu_Zhu2","~Weijia_Wu2","~Mike_Zheng_Shou1"]},"authors":{"value":["Zeyu Zhu","Weijia Wu","Mike Zheng Shou"]}},"version":2},{"content":{"summary":{"value":"This paper presents a method to adapt a video diffusion model (a text-to-video generative model) into a referral video segmenter (a text+video to masklet model).\nInstead of segmenting clearly defined objects as conventionally done in image and video object segmentation, the authors focus on segmenting dynamic concepts such as a wave or a cloud of smoke. In order to evaluate their model on this task, the authors introduce the Referral Video Process Segmentation (Ref-VPS) benchmark, consisting of 111 annotated Tiktok videos.\nBecause of the lack of training data for segmenting such dynamic \"entities\", the authors propose to use a video diffusion model, which has learned the mapping from linguistic descriptions to video on massive amounts of text-video pairs. The model is very minimally adapted to the Referral Video Segmentation task in order to conserve its capabilities acquired during pretraining."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"In order to match the pre-training input distribution of the model, the authors add noise to the input (as performed during pre-training).\nHowever, it corrupts the input image and reduces the information amount accessed by the model.\nHave the authors experimented with not adding noise to the input? That would simplify the approach further.\n\nHow many frames does the model process at once? (What is the value of T ?)\n\nMinor typos:\nL309: \"enities\"\nL386: \"UINNEXT\"\nL495: \"StableDiffsuion\"\nL523: \"changing a little as possible\""},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"The paper tackles an interesting problem in video segmentation.\nThe approach is straightforward and well-motivated, leveraging existing diffusion models.\nThe contributions are clearly articulated, and the paper is well-structured.\nThe authors commit to releasing code, models, and data, which supports reproducibility of their results by the community."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"Evaluation of Ref-VPS:\nDynamic concepts such as a wave or a cloud of smoke don't have clearly defined edges. Providing a single ground-truth contour for them is necessarily arbitrary. How then to evaluate? \nFor example, in Figure 1, the mask for “the smoke dissipating” arguably does not cover all the smoke. Similarly, for ”the wave crashing in the ocean”, the ground-mask does not include the wave foam on the left. A model prediction that includes these elements would be unjustly penalised.\nSegmentations on this kind of \"dynamic concepts\" are notoriously hard to evaluate but this issue is mostly not covered in this paper. The authors exclude the contour accuracy F metric but the J metric also suffers from this issue. There are possible solutions to alleviate this problem and their discussion is expected for a paper introducing a benchmark on segmenting \"Dynamic concepts\".\n\nRuntime:\nIn order to be practical, the approach presented in this paper should have a reasonable running time. The paper does not give details on this subject.\nFor example, how long does it take for the model to segment a video of 256 frames, using an A100 GPU?"}},"nonreaders":[],"tmdate":1731427282102,"tcdate":1730578179133,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission581/Reviewer_NxrS"],"signatures":["ICLR.cc/2025/Conference/Submission581/Reviewer_NxrS"],"forum":"Ir6JxcuP6H","number":4,"license":"CC BY 4.0","cdate":1730578179133,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission581/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427282102,"domain":"ICLR.cc/2025/Conference","replyto":"Ir6JxcuP6H","id":"U0GRnnm4QP","forumContent":{"venue":{"value":"Submitted to ICLR 2025"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["Referring Video Segmentation","Video Diffusion Models"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"We present REM, a framework for segmenting a wide variety of concepts in video that can be described through natural language. To achieve this level of generalization, our method capitalizes on visual-language representations learned by video diffusion models on Internet-scale datasets. A key insight of our approach is preserving as much of the generative model’s original representation as possible, while fine-tuning it on narrow-domain Referral Object Segmentation datasets. As a result, despite being exclusively trained on object masks from a limited set of categories, our framework is able to accurately segment and track both rare, unseen objects and non-object, dynamic concepts, such as waves crushing in the ocean. To better quantify the generalization capabilities of our model, we introduce a new benchmark for Referral Video Process Segmentation (RVPS), which captures dynamic phenomena that exist at the intersection of video and language. Our experiments show that REM performs comparably to state-of-the-art approaches on in-domain datasets while outperforming them by up to 28\\% out-of-domain, leveraging the power of Internet-scale pre-training."},"_bibtex":{"value":"@misc{\nbagchi2025refereverything,\ntitle={ReferEverything: Towards segmenting everything we can speak of in videos},\nauthor={Anurag Bagchi and Zhipeng Bao and Yu-Xiong Wang and Pavel Tokmakov and Martial Hebert},\nyear={2025},\nurl={https://openreview.net/forum?id=Ir6JxcuP6H}\n}"},"title":{"value":"ReferEverything: Towards segmenting everything we can speak of in videos"},"pdf":{"value":"/pdf/bd7d631ffb089b4987fd8dff3b7b73ecbf21660c.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Rejected_Submission"},"paperhash":{"value":"bagchi|refereverything_towards_segmenting_everything_we_can_speak_of_in_videos"},"authorids":{"value":["~Anurag_Bagchi1","~Zhipeng_Bao1","~Yu-Xiong_Wang1","~Pavel_Tokmakov2","~Martial_Hebert1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Anurag Bagchi","Zhipeng Bao","Yu-Xiong Wang","Pavel Tokmakov","Martial Hebert"]}},"version":2},{"content":{"venue":{"value":"ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)"},"pdf":{"value":"/pdf/0f77432a1b2686be1d437752103cb6ec086ab4b7.pdf"},"venueid":{"value":"OpenReview.net/Archive"},"paperhash":{"value":"lyu|cleargcd_mitigating_shortcut_learning_for_robust_generalized_category_discovery"},"authorids":{"value":["~Kailin_Lyu1","~Jianwei_He1","~Long_Xiao2","~Jianing_Zeng1","~Liang_Fan2","~Lin_Shu1","~Jie_Hao2"]},"html":{"value":"https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=11463966"},"abstract":{"value":"In open-world scenarios, Generalized Category Discovery (GCD) requires identifying both known and novel categories in unlabeled data, but existing methods often face prototype confusion from shortcut learning, weakening generalization and causing forgetting of known classes. We propose ClearGCD, a framework that suppresses reliance on non-semantic cues through two complementary mechanisms: Semantic View Alignment (SVA), generating strong augmentations via cross-class patch replacement while enforcing semantic consistency with weak augmentations, and Shortcut Suppression Regularization (SSR), maintaining an adaptive prototype bank that aligns known classes and separates potential novel ones. ClearGCD can be seamlessly integrated into parametric GCD methods and consistently outperforms state-of-the-art methods across multiple benchmarks."},"title":{"value":"ClearGCD: Mitigating Shortcut Learning for Robust Generalized Category Discovery"},"authors":{"value":["Kailin Lyu","Jianwei He","Long Xiao","Jianing Zeng","Liang Fan","Lin Shu","Jie Hao"]}},"tmdate":1791560279566,"pdate":1776722400000,"tcdate":1791560279566,"writers":["~Kailin_Lyu1","~Jianwei_He1","~Long_Xiao2","~Jianing_Zeng1","~Liang_Fan2","~Lin_Shu1","~Jie_Hao2"],"signatures":["~Liang_Fan2"],"forum":"iJTTChsTqB","license":"CC BY 4.0","number":55557,"cdate":1791560279566,"readers":["everyone"],"invitations":["OpenReview.net/Archive/-/Direct_Upload"],"mdate":1791560279566,"domain":"OpenReview.net/Archive","id":"iJTTChsTqB","version":2},{"content":{"comment":{"value":"I thank the authors for their careful replay. My first two questions have been addressed. I still concern about the basic assumption the frequent existence of \"one-to-many\" as mentioned in my question (3). I can see the same concern from Reviewer muWC.\n\nIn my understanding, the Uncertainty-Aware Alignment Mechanism module can be seen as a pseudo-label generator that is used to find matching video text pairs in the target domain. Even if the \"one-to-many\" assumption does not hold, it seems that the “top-2 UAM” algorithm helps to find more true-positive pairs (albeit with a potential for increased noise), thus allowing the model to learn more on target domain, while the basic “top-1 UAM” algorithm results in many true-positive pairs being filtered out. \n\n(1) In the authors' reply,  the authors calculate the similarities among 1000 text and pick more than 200 new pairs whose similarities are larger than 0.9. Can the authors show some pairs they found. By the way, it is different from how UAM finds multiple text pairs belonging to the same video, so I am still confused about how many new pairs will be found by the UAM. Top-2 UAM algorithm will obviously finds more pairs than top-1 UAM, so in the newly found pairs, how many of them belongs to the ground-truth pairs.\n\n(2) If \"one-to-many\" exists frequently in the datasets, does that mean the UAM should also be applied in the source domain.\n\nThe Uncertainty-Aware Alignment Mechanism module does show effectiveness. My concern relates to the author's explanation about why it is effective."}},"tmdate":1704271991419,"tcdate":1692258778979,"writers":["NeurIPS.cc/2023/Conference","NeurIPS.cc/2023/Conference/Submission1873/Reviewer_NzVt"],"signatures":["NeurIPS.cc/2023/Conference/Submission1873/Reviewer_NzVt"],"forum":"iQlK3VJxV7","number":1,"license":"CC BY 4.0","cdate":1692258778979,"mdate":1704271991419,"readers":["everyone"],"invitations":["NeurIPS.cc/2023/Conference/Submission1873/-/Official_Comment","NeurIPS.cc/2023/Conference/-/Edit"],"domain":"NeurIPS.cc/2023/Conference","replyto":"RzCl7BJZMh","id":"gIZnjhMFJb","forumContent":{"venue":{"value":"NeurIPS 2023 poster"},"keywords":{"value":["video-text retrieval; cross-domain;Unsupervised Domain Adaptation Video-text Retrieval;"]},"_bibtex":{"value":"@inproceedings{\nhao2023uncertaintyaware,\ntitle={Uncertainty-Aware Alignment  Network  for Cross-Domain Video-Text Retrieval},\nauthor={Xiaoshuai Hao and Wanqian Zhang},\nbooktitle={Thirty-seventh Conference on Neural Information Processing Systems},\nyear={2023},\nurl={https://openreview.net/forum?id=iQlK3VJxV7}\n}"},"title":{"value":"Uncertainty-Aware Alignment  Network  for Cross-Domain Video-Text Retrieval"},"paperhash":{"value":"hao|uncertaintyaware_alignment_network_for_crossdomain_videotext_retrieval"},"TLDR":{"value":"In this paper, we address the challenge task of Unsupervised Domain Adaptation Video-text Retrieval (UDAVR), assuming that training (source) data and testing (target) data are from different domains."},"abstract":{"value":"Video-text retrieval is an important but challenging research task in the multimedia community.  In this paper, we address the challenge task of Unsupervised Domain Adaptation Video-text Retrieval (UDAVR), assuming that training (source) data and testing (target) data are from different domains. Previous approaches are mostly derived from classification based domain adaptation methods, which are neither multi-modal nor suitable for retrieval task.  In addition, as to the pairwise misalignment issue in target domain, i.e., no pairwise annotations between target videos and texts, the existing method assumes that a video corresponds to a text. Yet we empirically find that in the real scene, one text usually corresponds to multiple videos and vice versa. To tackle this one-to-many issue, we propose a novel method named Uncertainty-aware Alignment Network (UAN). Specifically, we first introduce the multimodal mutual information module to balance the minimization of domain shift in a smooth manner. To tackle the multimodal uncertainties pairwise misalignment in target domain, we propose the Uncertainty-aware Alignment Mechanism (UAM) to fully exploit the semantic information of both modalities in target domain. Extensive experiments in the context of domain-adaptive video-text retrieval demonstrate that our proposed method consistently outperforms multiple baselines, showing a superior generalization ability for target data."},"pdf":{"value":"/pdf/8e69ebbe33b1ba360a90af65d4977177eb771a6c.pdf"},"venueid":{"value":"NeurIPS.cc/2023/Conference"},"authorids":{"value":["~Xiaoshuai_Hao1","~Wanqian_Zhang1"]},"authors":{"value":["Xiaoshuai Hao","Wanqian Zhang"]}},"version":2},{"content":{"data_release":{"value":"We authorize the release of our submission and author names to the public in the event of acceptance."},"venue":{"value":"ReALM-GEN 2026 - ICLR 2026 Workshop"},"email_sharing":{"value":"We authorize the sharing of all author emails with Program Chairs."},"pdf":{"value":"/pdf/3c8fea02a6506b71c230458a1df8509ae7d05025.pdf"},"keywords":{"value":["text-video diffusion","corruption-aware training","structured noise injection","multimodal robustness","temporal coherence"]},"venueid":{"value":"ICLR.cc/2026/Workshop/ReALM-GEN"},"paperhash":{"value":"maduabuchi|corruptionaware_training_of_latent_video_diffusion_models_for_robust_texttovideo_generation"},"authorids":{"value":["~Chika_Maduabuchi1","~Hao_Chen15","~Yujin_Han1","~Jindong_Wang4"]},"abstract":{"value":"Latent Video Diffusion Models (LVDMs) have achieved state-of-the-art generative quality for image and video generation; however, they remain brittle under noisy conditioning, where small perturbations in text or multimodal embeddings can cascade over timesteps and cause semantic drift. Existing corruption strategies from image diffusion (Gaussian, Uniform) fail in video settings because static noise disrupts temporal fidelity. In this paper, we propose CAT-Video, a corruption-aware training framework with structured, data-aligned noise injection tailored for video diffusion. Our two operators—Batch-Centered Noise Injection (BCNI) and Spectrum-Aware Contextual Noise (SACN)—align perturbations with batch semantics or spectral dynamics to preserve coherence. CAT-Video yields substantial gains: BCNI reduces FVD by 31.9% on WebVid-2M, MSR-VTT, and MSVD, while SACN improves UCF-101 by 12.3%, outperforming Gaussian, Uniform, and even large diffusion baselines like DEMO (2.3B) and Lavie (3B) despite training on \n5x less data. Ablations confirm the unique value of low-rank, data-aligned noise, and theory establishes why these operators tighten robustness and generalization bounds. CAT-Video thus sets a new framework for robust video diffusion, and our experiments show that it can also be extended to autoregressive generation and multimodal video understanding LLMs."},"title":{"value":"Corruption-Aware Training of Latent Video Diffusion Models for Robust Text-to-Video Generation"},"authors":{"value":["Chika Maduabuchi","Hao Chen","Yujin Han","Jindong Wang"]}},"tmdate":1774799810107,"pdate":1772442528397,"tcdate":1769948238024,"writers":["ICLR.cc/2026/Workshop/ReALM-GEN","ICLR.cc/2026/Workshop/ReALM-GEN/Submission14/Authors"],"signatures":["ICLR.cc/2026/Workshop/ReALM-GEN/Submission14/Authors"],"forum":"nDIA7FjQtX","license":"CC BY 4.0","number":14,"cdate":1769948238024,"readers":["everyone"],"invitations":["ICLR.cc/2026/Workshop/ReALM-GEN/-/Submission","ICLR.cc/2026/Workshop/ReALM-GEN/-/Submission_Change_Before_Bidding","ICLR.cc/2026/Workshop/ReALM-GEN/-/Submission_Change_Before_Reviewing","ICLR.cc/2026/Workshop/ReALM-GEN/-/Submission_Release","ICLR.cc/2026/Workshop/ReALM-GEN/Submission14/-/Camera_Ready_Revision"],"mdate":1774799810107,"odate":1772442528397,"domain":"ICLR.cc/2026/Workshop/ReALM-GEN","id":"nDIA7FjQtX","version":2},{"content":{"summary":{"value":"This paper presents a new task: grounded video caption generation, which aims to generate captions for videos while also providing bounding boxes for the objects mentioned in the captions. \nTo achieve this, the authors designed an automatic grounded video annotation pipeline. Based on the dataset constructed using this pipeline, the paper trains a model called VideoGLaMM, which performs well on the grounded video caption generation task."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"See Weaknesses."},"rating":{"value":3},"code_of_conduct":{"value":"Yes"},"presentation":{"value":1},"contribution":{"value":2},"strengths":{"value":"- An automated grounded video annotation pipeline was designed, significantly reducing annotation costs."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The authors claim to introduced a new task: grounded video caption generation. However, there is prior research on this task, such as PG-VIDEO-LLAVA[1]. How do the authors justify their assertion of novelty?\n\n[1] PG-Video-LLaVA: Pixel Grounding Large Video-Language Models.\n\n- For videos that contain two or more events, how should the automated annotation pipeline be applied? For example, consider a scenario where a man picks up a knife to chop vegetables and then puts the vegetables into a pot. If the knife and pot appear simultaneously in the frame but the action hasn't progressed to the pot yet, how can we avoid generating a bounding box for the pot prematurely?\n\n- The case studies (e.g., Figure 4) presented by the authors seem overly simplistic. Can VideoGLaMM be effectively applied to videos that involve multiple events?\n\n- There are very few comparative methods included in the study. The authors should compare more methods, such as the integration of video captioning models with object detection models, and PG-Video-LLaVA."}},"nonreaders":[],"tmdate":1731427749624,"tcdate":1730550141694,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission12873/Reviewer_LZxE"],"signatures":["ICLR.cc/2025/Conference/Submission12873/Reviewer_LZxE"],"forum":"xYzOkOGD96","number":3,"license":"CC BY 4.0","cdate":1730550141694,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/Submission12873/-/Official_Review","ICLR.cc/2025/Conference/-/Edit"],"mdate":1731427749624,"domain":"ICLR.cc/2025/Conference","replyto":"xYzOkOGD96","id":"oPDDeEleYQ","forumContent":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["vision-language models","VLM","LLM","video grounding","automatic annotation","pseudo-labeling"]},"primary_area":{"value":"datasets and benchmarks"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"We propose a new task, dataset and model for grounded video caption generation. This task unifies captioning and object grounding in video, where the objects in the caption are grounded in the video via temporally consistent bounding boxes.  We introduce the following contributions. First, we present a task definition and a manually annotated test dataset for this task, referred to as GROunded Video Caption Generation (GROC). Second, we introduce a large-scale automatic annotation method leveraging an existing model for grounded still image captioning together with an LLM for summarising frame-level captions into temporally consistent captions in video. \nFurthermore, we prompt the LLM to track by language – classifying noun phrases from the frame-level captions into noun phrases of the video-level generated caption. We apply this approach to videos from the HowTo100M dataset, which results in a new large-scale training dataset, called HowToGround, with automatically annotated captions and spatio-temporally consistent bounding boxes with coherent natural language labels. Third, we introduce a new grounded video caption generation model, called VideoGLaMM, and train the model on the new automatically annotated HowToGround dataset. Finally, results of our VideoGLaMM model set the state of the art for the new task of grounded video caption generation. We perform extensive ablations and demonstrate the importance of key technical contributions of our model."},"_bibtex":{"value":"@misc{\nkazakos2024grounded,\ntitle={Grounded Video Caption Generation},\nauthor={Evangelos Kazakos and Cordelia Schmid and Josef Sivic},\nyear={2024},\nurl={https://openreview.net/forum?id=xYzOkOGD96}\n}"},"title":{"value":"Grounded Video Caption Generation"},"pdf":{"value":"/pdf/7c363c0c76057bc0f14cf8ad8d057a1f6e8bc84b.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"kazakos|grounded_video_caption_generation"},"authorids":{"value":["~Evangelos_Kazakos2","~Cordelia_Schmid1","~Josef_Sivic1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Evangelos Kazakos","Cordelia Schmid","Josef Sivic"]}},"version":2},{"content":{"summary":{"value":"1. The paper proposes a joint audio-video generation model named JavisDiT, which is based on the Diffusion Transformer (DiT) architecture and aims to simultaneously generate high-quality and synchronized audio-video content from text prompts.\n2. The paper designs a core module, the Hierarchical Spatial-Temporal Synchronized Prior Estimator, which extracts both global coarse-grained semantic and local fine-grained spatio-temporal priors from text. These priors are then injected into the DiT module to guide the precise spatial and temporal synchronization of the audio and video.\n3. Addressing the lack of scene diversity and complexity in existing benchmarks, the authors constructed a new and more challenging benchmark dataset, JavisBench. If this dataset were to be made open-source, it would be beneficial to the community's development.\n4. To address the deficiencies of existing metrics (such as AV-Align) , the paper introduces JavisScore. This new metric uses a temporal-aware semantic alignment mechanism to robustly measure synchronization in complex real-world scenarios."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"1. The HiST-Sypo estimator uses ImageBind's text encoder, while the proposed JavisScore metric also relies on ImageBind to measure audio-visual synchronization. Does this not create a 'circular' evaluation, where the model is effectively trained to optimize for the feature space of the metric that is supposed to be judging it?\n2. Regarding the audio autoencoder, why did the authors opt for a Mel-spectrogram based VAE (which necessitates a vocoder) instead of a waveform-based VAE (e.g., Wave-VAE)? The choice of a Mel-VAE introduces an inherent quality loss from the vocoder, which could be a critical bottleneck for the final audio generation quality.\n3. The paper claims to support variable-length audio generation, possibly via dynamic temporal masking. How is this practically implemented during inference? Given that the main evaluations are on 4-second videos, what evidence is there that this mechanism can robustly scale to longer durations (e.g., 10 seconds) while maintaining quality and synchronization?\n4. In the ablation studies (Sec 5.3, line 795), the paper introduces normalized scores ($S_{AVQ}$, $S_{AVC}$, $S_{AVS}$) to simplify the results. What is the justification for the specific weighting factors used in these formulas? These equations appear arbitrary and lack a clear theoretical or empirical basis, making the ablation results difficult to interpret."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"The paper's motivation is clear and precisely targets the two core challenges in the field of Joint Audio-Video Generation (JAVG): ensuring high-quality generation of audio and video, and maintaining perfect synchronization between the two modalities. It is rare for a single work to contribute a model, dataset, and metric simultaneously. This represents a significant potential to advance the field and, as an academic work, is sufficiently compelling."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. There is a clear disconnect between the paper's core claims and the qualitative results provided in the supplementary demo. The paper claims to solve high-quality generation and perfect synchronization, but the demos fail on both counts. The visual quality is significantly lower than the level of current video generation models, and audio-video synchronization is not well demonstrated in these demos, often appearing as coarse \"scene-ambience\" rather than the claimed \"fine-grained\" alignment. Furthermore, the number of samples is insufficient to cover the breadth of capabilities claimed. Critically, the demos lack the most essential evidence: direct comparisons against baseline methods (e.g., \"T2V + V2A\" or \"T2A + A2V\") for the same prompt. While the authors provide objective metrics in Table A8, the corresponding qualitative examples to visually substantiate these claims are missing.\n2. The paper lacks sufficient subjective (human) evaluation. This is a necessary component for measuring video quality, audio quality, and audio-video alignment, and its absence is a significant omission.\n3. The authors' justification for their model's weaker audio performance (e.g., in line 1538) is unconvincing. Attributing this weakness to supporting \"variable-length audio generation\" is a flawed argument, as the baseline AudioLDM 2 also supports this feature. Additionally, while the authors' concern about AudioLDM 2 training on AudioCaps is valid, their use of the term \"data leakage\" is imprecise.\n4. The multi-stage training strategy is overly complex. Does the complex multi-stage training strategy require significant manual intervention and tuning when applied to new datasets.\n5. The paper's readability is poor due to dense prose. For example, the term \"Fine-Grained Spatial-Temporal Self-Attention Cross-Attention\" (line 89) is structurally confusing. This confusion is amplified by the inconsistent naming in Figure 2 (e.g., \"Fine-Grained ST-CrossAttn\" and \"ST-SelfAttn\"), which hinders a clear understanding of the architecture. While the authors state in the Appendix (A.4) that LLMs were used \"solely as writing assistants,\" we believe the paper's over-reliance on such tools has resulted in these convoluted descriptions and reduced overall readability."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762931358570,"tcdate":1761970861634,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission19437/Reviewer_Fny1"],"signatures":["ICLR.cc/2026/Conference/Submission19437/Reviewer_Fny1"],"forum":"y7HV7KT3Bd","number":3,"license":"CC BY 4.0","cdate":1761970861634,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission19437/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762931358570,"domain":"ICLR.cc/2026/Conference","replyto":"y7HV7KT3Bd","id":"2edlLR9MO1","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"TLDR":{"value":"We introduce JavisDiT, a novel Joint Audio-Video Diffusion Transformer designed for synchronized audio-video generation (JAVG) from open-ended user prompts."},"keywords":{"value":["Diffusion Transformer","Joint Audio-Video Generation","Text-to-Audio-Video Generation","Video Generation"]},"supplementary_material":{"value":"/attachment/6d0bd1487a58f9109bb452ac66b2d5b21654cfea.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"This paper introduces JavisDiT, a novel Joint Audio-Video Diffusion Trans- former designed for synchronized audio-video generation (JAVG). Based on the powerful Diffusion Transformer (DiT) architecture, JavisDiT simultaneously generates high-quality audio and video content from open-ended user prompts in a unified framework. To ensure audio-video synchronization, we introduce a fine-grained spatio-temporal alignment mechanism through a Hierarchical Spatial-Temporal Synchronized Prior (HiST-Sypo) Estimator. This module extracts both global and fine-grained spatio-temporal priors, guiding the synchronization between the visual and auditory components. Furthermore, we propose a new benchmark, JavisBench, which consists of 10,140 high-quality text-captioned sounding videos and focuses on synchronization evaluation in diverse and complex real-world scenarios. Further, we specifically devise a robust metric for measuring the synchrony between generated audio-video pairs in real-world content. Experimental results demonstrate that JavisDiT significantly outperforms existing methods by ensuring both high-quality generation and precise synchronization, setting a new standard for JAVG tasks. Our code, model, and data are available at https://javisverse.github.io/JavisDiT-page/."},"_bibtex":{"value":"@inproceedings{\nliu2026javisdit,\ntitle={JavisDiT: Joint Audio-Video Diffusion Transformer with Hierarchical Spatio-Temporal Prior Synchronization},\nauthor={Kai Liu and Wei Li and Lai Chen and Shengqiong Wu and Yanhao Zheng and Jiayi Ji and Fan Zhou and Jiebo Luo and Ziwei Liu and Hao Fei and Tat-Seng Chua},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=y7HV7KT3Bd}\n}"},"title":{"value":"JavisDiT: Joint Audio-Video Diffusion Transformer with Hierarchical Spatio-Temporal Prior Synchronization"},"pdf":{"value":"/pdf/07f1f045a6743ca22be0e5f8d9852115a1ba5d3f.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"liu|javisdit_joint_audiovideo_diffusion_transformer_with_hierarchical_spatiotemporal_prior_synchronization"},"authorids":{"value":["~Kai_Liu8","~Wei_Li111","~Lai_Chen1","~Shengqiong_Wu2","~Yanhao_Zheng1","~Jiayi_Ji1","~Fan_Zhou13","~Jiebo_Luo1","~Ziwei_Liu1","~Hao_Fei1","~Tat-Seng_Chua2"]},"authors":{"value":["Kai Liu","Wei Li","Lai Chen","Shengqiong Wu","Yanhao Zheng","Jiayi Ji","Fan Zhou","Jiebo Luo","Ziwei Liu","Hao Fei","Tat-Seng Chua"]}},"version":2},{"content":{"summary":{"value":"This paper presents InternVid, an large-scaled video-language dataset, consisting of a collection of 7M videos, 234M video clips and corresponding textual descriptions. They consider diversity and data quality by sourcing YouTube contents from various countries and languages. This new dataset enables pretraining a video-language model, named ViCLIP, exhibiting state-of-the-art scores in action recognition and video retrieval tasks. It also demonstrates a credible generative capability in text-to-video generation, supported by qualitative and quantitative assessment."},"soundness":{"value":"2 fair"},"confidence":{"value":"4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work."},"questions":{"value":"1. Consideration with Table 6\n\n- Comment of VideoCrafter, VideoFusion is absent. Does InternVID also improve VideoCrafter, VideoFusion in IS, FID, FVD, and CLIPSIM? Can you provide further analysis about the choice of InternVID-Aes-18M rather than exploiting the full InternVID-200M? \n\n2. Consideration with Section 5.2 and Figure 13~17\n\n- To insist that InternVID provides more powerful video-text aligned representations for video-centric dialogue systems, authors must compare the results from VideoChat-ViCLIP with that of vanilla VideoChat. Can you provide quantitative and qualitative comparisons, showing improvement of the results after replacement of the previous visual encoder with ViCLIP?"},"rating":{"value":"6: marginally above the acceptance threshold"},"details_of_ethics_concerns":{"value":"The data source is YouTube, where the content often includes personal recordings, potentially raising privacy concerns. The collectors should carefully address these issues and rigorous data selection criteria should be put in place to ensure strict adherence to privacy and other relevant guidelines."},"code_of_conduct":{"value":"Yes"},"presentation":{"value":"2 fair"},"contribution":{"value":"3 good"},"strengths":{"value":"1. Clarity in writing\n\n- The paper is well-written and easy to understand.\n\n2. Scale and impact\n\n- InternVid includes large-scaled 7.1M videos, making it a significant contribution to the video-language pretraining. Referring to table 1, InternVID comprises 234M 720p video clips with LLM generated captions which shows outperforming accuracy in description. The dataset’s substantial size can potentially have a profound impact on the field.\n\n3. Diversity and Quality\n\n- The effort to collect videos from various countries and languages is commendable. This diversity in data collection may address the issue of bias, a crucial aspect of large-scale datasets, as it mitigates the risk of training biased models.\n\n4. High performance\n\n- The suggested model, referred as ViCLIP, achieve the state-of-the-art in video-language benchmarks such as  action recognition and text-to-video retrieval.\n- InternVID improves text-to-video generation by empowering the previous models with its large scale video-text dataset. It shows quantitative and qualitative (Figure 12) improvement in generation quality."},"flag_for_ethics_review":{"value":["Yes, Discrimination / bias / fairness concerns","Yes, Privacy, security and safety"]},"weaknesses":{"value":"1. Lack of definitions, explanations\n\n- Definition of UMT-SIM is absent. In table 2, justification of the design choice of InternVid-10M-FLT is not explained. (At glance, the result of performance improvement after filtering the dataset with a high UMT-SIM score seems trivial.)\n- Figure 5 introduced 3 schemes of interleaving video clips, text, and ASR text. However, comparison between the methods or further analysis is not given.\n- It is doubtful that CLIP is a suitable choice for comparing their method in action recognition tasks. There are likely more appropriate video models to use in the benchmark for a fairer comparison.\n\n2. Lack of novelty\n\n- The proposed method for generating the dataset is not novel. It simply uses BLIP to generate captions for video frames, and exploits LLM to generate a summarized caption for the video. Also, description of the used LLM is not given.\n- ViCLIP is not a novel architecture and inherits the CLIP model without significant modifications. The paper predominantly focuses on introducing the dataset, which leaves room for exploring further architectural improvement and training techniques.\n\n3. Unclear model selection process\n\n- The paper lacks information on the criteria used to select the image captioning and language models.\n\n4. Absence of ablation study and potential impact\n\n- In contrast to previous models which are benchmarked in datasets mostly composed of English, InternVID consists of clips with diverse languages. Further analysis about the potential impact caused by this aspect should be considered.\n\n5. Minor comments on Figure 2\n\n- Similar colors like green and dark green are used, but this may cause confusion. Using more distinct and discrete colors, such as red and blue, could enhance the clarity."}},"nonreaders":[],"tmdate":1700553891753,"tcdate":1697874413041,"writers":["ICLR.cc/2024/Conference","ICLR.cc/2024/Conference/Submission4917/Reviewer_1UqW"],"signatures":["ICLR.cc/2024/Conference/Submission4917/Reviewer_1UqW"],"forum":"MLBdiWu4Fw","number":1,"license":"CC BY 4.0","cdate":1697874413041,"readers":["everyone"],"invitations":["ICLR.cc/2024/Conference/Submission4917/-/Official_Review","ICLR.cc/2024/Conference/-/Edit"],"mdate":1700553891753,"domain":"ICLR.cc/2024/Conference","replyto":"MLBdiWu4Fw","id":"H1HxdqN6Qp","forumContent":{"venue":{"value":"ICLR 2024 spotlight"},"TLDR":{"value":"This paper introduces InternVid, a large-scale video-centric multimodal dataset that enables learning powerful and transferable video-text representations for multimodal understanding and generation."},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["video-language dataset","video understanding","video generation","multimodal understanding","action recognition","video retrieval"]},"supplementary_material":{"value":"/attachment/ab6262d86bfdbdf3844255ccbd7ff6bdca41c17b.zip"},"primary_area":{"value":"datasets and benchmarks"},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2024/AuthorGuide."},"abstract":{"value":"This paper introduces InternVid, a large-scale video-centric multimodal dataset that enables learning powerful and transferable video-text representations for multimodal understanding and generation. InternVid contains over 7 million videos lasting nearly 760K hours, yielding 234M video clips accompanied by detailed descriptions of total 4.1B words. Our core contribution is to develop a scalable approach to autonomously build a high-quality video-text dataset with large language models (LLM), thereby showcasing its efficacy in learning video-language representation at scale. Specifically, we utilize a multi-scale approach to generate video-related descriptions. Furthermore, we introduce ViCLIP, a video-text representation learning model based on ViT-L. Learned on InternVid via contrastive learning, this model demonstrates leading zero-shot action recognition and competitive video retrieval performance. Beyond basic video understanding tasks like recognition and retrieval, our dataset and model have broad applications. They are particularly beneficial for generating interleaved video-text data for learning a video-centric dialogue system, advancing video-to-text and text-to-video generation research. These proposed resources provide a tool for researchers and practitioners interested in multimodal video understanding and generation."},"_bibtex":{"value":"@inproceedings{\nwang2024internvid,\ntitle={InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation},\nauthor={Yi Wang and Yinan He and Yizhuo Li and Kunchang Li and Jiashuo Yu and Xin Ma and Xinhao Li and Guo Chen and Xinyuan Chen and Yaohui Wang and Ping Luo and Ziwei Liu and Yali Wang and Limin Wang and Yu Qiao},\nbooktitle={The Twelfth International Conference on Learning Representations},\nyear={2024},\nurl={https://openreview.net/forum?id=MLBdiWu4Fw}\n}"},"title":{"value":"InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation"},"pdf":{"value":"/pdf/5355ce2fec3ff26dca65a969b767fd7b1102bb05.pdf"},"venueid":{"value":"ICLR.cc/2024/Conference"},"paperhash":{"value":"wang|internvid_a_largescale_videotext_dataset_for_multimodal_understanding_and_generation"},"authorids":{"value":["~Yi_Wang19","~Yinan_He1","~Yizhuo_Li1","~Kunchang_Li1","~Jiashuo_Yu1","~Xin_Ma3","~Xinhao_Li1","~Guo_Chen2","~Xinyuan_Chen1","~Yaohui_Wang1","~Ping_Luo2","~Ziwei_Liu1","~Yali_Wang1","~Limin_Wang1","~Yu_Qiao1"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors' identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Yi Wang","Yinan He","Yizhuo Li","Kunchang Li","Jiashuo Yu","Xin Ma","Xinhao Li","Guo Chen","Xinyuan Chen","Yaohui Wang","Ping Luo","Ziwei Liu","Yali Wang","Limin Wang","Yu Qiao"]}},"version":2},{"content":{"summary":{"value":"The paper introduces EgoNight, a benchmark focused on egocentric vision understanding in nighttime. It consists of three data sources: (i) EgoNight-Synthetic, 50 ideally aligned egocentric pairs with varying illumination levels using a simulator; (ii) EgoNight-Sofia, 20 pairs of real-world egocentric videos with spatially and temporally aligned day–night counterparts; (iii) EgoNight-Oxford,  20 nighttime videos from the Oxford Day-and-Night dataset"},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"- Could the authors explain the human verification process for the day-augmented annotation? Specifically, how were QA pairs handled if the answer was only derivable from the day video and not the night video?\n- Are there any strategies to address the synthetic-to-real gap in EgoNight-Synthetic?\n- In Figure5, is NonC in the egoNight-Syn bar plot a typo? Could the authors use the same color for the same task to improve readability?\n- What is the performance of reasoning on MLLM-generated captions of the videos?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"- While most egocentric datasets are overwhelmingly focused on well-lit, daytime scenarios, this paper provides valuable nighttime egocentric data and thus has high significance for real-world applications.\n- The proposed benchmark includes a comprehensive list of QA tasks, along with three critical tasks (day-night correspondence retrieval, temporal localization, egocentric depth estimation) that are unique for nighttime videos.\n- The paper is well-written and easy to follow with carefully designed figures."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- The EgoNight-Sofia (real-world) dataset is the most valuable and novel component in the proposed benchmark. However, it is quite small with only 20 video pairs. The vast majority of QA pairs (2029 of 3658) come from the synthetic dataset, which may not capture the full complexity of real-world sensor noise and lighting artifacts.\n- The day-augmented annotation pipeline could introduce a bias. It risks generating questions about details that are perfectly visible in the day video but are genuinely invisible or non-inferable in the night video. Is this considered in the human verification stage?\n- The evaluation uses an \"LLM-as-a-Judge\" (GPT-4.1) to score answers, which might introduce bias towards its own outputs. GPT4.1 achieving the highest accuracy under this metric is less convincing."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762919699101,"tcdate":1762143358479,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission7619/Reviewer_akCw"],"signatures":["ICLR.cc/2026/Conference/Submission7619/Reviewer_akCw"],"forum":"DKD4QbOKBN","number":4,"license":"CC BY 4.0","cdate":1762143358479,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission7619/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762919699101,"domain":"ICLR.cc/2026/Conference","replyto":"DKD4QbOKBN","id":"8jZ4EP8Re9","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Egocentric vision; Benchmark; MLLMs; VQA"]},"supplementary_material":{"value":"/attachment/d76f00441714388d92564268854c368aff0c8323.pdf"},"primary_area":{"value":"datasets and benchmarks"},"abstract":{"value":"Most existing benchmarks for egocentric vision understanding focus primarily on daytime scenarios, overlooking the low-light conditions that are inevitable in real-world applications. To investigate this gap, we present EgoNight, the first comprehensive benchmark for nighttime egocentric vision, with visual question answering (VQA) as the core task. A key feature of EgoNight is the introduction of day–night aligned videos, which enhance night annotation quality using the daytime data and reveal clear performance gaps between lighting conditions. To achieve this, we collect both synthetic videos rendered by Blender and real-world recordings, ensuring that scenes and actions are visually and temporally aligned. Leveraging these paired videos, we construct EgoNight-VQA, supported by a novel day-augmented night auto-labeling engine and refinement through extensive human verification. Each QA pair is double-checked by annotators for reliability. In total, EgoNight-VQA contains 3658 QA pairs across 90 videos, spanning 12 diverse QA types, with more than 300 hours of human work. Evaluations of the state-of-the-art multimodal large language models (MLLMs) reveal substantial performance drops when transferring from day to night, underscoring the challenges of reasoning under low-light conditions. Beyond VQA, EgoNight also introduces two auxiliary tasks, day–night correspondence retrieval and egocentric depth estimation at night, that further explore the boundaries of existing models. We believe EgoNight-VQA provides a strong foundation for advancing application-driven egocentric vision research and for developing models that generalize across illumination domains.  The code and data can be found in https://github.com/dehezhang2/EgoNight"},"_bibtex":{"value":"@inproceedings{\nzhang2026egonight,\ntitle={EgoNight: Towards Egocentric Vision Understanding at Night with a Challenging Benchmark},\nauthor={Deheng Zhang and Yuqian Fu and Runyi Yang and Yang Miao and Tianwen Qian and Xu Zheng and Guolei Sun and Ajad Chhatkuli and Xuanjing Huang and Yu-Gang Jiang and Luc Van Gool and Danda Pani Paudel},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=DKD4QbOKBN}\n}"},"title":{"value":"EgoNight: Towards Egocentric Vision Understanding at Night with a Challenging Benchmark"},"pdf":{"value":"/pdf/3f37297889894b8d99e35fc46e1a76d639fadaf7.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"zhang|egonight_towards_egocentric_vision_understanding_at_night_with_a_challenging_benchmark"},"authorids":{"value":["~Deheng_Zhang1","~Yuqian_Fu2","~Runyi_Yang1","~Yang_Miao2","~Tianwen_Qian1","~Xu_Zheng2","~Guolei_Sun2","~Ajad_Chhatkuli1","~Xuanjing_Huang1","~Yu-Gang_Jiang1","~Luc_Van_Gool1","~Danda_Pani_Paudel3"]},"authors":{"value":["Deheng Zhang","Yuqian Fu","Runyi Yang","Yang Miao","Tianwen Qian","Xu Zheng","Guolei Sun","Ajad Chhatkuli","Xuanjing Huang","Yu-Gang Jiang","Luc Van Gool","Danda Pani Paudel"]}},"version":2},{"content":{"venue":{"value":"ICANN (3) 2024"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-031-72338-4_14.pdf"},"venueid":{"value":"dblp.org/conf/ICANN/2024"},"paperhash":{"value":"yu|boundaryaware_noiseresistant_video_moment_retrieval"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Fengzhen_Yu:","~Xiaodong_Gu2"]},"html":{"value":"https://doi.org/10.1007/978-3-031-72338-4_14"},"_bibtex":{"value":"@inproceedings{DBLP:conf/icann/YuG24,\n  author={Fengzhen Yu and Xiaodong Gu},\n  title={Boundary-Aware Noise-Resistant Video Moment Retrieval},\n  year={2024},\n  cdate={1704067200000},\n  pages={193-206},\n  url={https://doi.org/10.1007/978-3-031-72338-4_14},\n  booktitle={ICANN (3)},\n  crossref={conf/icann/2024-3}\n}\n"},"abstract":{"value":"Video Moment Retrieval (VMR) is a critical task that aims to retrieve relevant video segments based on textual queries from untrimmed videos. This paper introduces a Boundary-aware variant of the 2D Temporal Adjacency Network (BA-TAN) to address VMR challenges. The proposed model leverages cross-modal interactions and emphasizes video moment boundary in the temporal network to improve retrieval accuracy. Attention mechanisms are integrated to suppress background interference and highlight query-relevant video features. Additionally, we introduce a Time Proximity Mask Convolution to learn video features that are adjacent to the current time segment. This enhances the model’s ability to capture temporal context and improves moment retrieval accuracy. Experimental results demonstrate the effectiveness of our approach BA-TAN in enhancing VMR performance. This work provides valuable insights for further research and applications in related fields."},"title":{"value":"Boundary-Aware Noise-Resistant Video Moment Retrieval"},"authors":{"value":["Fengzhen Yu","Xiaodong Gu"]}},"tmdate":1769329983977,"pdate":1735603200000,"externalIds":["dblp:conf/icann/YuG24"],"tcdate":1769329965808,"writers":["~"],"signatures":["~Xiaodong_Gu2"],"forum":"BPYrpXFGCr","license":"CC BY-SA 4.0","number":795777,"cdate":1704067200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1769329983977,"domain":"DBLP.org","id":"BPYrpXFGCr","version":2},{"content":{"summary":{"value":"This paper presents ST-SimDiff, a balanced framework that explicitly models the trade-off between spatial-temporal similarity and difference in video representations. The method aims to improve video understanding by aligning similar semantics while preserving meaningful diversity across frames. The authors introduce novel similarity–difference guided modules that can be integrated into existing video backbones with minimal modifications. Extensive experiments on standard benchmarks demonstrate consistent improvements over strong baselines, showing superior performance across multiple tasks with manageable computational overhead."},"soundness":{"value":3},"confidence":{"value":5},"questions":{"value":"Can the authors provide qualitative visualization results, especially highlighting where the similarity–difference mechanism fails or produces ambiguous interpretations?\nWhere are the performance boundaries of this framework? Are there specific video tasks/types that ST-SimDiff is ill-equipped to handle?\nHow does the framework solve the input misalignment problem after pruning?\nIf the model makes an incorrect judgment on a video, how can you attribute the failure to the similarity module versus the difference module?"},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"This well-executed, clearly written paper addresses the critical issue of balancing temporal consistency and diversity in video representation learning. The ST-SimDiff framework is elegant, conceptually sound, and easily integrates into existing architectures, delivering a significant performance boost without the need for additional supervision or costly retraining.\nThe paper effectively highlights how previous methods either overemphasize similarity or difference, while ST-SimDiff strikes a balanced, principled approach. The motivation is compelling, the technical formulation is clear, and the presentation flows smoothly.\nThe design of interpretable components effectively clarifies how spatial-temporal cues interact. The experimental validation is comprehensive, covering diverse datasets, multiple backbones, and various evaluation metrics, with ablations that isolate the contribution of each component. Results are strong and consistent."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"The paper lacks crucial qualitative visualizations. It does not show failure cases or ambiguous scenarios, making it difficult to understand the mechanism's practical limitations or how it behaves when it produces an incorrect interpretation.\nThe framework's performance boundaries are not explored. The paper fails to investigate video types or specific tasks where the proposed balancing paradigm might be suboptimal or ill-equipped.\nThe paper does not address the critical misalignment problem that arises from sequence pruning. While pruning/dropping is a common efficiency method, it creates a gap between the dense data the model was trained on and the sparse data it receives at inference, for which the model has no explicit handling mechanism.\nThe paper lacks a method for error attribution, making it difficult to determine whether a specific failure originates from the similarity module, the difference module, or their interaction."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762943158255,"tcdate":1761834811761,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission24671/Reviewer_f9gs"],"signatures":["ICLR.cc/2026/Conference/Submission24671/Reviewer_f9gs"],"forum":"he8kYNcoMA","number":3,"license":"CC BY 4.0","cdate":1761834811761,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission24671/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762943158255,"domain":"ICLR.cc/2026/Conference","replyto":"he8kYNcoMA","id":"phVmksTTrH","forumContent":{"TLDR":{"value":"We propose a training-free method that builds a spatio-temporal graph to efficiently select video tokens by incorporating similarity for redundancy reduction and difference for key event detection."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Video Understanding","Visual Token Reduction","Multimodal Large Language Models"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Multimodal Large Language Models (MLLMs) face significant computational overhead when processing long videos due to the massive number of visual tokens required. To improve efficiency, existing methods primarily reduce redundancy by pruning or merging tokens based on importance or similarity. However, these approaches largely overlook a critical dimension of video content, i.e., changes and turning points, and they lack a collaborative model for spatio-temporal relationships.\nTo address this, we propose a new perspective: similarity is for identifying redundancy, while difference is for capturing key events. Based on this, we designed a training-free framework named ST-SimDiff. We first construct a spatio-temporal graph from the visual tokens to uniformly model their complex associations. Subsequently, we employ a parallel dual-selection strategy: 1) similarity-based selection uses community detection to retain representative tokens, compressing static information; 2) temporal difference-based selection precisely locates content-changing points to preserve tokens that capture key dynamic shifts. This allows it to preserve both static and dynamic content with a minimal number of tokens. Extensive experiments show our method significantly outperforms state-of-the-art approaches while substantially reducing computational costs.\nOur code is available in [https://github.com/bingjunluo/ST-SimDiff](https://github.com/bingjunluo/ST-SimDiff)."},"_bibtex":{"value":"@inproceedings{\nluo2026stsimdiff,\ntitle={{ST}-SimDiff: Balancing Spatiotemporal Similarity and Difference for Efficient Video Understanding with {MLLM}s},\nauthor={Bingjun Luo and Tony Wang and Chaoqi Chen and Xinpeng Ding},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=he8kYNcoMA}\n}"},"title":{"value":"ST-SimDiff: Balancing Spatiotemporal Similarity and Difference for Efficient Video Understanding with MLLMs"},"pdf":{"value":"/pdf/b697cea955353f3b35057bb9055738474d679568.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"luo|stsimdiff_balancing_spatiotemporal_similarity_and_difference_for_efficient_video_understanding_with_mllms"},"authorids":{"value":["~Bingjun_Luo2","~Tony_Wang4","~Chaoqi_Chen2","~Xinpeng_Ding1"]},"authors":{"value":["Bingjun Luo","Tony Wang","Chaoqi Chen","Xinpeng Ding"]}},"version":2},{"content":{"summary":{"value":"This paper introduces Animal-Bench, a video question answering benchmark focused on animals, which is usually overlooked in previous video benchmarks. Animal-Bench is sourced from six datasets and includes 13 tasks. Eight video-language models are evaluated on the benchmark and the results reveal shortcomings in the models. Moreover, the paper evaluates the robustness of the model by simulating weather and shooting parameter changes, which are challenging in real-world application."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. Is the accuracy drop in Table 2 relative drop or absolute drop?"},"rating":{"value":7},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. While the current video question answering benchmarks focus on human activities, Animal-Bench is the first animal-centric benchmark. It can not only boost the AI application in animal studies, but also help the development of video-language models.\n2. Animal-Bench aligns the benchmark with real-world applications by 1) including domain-specific tasks, e.g., breeding monitoring; 2) simulating realistic scenarios - weather and shooting parameter changes - by video editing.\n3. The paper thoroughly evaluates recent video-language models."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"### Missing dataset statistics:\n\nThe paper provides many detailed dataset statistics about the number of videos and the long-tail distribution of animal categories. But some important statistics are still missing:\n\n1. The number of questions on each video. If there can be multiple questions generated from one video, both the number of videos and number of questions should be listed in Section 3.1 and Table 3.\n2. The distribution of the video durations.\n\n### Answer correctness after video outpainting: \n\nIn the robustness evaluation, outpainting may introduce new animals to the video, e.g., row 2 column 2 in Figure 11. It may change the answer of the question on tasks like object existence and object count. As a result, the robustness evaluation could be inaccurate. It would be better to manually check a subset of the outpainted videos to see how well outpainting can preserve the correct answer.\n\n### Minor weaknesses:\n\nWhen a pre-trained model is applied to a specific downstream domain, it is natural to improve the performance by fine-tuning it on the downstream domain. However, this benchmark only provides a test set without training and validation sets. It would have greater impact if a training set is included and the fine-tuned model performance is evaluated.\n\n### Writing:\n\n1. It is better to explain the task abbreviations in Table 1 (either in the text or in the table caption).\n\n2. Showing the accuracy drop in Table 2 is a great way to demonstrate the model robustness. But it would be better to also list the absolute model accuracy."},"limitations":{"value":"Limitations are discussed in Appendix F. No potential negative societal impacts are mentioned."}},"nonreaders":[],"tmdate":1730879965089,"tcdate":1719417200684,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission18302/Reviewer_Mynx"],"signatures":["NeurIPS.cc/2024/Conference/Submission18302/Reviewer_Mynx"],"forum":"DexM7d1H6e","number":1,"license":"CC BY 4.0","cdate":1719417200684,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission18302/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730879965089,"domain":"NeurIPS.cc/2024/Conference","replyto":"DexM7d1H6e","id":"iE0tp4q1zD","forumContent":{"venue":{"value":"NeurIPS 2024 poster"},"keywords":{"value":["Multimodal video model","Evaluation benchmark","Robustness testing"]},"primary_area":{"value":"evaluation"},"abstract":{"value":"With the emergence of large pre-trained multimodal video models, multiple benchmarks have been proposed to evaluate model capabilities. However, most of the benchmarks are human-centric, with evaluation data and tasks centered around human applications. Animals are an integral part of the natural world, and animal-centric video understanding is crucial for animal welfare and conservation efforts. Yet, existing benchmarks overlook evaluations focused on animals, limiting the application of the models. To address this limitation, our work established an animal-centric benchmark, namely Animal-Bench, to allow for a comprehensive evaluation of model capabilities in real-world contexts, overcoming agent-bias in previous benchmarks. Animal-Bench includes 13 tasks encompassing both common tasks shared with humans and special tasks relevant to animal conservation, spanning 7 major animal categories and 819 species, comprising a total of 41,839 data entries. To generate this benchmark, we defined a task system centered on animals and proposed an automated pipeline for animal-centric data processing. To further validate the robustness of models against real-world challenges, we utilized a video editing approach to simulate realistic scenarios like weather changes and shooting parameters due to animal movements. We evaluated 8 current multimodal video models on our benchmark and found considerable room for improvement. We hope our work provides insights for the community and opens up new avenues for research in multimodal video models. Our data and code will be released at https://github.com/PRIS-CV/Animal-Bench."},"_bibtex":{"value":"@inproceedings{\njing2024animalbench,\ntitle={Animal-Bench: Benchmarking Multimodal Video Models for Animal-centric Video Understanding},\nauthor={Yinuo Jing and Ruxu Zhang and Kongming Liang and Yongxiang Li and Zhongjiang He and Zhanyu Ma and Jun Guo},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=DexM7d1H6e}\n}"},"title":{"value":"Animal-Bench: Benchmarking Multimodal Video Models for Animal-centric Video Understanding"},"pdf":{"value":"/pdf/1385b842d017357564468802fe0fd83c961c61a1.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"jing|animalbench_benchmarking_multimodal_video_models_for_animalcentric_video_understanding"},"authorids":{"value":["~Yinuo_Jing1","~Ruxu_Zhang2","~Kongming_Liang2","~Yongxiang_Li2","~Zhongjiang_He1","~Zhanyu_Ma1","~Jun_Guo1"]},"authors":{"value":["Yinuo Jing","Ruxu Zhang","Kongming Liang","Yongxiang Li","Zhongjiang He","Zhanyu Ma","Jun Guo"]}},"version":2},{"content":{"summary":{"value":"This paper proposes modeling the physical realism of image editing using video generation models. By leveraging the strong temporal, physical, and motion consistency capabilities of video generation models, the approach achieves impressive editing results. Furthermore, the introduction of a temporal reasoning token to simulate intermediate video frames is a very intuitive idea, and the experimental results are remarkable."},"soundness":{"value":4},"confidence":{"value":5},"questions":{"value":"See the weakness."},"rating":{"value":8},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. Utilizing video generation model to model the image editing task, which introducing great physical prior, achieving great results.\n2. The Temporal Reasoning Token simulate the intermidiate step of video and it fills the gap of modeling intermediate changes in image editing, resulting in stronger interpretability.\n3. After distilling, the result is still great and the speed is nearly comparable with image editing model."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"I do not have many concerns regarding the content of the paper itself. However, I am curious about how the video generation based model would perform in multi-turn editing scenarios involving different user instructions, where each round would introduce an additional input frame, which is an interesting attempt for future development.\nAdditionally, I want to know the true performance of trained video generative model in video generation. This would provide insights into how demanding the requirements are for training video generation models to achieve high-quality image editing. Please show the result in video generation benchmark like VBench series."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762915654224,"tcdate":1760974075687,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission1003/Reviewer_kjZH"],"signatures":["ICLR.cc/2026/Conference/Submission1003/Reviewer_kjZH"],"forum":"MbMzoQ91Gk","number":1,"license":"CC BY 4.0","cdate":1760974075687,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission1003/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762915654224,"domain":"ICLR.cc/2026/Conference","replyto":"MbMzoQ91Gk","id":"mHDM9FKd1e","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["image editing","generative models"]},"supplementary_material":{"value":"/attachment/34a2821229a2f71de5d2675cd2cb374fc427d416.pdf"},"primary_area":{"value":"generative models"},"abstract":{"value":"Recent advances in large generative models have greatly enhanced both image editing and in-context image generation, yet a critical gap remains in ensuring physical consistency, where edited objects must remain coherent. This capability is especially vital for world simulation related tasks. In this paper, we present ChronoEdit, a framework that reframes image editing as a video generation problem. First, ChronoEdit treats the input and edited images as the first and last frames of a video, allowing it to leverage large pretrained video generative models that capture not only object appearance but also the implicit physics of motion and interaction through learned temporal consistency. Second, ChronoEdit introduces a temporal reasoning stage that explicitly performs editing at inference time. Under this setting, target frame is jointly denoised with reasoning tokens to imagine a plausible editing trajectory that constrains the solution space to physically viable transformations. The reasoning tokens are then dropped after a few steps to avoid the high computational cost of rendering a full video. To validate ChronoEdit, we introduce PBench-Edit, a new benchmark of image–prompt pairs for contexts that require physical consistency, and demonstrate that ChronoEdit surpasses state-of-the-art baselines in both visual fidelity and physical plausibility. Project page for code and models: https://research.nvidia.com/labs/toronto-ai/chronoedit"},"_bibtex":{"value":"@inproceedings{\nwu2026chronoedit,\ntitle={ChronoEdit: Towards Temporal Reasoning for In-Context Image Editing and World Simulation},\nauthor={Jay Zhangjie Wu and Xuanchi Ren and Tianchang Shen and Tianshi Cao and Kai He and Yifan Lu and Ruiyuan Gao and Enze Xie and Shiyi Lan and Jose M. Alvarez and Jun Gao and Sanja Fidler and Zian Wang and Huan Ling},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=MbMzoQ91Gk}\n}"},"title":{"value":"ChronoEdit: Towards Temporal Reasoning for In-Context Image Editing and World Simulation"},"pdf":{"value":"/pdf/e1212649c2d96c4e59c526a148e9c4dbc8b8a959.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"wu|chronoedit_towards_temporal_reasoning_for_incontext_image_editing_and_world_simulation"},"authorids":{"value":["~Jay_Zhangjie_Wu1","~Xuanchi_Ren1","~Tianchang_Shen1","~Tianshi_Cao1","~Kai_He5","~Yifan_Lu1","~Ruiyuan_Gao2","~Enze_Xie1","~Shiyi_Lan3","~Jose_M._Alvarez2","~Jun_Gao3","~Sanja_Fidler1","~Zian_Wang1","~Huan_Ling1"]},"authors":{"value":["Jay Zhangjie Wu","Xuanchi Ren","Tianchang Shen","Tianshi Cao","Kai He","Yifan Lu","Ruiyuan Gao","Enze Xie","Shiyi Lan","Jose M. Alvarez","Jun Gao","Sanja Fidler","Zian Wang","Huan Ling"]}},"version":2},{"content":{"venue":{"value":"ICLR 2027 Conference Submission"},"keywords":{"value":["Empirical Neural Tangent Kernel","eNTK","shortcut learning"]},"supplementary_material":{"value":"/attachment/20057892c54319a42ca04d8f1ffd46407d321431.zip"},"primary_area":{"value":"learning theory"},"abstract":{"value":"Deep neural networks frequently rely on easily accessible spurious correlations, driving a phenomenon known as shortcut learning. Remarkably, this reliance remains invisible on standard loss curves, which decrease smoothly as models build brittle representations. In this work, we uncover the mathematical root of this invisibility by analyzing optimization dynamics via the empirical Neural Tangent Kernel (eNTK) along the training trajectory. We establish the non-asymptotic **Dynamic Starvation Theorem**, proving that a dominant shortcut eigenvalue restricts causal parameter updates to a bounded regime $O(1/\\sigma_f)$ in the worst case of unbounded shortcut dominance. We further characterize how this starvation mechanism manifests under different loss functions: under **squared loss**, joint and ablated training yield identical predictions in output space, trapping the causal subnetwork in a degenerate null space as loss vanishes; under **cross-entropy loss**, self-correction is merely logarithmic. To resolve these failure modes, we propose **Farsighted Signal Optimization** (FSO), which reinjects counterfactual causal gradients into the training objective. We extend this to model-agnostic settings via GAP-FSO, proving that inter-group confidence gaps serve as exact proxies for shortcut parameters without explicit architectural splits. Unlike prior theoretical treatments of shortcut learning, which validate their predictions on simple architectures, we test GAP-FSO on ResNet-50 and DistilBERT across six vision and language benchmarks. Results are competitive with established baselines overall and clearly outperform them on the NLP benchmark."},"_bibtex":{"value":"@inproceedings{\nanonymous2026farsighted,\ntitle={Farsighted Correction of Shortcut Learning: From Dynamic Starvation to Confidence-Gap Optimization},\nauthor={Anonymous},\nbooktitle={Submitted to The Fifteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=JrKJaRbQpz},\nnote={under review}\n}"},"title":{"value":"Farsighted Correction of Shortcut Learning: From Dynamic Starvation to Confidence-Gap Optimization"},"pdf":{"value":"/pdf/3aac61fcba08525d869fcb3c0477a576ce8d300f.pdf"},"venueid":{"value":"ICLR.cc/2027/Conference/Submission"}},"tmdate":1791225960744,"tcdate":1789053326612,"writers":["ICLR.cc/2027/Conference","ICLR.cc/2027/Conference/Submission13342/Authors"],"signatures":["ICLR.cc/2027/Conference/Submission13342/Authors"],"forum":"JrKJaRbQpz","license":"CC BY 4.0","number":13342,"cdate":1789053326612,"readers":["everyone"],"invitations":["ICLR.cc/2027/Conference/-/Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Bidding","ICLR.cc/2027/Conference/Submission13342/-/Full_Submission","ICLR.cc/2027/Conference/-/Submission_Change_Before_Reviewing","ICLR.cc/2027/Conference/-/Edit"],"mdate":1791225960744,"odate":1791031328286,"domain":"ICLR.cc/2027/Conference","id":"JrKJaRbQpz","version":2},{"content":{"venue":{"value":"ICLR 2025 Conference Withdrawn Submission"},"code_of_ethics":{"value":"I acknowledge that I and all co-authors of this work have read and commit to adhering to the ICLR Code of Ethics."},"keywords":{"value":["fMRI; fMRI-to-Video; Dynamic-Aware; Video Reconstruction from Brain Activities;"]},"primary_area":{"value":"applications to neuroscience & cognitive science"},"reciprocal_reviewing":{"value":"I understand the reciprocal reviewing requirement as described on https://iclr.cc/Conferences/2025/CallForPapers. If none of the authors are registered as a reviewer, it may result in a desk rejection at the discretion of the program chairs. To request an exception, please complete this form at https://forms.gle/Huojr6VjkFxiQsUp6."},"submission_guidelines":{"value":"I certify that this submission complies with the submission instructions as described on https://iclr.cc/Conferences/2025/AuthorGuide."},"abstract":{"value":"Existing methods for fMRI-to-video reconstruction typically focus on accurately reconstructing visual content ($i.e.$, appearance), neglecting dynamic event information. However, as highlighted in cognitive neurology, these key dynamic events significantly influence brain signal changes during video perception. In this article, we introduce Mindgrapher, a two-stream framework designed to address this gap by enhancing the reconstruction of dynamic-aware videos from fMRI data. Mindgrapher comprises $i)$ a visual content reconstruction stream, that improves the accuracy of the reconstructed visual content from sparsely distributed fMRI data through a temporal dynamics enrichment approach and multi-moment multimodal contrastive learning; $ii)$ a dynamics injection stream, that firstly crafts dynamic-aware fMRI embeddings and then integrates them into the reconstruction process via a fine-grained approach, thereby producing videos that effectively perceive dynamic events. Moreover, to address the lack of suitable metrics for evaluating dynamic event information, we introduce a new evaluation metric named dynamic content fidelity (DCF), which measures how accurately dynamic events within the video are reconstructed. Upon evaluation with a publicly available fMRI dataset, Mindgrapher outperforms the state-of-the-arts on all metrics, $i.e.$, semantic classification accuracy, structural similarity index, and DCF. The reconstructed video results are available on our web page. Code shall be released."},"_bibtex":{"value":"@misc{\nquan2024mindgrapher,\ntitle={MindGrapher: Dynamic-Aware f{MRI}-to-Video Reconstruction},\nauthor={Ruijie Quan and Wensong Song and Liulei Li and Wenguan Wang and Yi Yang},\nyear={2024},\nurl={https://openreview.net/forum?id=JfKF7Pdigi}\n}"},"title":{"value":"MindGrapher: Dynamic-Aware fMRI-to-Video Reconstruction"},"pdf":{"value":"/pdf/4ddf08c68ba61444352b3a6245d4f2283624e38a.pdf"},"venueid":{"value":"ICLR.cc/2025/Conference/Withdrawn_Submission"},"paperhash":{"value":"quan|mindgrapher_dynamicaware_fmritovideo_reconstruction"},"authorids":{"value":["~Ruijie_Quan1","~Wensong_Song1","~Liulei_Li1","~Wenguan_Wang4","~Yi_Yang4"]},"anonymous_url":{"value":"I certify that there is no URL (e.g., github page) that could be used to find authors’ identity."},"no_acknowledgement_section":{"value":"I certify that there is no acknowledgement section in this submission for double blind review."},"authors":{"value":["Ruijie Quan","Wensong Song","Liulei Li","Wenguan Wang","Yi Yang"]}},"tmdate":1731469650138,"tcdate":1727163907659,"writers":["ICLR.cc/2025/Conference","ICLR.cc/2025/Conference/Submission3510/Authors"],"signatures":["ICLR.cc/2025/Conference/Submission3510/Authors"],"forum":"JfKF7Pdigi","license":"CC BY 4.0","number":3510,"cdate":1727163907659,"readers":["everyone"],"invitations":["ICLR.cc/2025/Conference/-/Submission","ICLR.cc/2025/Conference/-/Post_Submission","ICLR.cc/2025/Conference/Submission3510/-/Full_Submission","ICLR.cc/2025/Conference/-/Withdrawn_Submission"],"mdate":1731469650138,"odate":1728008565725,"domain":"ICLR.cc/2025/Conference","id":"JfKF7Pdigi","version":2},{"content":{"summary":{"value":"This paper proposes MIMIC, a novel two-stage image-to-video diffusion framework designed to generate realistic and controllable manipulation videos. The method first utilizes an Interaction-Motion-Aware (IMA) module to generate a sequence of semantic interaction masks guided by a reference video. Subsequently, it introduces a Pair Prompt Control mechanism that conditions the final video generation on both the predicted masks and the original reference video, effectively disentangling object motion from camera motion and enhancing controllability."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. The experiments only show in-domain cases. Have the authors tested cross-domain generalization? For example, can a video of a human hand be used to guide a robot arm?\n\n2. Figure 6 shows that lacking the IMA module leads to failure. With the IMA module, how well does it work if the reference motion is complex or the object is occluded?\n\n3. The paper mentions that the input of Stage I includes text, but it appears to be unused in the architecture (Figure 2) and the IMA module description. Could the authors clarify how exactly the text is used in Stage I?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"1. The authors' decomposition of the task into two stages, \"mask generation\" and \"video rendering,\" is a reasonable design. Furthermore, the authors demonstrate the necessity of this decomposition through ablation studies (such as \"Two-Stage vs. One-Stage\" and Figure 6) and clearly show the contributions of each key module (IMA and PPC).\n\n2. The authors identified the problem with using only masks for control. The Pair Prompt Control (PPC) part is a clever solution. Figure 6 (background drift vs. static background) shows how this part is effective at separating object motion from camera motion."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The IMA module just copies the motion from the reference video and does not understand the physics of the target scene. The predicted masks in Figure 6 support this, appearing to be a direct motion transfer that ignores the specific geometry and constraints of the target object.\n\n2. The examples in the paper are all very similar, showing only robot-to-robot or hand-to-hand tasks. This makes it unclear if the method can work for more general cases where the reference and target are different. For instance, it is not shown if a human hand video can be used to control a robot arm.\n\n3. The method's first stage requires reference manipulation masks as input, but the paper does not explain how these are obtained. Class-agnostic models like SAM are difficult to automatically and accurately prompt to segment only the manipulated object in complex videos. Meanwhile, class-specific semantic segmentation models (e.g., Mask R-CNN) would restrict the method to only predefined object categories, causing it to fail on any new objects."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762919749365,"tcdate":1761985529705,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission7691/Reviewer_xAXU"],"signatures":["ICLR.cc/2026/Conference/Submission7691/Reviewer_xAXU"],"forum":"COrUdVuInH","number":3,"license":"CC BY 4.0","cdate":1761985529705,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission7691/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762919749365,"domain":"ICLR.cc/2026/Conference","replyto":"COrUdVuInH","id":"FxXkwEUGqQ","forumContent":{"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["video diffusion model","manipulation video"]},"supplementary_material":{"value":"/attachment/b6d9b8cb3acbab21f3cbce260e6148e91bc5581b.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Embodied intelligence faces a fundamental bottleneck from limited large-scale interaction data. Video generation offers a scalable alternative, but manipulation videos remain particularly challenging, as they require capturing subtle, contact-rich dynamics. Despite recent advances, video diffusion models still struggle to balance semantic understanding with fine-grained visual details, restricting their effectiveness in manipulation scenarios. Our key insight is that reference videos provide rich semantic and motion cues that can effectively drive manipulation video generation. Building on this, we propose MIMIC, a two-stage image-to-video diffusion framework. (1) We first introduce an Interaction-Motion-Aware (IMA) module to fuse visual features from the reference video, producing coherent semantic masks that correspond to the target image. (2) then utilize these masks as semantic control signals to guide the video generation process. Moreover, considering the ambiguity of the motion attribution,  we introduce a Pair Prompt Control mechanism to disentangle object and camera motion by adding the reference video as an additional input. Extensive experiments demonstrate that MIMIC significantly outperforms existing methods, effectively preserves manipulation intent and motion details, even when handling diverse and deformable objects. Our findings underscore the effectiveness of reference-driven semantics for controllable and realistic manipulation video generation."},"_bibtex":{"value":"@inproceedings{\nchen2026mimic,\ntitle={{MIMIC}: Mask-Injected Manipulation Video Generation with Interaction Control},\nauthor={Tianxiao Chen and Jintao Rong and Huajin Chen and Jingya Wang and Tao Zhou and Jiming Chen and Qi Ye},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=COrUdVuInH}\n}"},"title":{"value":"MIMIC: Mask-Injected Manipulation Video Generation with Interaction Control"},"pdf":{"value":"/pdf/84bc45b7fa57b5fa895a891bd9cf268ab48a7972.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"chen|mimic_maskinjected_manipulation_video_generation_with_interaction_control"},"authorids":{"value":["~Tianxiao_Chen1","~Jintao_Rong2","~Huajin_Chen1","~Jingya_Wang3","~Tao_Zhou8","~Jiming_Chen1","~Qi_Ye2"]},"authors":{"value":["Tianxiao Chen","Jintao Rong","Huajin Chen","Jingya Wang","Tao Zhou","Jiming Chen","Qi Ye"]}},"version":2},{"content":{"venue":{"value":"LORI 2013"},"pdf":{"value":"https://link.springer.com/content/pdf/10.1007/978-3-642-40948-6_2.pdf"},"venueid":{"value":"dblp.org/conf/LORI/2013"},"paperhash":{"value":"alechina|minimal_preference_change"},"authorids":{"value":["~Natasha_Alechina1","~Fenrong_Liu1","~Brian_Logan1"]},"html":{"value":"https://doi.org/10.1007/978-3-642-40948-6_2"},"_bibtex":{"value":"@inproceedings{DBLP:conf/lori/AlechinaLL13,\n  author={Natasha Alechina and Fenrong Liu and Brian Logan},\n  title={Minimal Preference Change},\n  year={2013},\n  cdate={1356998400000},\n  pages={15-26},\n  url={https://doi.org/10.1007/978-3-642-40948-6_2},\n  booktitle={LORI},\n  crossref={conf/lori/2013}\n}\n"},"abstract":{"value":"We propose a novel approach to preference change. We treat a set of preferences as a special kind of theory, and define minimal change contraction and revision operations in the spirit of minimal change as advocated by the Alchourron, Gardenfors, and Makinson (AGM) theory of belief revision. We characterise minimal contraction of preference sets by a set of postulates and prove a representation theorem. We also give a linear time algorithm which implements minimal contraction by a single preference. We also define minimal contraction by a set of preferences, and for a significant special case state postulates, prove a representation theorem, and provide an efficient algorithm implementing minimal contraction by a set of preferences."},"title":{"value":"Minimal Preference Change"},"authors":{"value":["Natasha Alechina","Fenrong Liu","Brian Logan"]}},"tmdate":1789969453131,"pdate":1356998400000,"tcdate":1722926881346,"writers":["~"],"signatures":["~Natasha_Alechina1"],"forum":"5QK7zPPY6t","license":"CC BY-SA 4.0","number":51932,"cdate":1356998400000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1789969453131,"domain":"DBLP.org","id":"5QK7zPPY6t","version":2},{"content":{"summary":{"value":"The paper introduces GenCompositor, a diffusion-transformer-based framework for generative video compositing, enabling insertion, removal, and harmonization of dynamic foreground objects within videos. The method takes as input a masked background video, a mask video, and a foreground video, along with user-defined trajectory and scale controls. The model architecture follows a  MMDiT, similar to those used in FLUX, SD3, or HunyuanVideo, with full self-attention, and includes:\n- A background preservation branch to maintain scene consistency,\n- A DiT fusion block, which concatenates background and foreground tokens for fusion instead of using cross-attention, and\n- An Extended Rotary Position Embedding (ERoPE) to mitigate layout misalignment artifacts.\n\nFurthermore, a new dataset, VideoComp, containing 61K video triplets, is curated for training. The qualitative results are visually strong and demonstrate convincing object insertion and harmonization, though the quantitative metrics show only modest improvements over existing methods."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"- The results without the fusion block are much worse. Could this be due to the limited representational power of the VAE encoder? I would like to see additional results where a stronger feature extractor, such as DINO or CLIP, is used for the cross-attention variant, as this could provide more meaningful semantic conditioning.\n- The model depends strongly on mask and trajectory inputs. How flexible is it when these are imprecise or unavailable? Can it generalize to more unconstrained compositional setups?\n- How well does it handle multiple objects or occlusion? \n- The paper mentions luminance augmentation for harmonization. Is it sufficient for handling complex illumination changes, or does it fail in extreme lighting conditions?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- **Strong engineering effort:** The system is well-implemented and technically solid. The architecture is clean and the pipeline is carefully designed.\n- **Good visual quality:** Qualitative and video results are impressive, showing smooth motion and coherent integration of foreground and background.\n- **Clear writing and presentation:** The paper is well-organized, easy to follow, and supported by clear figures and a detailed video presentation.\n- **Practical relevance:** The task aligns well with real-world video editing workflows and could be useful for production or creative pipelines."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. **Limited conceptual novelty**: \n    - The core components are incremental modifications of existing DiT architectures.\n    - The background preservation branch resembles ControlNet-style conditioning, which is also stated in the paper.\n    - ERoPE is a simple positional embedding shift; while useful, it is not conceptually innovative.\nOverall, the work feels like a strong system-level integration rather than a conceptual contribution.\n\n2. **Inadequate quantitative evaluation**:\n    - The paper reports only frame-level metrics (PSNR, SSIM, LPIPS, CLIP). These are not sufficient for video generation; video-level metrics such as FVD or KVD are missing.\n    - Quantitative improvements are relatively small, making it difficult to assess significance.\n\n3. **Weak baseline comparisons**:\n    - The compared baselines (Harmonizer, VideoTripletTransformer) are both video harmonization methods, not true generative compositing or conditional video generation systems.\n    - DynVFX can be used for comparisons. Although it cannot add directly the reference objects, it can add objects through text-prompts, which will also strengthen the paper's results.\n\n4. **Dataset concerns**:\n    - The construction of the new VideoComp dataset is clearly described, combining 409K source videos into 61K compositing triplets via a semi-automatic pipeline (SAM2 + CogVLM + Qwen). While this is a solid contribution, the paper provides limited analysis of dataset characteristics such as category diversity, motion distribution, or domain coverage, which limits understanding of how well the dataset covers diverse motion and scene types. Including a brief quantitative summary would strengthen this part.\n\n5. **Conceptual framing and scope**:\n    - The paper frames generative video compositing as a new task, but it heavily overlaps with existing controllable video generation and video inpainting setups."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762915818741,"tcdate":1761381406458,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission1570/Reviewer_irsQ"],"signatures":["ICLR.cc/2026/Conference/Submission1570/Reviewer_irsQ"],"forum":"ynim5u2N4i","number":2,"license":"CC BY 4.0","cdate":1761381406458,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission1570/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762915818741,"domain":"ICLR.cc/2026/Conference","replyto":"ynim5u2N4i","id":"vSIcUHzLZo","forumContent":{"TLDR":{"value":"GenCompositor is capable of effortlessly compositing different videos guided by user-specified trajectories and scales."},"venue":{"value":"ICLR 2026 Poster"},"keywords":{"value":["Diffusion Models","Video Editing","Video Compositing"]},"supplementary_material":{"value":"/attachment/27f18c510c51d9af8b342e392cb17ce1d09641cc.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Video compositing combines live-action footage to create video production, serving as a crucial technique in video creation and film production. Traditional pipelines require intensive labor efforts and expert collaboration, resulting in lengthy production cycles and high manpower costs. To address this issue, we automate this process with generative models, called generative video compositing. This new task strives to adaptively inject identity and motion information of foreground video to the target video in an interactive manner, allowing users to customize the size, motion trajectory, and other attributes of the dynamic elements added in final video. Specifically, we designed a novel Diffusion Transformer (DiT) pipeline based on its intrinsic properties. To maintain consistency of the target video before and after editing, we revised a light-weight DiT-based background preservation branch with masked token injection. As to inherit dynamic elements from other sources, a DiT fusion block is proposed using full self-attention, along with a simple yet effective foreground augmentation for training. Besides, for fusing background and foreground videos with different layouts based on user control, we developed a novel position embedding, named Extended Rotary Position Embedding (ERoPE). Finally, we curated a dataset comprising 61K sets of videos for our new task, called VideoComp. This data includes complete dynamic elements and high-quality target videos. Experiments demonstrate that our method effectively realizes generative video compositing, outperforming existing possible solutions in fidelity and consistency. Project is available at https://gencompositor.github.io/"},"_bibtex":{"value":"@inproceedings{\nyang2026gencompositor,\ntitle={GenCompositor: Generative Video Compositing with Diffusion Transformer},\nauthor={Shuzhou Yang and Xiaoyu Li and Xiaodong Cun and Guangzhi Wang and Lingen Li and Ying Shan and Jian Zhang},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=ynim5u2N4i}\n}"},"title":{"value":"GenCompositor: Generative Video Compositing with Diffusion Transformer"},"pdf":{"value":"/pdf/cfd1ebc349fc91f9a0fce0b4ac6a434601465415.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"yang|gencompositor_generative_video_compositing_with_diffusion_transformer"},"authorids":{"value":["~Shuzhou_Yang1","~Xiaoyu_Li2","~Xiaodong_Cun1","~Guangzhi_Wang1","~Lingen_Li1","~Ying_Shan2","~Jian_Zhang22"]},"authors":{"value":["Shuzhou Yang","Xiaoyu Li","Xiaodong Cun","Guangzhi Wang","Lingen Li","Ying Shan","Jian Zhang"]}},"version":2},{"content":{"venue":{"value":"IEEE Trans. Circuits Syst. Video Technol. 2017"},"pdf":{"value":"https://ieeexplore.ieee.org/iel7/76/7891656/07373614.pdf"},"venueid":{"value":"dblp.org/journals/TCSV/2017"},"paperhash":{"value":"li|realtime_featurebased_video_stabilization_on_fpga"},"authorids":{"value":["~Jianan_Li1","https://dblp.org/search/pid/api?q=author:Tingfa_Xu:","https://dblp.org/search/pid/api?q=author:Kun_Zhang:"]},"html":{"value":"https://doi.org/10.1109/TCSVT.2016.2515238"},"_bibtex":{"value":"@article{DBLP:journals/tcsv/LiXZ17,\n  author={Jianan Li and Tingfa Xu and Kun Zhang},\n  title={Real-Time Feature-Based Video Stabilization on FPGA},\n  year={2017},\n  cdate={1483228800000},\n  journal={IEEE Trans. Circuits Syst. Video Technol.},\n  volume={27},\n  number={4},\n  pages={907-919},\n  url={https://doi.org/10.1109/TCSVT.2016.2515238}\n}\n"},"abstract":{"value":"Digital video stabilization is an important video enhancement technology that aims to remove unwanted camera vibrations from video sequences. Trading off between stabilization performance and real-time hardware implementation feasibility, this paper presents a feature-based full-frame video stabilization method and a novel complete fully pipelined architectural design to implement it on field-programmable gate array (FPGA). In the proposed method, feature points are first extracted with the oriented features from accelerated segment test and rotated binary robust independent elementary features algorithm and matched between consecutive frames. Next, the matched point pairs are fitted to the affine transformation model using a random-sample consensus-based approach to estimate inter-frame motion robustly. Then, the estimated results are accumulated to compute the cumulative motion parameters between the current and reference frames, and the translational components are smoothed by a Kalman filter representing intentional camera movement. Finally, a mosaicked image is constructed based on cumulative motion parameters using an image mosaicking technique, and then a display window is created with the desired frame size according to the computed intentional camera movement to obtain a full motion-compensated frame. Using pipelining and parallel processing strategies, the whole process has been designed using a novel complete fully pipelined architecture and implemented on Altera's Cyclone III FPGA to build a real-time stabilization system. The experimental results have shown that the proposed system can deal with standard PAL video input including arbitrate translation and rotation and can produce full-frame stabilized output providing a better viewing experience at 22.37 ms/frame, thus achieving real-time processing performance."},"title":{"value":"Real-Time Feature-Based Video Stabilization on FPGA"},"authors":{"value":["Jianan Li","Tingfa Xu","Kun Zhang"]}},"tmdate":1731487428124,"pdate":1483228800000,"tcdate":1731484186819,"writers":["~"],"signatures":["~Jianan_Li1"],"forum":"udNAZhQxD5","license":"CC BY-SA 4.0","number":213905,"cdate":1483228800000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1731487428124,"domain":"DBLP.org","id":"udNAZhQxD5","version":2},{"content":{"summary":{"value":"- The paper proposes a model for video-to-audio generation (main task) and video-to-audio captioning (auxilliary task). \n- The text is optionally used to control audio generation for ambiguous video cases. \n- They use two step approach i.e. video-to-caption stage and video+text-to-audio stage\n- Stage 1: video-to-caption\n  - This stage basically converts videos and text to audio relevant captions\n  - They train a VATT converter that uses LoRA finetuning to get audio captions. \n  - This stage uses AudioSet and VGGSound and captions generated using LTU model.\n\n- Stage 2: video+text-to-audio \n  - This stage generates audio tokens, given video+text as input. The video+text is converted to audio relevant captions using stage-1\n  - They use LM based audio token generation. The VATT Audio decoder learns to predict the masked audio tokens. During inference, iterating parallel decoding is used.\n  - VGGSound is used to train this stage.\n\n- There model shows comparative and better performance on video-to-audio(V2A) generation task and text to audio generation task. Add text to V2A task further improves KLD score.\n- Stage 1 when compared to baselines can perform better on video-to-audio captioning.\n- Qualitative and human evaluation shows preference towards their work."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"My questions are based on the arguments mentioned in the weakness section:\n\n- Can the correctness of the LTU-generated captions be verified? Since they are used as the ground-truth for several tasks?\n\n- Can the baseline text-to-audio(TTA) generation models be trained on VGGSound+LTU captions? Can VATT be compared on AudioCaps?\n\n- Can VATT be compared on other video captioning datasets? Or the existing video-captioning models be finetuned/trained from scratch on this LTU data?\n\n- Why are some metrics different from the numbers quoted in the original paper? For eg. Diff-foley alignment score."},"rating":{"value":6},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":3},"strengths":{"value":"- The paper proposes a new problem of utilizing text for video-to-audio generation.\n- Several novel and interesting technical contributions. i.e. \n - Video+text conditioning to generate audio by converting this condition to joint audio-caption.\n - Using synthetic captions for training and LoRA for MLLM finetuning.\n - Using iterative parallel decoding and bi-directional self-attention. \n- Their approach improves on video-to-audio generation, and text-to-audio generation both quantitatively and qualitatively.\n- The results also show improvement in video-to-caption generation."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"- Text captions generated by LTU are used as ground truth for video-to-audio generation, text-to-audio generation, and video-to-audio captioning. The authors should verify the correctness of the captions. \n\n- Table 2 results for Text-to-audio generation seem slightly unfair for the baselines. \n      - The LTU-generated caption may have some biases or patterns of caption generation that VATT has seen during training while the \n         baselines have not.\n      - Though AudioGen has seen VGGSound tags, it has not seen the captions generated by LTU (which VATT has). \n      - AudioLDM2 has not even seen VGGSound data for training. \n      - A fair comparison would be finetuning/training from scratch Text-to-audio models on VGGSound with LTU-generated captions. Or VATT should be compared on AudioCaps (with videos downloaded from YouTube). \n\n- Table 3 results are also a little unfair:\n  - Similar to the argument above LLAVA has not seen LTU-generated captions for VGGSOUND.\n  - Also LLAVA takes only 1 frame while VATT takes more frames. \n  - Maybe a fair comparison would be to train some captioning models on LTU-generated data.\n\nMinor comments:\n- The captions for the tables should come above the table. For eg. Table 1,2,3\n- Missing reference: A previous work by Mo. et al [1] that utilizes text+video-to-audio generation. This should be mentioned \n- Align-acc for diff-foley is 82.47 in the paper vs 94.05 in the original paper\n\n[1] DiffAVA: Personalized text-to-audio generation with visual alignment"},"limitations":{"value":"Yes the authors adequately addressed the limitations."}},"nonreaders":[],"tmdate":1730878736509,"tcdate":1720852231723,"writers":["NeurIPS.cc/2024/Conference","NeurIPS.cc/2024/Conference/Submission1915/Reviewer_5yrj"],"signatures":["NeurIPS.cc/2024/Conference/Submission1915/Reviewer_5yrj"],"forum":"kr7eN85mIT","number":4,"license":"CC BY 4.0","cdate":1720852231723,"readers":["everyone"],"invitations":["NeurIPS.cc/2024/Conference/Submission1915/-/Official_Review","NeurIPS.cc/2024/Conference/-/Edit"],"mdate":1730878736509,"domain":"NeurIPS.cc/2024/Conference","replyto":"kr7eN85mIT","id":"OcZPuPcO6d","forumContent":{"TLDR":{"value":"A novel multi-modal generation framework for text guided video-to-audio generation and video-to-audio captioning."},"venue":{"value":"NeurIPS 2024 poster"},"keywords":{"value":["multi-modal learning","audio-visual learning","multi-modal large-language-model","text-guided video-to-audio generation","video-to-audio captioning"]},"primary_area":{"value":"speech_and_audio"},"flagged_for_ethics_review":{"value":true},"abstract":{"value":"The content of visual and audio scenes is multi-faceted such that a video stream can\nbe paired with various audio streams and vice-versa. Thereby, in video-to-audio\ngeneration task, it is imperative to introduce steering approaches for controlling the\ngenerated audio. While Video-to-Audio generation is a well-established generative\ntask, existing methods lack such controllability. In this work, we propose VATT, a\nmulti-modal generative framework that takes a video and an optional text prompt\nas input, and generates audio and optional textual description (caption) of the\naudio. Such a framework has two unique advantages: i) Video-to-Audio generation\nprocess can be refined and controlled via text which complements the context\nof the visual information, and ii) The model can suggest what audio to generate\nfor the video by generating audio captions. VATT consists of two key modules:\nVATT Converter, which is an LLM that has been fine-tuned for instructions and\nincludes a projection layer that maps video features to the LLM vector space, and\nVATT Audio, a bi-directional transformer that generates audio tokens from visual\nframes and from optional text prompt using iterative parallel decoding. The audio\ntokens and the text prompt are used by a pretrained neural codec to convert them\ninto a waveform. Our experiments show that when VATT is compared to existing\nvideo-to-audio generation methods in objective metrics, such as VGGSound audiovisual dataset, it achieves competitive performance when the audio caption is\nnot provided. When the audio caption is provided as a prompt, VATT achieves\neven more refined performance (with lowest KLD score of 1.41). Furthermore,\nsubjective studies asking participants to choose the most compatible generated\naudio for a given silent video, show that VATT Audio has been chosen on average\nas a preferred generated audio than the audio generated by existing methods. VATT\nenables controllable video-to-audio generation through text as well as suggesting\ntext prompts for videos through audio captions, unlocking novel applications such\nas text-guided video-to-audio generation and video-to-audio captioning."},"_bibtex":{"value":"@inproceedings{\nliu2024tell,\ntitle={Tell What You Hear From What You See - Video to Audio Generation Through Text},\nauthor={Xiulong Liu and Kun Su and Eli Shlizerman},\nbooktitle={The Thirty-eighth Annual Conference on Neural Information Processing Systems},\nyear={2024},\nurl={https://openreview.net/forum?id=kr7eN85mIT}\n}"},"title":{"value":"Tell What You Hear From What You See - Video to Audio Generation Through Text"},"pdf":{"value":"/pdf/7f048a863a36075771abbaa28445e76c2233bd97.pdf"},"venueid":{"value":"NeurIPS.cc/2024/Conference"},"paperhash":{"value":"liu|tell_what_you_hear_from_what_you_see_video_to_audio_generation_through_text"},"authorids":{"value":["~Xiulong_Liu1","~Kun_Su1","~Eli_Shlizerman1"]},"authors":{"value":["Xiulong Liu","Kun Su","Eli Shlizerman"]}},"version":2},{"content":{"summary":{"value":"The paper describes a dataset consisting of Question-Answer pairs about object-based tasks (such as counting, discances, movements, etc.) that were automatically generated from existing video dataset (TAO, ScanNet and KITTI). It also describes a reward design for RL-based training, based on the positions of objects as well as their positions relative to one another. It then shows that together, these two ingredients make it possible to fine-tune Qwen2.5-VL-7B to achieve relatively high recognition performance on several video understanding benchmarks (STI-bench, V-STaR, etc.). Ablations furthermore show that all the ingredients contribute positively to performance improvements on these benchmarks."},"soundness":{"value":3},"confidence":{"value":3},"questions":{"value":"After fine-tuning the model on graph-based spatio-temporal reasoning, how is it applied to the downstream benchmark tasks? Is there additional fine-tuning involved? Any task-specific (few-shot?) prompting?  \n\nThe work is reminiscent of older works on graph-based grounding, such as “Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations”, Krishna et al 2016, and many related and follow-up papers). It would be good to see a discussion on how it is related to this line of research. \n\n(small comment:) The frames in the figures (for example, Figure 1, Figure 2) are impossible to see properly or understand in a print-out version of the paper."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"The paper is quite easy to follow. It describes an (automatically generated) dataset, a training method and a model (based on fine-tuning Qwen2.5-VL-7B), all of which will be publicly released. The model yields fairly good performance on some video understanding benchmarks, specifically those that test for an understanding of object positions and their relations."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"I have the concern that the object-centric data and reward designs are benefitting specifically tasks that measure the ability to understand object positions and relations, without enhancing the model’s visual understanding overall. The results suggest that this is indeed the case, as model performance on tasks like STI-bench and V-STaR is comparably high, while this is not nearly as much the case for tasks like Video-MME (where the model is in fact significantly behind the state of the art). \n\nOverall, this paper describes an engineering effort, based on automatically generated data and RL-training with object-centric rewards, to mildly improve performance on video-understanding benchmarks (and most significantly on object-centric benchmarks). I am not quite sure how this advances our understanding of visual capabilities, models and model limitations overall. I am also missing a view of the limitations of the object-centric design. For example, are there any tasks where performance degrades rather than benefits from this design?"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762919960451,"tcdate":1761849400559,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission7938/Reviewer_q9jw"],"signatures":["ICLR.cc/2026/Conference/Submission7938/Reviewer_q9jw"],"forum":"D6v3B6oTDA","number":1,"license":"CC BY 4.0","cdate":1761849400559,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission7938/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762919960451,"domain":"ICLR.cc/2026/Conference","replyto":"D6v3B6oTDA","id":"WpfcZjmm3X","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Multi-modal large language odel","reinforcement learning with verifiable rewards","spatio-temporal reasoning","video understanding"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Recent progress in Multimodal Large Language Models (MLLMs) has demonstrated strong semantic understanding capabilities, but struggles to perform precise spatio-temporal understanding. Existing spatio-temporal methods primarily focus on the video itself, while overlooking the physical information within the video, such as multi-object layouts and motion. Such limitations restrict the use of MLLMs in downstream applications that demand high precision, including embodied intelligence and VR. To address this issue, we present Video-STR, a novel graph-based reinforcement method for precise Video Spatio-Temporal Reasoning. Building upon the capacity of Reinforcement Learning with Verifiable Reward (RLVR) to improve model abilities, we introduce a reasoning mechanism using graph representation based on Group Relative Policy Optimization (GRPO) to guide the model in inferring the underlying spatio-temporal topology of scenarios during the thinking process.To resolve the lack of spatio-temporal training data, we construct the STV-205k dataset with 205k question-answering pairs, covering dynamic multi-object scenes in both indoor and outdoor environments, to support the model training. Experiments show that Video-STR achieves state-of-the-art results on various benchmarks, outperforming the base model by 13% on STI-Bench, and demonstrating the effectiveness of our approach and dataset. Code, model, and data will be released."},"_bibtex":{"value":"@misc{\nwang2026videostr,\ntitle={Video-{STR}: Reinforcing {MLLM}s in Video Spatio-Temporal Reasoning with Relation Graph},\nauthor={Wentao Wang and Heqing Zou and Tianze Luo and Guiyang Xie and Rui Huang and Yutian Zhao and Zhuochen Wang and Hansheng Zhang and Chengwei Qin and Yan Wang and Lin Zhao and Zhang huaijian},\nyear={2026},\nurl={https://openreview.net/forum?id=D6v3B6oTDA}\n}"},"title":{"value":"Video-STR: Reinforcing MLLMs in Video Spatio-Temporal Reasoning with Relation Graph"},"pdf":{"value":"/pdf/44d5d0d35e4beff6b3afc46e66772eed521c69d1.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"wang|videostr_reinforcing_mllms_in_video_spatiotemporal_reasoning_with_relation_graph"},"authorids":{"value":["~Wentao_Wang9","~Heqing_Zou1","~Tianze_Luo1","~Guiyang_Xie1","~Rui_Huang23","~Yutian_Zhao3","~Zhuochen_Wang1","~Hansheng_Zhang1","~Chengwei_Qin1","~Yan_Wang12","~Lin_Zhao3","~Zhang_huaijian1"]},"authors":{"value":["Wentao Wang","Heqing Zou","Tianze Luo","Guiyang Xie","Rui Huang","Yutian Zhao","Zhuochen Wang","Hansheng Zhang","Chengwei Qin","Yan Wang","Lin Zhao","Zhang huaijian"]}},"version":2},{"content":{"summary":{"value":"This paper introduces EditVerse, a novel unified framework designed to perform both image and video generation and editing within a single model. The authors identify two primary bottlenecks in video editing: (1) architectural limitations of existing models, which are often task-specific, and (2) the scarcity of high-quality, instruction-based video editing data."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"1. Could the authors provide more details on the scaling properties of EditVerse? How do inference time and VRAM usage scale with (a) video duration (e.g., 3s vs. 10s vs. 30s) and (b) resolution (360p vs. 720p)? Is the full self-attention approach feasible beyond the short, low-resolution clips shown?\n2. The dimension allocations for the 4D RoPE (56H, 56W, 12Seq, 4Temp) are specific. Could the authors provide a brief rationale or ablation for this design choice? For example, why is the sequential dimension (12) given more embedding space than the temporal dimension (4)?\n3. How does the model perform on editing tasks or concepts that are *not* present in its synthetic \"teacher\" models? For instance, if VACE is poor at \"making an object transparent,\" can EditVerse still learn this concept purely from the image-editing data and apply it successfully to video, or is its video performance on this task limited by VACE?\n4. The \"wrong position\" failure case in Fig 9a suggests limitations in complex spatial reasoning. The VLM evaluation (Table 2) is high, but this metric averages 3 frames. How well does this VLM metric capture these kinds of high-level logical or instruction-misalignment failures, as opposed to per-frame artifacts or quality?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":4},"contribution":{"value":3},"strengths":{"value":"1. The core design of an interleaved 1D token sequence for mixed image/video data, combined with a full self-attention mechanism and the novel 4D RoPE, is a powerful and original solution for multimodal in-context learning.\n2. The paper's standout strength is its clear demonstration that a unified model can learn video editing from data-abundant *image* editing datasets. The ablation in Table 4 (showing the model *can* edit video with zero video-edit data, albeit poorly) and the major quality drop from removing image-edit data (Fig 8) provide powerful evidence for this hypothesis.\n3. The authors address the ecosystem problem by not only building a model but also creating the data (232K video-edit samples) and the evaluation tools (EditVerseBench) necessary to make progress, and they are releasing the benchmark."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The most significant weakness, which the authors acknowledge in the appendix, is the computational cost. Using full self-attention on a 1D-tokenized video (where the sequence length $L$ includes all frames) leads to $O(L^2)$ complexity. The reported 118 seconds for a single 360p video on an A100 is very slow and will likely scale quadratically or worse with resolution and duration, making it impractical for real-world use on high-resolution or long-form videos.\n2. The video editing data is synthetically generated by a pipeline of specialist \"teacher\" models (VACE, DiffuEraser, etc.). While this is a clever solution, it means the model's performance may be capped by the quality and biases of these teachers. It's unclear if the model can perform edits that its teacher models are incapable of."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1764359019292,"tcdate":1761666220812,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission8067/Reviewer_3VqN"],"signatures":["ICLR.cc/2026/Conference/Submission8067/Reviewer_3VqN"],"forum":"blJXE07r7I","number":2,"license":"CC BY 4.0","cdate":1761666220812,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission8067/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1764359019292,"domain":"ICLR.cc/2026/Conference","replyto":"blJXE07r7I","id":"rvcL0IKCyS","forumContent":{"venue":{"value":"ICLR 2026 Oral"},"keywords":{"value":["Video Editing","Content Generation","Artificial Intelligence"]},"supplementary_material":{"value":"/attachment/54f1a6cabd37b58af95d8fc7a1065d3ac273ac39.zip"},"primary_area":{"value":"generative models"},"abstract":{"value":"Recent advances in foundation models highlight a clear trend toward unification and scaling, showing emergent capabilities across diverse domains. While image generation and editing have rapidly transitioned from task-specific to unified frameworks, video generation and editing remain fragmented due to architectural limitations and data scarcity. In this work, we introduce EditVerse, a unified framework for image and video generation and editing within a single model. By representing all modalities, i.e., text, image, and video, as a unified token sequence, EditVerse leverages self-attention to achieve robust in-context learning, natural cross-modal knowledge transfer, and flexible handling of inputs and outputs with arbitrary resolutions and durations. To address the lack of video editing training data, we design a scalable data pipeline that curates 232K video editing samples and combines them with large-scale image and video datasets for joint training. Furthermore, we present EditVerseBench, the first benchmark for instruction-based video editing covering diverse tasks and resolutions. Extensive experiments and user studies demonstrate that EditVerse achieves state-of-the-art performance, surpassing existing open-source and commercial models, while exhibiting emergent editing and generation abilities across modalities."},"_bibtex":{"value":"@inproceedings{\nju2026editverse,\ntitle={EditVerse: Unifying Image and Video Editing and Generation with In-Context Learning},\nauthor={Xuan Ju and Tianyu Wang and Yuqian Zhou and He Zhang and Qing Liu and Nanxuan Zhao and Zhifei Zhang and Yijun Li and Yuanhao Cai and Shaoteng Liu and Daniil Pakhomov and Zhe Lin and Soo Ye Kim and Qiang Xu},\nbooktitle={The Fourteenth International Conference on Learning Representations},\nyear={2026},\nurl={https://openreview.net/forum?id=blJXE07r7I}\n}"},"title":{"value":"EditVerse: Unifying Image and Video Editing and Generation with In-Context Learning"},"pdf":{"value":"/pdf/b94811b83a0530da511e76edf48c05b2bbbd4725.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference"},"paperhash":{"value":"ju|editverse_unifying_image_and_video_editing_and_generation_with_incontext_learning"},"authorids":{"value":["~Xuan_Ju1","~Tianyu_Wang3","~Yuqian_Zhou2","~He_Zhang16","~Qing_Liu1","~Nanxuan_Zhao1","~Zhifei_Zhang2","~Yijun_Li2","~Yuanhao_Cai1","~Shaoteng_Liu1","~Daniil_Pakhomov2","~Zhe_Lin1","~Soo_Ye_Kim1","~Qiang_Xu1"]},"authors":{"value":["Xuan Ju","Tianyu Wang","Yuqian Zhou","He Zhang","Qing Liu","Nanxuan Zhao","Zhifei Zhang","Yijun Li","Yuanhao Cai","Shaoteng Liu","Daniil Pakhomov","Zhe Lin","Soo Ye Kim","Qiang Xu"]}},"version":2},{"content":{"summary":{"value":"The paper introduces Single-Frame Video set Distillation (SFVD), a method that compresses large video datasets into a small set of learnable key frames. Each distilled frame is interpolated into video sequences and refined using a Temporal Reshaping Network to retain temporal cues. Across diverse benchmarks, SFVD achieves consistent accuracy gains over prior video distillation methods while being far more efficient."},"soundness":{"value":3},"confidence":{"value":2},"questions":{"value":"1,2: See wealkness\n\n3. Following last question, I consider SFVD a more semantics driven method. It presents promising results even without temporal components. I am curious whether semantics are dominant factors (compared to motions) in video data distillation and whether we need more motion heavy benchmarks for this task? Look forward to authors' insight and discussions."},"rating":{"value":6},"details_of_ethics_concerns":{"value":"N/A"},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. Novel and simple formulation: SFVD identify the redundant information issues in video data and reformulates video distillation as a single-frame optimization problem.\n2. Strong empirical results with high efficiency: SFVD demonstrates sota performance across multiple benchmarks  with significantly lower computation and storage costs.\n3. Comprehensive ablations: Detailed ablations and cross-architecture experiments support SFVD design choices and show consistent improvements over prior baselines."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Motion capture: Although the proposed TRN can recover motion patterns, the distilled representation is still mainly derived from static single frames, which may fail to capture complex temporal dependencies such as long-range interactions or causal motion dynamics. Some analysis or discussion on whether motion is effectively captured should be provided.\n2. Assumption limitation: The SFVD build on the observation that video data usually have redundant information, which is reasonable. However, the assumption that one image can contain sufficient information within a video seem limited (by video duration and dynamics) and has no support. The ablation and main results are promising, however I am curious whether such method can be applied to longer video data or videos with more dramatic motions."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762925477843,"tcdate":1761951099539,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission15168/Reviewer_EBp7"],"signatures":["ICLR.cc/2026/Conference/Submission15168/Reviewer_EBp7"],"forum":"U4iaubcD7N","number":2,"license":"CC BY 4.0","cdate":1761951099539,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission15168/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762925477843,"domain":"ICLR.cc/2026/Conference","replyto":"U4iaubcD7N","id":"pwzjjIcRcr","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"keywords":{"value":["Dataset Distillation","Video Recognition"]},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Dataset distillation aims to synthesize compact yet informative datasets that allow models trained on them to achieve performance comparable to training on the full dataset. While this approach has shown promising results for image data, extending dataset distillation methods to video data has proven challenging and often leads to suboptimal performance. In this work, we first identify the core challenge in video set distillation as the substantial increase in learnable parameters introduced by the temporal dimension of video, which complicates optimization and hinders convergence. To address this issue, we observe that a single frame is often sufficient to capture the discriminative semantics of a video. Leveraging this insight, we propose Single-FrameVideo set Distillation (SFVD), a framework that distills videos into highly informative frames for each class. Our method focuses on distilling videos into highly informative frames for each class for effective optimization during distillation, a framework that distills videos into highly informative frames for each class. Using differentiable interpolation, these frames are transformed into video sequences and matched with the original dataset, while updates are restricted to the frames themselves for improved optimization efficiency. To further incorporate temporal information, the distilled frames are combined with sampled real videos from real videos during the matching process through a temporal reshaping network. Extensive experiments on multiple benchmarks demonstrate that SFVD substantially outperforms prior methods, achieving improvements of up to 5.3% on MiniUCF, thereby offering a more effective solution for video dataset distillation."},"_bibtex":{"value":"@misc{\nzhao2025distilling,\ntitle={Distilling Video Datasets into Images},\nauthor={Zhenghao Zhao and Haoxuan Wang and Kai Wang and Yuzhang Shang and Yuan Hong and Yan Yan},\nyear={2025},\nurl={https://openreview.net/forum?id=U4iaubcD7N}\n}"},"title":{"value":"Distilling Video Datasets into Images"},"pdf":{"value":"/pdf/55f73c960bf152b1bf596db27a6480c65292ee43.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"zhao|distilling_video_datasets_into_images"},"authorids":{"value":["~Zhenghao_Zhao1","~Haoxuan_Wang1","~Kai_Wang8","~Yuzhang_Shang1","~Yuan_Hong1","~Yan_Yan6"]},"authors":{"value":["Zhenghao Zhao","Haoxuan Wang","Kai Wang","Yuzhang Shang","Yuan Hong","Yan Yan"]}},"version":2},{"content":{"summary":{"value":"This paper proposes VideoAgent, an agent-based framework for automated video editing. By introducing a global-aware video shot creation mechanism and a self-reflective agent graph orchestration strategy, VideoAgent demonstrates promising results. Nevertheless, the paper still has several aspects that could be further improved."},"soundness":{"value":3},"confidence":{"value":4},"questions":{"value":"Please refer to Weaknesses."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. The paper focuses on the task of video editing and content creation, which holds significant practical value in real-world applications.\n\n2. The paper is well-written and easy to follow, with a comprehensive appendix that provides detailed explanations of the technical aspects of the proposed work."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. The definition and research scope of the task are not clearly articulated. Video editing is a highly broad concept, and the authors should explicitly specify which sub-tasks are covered by this work.\n\n2. In Section 2.3.1, the authors mention functionalities such as face swapping and lip synchronization, yet there appears to be no corresponding agent described in Appendix A.5.\n\n3. The paper lacks methodological novelty and sufficient contribution; the proposed system is largely built upon existing techniques and relies heavily on prompt engineering rather than introducing new algorithmic insights.\n\n4. The overall framework appears redundant and overly complicated. Constructing a dedicated dataset to train a more compact and unified model would likely be more effective.\n\n5. The paper makes extensive use of LLMs, but does not include a dedicated section “Usage of LLMs”."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762919017546,"tcdate":1761966546881,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission6731/Reviewer_oTRV"],"signatures":["ICLR.cc/2026/Conference/Submission6731/Reviewer_oTRV"],"forum":"cTqGsLYkRl","number":3,"license":"CC BY 4.0","cdate":1761966546881,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission6731/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762919017546,"domain":"ICLR.cc/2026/Conference","replyto":"cTqGsLYkRl","id":"mQisVrLNbk","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Multi-Agent Systems; Multimodal Content Editing; Agentic AI"]},"supplementary_material":{"value":"/attachment/026381dc502ed23824895d4016100d97ae3b3431.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Video editing has become essential in digital media creation, yet existing automated systems are restricted to short segment processing and domain-specific tasks. They face two critical limitations: i) inability to handle diverse video comprehension and editing operations, and ii) lack of long-video understanding for coherent narrative creation. We propose VideoAgent, an all-in-one agentic framework addressing these challenges through two key innovations. First, we develop automated video shot creation with shot planning agents for coherent narratives and cross-modal retrieval for aligned visual content. Second, we design a multi-agent orchestration framework integrating over thirty specialized editing agents. Intent parsing filters relevant tools while self-reflective graph orchestration assembles complex editing pipelines. Extensive experiments on our newly-proposed VideoEdit benchmark and public datasets demonstrate VideoAgent's superiority over existing multimodal LLMs and agentic systems. VideoAgent achieves 87-98% orchestration success rates while reducing API costs by 60%. Human evaluation across six video categories shows VideoAgent produces professional-quality content approaching human-level performance, with ratings only 4% below human-created videos."},"_bibtex":{"value":"@misc{\nzhou2026videoagent,\ntitle={VideoAgent: All-in-One Agentic Framework for Video Understanding and Editing},\nauthor={Hengji Zhou and Lingxuan Huang and KunpengTan and Si Wu and Lianghao Xia and Chao Huang},\nyear={2026},\nurl={https://openreview.net/forum?id=cTqGsLYkRl}\n}"},"title":{"value":"VideoAgent: All-in-One Agentic Framework for Video Understanding and Editing"},"pdf":{"value":"/pdf/d2174487eb6853747d5a8f5d7a4743a4ef1fd6b8.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"zhou|videoagent_allinone_agentic_framework_for_video_understanding_and_editing"},"authorids":{"value":["~Hengji_Zhou1","~Lingxuan_Huang1","~KunpengTan1","~Si_Wu5","~Lianghao_Xia1","~Chao_Huang4"]},"authors":{"value":["Hengji Zhou","Lingxuan Huang","KunpengTan","Si Wu","Lianghao Xia","Chao Huang"]}},"version":2},{"content":{"summary":{"value":"This paper proposes MomentSeg, a novel framework for referring video object segmentation (RVOS) that addresses the limitations of sparse and fixed sampling in existing methods. The core contribution is the Moment-Centric Sampling (MCS) mechanism, which intelligently selects frames most relevant to the referring expression to improve temporal grounding and visual context. Furthermore, the model employs a Bidirectional Anchor-updated Propagation (BAP) strategy to ensure robust and temporally consistent mask tracking across the selected video segment. The experimental results, validated across multiple challenging datasets, demonstrate that MomentSeg achieves state-of-the-art or highly competitive performance, confirming the effectiveness of the proposed sampling and propagation strategies."},"soundness":{"value":2},"confidence":{"value":4},"questions":{"value":"See the weakness."},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"1. Clear Motivation: The work provides a clear and compelling motivation, effectively highlighting the inherent limitations of conventional fixed or simple keyframe sampling approaches when dealing with moment-centric video-language tasks.\n2. Comprehensive Experiments and Strong Results: The experimental evaluation is notably comprehensive, providing ample technical detail and achieving excellent, state-of-the-art results across a diverse set of benchmark datasets.\n3. Well-Structured and Accessible Methodology: The overall methodology and presentation of the paper are clear, logical, and easy to follow, making the proposed MomentSeg architecture and techniques accessible to the broader reader community."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"1. Incremental Gain of MCS: The core idea of \"keyframe mining\" appears somewhat incremental, and the reported empirical gains are small; Figure 2 suggests the impact of the sampling strategy is marginal (around 2 points), and Table 7 confirms that the gain provided by MCS over the base keyframe methods is minimal.\n2. Similarity to Prior Grounding Work: The approach of using the [FIND] token to localize keyframes is highly similar to the established \"streaming EOS prediction\" mechanism found in related work, such as VideoLLM-Online, which diminishes the technical novelty of this specific component.\n3. Limited Applicability to Online Scenarios: Key components of the proposed method, including Moment-Centric Sampling (MCS) and Bidirectional Anchor-updated Propagation (BAP), fundamentally require access to the entire video's context, severely limiting their applicability to real-world online or streaming environments.\n4. Novelty of Bidirectional Propagation: While effective within this framework, the concept of bidirectional propagation for ensuring temporal consistency is not a novel technique in the broader Video Object Segmentation (VOS) literature and represents mainly an adaptation to the SAM model rather than a core innovation.\n5. Missing Robustness Analysis: The paper lacks a deep-dive analysis on the method's robustness, particularly how BAP performs when anchor frames contain significant noise or the model encounters high degrees of occlusion, which are common challenges in propagation-based models.\n\n[*] VideoLLM-online: Online Video Large Language Model for Streaming Video, CVPR2024."}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762915731920,"tcdate":1761835349804,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission1310/Reviewer_CUfj"],"signatures":["ICLR.cc/2026/Conference/Submission1310/Reviewer_CUfj"],"forum":"CxpKHWqT1n","number":3,"license":"CC BY 4.0","cdate":1761835349804,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission1310/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762915731920,"domain":"ICLR.cc/2026/Conference","replyto":"CxpKHWqT1n","id":"neYLZHosiu","forumContent":{"venue":{"value":"Submitted to ICLR 2026"},"keywords":{"value":["Large multi-modal model","Referring Video Object Segmentation","Temporal Sentence Grounding","Key frame sampling"]},"supplementary_material":{"value":"/attachment/cbd5fb8e44d6f3a43215cd22f6c96fd587802f7d.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"Referring Video Object Segmentation (RefVOS) seeks to segment target objects in videos guided by natural language descriptions, demanding both temporal reasoning and fine-grained visual comprehension. Existing sampling strategies for LLM-based approaches typically rely on either handcrafted heuristics or external keyframe models. The former often overlooks essential temporal cues, while the latter increases system complexity. To address this, we propose a unified framework that jointly optimizes Temporal Sentence Grounding (TSG) and RefVOS, naturally incorporating key moment grounding capability. During training, we introduce a novel TSG paradigm that employs a dedicated $\\texttt{[FIND]}$ token for key moment identification through temporal token similarity matching, thereby avoiding the need for external timestamp encodings. For inference, we design a Moment-Centric Sampling (MCS) strategy that densely samples informative moments while sparsely sampling non-essential frames, preserving both motion details and global context. To further enhance tracking stability, we develop Bidirectional Anchor-updated Propagation (BAP), which leverages the most relevant moment as start point for high-quality mask initialization and dynamically updates at sampled points to mitigate accumulated errors."},"_bibtex":{"value":"@misc{\ndai2026momentseg,\ntitle={MomentSeg: Moment-Centric Sampling for Enhanced Video Pixel Understanding},\nauthor={Ming Dai and Sen Yang and Boqiang Duan and Wankou Yang and Jingdong Wang},\nyear={2026},\nurl={https://openreview.net/forum?id=CxpKHWqT1n}\n}"},"title":{"value":"MomentSeg: Moment-Centric Sampling for Enhanced Video Pixel Understanding"},"pdf":{"value":"/pdf/1d4b3e8aecf923ccbfe8226bf27392c142b88f34.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Rejected_Submission"},"paperhash":{"value":"dai|momentseg_momentcentric_sampling_for_enhanced_video_pixel_understanding"},"authorids":{"value":["~Ming_Dai3","~Sen_Yang3","~Boqiang_Duan1","~Wankou_Yang1","~Jingdong_Wang1"]},"authors":{"value":["Ming Dai","Sen Yang","Boqiang Duan","Wankou Yang","Jingdong Wang"]}},"version":2},{"content":{"venue":{"value":"NeurIPS 2022"},"pdf":{"value":"https://proceedings.neurips.cc/paper_files/paper/2022/file/536d643875321d6c3282ee8c7ea5eb6a-Paper-Conference.pdf"},"venueid":{"value":"dblp.org/conf/NIPS/2022"},"paperhash":{"value":"zheng|causally_motivated_multishortcut_identification_and_removal"},"authorids":{"value":["~Jiayun_Zheng1","~Maggie_Makar1"]},"html":{"value":"http://papers.nips.cc/paper_files/paper/2022/hash/536d643875321d6c3282ee8c7ea5eb6a-Abstract-Conference.html"},"_bibtex":{"value":"@inproceedings{DBLP:conf/nips/ZhengM22,\n  author={Jiayun Zheng and Maggie Makar},\n  title={Causally motivated multi-shortcut identification and removal},\n  year={2022},\n  cdate={1640995200000},\n  url={http://papers.nips.cc/paper_files/paper/2022/hash/536d643875321d6c3282ee8c7ea5eb6a-Abstract-Conference.html},\n  booktitle={NeurIPS},\n  crossref={conf/nips/2022}\n}\n"},"abstract":{"value":"For predictive models to provide reliable guidance in decision making processes, they are often required to be accurate and robust to distribution shifts. Shortcut learning--where a model relies on spurious correlations or shortcuts to predict the target label--undermines the robustness property, leading to models with poor out-of-distribution accuracy despite good in-distribution performance. Existing work on shortcut learning either assumes that the set of possible shortcuts is known a priori or is discoverable using interpretability methods such as saliency maps, which might not always be true. Instead, we propose a two step approach to (1) efficiently identify relevant shortcuts, and (2) leverage the identified shortcuts to build models that are robust to distribution shifts. Our approach relies on having access to a (possibly) high dimensional set of auxiliary labels at training time, some of which correspond to possible shortcuts. We show both theoretically and empirically that our approach is able to identify a sufficient set of shortcuts leading to more efficient predictors in finite samples."},"title":{"value":"Causally motivated multi-shortcut identification and removal"},"authors":{"value":["Jiayun Zheng","Maggie Makar"]}},"tmdate":1762256183187,"pdate":1640995200000,"tcdate":1727781684624,"writers":["~"],"signatures":["~Maggie_Makar1"],"forum":"WV6zdzO3VS","license":"CC BY-SA 4.0","number":131082,"cdate":1640995200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit","DBLP.org/-/Author_Coreference"],"mdate":1762256183187,"domain":"DBLP.org","id":"WV6zdzO3VS","version":2},{"content":{"venue":{"value":"3DV 2026 Poster"},"keywords":{"value":["Diffusion Models","Camera-controlled Generation","Context-aware Generation"]},"supplementary_material":{"value":"/attachment/8647d9be40d85d87a935a3dbc80391484017debc.zip"},"abstract":{"value":"Recently, image-to-video (I2V) diffusion models have demonstrated impressive scene understanding and generative quality, incorporating image conditions to guide generation. However, these models primarily animate static images without extending beyond their provided context. Introducing additional constraints, such as camera trajectories, can enhance diversity but often degrade visual quality, limiting their applicability for tasks requiring faithful scene representation. We propose CamC2V, a context-to-video (C2V) model that integrates multiple image conditions as context with 3D constraints alongside camera control to enrich both global semantics and fine-grained visual details. This enables more coherent and context-aware video generation. Moreover, we motivate the necessity of temporal awareness for an effective context representation. Our comprehensive study on the RealEstate10K dataset demonstrates a $24.09\\%$ (FVD) improvement in visual quality and camera controllability. Our code is publicly available at: https://github.com/LDenninger/CamC2V."},"_bibtex":{"value":"@inproceedings{\ndenninger2026camcv,\ntitle={CamC2V: Context-aware Controllable Video Generation},\nauthor={Luis Denninger and Sina Mokhtarzadeh Azar and Juergen Gall},\nbooktitle={Thirteenth International Conference on 3D Vision},\nyear={2026},\nurl={https://openreview.net/forum?id=09WlFpCoc3}\n}"},"title":{"value":"CamC2V: Context-aware Controllable Video Generation"},"pdf":{"value":"/pdf/45b62ae2ecdab45bf7b954f23929f5d9180cfd3b.pdf"},"venueid":{"value":"3DV/2026/Conference"},"paperhash":{"value":"denninger|camc2v_contextaware_controllable_video_generation"},"authorids":{"value":["~Luis_Denninger1","~Sina_Mokhtarzadeh_Azar1","~Juergen_Gall1"]},"authors":{"value":["Luis Denninger","Sina Mokhtarzadeh Azar","Juergen Gall"]}},"tmdate":1769776566568,"pdate":1762379899362,"tcdate":1754846407951,"writers":["3DV/2026/Conference","3DV/2026/Conference/Submission137/Authors"],"signatures":["3DV/2026/Conference/Submission137/Authors"],"forum":"09WlFpCoc3","license":"CC BY 4.0","number":137,"cdate":1754846407951,"readers":["everyone"],"invitations":["3DV/2026/Conference/-/Submission","3DV/2026/Conference/-/Post_Submission","3DV/2026/Conference/Submission137/-/Supplementary_Material","3DV/2026/Conference/-/Edit","3DV/2026/Conference/Submission137/-/Post-Decision_Revision"],"mdate":1769776566568,"odate":1769776566534,"domain":"3DV/2026/Conference","id":"09WlFpCoc3","version":2},{"content":{"venue":{"value":"IEEE Trans. Circuits Syst. Video Technol. 2024"},"pdf":{"value":"https://ieeexplore.ieee.org/iel8/76/10811772/10621651.pdf"},"venueid":{"value":"dblp.org/journals/TCSV/2024"},"paperhash":{"value":"zhang|entity_dependency_learning_network_with_relation_prediction_for_video_visual_relation_detection"},"authorids":{"value":["https://dblp.org/search/pid/api?q=author:Guoguang_Zhang:","https://dblp.org/search/pid/api?q=author:Yepeng_Tang:","https://dblp.org/search/pid/api?q=author:Chunjie_Zhang:","~Xiaolong_Zheng4","https://dblp.org/search/pid/api?q=author:Yao_Zhao_0001:"]},"html":{"value":"https://doi.org/10.1109/TCSVT.2024.3437437"},"_bibtex":{"value":"@article{DBLP:journals/tcsv/ZhangTZZZ24,\n  author={Guoguang Zhang and Yepeng Tang and Chunjie Zhang and Xiaolong Zheng and Yao Zhao},\n  title={Entity Dependency Learning Network With Relation Prediction for Video Visual Relation Detection},\n  year={2024},\n  month={December},\n  cdate={1733011200000},\n  journal={IEEE Trans. Circuits Syst. Video Technol.},\n  volume={34},\n  number={12},\n  pages={12425-12436},\n  url={https://doi.org/10.1109/TCSVT.2024.3437437}\n}\n"},"abstract":{"value":"Video Visual Relation Detection (VidVRD) is a pivotal task in the field of video analysis. It involves detecting object trajectories in videos, predicting potential dynamic relation between these trajectories, and ultimately representing these relationships in the form of <subject, predicate, object> triplets. Correct prediction of relation is vital for VidVRD. Existing methods mostly adopt the simple fusion of visual and language features of entity trajectories as the feature representation for relation predicates. However, these methods do not take into account the dependency information between the relation predication and the subject and object within the triplet. To address this issue, we propose the entity dependency learning network(EDLN), which can capture the dependency information between relation predicates and subjects, objects, and subject-object pairs. It adaptively integrates these dependency information into the feature representation of relation predicates. Additionally, to effectively model the features of the relation existing between various object entities pairs, in the context encoding phase for relation predicate features, we introduce a fully convolutional encoding approach as a substitute for the self-attention mechanism in the Transformer. Extensive experiments on two public datasets demonstrate the effectiveness of the proposed EDLN."},"title":{"value":"Entity Dependency Learning Network With Relation Prediction for Video Visual Relation Detection"},"authors":{"value":["Guoguang Zhang","Yepeng Tang","Chunjie Zhang","Xiaolong Zheng","Yao Zhao"]}},"tmdate":1747151592589,"pdate":1704067200000,"tcdate":1747151589516,"writers":["~"],"signatures":["~Xiaolong_Zheng4"],"forum":"2xdLrbeQon","license":"CC BY-SA 4.0","number":448955,"cdate":1733011200000,"readers":["everyone"],"invitations":["DBLP.org/-/Record","DBLP.org/-/Edit"],"mdate":1747151592589,"domain":"DBLP.org","id":"2xdLrbeQon","version":2},{"content":{"summary":{"value":"This paper proposes Any2Caption, a framework that turns a vision–language model into a structured re-captioning model conditioned on multiple modalities (depth, pose, camera, identity, etc.). The method generates six-part structured captions (dense, object, background, action, style, camera) from arbitrary multimodal inputs, which can then be fed into existing text-to-video models to improve controllability and visual quality."},"soundness":{"value":2},"confidence":{"value":5},"questions":{"value":"- How is CLIP-T for long structured caption calculated (As CLIP textual encoder can not encode long sequence without losing information)? Is it still short-caption-video similarity score?"},"rating":{"value":4},"code_of_conduct":{"value":"Yes"},"presentation":{"value":3},"contribution":{"value":2},"strengths":{"value":"- The paper is well-written and clearly explains motivation and method.\n\n - The proposed structured-captioning paradigm is intuitive yet effective, successfully turning a general VLM into a condition-aware re-captioning model.\n\n- The implementation is efficient, requiring no modification to downstream video generators.\n\n- The results on multiple generators show that longer, structured captions can improve controllability and consistency to some extent."},"flag_for_ethics_review":{"value":["No ethics review needed."]},"weaknesses":{"value":"**Distribution Shift and Potential Suboptimal Improvement**\n\nThe paper evaluates multiple video generation models in a zero-shot setting, but the structured caption model was likely never exposed to them during training.\nSince different generators prefer different caption styles, the improvement may simply come from length improvement rather than genuine interpretive ability(as many video generators are trained on dense captions, and Any2Caption’s outputs are also long and verbose, which could coincidentally fit those models better).\nThis suggests that the observed improvement might be partly artefactual and potentially suboptimal rather than reflecting a true generalization capability.\n\n**Limited Necessity Demonstration**\n\nIn RQ1 (“Is the structured caption necessary?”), the authors intentionally choose the multi-ID condition, which is a case where structured rewriting may help the most due to better seperation of semantic entities among input identities (though such seperation can potentially be achieved without structured caption).\n\nThis makes this ablation reasonable but also biased: it only shows necessity for this task, potentially most favorable setting.\n\nThe same necessity is not demonstrated for other settings, where structured captions may be unnecessary or even redundant.\n\nThus, RQ1 provides only partial evidence and does not establish the claimed general necessity of structured captioning.\n\n\n**Inadequate Baselines and Limited Practical Impact**\n\nAs a paper essentially proposing a prompt enhancer, Any2Caption should be compared not only with short prompts but also with existing prompt enhancers. Most well-known open-sourced video generation models/projects already include prompt enhancers, often large LLMs fine-tuned or prompted via In-context learning to generate model-preferred prompts (e.g., CogVideoX’, Hunyuan’s, etc.). Additionally, as author mentioned, many video recaptioning methods are also proposed to enhance video generation ( ShareGPT4Video, InstanceCap, etc.)\nWithout such baselines, it remains unclear whether Any2Caption provides benefits over existing prompt optimization strategies.\n\nFrom a more practical standpoint, for instance, if a user is already employing models such as HunyuanVideo or CogVideoX, both of which feature built-in prompt enhancers optimized for their respective training data, it is not obvious why one would replace them with Any2Caption. In the absence of clear evidence of superior generalization, adaptability, or usability, the practical contribution and real-world impact of this work appear limited.\n\n**Unreliable Evaluation**\n\nThe evaluation pipeline is also not convincing to me.\n\nFor text/caption quality, Table 3 already shows that higher caption metrics do not correlate with better video results, undermining the relevance of Table 4.\n\nAlthough the appendix includes VDC benchmark results for captioning evaluation, the compared baselines are largely outdated, and notably, AuroraCap (introduced in the VDC paper itself and reported results better than Any2Caption) is missing. Thus VDC results is still not meaningful to me.\n\nFor video quality, most of the reported gains are minor and may fall within metric variance (many of these metrics are known to be not stable, e.g. aesthetic quality, etc), while qualitative examples in the paper are relatively limited considering the large amount of tasks the method claimed to tackle.\n\nI also checked the provided videos in the supplementary materials, but in many cases I can't find significant differences/improvements comparing the short caption version and the structured caption version (e.g, camera to video, ids + depth to video).\n\nGiven these issues, the evaluation is relatively weak and potentially misleading. A proper human study (e.g. voting) comparing videos from short prompts, existing enhancers, and Any2Caption-generated captions would provide much more credible evidence, and a more qualitative comparison is also needed per task.\n\n**Minor Issues**\n- The camera pose visualization uses overly large frustum cones, making motion changes almost invisible.\n\n- Several typos:\n    - Fig 1 “normal bae” → “normal base”\n    - Table 4 “METER” → “METEOR”\n    - L307 “access” → “assess”?\n    - Table 2 “vieo” → “video”"}},"parentInvitations":"ICLR.cc/2026/Conference/-/Official_Review","nonreaders":[],"tmdate":1762921774298,"tcdate":1761808733600,"writers":["ICLR.cc/2026/Conference","ICLR.cc/2026/Conference/Submission10483/Reviewer_WxPH"],"signatures":["ICLR.cc/2026/Conference/Submission10483/Reviewer_WxPH"],"forum":"XoT51yzqz7","number":2,"license":"CC BY 4.0","cdate":1761808733600,"readers":["everyone"],"invitations":["ICLR.cc/2026/Conference/Submission10483/-/Official_Review","ICLR.cc/2026/Conference/-/Edit"],"mdate":1762921774298,"domain":"ICLR.cc/2026/Conference","replyto":"XoT51yzqz7","id":"LEslfkbJiv","forumContent":{"venue":{"value":"ICLR 2026 Conference Withdrawn Submission"},"TLDR":{"value":"we present Any2Caption, a novel framework for controllable video generation from any condition."},"keywords":{"value":["Mulit-Modal Understanding","Video Caption","Controllable Video Generation"]},"supplementary_material":{"value":"/attachment/c4e52f0620c3c80aa9c10467c519ffdf4fd194e1.zip"},"primary_area":{"value":"applications to computer vision, audio, language, and other modalities"},"abstract":{"value":"To address the bottleneck of accurately interpreting user intent within the current video generation community, we present Any2Caption, a novel framework for controllable video generation from any condition. The key idea is decoupling various condition interpretation steps from the video synthesis step. By leveraging modern multimodal large language models (MLLMs), Any2Caption interprets diverse inputs—text, images, videos, and specialized cues such as region, motion, and camera poses—into dense, structured captions that offer backbone video generators with better guidance. We also introduce Any2CapIns, a large-scale dataset with 337K instances and 407K annotations for any-condition-to-caption instruction tuning. Comprehensive evaluations demonstrate significant improvements of our system in controllability and video quality across various aspects of existing video generation models."},"_bibtex":{"value":"@misc{\nwu2025interpreting,\ntitle={Interpreting Any Condition to Caption for Controllable Video Generation},\nauthor={Shengqiong Wu and Weicai Ye and Jiahao Wang and Quande Liu and Xintao Wang and Pengfei Wan and Di ZHANG and Kun Gai and Shuicheng YAN and Hao Fei and Tat-Seng Chua},\nyear={2025},\nurl={https://openreview.net/forum?id=XoT51yzqz7}\n}"},"title":{"value":"Interpreting Any Condition to Caption for Controllable Video Generation"},"pdf":{"value":"/pdf/6007d727cfcd1e79dc015c7c99c6ee43191f837a.pdf"},"venueid":{"value":"ICLR.cc/2026/Conference/Withdrawn_Submission"},"paperhash":{"value":"wu|interpreting_any_condition_to_caption_for_controllable_video_generation"},"authorids":{"value":["~Shengqiong_Wu2","~Weicai_Ye3","~Jiahao_Wang15","~Quande_Liu1","~Xintao_Wang1","~Pengfei_Wan1","~Di_ZHANG3","~Kun_Gai1","~Shuicheng_YAN3","~Hao_Fei1","~Tat-Seng_Chua2"]},"authors":{"value":["Shengqiong Wu","Weicai Ye","Jiahao Wang","Quande Liu","Xintao Wang","Pengfei Wan","Di ZHANG","Kun Gai","Shuicheng YAN","Hao Fei","Tat-Seng Chua"]}},"version":2}],"count":10000}